Automatic labeling method and system
By using the Patch Fusion model in the automatic labeling method to fuse the depth estimation results of the image and generate depth prompts, the problem of SAM model relying on manual labeling in the full-graphic semantic segmentation task is solved, and the effect of automatic labeling is achieved, ensuring the accuracy and completeness of labeling.
Patent Information
- Application Number
- CN202510096379.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-27
AI Technical Summary
In the semantic segmentation task of all objects in the whole image, the existing SAM model relies on automatic configuration of fixed and uniform dot matrix, resulting in missing potential connections of objects in the image and requiring manual participation for labeling.
Automatic annotation is completed by fusing the depth estimation results of high-resolution and low-resolution images using the Patch Fusion model in the automatic annotation method, a depth prompt is generated, and the image is embedded and the depth prompt is input into the mask decoder of the SAM model.
The purpose of automatic labeling is achieved, avoiding the problem of time-consuming and costly labeling information manually, ensuring that the complete object is not divided into multiple masks, and that multiple objects are not divided into the same block mask.
Smart Images

Figure CN120047946A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and particularly to an automatic annotation method and system. Background Art
[0002] The segmentation technology of autonomous driving vehicles mainly belongs to the field of computer vision, usually known as image segmentation or semantic segmentation, and sometimes also involves instance segmentation and panoramic segmentation. These technologies enable autonomous driving systems to parse meaningful information from visual data, such as different categories like roads, pedestrians, vehicles, traffic signs, etc.
[0003] Autonomous driving full-scene segmentation is crucial for ensuring vehicle safety, accurate navigation, and understanding complex traffic environments. It enables autonomous driving vehicles to identify and adapt to various situations on the road, including different objects, obstacles, and road surface conditions, which is the basis for achieving reliable autonomous driving. In short, full-scene segmentation is an important technology for autonomous vehicles to "see" and "understand" the world.
[0004] In autonomous driving, image segmentation technology is used to understand and interpret the environment around the vehicle. The data captured by the sensor cameras used in autonomous driving vehicles needs to be processed quickly and accurately to ensure safe and effective navigation. With the development of deep learning technology, especially the successful application of convolutional neural network CNN and transformer in the field of image recognition and processing, the image segmentation technology in the field of autonomous driving has also made significant progress. Modern autonomous driving systems usually adopt end-to-end deep learning models to perform the above-mentioned image segmentation tasks, and these models can be trained through large-scale datasets to improve their accuracy and robustness.
[0005] With the development of technologies such as large models like SAM (Segment Anything Model), deep learning models have made significant progress in processing complex visual tasks such as image segmentation. Through the training of a large amount of data, these models can better understand the scene and perform precise pixel-level classification, thus providing more reliable perception capabilities for applications such as autonomous driving systems. However, although the automation level of these models is very high, they still rely on human participation in some cases, such as providing annotation information during the training process, such as point box, mask, etc.
[0006] Figure 1 Shows the structural schematic diagram of the existing SAM model. As Figure 1As shown in the figure, the existing SAM model consists of three parts: a prompt encoder, an image encoder, and a lightweight mask decoder. Among them, the prompt encoder can accept four types of prompts, namely point sets, object bounding boxes, text, and masks, and output the final mask specifically through this prompt to complete the open-set semantic segmentation task. The input of the SAM model is divided into two parts: an image and a prompt. The image passes through the image encoder to obtain basic image features, called Image Embedding, while the prompt, according to the different types of input prompts, passes through different prompt encoders to obtain Prompt Embedding. Finally, the two inputs of the image and the prompt are sent into the mask decoding module together to complete the prediction of the final mask.
[0007] The existing SAM model is more suitable for open-set semantic segmentation tasks for partial objects, such as providing pre-detected object bounding boxes or completing segmentation specifically through text. For the semantic segmentation of all objects in the whole image, the existing SAM model is completed by automatically configuring a fixed and uniform dot matrix, which has nothing to do with the image content and misses the potential connections between objects in the image. Summary of the Invention
[0008] In view of some or all of the problems in the prior art, the present invention provides an automatic annotation method, which includes the following steps:
[0009] Input the collected original autonomous driving image;
[0010] Input the original autonomous driving image into the image encoder to obtain an image embedding;
[0011] Preprocess the original autonomous driving image to obtain a high-resolution image and a low-resolution image, where the resolution of the high-resolution image is greater than that of the low-resolution image;
[0012] Use the Patch Fusion model to fuse the high-resolution image and the low-resolution image to obtain a fused depth estimation result;
[0013] Input the fused depth estimation result into the depth prompt encoder to obtain a depth prompt; and
[0014] Input the image embedding and the depth prompt into the mask decoder of the SAM model to obtain the final mask, thereby completing automatic annotation.
[0015] Further, the preprocessing includes:
[0016] Image preparation, upsample the original autonomous driving image using the bilinear interpolation method to obtain a high-resolution image, and downsample the original autonomous driving image to obtain a low-resolution image; and
[0017] Image segmentation, according to the set slice size, perform horizontal and vertical slicing on the high-resolution image.
[0018] Further, using the Patch Fusion model, fuse the high-resolution image and the low-resolution image to obtain the depth estimation result, including:
[0019] Globally scale-aware depth estimation, use the global depth estimation network Nc to process the low-resolution image to obtain a rough depth estimation result Dc and the features Fc of each layer;
[0020] Chunk and crop the high-resolution image, input the cropped image into the image chunk depth estimation network Nf to obtain a fine depth estimation result Df and the features Ff of each layer; and
[0021] Input the rough depth estimation result Dc and the fine depth estimation result Df into the fusion network Ng to obtain a fused depth estimation result Dg and the features Fg of each layer.
[0022] Further, the input of the depth prompt encoder is the fused depth estimation result, and the tensor shape of the fused depth estimation result is 1xHxW.
[0023] Further, inputting the image embedding and the depth prompt into the mask decoder of the SAM model includes:
[0024] The depth prompt is input into the prompt encoder, and the prompt encoder inputs the aggregated prompt into the mask decoder.
[0025] Further, the aggregated prompt includes depth prompt, mask prompt, point set prompt, target box prompt, and text prompt.
[0026] The present invention also provides an automatic annotation system, which includes the following modules:
[0027] An image input module, configured to input the collected original autonomous driving image;
[0028] An image embedding module, configured to input the original autonomous driving image into an image encoder to obtain an image embedding;
[0029] A preprocessing module, configured to preprocess the original autonomous driving image to obtain a high-resolution image and a low-resolution image;
[0030] A depth estimation module, configured to use the Patch Fusion model to fuse the high-resolution image and the low-resolution image to obtain a fused depth estimation result;
[0031] A depth hint module, configured to input the fused depth estimation result into a depth hint encoder to obtain a depth hint; and
[0032] A mask output module, configured to input the image embedding and the depth hint into the mask decoder of the SAM model to obtain a final mask, thereby completing automatic annotation.
[0033] Furthermore, the depth hint is input into the hint encoder, and the hint encoder inputs the aggregated hints into the mask decoder;
[0034] The aggregated hints include depth hints, mask hints, point set hints, target box hints, and text hints.
[0035] The present invention also provides a computer system, including:
[0036] A processor configured to execute machine-readable instructions;
[0037] A graphics card with an artificial intelligence chip, configured to train the automatic annotation method; and
[0038] A memory configured to store machine-readable instructions, and the machine-readable instructions execute the steps of the automatic annotation method when executed by the processor and / or the graphics card.
[0039] The present invention also provides a computer-readable storage medium, on which machine-readable instructions are stored, and the machine-readable instructions execute the steps of the automatic annotation method when executed by the processor.
[0040] The technical solution provided by the present invention has the following advantages:
[0041] 1. The automatic annotation method proposed by the present invention successfully overcomes the problems of time-consuming, laborious, and high-cost of manual annotation information. With the help of automated annotation, the image set can be quickly customized and expanded according to actual needs, providing a rich sample source for model training.
[0042] 2. The automatic annotation method proposed by the present invention gives the SAM model a certain degree of object information by adding depth hints to the hint decoder, so that complete objects are no longer cut into multiple masks, and multiple objects are not segmented into the same mask, achieving the purpose of automatic annotation. Description of the Drawings
[0043] To further clarify the above and other advantages and features of the embodiments of the present invention, a more specific description of the embodiments of the present invention will be presented with reference to the accompanying drawings. It can be understood that these drawings only depict typical embodiments of the present invention and thus will not be considered as limiting its scope. In the drawings, for clarity, the same or corresponding components will be denoted by the same or similar reference numerals.
[0044] Figure 1 The structural schematic diagram of the existing SAM model is shown;
[0045] Figure 2 The flowchart of the automatic annotation method according to an embodiment of the present invention is shown;
[0046] Figure 3 The principle schematic diagram of the automatic annotation method according to an embodiment of the present invention is shown;
[0047] Figure 4 The structural schematic diagram of the Patch Fusion model according to an embodiment of the present invention is shown;
[0048] Figure 5 The structural schematic diagram of the guiding fusion network according to an embodiment of the present invention is shown;
[0049] Figure 6 The structural schematic diagram of the depth hint encoder according to an embodiment of the present invention is shown;
[0050] Figure 7 The autonomous driving image of a scene is shown;
[0051] Figure 8 Shows the existing SAM model for Figure 7 The schematic diagram of the annotation result of the shown image;
[0052] Figure 9 Shows the fusion depth estimation result of the automatic annotation method provided by the present invention for Figure 7 the shown image;
[0053] Figure 10 Shows the automatic annotation method provided by the present invention for Figure 7 the schematic diagram of the annotation result of the shown image; and
[0054] Figure 11 The schematic diagram of the automatic annotation system according to an embodiment of the present invention is shown. Detailed implementation manners
[0055] In the following description, the present invention is described with reference to various embodiments. However, those skilled in the art will recognize that the embodiments can be implemented without one or more specific details or in conjunction with other alternative and / or additional methods or components. In other instances, well-known structures or operations are not shown or described in detail so as not to obscure the inventive aspects of the present invention. Similarly, for purposes of explanation, specific numbers and configurations are set forth in order to provide a thorough understanding of the embodiments of the present invention. However, the present invention is not limited to these specific details.
[0056] In this specification, the reference to "an embodiment" or "the embodiment" means that the particular features, structures, or characteristics described in connection with the embodiment are included in at least one embodiment of the present invention. The phrase "in an embodiment" appearing throughout this specification does not necessarily all refer to the same embodiment.
[0057] It should be noted that the embodiments of the present invention describe the method steps in a specific order. However, this is only for the purpose of illustrating the specific embodiment and does not limit the order of the steps. On the contrary, in different embodiments of the present invention, the order of the steps can be adjusted according to the actual requirements.
[0058] In the present invention, the various modules of the system according to the present invention can be implemented using software, hardware, firmware, or a combination thereof. When a module is implemented using software, the function of the module can be realized through a computer program process. For example, the module can be realized by a code segment (such as a code segment in languages like C, C++) stored in a storage device (such as a hard disk, memory, etc.), where when the code segment is executed by a processor, the corresponding function of the module can be realized. When a module is implemented using hardware, the function of the module can be realized by setting up the corresponding hardware structure. For example, the function of the module can be realized by hardware programming of a programmable device such as a field programmable gate array (FPGA), or by designing an application specific integrated circuit (ASIC) including multiple electronic devices such as transistors, resistors, and capacitors. When a module is implemented using firmware, the function of the module can be written in a read-only memory such as an EPROM or EEPROM of the device in the form of program code, and when the program code is executed by a processor, the corresponding function of the module can be realized. Additionally, certain functions of the module may need to be realized by separate hardware or in cooperation with the hardware. For example, the detection function is realized by a corresponding sensor (such as a proximity sensor, an acceleration sensor, a gyroscope, etc.), the signal emission function is realized by a corresponding communication device (such as a Bluetooth device, an infrared communication device, a baseband communication device, a Wi-Fi communication device, etc.), the output function is realized by a corresponding output device (such as a display, a speaker, etc.), and so on.
[0059] In the present invention, the resolution of the high-resolution image is 2160x3840 (HxW), and the resolution of the low-resolution image is 392x518 (HxW).
[0060] Aiming at the problem that the existing SAM model still relies on manual participation in some cases, such as manual annotation of information, which is time-consuming, laborious and costly, the present invention provides an automatic annotation method. By adding depth prompts to the prompt decoder, the SAM model is given certain object information to a certain extent, so that a complete object is no longer cut into multiple masks, and multiple objects are not segmented into the same mask, and automatic annotation can be achieved.
[0061] Figure 2 The flowchart of the automatic annotation method according to an embodiment of the present invention is shown. Figure 3 The schematic diagram of the principle of the automatic annotation method according to an embodiment of the present invention is shown. The following combines Figure 2 and Figure 3 , to illustrate the automatic annotation method proposed by the present invention. In an embodiment of the present invention, the automatic annotation method can be executed by a computer. The automatic annotation method includes the following steps:
[0062] First, input the collected original autonomous driving image. In an embodiment of the present invention, the original autonomous driving image can be obtained through a mass-produced autonomous driving system, where the mass-produced autonomous driving system is configured with sensor devices, and the original autonomous driving image is constructed by receiving the information collected by the sensor devices. Among them, the sensor devices include image sensors, radar sensors, GPS / IMU, etc. In an embodiment of the present invention, the radar sensor includes millimeter-wave radar, lidar, etc.
[0063] Next, input the original autonomous driving image into the image encoder to obtain an image embedding.
[0064] Next, preprocess the original autonomous driving image to obtain a high-resolution image and a low-resolution image, where the resolution of the high-resolution image is greater than that of the low-resolution image. The preprocessing includes image preparation and image segmentation. Image preparation is: using the bilinear interpolation method to upsample the original autonomous driving image to obtain a high-resolution image for subsequent slicing, and downsampling the original autonomous driving image to obtain a low-resolution image for rough depth estimation; image segmentation is: according to the set slice size, the high-resolution image is horizontally and vertically sliced.
[0065] Next, use the Patch Fusion model to fuse the high-resolution image and the low-resolution image to obtain a fused depth estimation result.
[0066] The Patch Fusion model is a deep learning architecture for monocular depth estimation. Its idea is to fuse depth estimation maps at high and low resolutions using global context information, thereby achieving fine depth estimation at high resolution. Figure 4 FIG. shows a schematic structural diagram of the Patch Fusion model according to an embodiment of the present invention. As Figure 4 shown, the Patch Fusion model consists of three parts: a global depth estimation network Nc for obtaining global context information, a patch depth estimation network Nf for obtaining refined depth estimation results on each image patch, and a fusion network Ng with global guidance information.
[0067] Using the Patch Fusion model to fuse the high-resolution image and the low-resolution image, the depth estimation result includes:
[0068] Globally scale-aware depth estimation. Process the low-resolution image according to the model input size, use the global depth estimation network Nc to process the low-resolution image to obtain a rough depth estimation result Dc, and at the same time output the features Fc of each layer for subsequent processing. This result lacks a lot of high-frequency detail information but retains good global information;
[0069] Cut the high-resolution image into blocks according to the model input size, input the cut image into the patch depth estimation network Nf to obtain the depth estimation result of each block, that is, the refined depth estimation result Df, and at the same time output the features Ff of each layer. This result has rich details in the image texture, but due to the inference being performed in blocks, there will be a phenomenon of depth discontinuity between blocks;
[0070] Input the rough depth estimation result Dc and the refined depth estimation result Df into the fusion network Ng to obtain the fused depth estimation result Dg and the features Fg of each layer.
[0071] The fusion network Ng is a fusion network with fusion and guidance, including a G2L (Global-to-Local) module and a guided fusion network. The G2L module uses the STL (Swin Transformer Layer) layer to process the rough global features Fc. The STL can be divided into W-SA (localized windows for self-attention) and SW_SA (shifted-window attention).
[0072] FIG. 5 shows a schematic structural diagram of the guided fusion network according to an embodiment of the present invention. As Figure 5As shown in the figure, the input of the guiding fusion network is the rough depth estimation result Dc and the fine depth estimation result Df of the corresponding patches of the cropped image. The guiding fusion network is divided into three components. Component A is a continuous encoding layer, component B is a skip-connection module, and component C is an upsampling layer.
[0073] Next, the fused depth estimation result is input into the depth prompt encoder to obtain a depth prompt.
[0074] As Figure 2 shown in the figure, the present invention adds a prompt encoder, namely a depth prompt encoder, to the SAM model. The present invention improves the semantic segmentation effect of the SAM model for the whole image by introducing a monocular depth estimation map, thereby completing the automatic annotation task. The structure of the depth prompt encoder is the same as that of the original mask prompt encoder, except for the input part. The input of the mask prompt encoder is usually a single-channel binary image with the same size as the image, and the tensor shape is 1xHxW, and the values it contains are only 0 and 1. The input of the depth prompt encoder is the depth estimation result of the above Patch Fusion model, that is, the fused depth estimation result, and the tensor shape of the fused depth estimation result is 1xHxW, and the values it contains are represented by 32-bit floating-point numbers. Figure 6 The structural schematic diagram of the depth prompt encoder according to an embodiment of the present invention is shown. As Figure 6 shown in the figure, the depth prompt encoder consists of three convolutional layers and corresponding normalization and activation layers. LayerNorm is used for normalization, and GELU is used for activation. The first two of the three convolutional layers are 3x3 convolutions, which output 1024 channels and 256 channels respectively. The last convolutional layer is a 1x1 convolution, which outputs 1024 channels, and the output result is reshaped into a one-dimensional vector.
[0075] Finally, the image embedding and the depth prompt are input into the mask decoder of the SAM model to obtain the final mask, thereby completing the automatic annotation.
[0076] In order to enable the mask decoder of the SAM model to accept the depth estimation result, the present invention fine-tunes the existing SAM model, and the training process is consistent with the inference. As Figure 3As shown, for the existing SAM model, the rest remains unchanged, and depth cues are added, causing the prompt decoder and mask output to change after training. The depth cues are input into the prompt encoder, which then inputs the aggregated cues into the mask decoder. The aggregated cues include depth cues, mask cues, point set cues, object bounding box cues, and text cues. The original autonomous driving image is respectively input into the image encoder and the Patch Fusion model to obtain the image embedding and the fused depth estimation result. Then, the fused depth estimation result is fed into the depth cue encoder to obtain the depth cues. Finally, the depth cues and the image embedding are together fed into the mask decoder of the SAM model. During fine-tuning training, only the weights of the prompt encoder are updated, and the weights of the other parts of the SAM model are not involved in the update. At the same time, the loss function and other hyperparameters do not need to be changed.
[0077] Compared with the existing SAM model, the fine-tuned SAM model, due to the addition of depth cues, gives the model object information to a certain extent, so that a complete object is no longer cut into multiple masks, and multiple objects will not be segmented into the same mask, enabling the purpose of automatic annotation.
[0078] The effect of the fault diagnosis method provided by the present invention can be further illustrated by the following experimental results.
[0079] Figure 7 An autonomous driving image of a scene is shown. Figure 8 Shows the existing SAM model for Figure 7 The schematic diagram of the annotation result of the shown image. Figure 9 Shows the fused depth estimation result schematic diagram of the image by the automatic annotation method provided by the present invention for Figure 7 the shown image. Figure 10 Shows the automatic annotation method provided by the present invention for Figure 7 the schematic diagram of the annotation result of the shown image. As Figure 8 shown, the existing SAM model cuts a complete object into multiple masks, such as buildings, roads, and vehicles, etc., thus unable to accurately annotate the object. As Figure 9 shown, the scale on the right represents the normalized depth value. The smaller the depth value, the closer the depth distance from the shooting camera, and the closer the color on the graph is to blue; the larger the depth value, the farther the depth distance from the shooting camera, and the closer the color on the graph is to yellow. The formula for the normalized depth is: normalized depth = output depth / output maximum depth. The output depth of the model is the actual scale in the real world, with the unit of meter. As Figure 10As shown, the fine-tuned SAM model provided by the present invention, due to the addition of depth cues obtained from the fused depth estimation results, gives the model a certain degree of object information, so that a complete object is no longer cut into multiple masks, and multiple objects will not be segmented into the same mask, and the purpose of automatic annotation can be achieved.
[0080] The automatic annotation method proposed by the present invention overcomes the problems of time-consuming, laborious and high cost of manual annotation information; by adding depth cues to the prompt decoder, it gives the SAM model a certain degree of object information, so that a complete object is no longer cut into multiple masks, and multiple objects will not be segmented into the same mask, and the purpose of automatic annotation can be achieved.
[0081] In an embodiment of the present invention, the present invention also provides an automatic annotation system. Figure 11 The schematic diagram of the automatic annotation system according to an embodiment of the present invention is shown. As Figure 11 shown, the system includes the following modules:
[0082] An image input module, configured to input the collected original autonomous driving image;
[0083] An image embedding module, configured to input the original autonomous driving image into an image encoder to obtain an image embedding;
[0084] A preprocessing module, configured to preprocess the original autonomous driving image to obtain a high-resolution image and a low-resolution image;
[0085] A depth estimation module, configured to use a Patch Fusion model to fuse the high-resolution image and the low-resolution image to obtain a fused depth estimation result;
[0086] A depth cue module, configured to input the fused depth estimation result into a depth cue encoder to obtain a depth cue; and
[0087] A mask output module, configured to input the image embedding and the depth cue into the mask decoder of the SAM model to obtain a final mask, thereby completing automatic annotation.
[0088] In an embodiment of the present invention, the depth cue is input into a cue encoder, and the cue encoder inputs the aggregated cues into the mask decoder; the aggregated cues include depth cues, mask cues, point set cues, target box cues, and text cues.
[0089] In an embodiment of the present invention, the present invention further provides a computer system, which includes a processor, a graphics card with an artificial intelligence chip, and a memory. The memory is configured to store machine-readable instructions, the graphics card is configured to train the automatic annotation method, and the processor is configured to execute the machine-readable instructions. When the processor and / or the graphics card execute the machine-readable instructions, the following processing steps are implemented: input the collected original autonomous driving image; input the original autonomous driving image into an image encoder to obtain an image embedding; preprocess the original autonomous driving image to obtain a high-resolution image and a low-resolution image, where the resolution of the high-resolution image is greater than that of the low-resolution image; use the PatchFusion model to fuse the high-resolution image and the low-resolution image to obtain a fused depth estimation result; input the fused depth estimation result into a depth prompt encoder to obtain a depth prompt; input the image embedding and the depth prompt into the mask decoder of the SAM model to obtain a final mask, thereby completing the automatic annotation.
[0090] The graphics card may preferably be a graphics card with a GPU computing power higher than model 5.0. Since the amount of data to be trained is large, providing the graphics card configuration can significantly improve the training speed.
[0091] The memory includes: various media that can store machine-readable instructions, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disc.
[0092] It can be understood that in addition to the memory and the processor described above, the above computer system further includes other software and hardware components not listed in this specification. Specifically, it can be determined according to the model of the specific data processing device in different application scenarios, and this specification will not list and elaborate one by one.
[0093] In an embodiment of the present invention, the present invention further provides a computer-readable storage medium, on which machine-readable instructions are stored. When the machine-readable instructions are executed by a processor, the following processing steps are implemented: input the collected original autonomous driving image; input the original autonomous driving image into an image encoder to obtain an image embedding; preprocess the original autonomous driving image to obtain a high-resolution image and a low-resolution image, where the resolution of the high-resolution image is greater than that of the low-resolution image; use the Patch Fusion model to fuse the high-resolution image and the low-resolution image to obtain a fused depth estimation result; input the fused depth estimation result into a depth prompt encoder to obtain a depth prompt; input the image embedding and the depth prompt into the mask decoder of the SAM model to obtain a final mask, thereby completing the automatic annotation.
[0094] Although the embodiments of the present invention have been described above, it should be understood that they are presented by way of example only and not as a limitation. It will be apparent to those skilled in the relevant art that various combinations, variations, and changes can be made thereto without departing from the spirit and scope of the present invention. Therefore, the breadth and scope of the present invention disclosed herein should not be limited by the exemplary embodiments disclosed above, but should be defined in accordance with the technical solutions of the present invention and their equivalents.
Claims
1. An automatic labeling method, characterized in that: The steps include: Input the collected original autonomous driving images; Inputting the original autonomous driving image into an image encoder to obtain an image embedding; Preprocessing the original autonomous driving image to obtain a high-resolution image and a low-resolution image, wherein a resolution of the high-resolution image is greater than a resolution of the low-resolution image; Using a Patch Fusion model, the high-resolution image and the low-resolution image are fused to obtain a fused depth estimation result; Inputting the fused depth estimation result into a depth cue encoder to obtain a depth cue; as well as The image embedding and the depth cue are input into the mask decoder of the SAM model to obtain the final mask, thereby completing the automatic labeling.
2. The automatic marking method according to claim 1, characterized in that: The pre-processing comprises: Image preparation, upsampling the original autonomous driving image using a bilinear interpolation method to obtain a high-resolution image, and downsampling the original autonomous driving image to obtain a low-resolution image; and Image segmentation: According to the set slice size, the high-resolution image is segmented horizontally and vertically.
3. The automatic marking method according to claim 1, characterized in that: Using the Patch Fusion model, the high-resolution image and the low-resolution image are fused to obtain a depth estimation result including: Global scale-aware depth estimation, using a global depth estimation network Nc to process the low-resolution image to obtain a rough depth estimation result Dc and features Fc of each layer; Cut the high-resolution image into blocks, input the cut images into an image block depth estimation network Nf, and obtain a fine depth estimation result Df and a feature Ff of each layer; and The rough depth estimation result Dc and the fine depth estimation result Df are input into the fusion network Ng to obtain the fused depth estimation result Dg and the feature Fg of each layer.
4. The automatic marking method according to claim 1, characterized in that: The input of the depth hint encoder is the fused depth estimation result, and the tensor shape of the fused depth estimation result is 1xHxW.
5. The automatic marking method according to claim 1, characterized in that: The image embedding and the depth cue are fed into the mask decoder of the SAM model including: The depth cues are input to a cue encoder, which inputs the summarized cues to a mask decoder.
6. The automatic marking method according to claim 5, characterized in that: The summarized prompts include depth prompts, mask prompts, point set prompts, target box prompts and text prompts.
7. A system for implementing the automatic marking method according to any one of claims 1 to 6, characterized in that: include: An image input module, configured to input a collected original autonomous driving image; An image embedding module is configured to input the original autonomous driving image into an image encoder to obtain an image embedding; A preprocessing module is configured to preprocess the original autonomous driving image to obtain a high-resolution image and a low-resolution image; A depth estimation module is configured to use a Patch Fusion model to fuse the high-resolution image and the low-resolution image to obtain a fused depth estimation result; A depth cue module, configured to input the fused depth estimation result into a depth cue encoder to obtain a depth cue; as well as The mask output module is configured to input the image embedding and the depth hint into the mask decoder of the SAM model to obtain a final mask, thereby completing automatic labeling.
8. The automatic labeling system according to claim 7, characterized in that: The depth hint is input into a hint encoder, and the hint encoder inputs the summarized hint into a mask decoder; The summarized prompts include depth prompts, mask prompts, point set prompts, target box prompts and text prompts.
9. A computer system, characterized in that: include: a processor configured to execute machine-readable instructions; A graphics card with an artificial intelligence chip configured to train an automatic labeling method; as well as A memory configured to store machine-readable instructions, wherein the machine-readable instructions, when executed by a processor and / or a graphics card, perform the steps of the method according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that: Machine-readable instructions are stored thereon, and when the machine-readable instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are performed.