Three-dimensional object detection method and device
By combining the image backbone, view transformer, and BEV encoder, the problems of high cost and insufficient accuracy of single view of LiDAR sensor are solved, and efficient and accurate 3D object detection is achieved.
Patent Information
- Application Number
- CN202510545250.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-07-26
- Filing Date
- 2025-04-28
- Publication Date
- 2025-11-18
AI Technical Summary
Existing 3D object detection technologies mainly rely on expensive LiDAR sensors or single-view estimation methods, resulting in high costs and insufficient accuracy of depth information.
The method extracts 2D image features from images using an image backbone, performs relative depth normalization and photometric matching using a view transformer, and combines a BEV encoder and a detection head to improve the accuracy and generalization ability of 3D object detection through multi-view image processing.
It reduces the cost of 3D object detection, improves the accuracy of depth information and the generalization performance of the model, and enhances the ability to recognize objects under different conditions.
Smart Images

Figure CN120976288A_ABST
Abstract
Description
Cross-reference to related applications
[0001] This application claims the benefit of Korean Patent Application No. 10-2024-0064025, filed on May 16, 2024, and Korean Patent Application No. 10-2024-0099581, filed on July 26, 2024, both of which have been filed with the Korean Intellectual Property Office, the entire disclosure of which is incorporated herein by reference for all purposes. Technical Field
[0002] The following description relates to methods and apparatus for three-dimensional object detection. Background Technology
[0003] 3D object detection typically involves using sensors (such as multiple cameras or light detection and ranging (LiDAR)) to collect 3D information about the surrounding environment and detecting objects based on the collected 3D information. By identifying other vehicles, pedestrians, obstacles, etc., 3D object detection is crucial for the safe operation of autonomous vehicles or robots.
[0004] Recent 3D object detection techniques primarily utilize expensive sensors such as LiDAR, or methods that estimate 3D information based on a single view. However, LiDAR is costly and involves complex data processing, while a single view can reduce the accuracy of depth information. Therefore, 3D object detection methods using multi-view images can help address these issues. Summary of the Invention
[0005] This summary is intended to introduce some concepts in a simplified form, which will be further described in the detailed description below. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to assist in determining the scope of the claimed subject matter.
[0006] In one general aspect, a method for detecting three-dimensional (3D) objects includes: extracting two-dimensional (2D) image features from an image using an image backbone; extracting a 3D feature map reflecting depth prediction information from the 2D image features using a view transformer configured to perform domain generalization; extracting BEV features from the 3D feature map using a bird's-eye view (BEV) encoder; and using a detection head to predict the location and category of the 3D object based on the BEV features.
[0007] Extracting 3D feature maps can include: predicting depth output from 2D image features using DepthNet, and inputting the outer product of DepthNet's depth output and 2D image features into the BEV pool.
[0008] The view transformer can be configured to perform a relative depth normalization method that minimizes depth and position prediction errors caused by differences in intrinsic / extrinsic parameters of the camera providing one of the images.
[0009] Multiple cameras, including the aforementioned cameras, can provide corresponding images, and the relative depth normalization method can include calculating a transformation matrix by which geometric transformations are performed between adjacent camera pairs based on intrinsic / extrinsic parameters.
[0010] The relative depth normalization method can obtain the relative depth by projecting image features onto neighboring image features using depth prediction information and transformation matrix, and then minimizing the relative depth loss based on the depth loss function.
[0011] The view transformer can be configured to perform a photometric matching method using depth prediction to optimize the alignment between the image and its neighboring images based on the photometric matching method.
[0012] The image backbone, view transformer, BEV encoder, and / or detection head may include corresponding domain adaptive adapters.
[0013] Each domain adaptive adapter can be added in parallel to the computation block to enable fine-tuning of parameters.
[0014] Each domain adaptive adapter can be configured to perform a skip connection in which features input to the view transformer, BEV encoder, and / or detector head are received, computed, and summed to update the gradient.
[0015] The method may also include: enhancing the 3D feature map by performing a generalization method based on decoupled image depth estimation.
[0016] In another general aspect, an electronic device includes: a memory storing instructions; and one or more processors, wherein, when the instructions are executed by the one or more processors, the one or more processors cause the one or more processors to: extract two-dimensional (2D) image features from an image using an image backbone, extract a 3D feature map reflecting depth prediction information from the 2D image features using a view transformer, extract BEV features from the 3D feature map using a bird's-eye view (BEV) encoder, and predict the location and category of an object based on the BEV features using a detection head.
[0017] When the instructions are executed by one or more processors, the processors may perform the following operations to extract 3D feature maps: the DepthNet predicts the depth output based on the 2D image features, and the outer product of the DepthNet's depth output and the 2D image features is input into the BEV pool.
[0018] The view transformer can be configured to perform a relative depth normalization method that minimizes depth and position prediction errors caused by differences in intrinsic / extrinsic parameters of the camera providing one of the images.
[0019] Multiple cameras, including the aforementioned cameras, can provide corresponding images, and the relative depth normalization method can include calculating a transformation matrix by which geometric transformations are performed between adjacent camera pairs based on intrinsic / extrinsic parameters.
[0020] The relative depth normalization method can obtain the relative depth by projecting image features onto adjacent image features using the depth prediction information and the transformation matrix, and by minimizing the relative depth loss based on the depth loss function.
[0021] The view transformer can be configured to perform a photometric matching method using depth prediction to optimize the alignment between the image and its neighboring images based on the photometric matching method.
[0022] The image backbone, view transformer, BEV encoder, and / or detection head can have corresponding domain adaptive adapters.
[0023] Domain adaptive adapters can be added in parallel to layers in the image backbone, view transformer, BEV encoder, and / or detection head.
[0024] Each domain adaptive adapter can be configured to perform a skip connection in which features input to the view transformer, BEV encoder, and / or detector head are received, computed, and summed to update the gradient.
[0025] When the instructions are executed by one or more processors, they can enable one or more processors to augment 3D feature maps by executing a generalization method based on decoupled image depth estimation.
[0026] Other features and aspects will become clear from the following detailed description, drawings and claims. Attached Figure Description
[0027] Figure 1 An example of a three-dimensional (3D) object detection method according to one or more embodiments is shown.
[0028] Figure 2 An example of a 3D object detection device according to one or more embodiments is shown.
[0029] Figure 3 and Figure 4 Each illustrates the operation of a view transformer according to one or more embodiments.
[0030] Figure 5An example of an adapter according to one or more embodiments is shown.
[0031] Figure 6 Example operation of an adapter according to one or more embodiments is shown.
[0032] Figure 7 An example of a generalized method for decoupled image depth estimation based on one or more embodiments is shown.
[0033] Figure 8 An example of an electronic device according to one or more embodiments is shown.
[0034] Throughout the accompanying drawings and detailed description, unless otherwise described or provided, the same or similar reference numerals shall be understood to refer to the same or similar elements, features, and structures. The drawings may not be drawn to scale, and for clarity, illustration, and convenience, the relative sizes, proportions, and depictions of elements in the drawings may be exaggerated. Detailed Implementation
[0035] The following detailed description is intended to help the reader fully understand the methods, apparatus, and / or systems described herein. However, various variations, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding the disclosure of this application. For example, the order of operations described herein is merely an example and is not limited to the order described herein, but can be changed as will become apparent upon understanding the disclosure of this application, except where the operations must occur in a specific order. Furthermore, for clarity and conciseness, descriptions of features known upon understanding the disclosure of this application may be omitted.
[0036] The features described herein may be embodied in various forms and should not be construed as being limited to the examples described herein. Rather, the examples described herein are merely illustrative of some of the many possible ways in which the methods, apparatus, and / or systems described herein will become apparent upon understanding the disclosure of this application.
[0037] The terminology used herein is for the purpose of describing various examples only and is not intended to limit this disclosure. The articles “a,” “an,” and “the” are also intended to include plural forms unless the context clearly indicates otherwise. The term “and / or” as used herein includes any one and any combination of any two or more of the related listed items. As a non-limiting example, the terms “comprising,” “including,” and “having” indicate the presence of the described features, numbers, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, numbers, operations, components, elements, and / or combinations thereof.
[0038] Throughout this specification, when a component or element is described as being "connected to," "coupled to," or "joined to" another component or element, that component or element may be directly "connected to," "coupled to," or "joined to" that other component or element, or one or more other components or elements may reasonably be present between them. When a component or element is described as being "directly connected to," "directly coupled to," or "directly joined to" another component or element, there are no other elements between them. Similarly, expressions such as "between" and "immediately following" and "adjacent" and "next to" may also be interpreted as described above.
[0039] While this document may use terms such as “first,” “second,” and “third,” or A, B, (a), (b), to describe various components, assemblies, regions, layers, or parts, these components, assemblies, regions, layers, or parts are not limited by these terms. Each such term is not intended to define the nature, order, or sequence of the corresponding component, assembly, region, layer, or part, but only to distinguish the corresponding component, assembly, region, layer, or part from other components, assemblies, regions, layers, or parts. Therefore, the first component, assembly, region, layer, or part mentioned in the examples described herein may also be referred to as the second component, assembly, region, layer, or part without departing from the teachings of the examples.
[0040] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains based on an understanding of the disclosure of this application. Terms (such as those defined in common dictionaries) should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology and the disclosure of this application, and should not be interpreted in an idealized or overly formal sense unless explicitly defined herein. The use of the term "may" (e.g., regarding what an example or embodiment may include or implement) in connection with an example or embodiment means that there exists at least one example or embodiment that includes or implements such a feature, but not all examples are limited thereto.
[0041] Figure 1 An example of a three-dimensional (3D) object detection method according to one or more embodiments is shown.
[0042] Operations 110 to 140 can be performed by Figure 8 The electronic device 800 shown, or any other suitable electronic device in any suitable system, shall perform this action.
[0043] Electronic device 800 may include 3D object detection device 200. (See reference...) Figure 1 Describe operations 110 to 140.
[0044] Figure 2An example of a 3D object detection device according to one or more embodiments is shown.
[0045] Reference Figure 2 One or more blocks and combinations thereof may be implemented by a computer based on dedicated hardware that performs a predetermined function and / or by a combination of computer instructions and general-purpose hardware.
[0046] Simultaneously refer to Figure 1 and Figure 2 The electronic device 800 (e.g., 3D object detection device 200) may include an image backbone 210, a view transformer 220, a bird's-eye view (BEV) encoder 230, and a detection head 240. The view transformer 220 may include a DepthNet 221 and a BEV pool 222. The image backbone 210, view transformer 220, BEV encoder 230, and detection head 240 may be implemented as corresponding neural network models.
[0047] In operation 110, the 3D object detection device 200 / electronic device 800 can use the image backbone 210 to extract 2D image features 211-1 from the image 201 received by the corresponding camera. The image 201 can come from multiple viewpoints of the corresponding camera. For example, the image 201 may include images from the front, left front, right front, rear, left rear, right rear, or other camera viewpoints.
[0048] The 3D object detection device 200 / electronic device 800 can perform camera parameter enhancement, which can address the problem of intrinsic / extrinsic camera parameter biases of any camera. For any image from the camera, the image scale, image parameters, and image bounding box scale can be randomly transformed during data enhancement. It is a set of matrices containing the intrinsic (internal) parameters of n corresponding arbitrary cameras. The i-th element (for the i-th camera) is represented by Equation 1.
[0049] Equation 1
[0050] In equation 1, These represent the focal lengths in the x and y directions, respectively. Let u and v represent the center pixel coordinates of the i-th camera, and u and v represent the pixel coordinates in the image plane. Electronic device 800 can multiply the camera's intrinsic parameter matrix by a scaling factor expressed in homogeneous coordinates. ,Will Convert to a randomly scaled matrix ,Right now Camera intrinsic parameters represent the characteristics of the camera itself (e.g., these characteristics are generally the same regardless of how the camera is mounted in a vehicle).
[0051] The set of camera extrinsic parameter matrices (e.g., This includes a camera extrinsic parameter matrix, each matrix being... The extrinsic parameters are those that can vary for a given camera (e.g., they can vary depending on the mounting method on each vehicle). Here, R represents rotation, and t represents translation. Electronics 800 can randomly apply rescaling and / or shifting to parameters related to the camera mounting. (yaw, pitch, roll) and / or (height) is used to perform data augmentation on the camera's external parameters or external information. In short, camera external parameters can represent the camera's position and orientation.
[0052] The transformation matrix of the i-th camera is discussed below with reference to Equation 3. This matrix is based on its intrinsic matrix. and external matrix .
[0053] By using data augmented to train an object recognition model, the model can learn camera information under different conditions, thereby improving its generalization performance and adaptability.
[0054] According to an embodiment, the image backbone 210 can be an image feature extractor that receives image 201 and extracts 2D image features 211-1. The 2D image features 211-1 can include visual information that can be used to detect objects. Here, image features can be inferred jointly from various images, but processed in a way that combines them into a unified representation. Specifically, each image can contribute its individual features (e.g., features extracted using the image backbone), and these features can then be aggregated or transformed to represent relationships between instances of the same object in different views or images. This approach enables the detection of objects in multiple images while preserving their contextual and spatial information.
[0055] In operation 120, the electronic device 800 can extract a 3D feature map 222-1 that reflects the predicted depth information. The 3D feature map 222-1 can be extracted from the 2D image features 211-1 using a view transformer 220 that provides domain generalization. The view transformer 220 can perform domain generalization through a relative depth normalization method and 2D red, green, and blue (RGB) matching.
[0056] According to some embodiments, the view transformer 220 can extract / infer 3D feature maps 222-1 by: (a) predicting depth information (e.g., depth distribution prediction result 212-1) using DepthNet 221 (which predicts depth information based on 2D image features 211-1); and (b) inputting the result of the outer product of (i) the output (depth information) of DepthNet 221 and (ii) the corresponding 2D image features 211-1 into the BEV pool 222. Details of the view transformer 220 will now be described.
[0057] Bird's-eye view (BEV) refers to a visualization method (or data form) commonly used for vehicles or robots, and can involve projecting 3D information onto a 2D plane as if viewing 3D information from above. This can be done using data collected from cameras or sensors.
[0058] In some embodiments, DepthNet 221 can predict depth by receiving 2D image features 211-1 as input, calculate depth information of the image accordingly, and generate a depth distribution prediction result 212-1. View transformer 220 can generate a 3D feature map 222-1 by passing the outer product of (i) the 2D image features 211-1 and (ii) the depth distribution prediction result 212-1 to BEV pooling 222. BEV pooling 222 can generate the 3D feature map 222-1 by projecting the outer product result onto 3D space.
[0059] In some implementations, the view transformer 220 may perform a relative depth normalization method that minimizes depth and position prediction errors caused by differences in the intrinsic / extrinsic parameters of the camera providing image 201. In this case, the relative depth normalization method may involve: calculating a transformation matrix based on input intrinsic / extrinsic parameters, and performing a geometric transformation between adjacent cameras using the transformation matrix. Furthermore, the relative depth normalization method may involve: obtaining the relative depth after projecting image features onto adjacent image features using depth prediction information and the transformation matrix, and minimizing the relative depth loss based on a depth loss function.
[0060] The view transformer 220 can perform a photometric matching method using depth prediction to optimize the alignment between adjacent images based on the photometric matching method. The optical matching method may include an RGB matching method and / or a photometric matching method.
[0061] Electronic device 800 can enhance 3D feature maps 222-1 by performing a generalization method based on decoupled image depth estimation. The generalization method based on decoupled image depth estimation may require the use of camera extrinsic parameters, and the generalization performance of the model can be improved by performing consistent depth predictions on both the original image and the image after view transformation, respectively.
[0062] In operation 130, the electronic device 800 can extract BEV features from the 3D feature map 222-1 using the BEV encoder 230. The BEV encoder 230 can encode the 3D feature map 222-1 into BEV features, which can be output from the BEV encoder 230.
[0063] In operation 140, electronic device 800 can receive BEV features and can use detection head 240 to predict the position and category of objects based on the BEV features. Detection head 240 can predict the position and category of objects based on BEV features; multiple objects can be detected / predicted in this way. Therefore, detection head 240 can ultimately generate object detection result 250. The object detection result can be the final output obtained by performing object detection in 3D space. The object detection result can include the object's position coordinates (e.g., x, y, and z coordinates), the dimensions of the object or its bounding box (e.g., width, length, and / or height), the object's orientation (e.g., direction of travel), and / or the object's category.
[0064] Electronic device 800 can predict the location and category of objects through a Region of Interest Alignment (RoIAlign) operation in 3D space. The RoIAlign operation can align RoIs to achieve accurate object classification and location detection.
[0065] The image backbone 210, view transformer 220, BEV encoder 230, and detection head 240 (“network components”) may include corresponding domain adaptive adapters (or be supplemented by corresponding domain adaptive adapters) (in some embodiments, only one or more network components may have adapters). Electronic device 800 can perform domain generalization via the adapters. Adapters can be added in parallel to computational blocks (e.g., layers included in the network components) and their parameters can be fine-tuned (e.g., see [link to documentation]). Figure 6 (Adapter 500 in the network component). The adapter included in the network component can perform skip connections, in which features input to the network component are received, processed, and summed to update gradients.
[0066] Figure 3 and Figure 4 The operation of the view transformer according to one or more embodiments is illustrated.
[0067] Reference Figure 1 and Figure 2 The description provided is generally applicable Figure 3 and Figure 4 .
[0068] Reference Figure 3 The view transformer 220 can receive the image 201 (about the image backbone 210) from the image 201. I represents the set of images i (height H, width W, number of channels 3 (e.g., RGB)), and n represents the batch index, where each element i k This represents the 2D image features extracted from N images (height H, width W, number of channels 3) in the k-th batch. (211-1) , This represents a 2D image feature extractor. This represents the 2D image features 211-1 extracted through the image backbone 210 (feature extractor) to generate a 3D feature map 222-1 that reflects depth prediction information.
[0069] View transformer 220 can extract depth distribution prediction results 212-1 by using DepthNet 221 to predict depth based on 2D image features 211-1 (about...). , This indicates the depth prediction result. Let 221 represent DepthNet 221, and F represent image features 211-1). The view transformer 220 can predict results 212-1 based on (i) 2D image features 211-1 F and (ii) depth distribution. The outer product is used to obtain the 3D volume. (Refer to...) Figure 3 Some mathematical symbols in This represents the 3D feature map 222-1 (the outer product projected onto the 3D BEV space). This refers to the projection operation described above. This represents a matrix related to the camera's intrinsic parameters, and This represents a matrix related to the camera's extrinsic parameters. The view transformer 220 can extract / generate 3D feature maps 222-1 by inputting the outer product result into the BEV pool 222.
[0070] A 3D volume can be a data representation of 3D space generated by combining depth information with 2D image features 211-1. Here, depth information can be a value representing how far a corresponding pixel is in 3D space. The position of each pixel or each feature in 3D space can be represented in the 3D volume. This information can be used for object recognition.
[0071] Reference Figure 3 and Figure 4When integrating / combining the depth distribution prediction result 212-1 with the 2D image features 211-1, the view transformer 220 can minimize the depth and prediction errors caused by the differences in the intrinsic / extrinsic parameters of adjacent cameras, and this can be achieved through a relative depth normalization method; the depth prediction can be used to optimize the alignment between adjacent images through a photometric matching method.
[0072] The view transformer 220 can perform operations using a transformation matrix based on the input of intrinsic / extrinsic parameters, and perform geometric transformations between adjacent cameras (represented in the notation below as the i-th camera and the j-th camera) through the transformation matrix, as shown in Equation 2.
[0073] Equation 2
[0074] Equation 3
[0075] The bolded part of equation 3, namely ( ), which can be the transformation matrix mentioned above.
[0076] Equation 3 uses spatially and temporally adjacent views to calculate the corresponding depth. The view transformer 220 can use the camera's intrinsic parameters by performing the operation of Equation 3. and camera external parameters Coordinate transformations are performed between views from different cameras. In Equation 3, and This represents the pixels that correspond to each other in adjacent views. This indicates deep prediction, and This indicates the depth prediction made using the corresponding pixels.
[0077] here, Represents the pixel coordinates of camera i. (Matrix) It's a camera. The intrinsic parameter matrix. This can include information such as camera focal length and camera center coordinates. Matrix It is the intrinsic parameter transformation matrix used to transform the intrinsic parameters of camera j to the intrinsic parameters of camera i. Matrix This is the extrinsic parameter transformation matrix used to transform the extrinsic parameters of camera j to the extrinsic parameters of camera i. This can include rotation matrices and translation vectors. Extrinsic parameter matrices. This indicates the relative position and orientation / orientation between cameras. Electronic device 800 can use an extrinsic parameter matrix to... Observed 3D points are converted to camera The coordinate system. Indicates camera The depth value of the corresponding pixel, and Indicates camera The coordinates of the corresponding pixel. It's a camera. The inverse of the intrinsic parameter matrix. It can be used to mount a camera The pixel coordinates are transformed into the normalized camera coordinate system.
[0078] The view transformer 220 can minimize the depth prediction differences between cameras using a depth loss function. Depth loss function This is expressed by Equation 4 below. The depth loss function can minimize the difference between the two cameras. and The view transformer 220 minimizes the difference between depth predictions. In this way, it can obtain consistent depth information from different viewpoints. In other words, the view transformer 220 can minimize the depth prediction difference using a depth loss function to minimize the corresponding depth prediction result from the corresponding pixel. With depth prediction results The differences between them.
[0079] Equation 4
[0080] here, This represents the Euclidean distance. The view transformer 220 can minimize the error by calculating the difference between the depth distribution predictions for each camera pair using a depth loss function. In the context of this section, the depth loss function plays a role in ensuring consistency in depth predictions across multiple viewpoints. By minimizing the loss calculated for each camera pair, the model aligns the depth distribution predictions for greater accuracy and reliability. In one implementation, the depth loss function can be configured to compare camera pairs. While these pairs may simply be adjacent pairs, the implementation is not limited to this. Adjacent viewpoints can be used because they provide geometrically relevant depth information. However, depending on the use case and computational resources, the implementation can be extended to all possible pairs.
[0081] View transformer 220 can minimize depth distribution prediction normalization through a depth loss function. The depth loss function compares depth distribution predictions obtained from multiple viewpoints to obtain consistent depth distribution predictions for view transformer 220. The depth loss function enables consistent recognition of the same object from various viewpoints. Furthermore, regarding the use of the depth loss function, the computational differences quantified by the depth loss function are used to iteratively update the model's weights during training. Specifically, the depth loss function evaluates the consistency of depth predictions across multiple viewpoints, aiming to minimize the depth estimation differences of the same object observed from different angles. The calculated loss is used as an error signal, which is backpropagated through the neural network to adjust the model's weights. The adjustment of each weight may be proportional to its contribution to the loss, thus ensuring that the network learns effectively from the error signal. In this context, the depth loss function is crucial for achieving consistent depth distribution predictions across various camera viewpoints. By minimizing this loss, the model can be configured to ensure geometric consistency and reliable depth estimation, which is beneficial for tasks such as 3D object detection and recognition.
[0082] View transformer 220 maintains geometric consistency through a depth loss function. View transformer 220 normalizes the geometric relationships between cameras through a transformation matrix, thereby achieving consistent depth prediction even in images captured from various angles. In summary, view transformer 220 provides domain generalization. Even in finite data environments, view transformer 220 maintains consistent performance across domains by preventing model overfitting.
[0083] For example, suppose there are two cameras and Capture the same scene from different angles. The camera can generate different images by predicting the scene's depth information. and In this case, the camera can be transformed using a transformation matrix. Depth prediction Transform to coordinate system Then, the view transformer 220 can change the camera... Predicted depth distribution values and transformed camera The loss value is calculated by comparing the depth distribution predictions with the predicted values.
[0084] According to an embodiment, the view transformer 220 can perform the RGB matching method in the photometric matching method. The RGB matching method can minimize the image differences between different cameras using an RGB loss function. The view transformer 220 can minimize the differences between images based on depth information using an RGB loss function. RGB loss function This is represented by Equation 5 below.
[0085] Equation 5
[0086] here, and These are cameras and The image. It's a camera. The intrinsic parameter matrix, and It's a camera. The depth prediction value. From the camera To the camera The intrinsic parameter transformation matrix of the transformation, and It's a camera. The depth prediction value. The view transformer 220 can calculate and minimize the difference between images of each camera pair using the RGB loss function.
[0087] View transformer 220 can maintain consistent visual information by comparing images from different cameras using the RGB loss function. In summary, view transformer 220 can provide domain generalization through the RGB loss function.
[0088] For example, suppose there are two cameras and The same scene is captured from different viewpoints. Since each camera captures the same object from a different angle, there will be differences between their two images. The view transformer 220 can minimize position-related errors by performing alignment between 2D images through 2D RGB matching by comparing the actual image with the depth information predicted by each camera through an RGB loss function.
[0089] In some embodiments, the view transformer 220 can perform a photometric matching method. The view transformer 220 can minimize image differences between different cameras using a photometric loss function. The view transformer 220 can also minimize differences between images based on depth information using a photometric loss function. Photometric loss function This is represented by Equation 6 below.
[0090] Equation 6
[0091] here, Representing an image Point clouds, This represents the photometric error calculated using the Structural Similarity Index (SSIM) measure. SSIM represents bilinear sampling in an image. It measures the similarity between two images and is primarily used to evaluate the similarity between the original and compressed images. SSIM mimics how the human visual system recognizes structural information in images, and the SSIM index has values between 0 and 1, where 0 indicates dissimilarity and 1 indicates similarity.
[0092] Reference Figure 4 The view transformer 220 can use the depth and geometric information between images to perform depth projection and photometric reprojection on adjacent views.
[0093] For depth projection onto adjacent views, the central view features This can include depth information and 2D image features extracted through the image backbone 210. The view transformer 220 can use the depth information in the central view 403 to perform projection onto adjacent views. The first adjacent view projection 410 is based on features from the central view. Extracted depth information for features of the first neighboring view The projection of the second neighboring view 420 is performed. During this process, the depth value of each pixel is converted into coordinates of the first neighboring view 401 and then projected. The second neighboring view projection 420 is based on features from the center view. Extracted depth information for features of the second neighboring view The projection is performed. During this process, the depth value of each pixel can be converted to coordinates of the second adjacent view 402 and then projected. The view transformer 220 can perform a geometric transformation to the adjacent view based on the depth value of each pixel by projecting the depth onto the adjacent view. Depth projection onto the adjacent view can be performed using camera parameters, and depth information can be propagated to the adjacent view through geometric transformation.
[0094] For photometric reprojection, the central image This can be the actual image observed in the central view 403. The first adjacent view reprojection 410-1 is from the central image. To the first adjacent image The reprojection of the second neighboring view 420-1 is performed from the center image. During this process, depth information can be used to transform the coordinates of each pixel into the first neighboring view 401. To the second adjacent image The reprojection is performed. During this process, depth information can be used to transform the coordinates of each pixel into the second adjacent view 402. Photometric reprojection can transform the pixel values of the actual image into the adjacent views based on the depth projection results. The view transformer 220 can maintain pixel value consistency while performing photometric transformation on adjacent views through photometric reprojection.
[0095] View transformer 220 can be obtained through the final loss function ( or (e.g., Equation 7 or 8) to improve the accuracy of 3D object detection. This function integrates the above loss functions.
[0096] Equation 7
[0097] Equation 8
[0098] Here, in order to optimize , and Grid search can be used , and ,in This indicates the loss in the detection task. In summary, the view transformer 220 can mitigate the differences between time points by constraining the corresponding depths between multiple views.
[0099] Figure 5 An example of an adapter according to one or more embodiments is shown.
[0100] Reference Figures 1 to 4 The description provided is generally applicable Figure 5 .
[0101] Figure 5 illustrates an adapter 500 for Label-Efficient Domain Adaptation (LEDA). LEDA can improve domain adaptation performance by applying the PEFT (Parameter-Efficient Fine-Tuning) method, which fixes the parameters of the pre-trained network (e.g., image backbone 210, view transformer 220 (or DepthNet 221), BEV encoder 230, and detector head 240) and fine-tunes only a relatively small number of parameters of the additionally provided adapter 500. The adapter 500 can be implemented as a plug-in, preventing catastrophic forgetting of the pre-trained weights.
[0102] For example, PEFT can effectively fine-tune a small number of parameters while largely maintaining the parameters of a large language model (LLM).
[0103] Instances of adapter 500 can be connected in parallel to each computation block to replace / update the parameters of image backbone 210, view transformer 220, BEV encoder 230, and / or detector head 240. Adapter 500 has a bottleneck structure formed by project-up / project-down and can be updated by adding the values calculated by the computation blocks to the values calculated by the previous computation blocks. In this case, the parameters of the computation blocks are fixed and therefore omitted during gradient update. Gradient updates can be performed only on the parameters corresponding to adapter 500.
[0104] During fine-tuning of the adapter 500, catastrophic forgetting can be suppressed. Fine-tuning can overwrite the weights of the data and previous tasks during training on new data and on new tasks. PEFT can preserve the pre-trained parameter values because newly added (functionally, such as activation, instantiation, etc.) adapter 500s are connected in a plug-in manner, thus preserving the weights of previous parameters. As a result, the PEFT method can fine-tune a small number of parameters with a small amount of data, thereby not only improving adaptability to new domains but also maintaining stable predictions on the pre-trained domain.
[0105] The adapter 500 can be implemented for domain adaptation. Adapter 500 has a bottleneck (dimensionality reduction) structure formed by up-projection layers and down-projection layers, and can use a skip connection method that receives, computes, and uses the same features as the pre-trained network (or network components). Domain adaptation performance can be progressively improved by performing gradient updates only on adapter 500 while fixing (not changing) previous computed blocks.
[0106] The adapter 500 is represented by Equation 7 below. The module is constructed in parallel with pre-trained operational blocks B (e.g., convolutional blocks, linear blocks, or MLPs (multilayer perceptrons)).
[0107] Equation 9
[0108] here, and These represent the downward projection layer and the upward projection layer, respectively. Indicates activation function ( Figure 5 (In "Act. Func"). BN represents batch normalization. First, input... It is input into the down-projection layer and compressed into It can be restored to its original state using an upward projection layer. Then, the outputs can be fused using skip connections. Adapter 500 itself is scalable, capable of learning high-resolution details in the corresponding space while reducing network complexity and computational cost. Specifically, adapter 500 can be initialized with nearly identical functionality to preserve pre-trained weights. Finally, the parallel plug-in adapter framework achieves stable general domain adaptation (GDA) performance across all possible source and target domains, and can gradually adapt to unfamiliar domains while preserving prior knowledge. Regarding adapter 500 being initialized with nearly identical functionality to preserve pre-trained weights, the term "functionality" here refers to preserving the functionality of the computational blocks to which adapter 500 is connected in parallel. More specifically, adapter 500 can be designed to complement the existing functionality of the computational blocks while retaining pre-trained weights. This configuration ensures that the adapter operates as an extension of the computational blocks, enhancing their capabilities without compromising their original purpose. The computational blocks can be convolutional blocks, MLP blocks, or similar structures, performing specific tasks such as feature extraction or transformation. By connecting adapter 500 in parallel, the network achieves the additional flexibility of domain adaptation while maintaining the stability of pre-trained weights. This allows the model to adapt effectively to new domains, ensuring stable performance in both familiar and unfamiliar domains. In summary, the adapter 500 retains the functionality of the connected computational blocks while introducing the ability to fine-tune parameters in a controlled manner, thus achieving a balance between adaptability and stability.
[0109] Figure 6 An example of the operation of an adapter according to one or more embodiments is shown.
[0110] Reference Figures 1 to 5 The description provided is generally applicable Figure 6 .
[0111] Figure 6 A parallel structure is shown, consisting of instances of computational blocks (e.g., Conv. blocks, MLP blocks, etc.) in the BEV encoder 230 and the detector head 240, and an adapter 500. The Conv. blocks can extract spatial features of the image by performing convolution operations, while the adapter 500 can perform additional fine-tuning on the parameters input to the Conv. blocks.
[0112] MLP blocks can learn complex relationships between feature vectors by performing MLP operations, while adapter 500 can improve domain adaptability by further fine-tuning the parameters input to the MLP block. However, Conv. blocks and MLP blocks are just examples. Electronic device 800 can include various other types of operational blocks (network components), and their order can also be changed.
[0113] Furthermore, the adapter 500, which can be added in parallel to the corresponding computation blocks, can be added individually to each computation block, and can also be added in parallel to multiple computation blocks. The Conv. block and MLP block in the BEV encoder 230 and the detector head 240 are non-limiting examples. The adapter 500 can be applied to multiple computation blocks included in the general BEV encoder 230 and detector head 240. Instances of the adapter 500 can also be applied to multiple corresponding computation blocks included in the image backbone 210 and the view transformer 220.
[0114] Figure 7 An example of a generalized method for decoupled image depth estimation based on one or more embodiments is shown.
[0115] Reference Figures 1 to 6 The description provided is generally applicable Figure 7 .
[0116] Reference Figure 7 The electronic device 800 uses an image depth estimation generalization method based on camera extrinsic parameters to generate similar 3D feature maps by viewpoint transformation of related images (e.g., images captured simultaneously), thereby enabling consistent depth prediction from different viewpoints.
[0117] Electronic device 800 can receive multi-view original image 701 as input. Electronic device 800 can process original image 701. The 3D Gaussian Splatting (3DGS) method is applied, and decoupled images 702 of view transformations can be generated. 3DGS methods can include generating images from different viewpoints by transforming the viewpoint of the input image.
[0118] Image backbone 210 can process the original input image 701 and the decoupled image 702, and can extract the original image features 211-1 and the decoupled image features 211-2 respectively. DepthNet 221 of view transformer 220 can process the original image features 211-1 and the decoupled image features 211-2, and can generate the original depth distribution prediction result 212-1 and the decoupled depth distribution prediction result 212-2 respectively.
[0119] The view transformer 220 can input the original outer product of the original image features 211-1 and the original depth distribution prediction result 212-1, as well as the decoupled outer product of the decoupled image features 211-1 and the decoupled depth distribution prediction result 212-2, into the BEV pool 222.
[0120] BEV pooling 222 can take the original outer product result and the decoupled outer product result as input to generate the original 3D feature map 222-1 Decoupling 3D Feature Map 222-2 .
[0121] BEV encoder 230 can extract BEV features by encoding the 3D feature map sent from view transformer 220.
[0122] The detection head 240 can use BEV features sent from the BEV encoder 230 to predict the location and category of an object. The electronic device 800 can then generate a 3D object detection result 750 using the detection head 240.
[0123] Original image 701 and decoupled image 702 The image backbone 210 and view transformer 220 can be used to convert the image into the original 3D feature map 222-1, respectively. Decoupling 3D Feature Map 222-2 That is, 3D feature map.
[0124] Electronic device 800 can perform consistent predictions from different viewpoints by calculating the cosine similarity loss between 3D feature maps.
[0125] Cosine similarity loss function It is represented by the following equation 10.
[0126] Equation 10
[0127] Electronic device 800 can perform consistent predictions in images with view transformations by measuring the similarity between the original 3D feature map 222-1 and the decoupled 3D feature map 222-2 using a cosine similarity loss function.
[0128] Using cosine similarity, for two vectors, when their similarity value is close to 1, they have the same direction, while when their similarity value is close to -1, they have different directions. The cosine similarity loss function can be designed to minimize the loss value as the similarity increases by maximizing the similarity.
[0129] A high similarity indicates that the 3D feature maps of the original image 701 and the decoupled image 702 are similar. This suggests that the electronic device 800 performs consistent predictions from both viewpoints. Therefore, the electronic device 800 can perform consistent predictions from various viewpoints.
[0130] On the other hand, low similarity can indicate that the 3D feature maps between the original image 701 and the decoupled image 702 are dissimilar. This could suggest that the electronic device 800 is not making consistent predictions from both viewpoints. In this case, the value of the loss function increases. Therefore, the electronic device 800 can adjust its parameters or increase the similarity of the network.
[0131] Figure 8 An example of an electronic device according to one or more embodiments is shown.
[0132] Reference Figure 8 According to an embodiment, electronic device 800 (e.g., an autonomous vehicle and 3D object detection device 200) may include a processor 830, a memory 850, and an output device 870 (e.g., a display). The processor 830, memory 850, and output device 870 may be connected to each other via a communication bus 805. In the above processing, for ease of description, electronic device 800 may include processor 830 for executing at least one of the above methods or an algorithm corresponding to that at least one method.
[0133] Output device 870 can display a user interface related to the 3D object detection method provided by processor 830.
[0134] The memory 850 can store data obtained by the processor 830 executing the 3D object detection method. Furthermore, the memory 850 can store various information generated during the processing of the processor 830 as described above. Additionally, the memory 850 can store various data, programs, etc. The memory 850 may include volatile memory or non-volatile memory. The memory 850 may include a high-capacity storage medium (such as a hard disk) to store various types of data.
[0135] In addition, the processor 830 can execute references Figures 1 to 7 At least one method or an algorithm corresponding to said at least one method is described. Processor 830 may be a hardware-implemented data processing device having circuitry physically configured to perform desired operations. For example, the desired operations may include code or instructions in a program. Processor 830 may be implemented as, for example, a central processing unit (CPU), a graphics processing unit (GPU), or a neural network processing unit (NPU), or a combination thereof. Hardware-implemented electronic device 800 may include, for example, a microprocessor, CPU, processor core, multi-core processor, multiprocessor, application-specific integrated circuit (ASIC), and field-programmable gate array (FPGA).
[0136] The processor 830 can execute programs and control the electronic device 800. The program code to be executed by the processor 830 can be stored in the memory 850.
[0137] The training process and computational algorithms described above can be executed on a server and applied to autonomous vehicles, or executed inside an autonomous vehicle.
[0138] For example, a server can receive 2D images from an autonomous vehicle, use the received 2D images to perform 3D object detection, and send the 3D object detection results to the autonomous vehicle.
[0139] As another example, autonomous vehicles may include electronics and processors for 3D object detection, and 3D object detection can be performed by receiving 2D images from the autonomous vehicle's cameras.
[0140] The units described herein can be implemented using hardware components, software components, and / or combinations thereof. The processing device can be implemented using one or more general-purpose or special-purpose computers, such as processors, controllers and arithmetic logic units (ALUs), digital signal processors (DSPs), microcomputers, FPGAs, programmable logic units (PLUs), microprocessors, or any other device capable of responding to and executing instructions in a defined manner. The processing device can run an operating system (OS) and one or more software applications running on the OS. The processing device can also access, store, manipulate, process, and generate data in response to the execution of software. For simplicity, the processing device is described in the singular; however, those skilled in the art will understand that the processing device can include multiple processing elements and various types of processing elements. For example, a processing device can include multiple processors or include one processor and one controller. Furthermore, different processing configurations, such as parallel processors, are also possible.
[0141] Software can include computer programs, code, instructions, or combinations thereof, to independently or uniformly instruct or configure a processing device to operate as intended. Software and data can be permanently or temporarily embodied in any type of machine, component, physical or virtual device, computer storage medium, or device, or embodied in propagated signal waves capable of providing instructions or data to or being interpreted by the processing device. Software can also be distributed across network-connected computer systems for distributed storage and execution. Software and data can be stored on one or more non-transitory computer-readable recording media.
[0142] The methods described in the examples above can be recorded in a non-transitory computer-readable medium, which includes program instructions for implementing the various operations of the example embodiments described above. The medium may also include data files, data structures, etc., alone or in combination with the program instructions. The program instructions recorded on the medium may be program instructions specifically designed and constructed for the purposes of the example embodiments, or program instructions well known and usable by those skilled in the art of computer software. Examples of non-transitory computer-readable media include: magnetic media, such as hard disks, floppy disks, and magnetic tapes; optical media, such as CD-ROMs and DVDs; magneto-optical media, such as optical discs; and hardware devices specifically configured to store and execute program instructions, such as read-only memory (ROM), random access memory (RAM), flash memory, etc. Examples of program instructions include both machine code (such as code generated by a compiler) and files containing high-level code that a computer can execute using an interpreter.
[0143] Computing devices, vehicles, electronic devices, processors, memory, image sensors, vehicle / operational hardware, displays, information output systems and hardware, storage devices, and references herein. Figures 1-8Other devices, apparatuses, units, modules, and components described are implemented by or represent hardware components. Where appropriate, examples of hardware components that can be used to perform the operations described in this application include controllers, sensors, generators, drivers, memories, comparators, arithmetic logic units, adders, subtractors, multipliers, dividers, integrators, and any other electronic components configured to perform the operations described in this application. In other examples, one or more hardware components performing the operations described in this application are implemented by computing hardware, such as one or more processors or computers. A processor or computer may be implemented by one or more processing elements, such as logic gate arrays, controllers and arithmetic logic units, digital signal processors, microcomputers, programmable logic controllers, field-programmable gate arrays, programmable logic arrays, microprocessors, or any other device or combination of devices configured to respond to and execute instructions in a defined manner to achieve desired results. In one example, the processor or computer includes or is connected to one or more memories that store instructions or software executed by the processor or computer. Hardware components implemented by a processor or computer can execute instructions or software, such as an operating system (OS) and one or more software applications running on the OS, to perform the operations described in this application. Hardware components can also access, manipulate, process, create, and store data in response to the execution of instructions or software. For simplicity, the singular terms “processor” or “computer” may be used in the description of the examples described in this application, but in other examples, multiple processors or computers may be used, or a processor or computer may include multiple processing elements, or multiple types of processing elements, or both. For example, a single hardware component or two or more hardware components may be implemented by a single processor, or by two or more processors, or by a processor and a controller. One or more hardware components may be implemented by one or more processors, or by a processor and a controller, and one or more other hardware components may be implemented by one or more other processors, or by another processor and another controller. One or more processors, or a processor and a controller, may implement a single hardware component, or two or more hardware components. Hardware components may have any one or more different processing configurations, examples of which include a single processor, a standalone processor, a parallel processor, a single-instruction single-data (SISD) multiprocessing, a single-instruction multiple-data (SIMD) multiprocessing, a multiple-instruction single-data (MISD) multiprocessing, and a multiple-instruction multiple-data (MIMD) multiprocessing.
[0144] Figures 1-8The methods for performing the operations described in this application, as shown, are executed by computing hardware, such as one or more processors or computers, which are implemented as described above to implement instructions or software to perform the operations performed by the methods described in this application. For example, a single operation or two or more operations may be executed by a single processor, or by two or more processors, or by a processor and a controller. One or more operations may be executed by one or more processors, or by a processor and a controller, and one or more other operations may be executed by one or more other processors, or by another processor and another controller. One or more processors, or a processor and a controller, may execute a single operation, or two or more operations.
[0145] Instructions or software for controlling computing hardware (e.g., one or more processors or computers) to implement hardware components and perform the methods described above can be written as computer programs, code segments, instructions, or any combination thereof, for individually or collectively instructing or configuring one or more processors or computers to operate as machines or special-purpose computers to perform operations performed by the hardware components and methods described above. In one example, the instructions or software include machine code, such as machine code generated by a compiler, that is directly executed by one or more processors or computers. In another example, the instructions or software include higher-level code that is executed by one or more processors or computers using an interpreter. The instructions or software can be written in any programming language based on the block diagrams and flowcharts shown in the accompanying drawings and the corresponding description herein, which disclose algorithms for performing operations performed by the hardware components and methods described above.
[0146] Instructions or software used to control computing hardware (such as one or more processors or computers) to implement hardware components and perform the methods described above, as well as any associated data, data files, and data structures, may be recorded, stored, or fixed on or in one or more non-transitory computer-readable storage media. Examples of non-transitory computer-readable storage media include read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc storage, hard disk drive (HDD), solid-state drive (SSD), flash memory, card-type storage (such as multimedia card micro or card (e.g., Secure Digital (SD) or Extreme Digital (XD)), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, and any other device configured to store instructions or software and any associated data, data files, and data structures in a non-transitory manner, and to provide instructions or software and any associated data, data files, and data structures to one or more processors or computers, enabling the one or more processors or computers to execute the instructions. In one example, the instructions or software and any associated data, data files, and data structures are distributed across a network-connected computer system, such that the instructions and software and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.
[0147] While this disclosure includes specific examples, it will be clear upon understanding this disclosure that various changes in form and detail may be made in these examples without departing from the spirit and scope of the claims and their equivalents. The examples described herein should be considered descriptive only and not for limiting purposes. The description of a feature or aspect in each example should be considered applicable to similar features or aspects in other examples. Suitable results may be achieved even if the described techniques are performed in a different order, and / or the components in the described system, architecture, device, or circuit are combined in different ways and / or replaced or supplemented by other components or their equivalents.
[0148] Therefore, in addition to the above disclosure, the scope of this disclosure may also be defined by the claims and their equivalents, and all variations within the scope of the claims and their equivalents should be interpreted as being included in this disclosure.
Claims
1. A method for detecting a three-dimensional (3D) object, the method comprising: Extract two-dimensional (2D) image features from an image using the image backbone; A 3D feature map reflecting depth prediction information is extracted from the 2D image features using a view transformer configured to perform domain generalization; The BEV features are extracted from the 3D feature map using a bird's-eye view BEV encoder; and The detection head is used to predict the location and category of the 3D object based on the BEV features.
2. The method according to claim 1, wherein, Extracting the 3D feature map includes: predicting the depth output by DepthNet based on the 2D image features, and inputting the outer product of the depth output of DepthNet and the 2D image features into the BEV pool.
3. The method according to claim 1, wherein, The view transformer is configured to perform a relative depth normalization method that minimizes depth and position prediction errors caused by differences in intrinsic / extrinsic parameters of the camera providing one of the images.
4. The method according to claim 3, wherein, Multiple cameras, including the aforementioned camera, provide corresponding images, and the relative depth normalization method includes calculating a transformation matrix based on the intrinsic / extrinsic parameters, wherein the transformation matrix is used to perform geometric transformations between adjacent camera pairs.
5. The method according to claim 4, wherein, The relative depth normalization method uses the depth prediction information and the transformation matrix to project image features onto adjacent image features, and then minimizes the relative depth loss based on the depth loss function to obtain the relative depth.
6. The method according to claim 1, wherein, The view transformer is configured to perform a photometric matching method using depth prediction to optimize the alignment between the image and neighboring images based on the photometric matching method.
7. The method according to claim 1, wherein, The image backbone, the view transformer, the BEV encoder, and / or the detection head include corresponding domain adaptive adapters.
8. The method according to claim 7, wherein, Each domain adaptive adapter is added in parallel to the computation block to enable fine-tuning of the parameters.
9. The method according to claim 7, wherein, Each domain adaptive adapter is configured to perform a skip connection, in which features input to the view transformer, the BEV encoder, and / or the detection head are received, computed, and summed to update the gradient.
10. The method according to claim 1, further comprising: The 3D feature map is enhanced by performing a generalization method based on decoupled image depth estimation.
11. An electronic device, comprising: Memory, storing instructions; as well as One or more processors, Wherein, when the instruction is executed by the one or more processors, the one or more processors: Extracting 2D image features from an image using an image backbone. A view transformer is used to extract a 3D feature map reflecting depth prediction information from the 2D image features. The BEV features are extracted from the 3D feature map using a bird's-eye view BEV encoder, and The detection head is used to predict the location and category of the object based on the BEV features.
12. The electronic device according to claim 11, wherein, When the instruction is executed by the one or more processors, the one or more processors perform the following operations to extract the 3D feature map: predicting a depth output by DepthNet based on the 2D image features, and inputting the outer product of the depth output of DepthNet and the 2D image features into the BEV pool.
13. The electronic device according to claim 11, wherein, The view transformer is configured to perform a relative depth normalization method that minimizes depth and position prediction errors caused by differences in intrinsic / extrinsic parameters of the camera providing one of the images.
14. The electronic device according to claim 13, wherein, Multiple cameras, including the aforementioned camera, provide corresponding images, and the relative depth normalization method includes calculating a transformation matrix based on the intrinsic / extrinsic parameters, wherein the transformation matrix is used to perform geometric transformations between adjacent camera pairs.
15. The electronic device according to claim 14, wherein, The relative depth normalization method uses the depth prediction information and the transformation matrix to project image features onto adjacent image features, and then minimizes the relative depth loss based on the depth loss function to obtain the relative depth.
16. The electronic device according to claim 11, wherein, The view transformer is configured to perform a photometric matching method using depth prediction to optimize the alignment between the image and neighboring images based on the photometric matching method.
17. The electronic device according to claim 11, wherein, The image backbone, the view transformer, the BEV encoder, and / or the detection head have corresponding domain adaptive adapters.
18. The electronic device according to claim 17, wherein, The domain adaptive adapters are added in parallel to the layers in the image backbone, the view transformer, the BEV encoder, and / or the detection head.
19. The electronic device according to claim 17, wherein, Each domain adaptive adapter is configured to perform a skip connection, in which features input to the view transformer, the BEV encoder, and / or the detection head are received, computed, and summed to update the gradient.
20. The electronic device according to claim 11, wherein, When the instructions are executed by the one or more processors, the one or more processors enhance the 3D feature map by performing a generalization method based on decoupled image depth estimation.
Citation Information
Patent Citations
Optical connector module and method for manufacturing optical waveguide substrate
KR1020240064025A
Method of predicting pressure of heat pump system for vehicle
KR1020240099581A