Method and system for extracting bird's eye view information
By fusing and super-resolving low-resolution BEV features from multiple sensors, the method enhances computing efficiency and environmental detail in autonomous driving systems, ensuring safe navigation.
Patent Information
- Application Number
- PCT/KR2025/003019
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-07
- Filing Date
- 2025-03-07
- Publication Date
- 2025-09-11
AI Technical Summary
Existing autonomous driving systems face challenges in efficiently extracting high-resolution bird's eye view (BEV) information using limited computing resources, which is crucial for safe navigation and environmental recognition.
A method and system that fuse low-resolution BEV features from multiple sensors, such as cameras and LiDAR, using a low-resolution BEV feature fusion model and super-resolution model to generate high-resolution BEV features, optimizing computing resource usage and enhancing environmental detail.
This approach allows for efficient extraction of high-resolution BEV features, providing detailed environmental information while conserving computing resources, enabling accurate navigation and safe driving operations.
Smart Images

Figure KR2025003019_12092025_PF_FP_ABST
Abstract
Description
Method and system for extracting bird's eye view information
[0001] The present disclosure relates to a method and system for extracting bird's eye view information, and more particularly, to a method and system for extracting high-resolution bird's eye view information.
[0002] Autonomous driving technology refers to a technology that enables a vehicle to drive autonomously with minimal or no human intervention by recognizing its surroundings using radar, LiDAR (Light Detection and Ranging), GPS, cameras, and other technologies. To ensure safe driving, a vehicle equipped with autonomous driving technology must be able to identify its own location, areas within its environment where it can drive, where it must stop, and areas where it cannot drive.
[0003] Bird's Eye View (BEV) information can comprehensively represent a wide area, making it ideal for driving devices to perceive the surrounding environment. Due to these advantages, bird's eye view information is widely used to ensure the safe operation of autonomous driving devices.
[0004] The present disclosure provides a method and system (device) for extracting high-resolution bird's eye view information.
[0005] The present disclosure may be implemented in various ways, including as a method, a device (system), or a computer program stored on a readable storage medium.
[0006] According to one embodiment of the present disclosure, a method for extracting bird's eye view information, performed by at least one processor, includes the steps of: receiving first sensor data and second sensor data acquired by a first sensor and a second sensor mounted on a driving device at a specific point where the driving device is located; extracting a first sensor BEV feature and a second sensor BEV feature, which are bird's eye view (BEV) information for a region of interest, based on the first sensor data and the second sensor data; generating a first fused BEV feature of a first resolution based on the first sensor BEV feature and the second sensor BEV feature using a low-resolution BEV feature fusion model; and converting the first fused BEV feature into a second fused BEV feature of a second resolution higher than the first resolution using a super-resolution model, wherein the region of interest includes an area within a predefined range based on the specific point.
[0007] A computer program stored in a computer-readable recording medium is provided for executing a method according to one embodiment of the present disclosure on a computer.
[0008] An information processing system according to one embodiment of the present disclosure includes a memory and at least one processor connected to the memory and configured to execute at least one computer-readable program included in the memory, wherein the at least one program includes instructions for receiving first sensor data and second sensor data acquired by a first sensor and a second sensor mounted on a driving device at a specific point where the driving device is located, extracting a first sensor BEV feature and a second sensor BEV feature, which are bird's-eye view information for a region of interest, based on the first sensor data and the second sensor data, generating a first fused BEV feature of a first resolution based on the first sensor BEV feature and the second sensor BEV feature using a low-resolution BEV feature fusion model, and converting the first fused BEV feature into a second fused BEV feature of a second resolution higher than the first resolution using a super-resolution model, wherein the region of interest includes an area within a predefined range based on the specific point.
[0009] According to some embodiments of the present disclosure, BEV feature fusion, which can cause a bottleneck when data sizes are large, is performed at low resolution, thereby efficiently extracting fused BEV features using limited computing resources. Furthermore, by super-resolving the BEV features fused at low resolution, high-resolution fused BEV features can be extracted, providing more detailed information about the surrounding environment. In other words, both computing resource savings and detailed environmental recognition can be achieved simultaneously.
[0010] According to some embodiments of the present disclosure, high-quality target information based on high-resolution BEV features that provide detailed information about the surrounding environment for an area close to the driving device is provided, and target information based on low-resolution BEV features is provided for an area not close to the driving device, thereby efficiently providing information under limited resources.
[0011] The effects of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs (referred to as “one skilled in the art”) from the description of the claims.
[0012] Embodiments of the present disclosure will be described below with reference to the accompanying drawings, wherein like reference numerals represent similar elements, but are not limited thereto.
[0013] FIG. 1 is a diagram illustrating an example of a method for extracting bird's eye view information according to one embodiment of the present disclosure.
[0014] FIG. 2 is a block diagram showing the internal configuration of an information processing system according to one embodiment of the present disclosure.
[0015] FIG. 3 is a diagram showing the internal configuration of a processor of an information processing system according to one embodiment of the present disclosure.
[0016] FIG. 4 is a diagram illustrating an example of first sensor data according to one embodiment of the present disclosure.
[0017] FIG. 5 is a diagram illustrating an example of second sensor data according to one embodiment of the present disclosure.
[0018] FIG. 6 is a diagram illustrating an example of extracting BEV features for each sensor according to one embodiment of the present disclosure.
[0019] FIG. 7 is a diagram illustrating an example of extracting high-resolution fused BEV features according to one embodiment of the present disclosure.
[0020] FIG. 8 is a diagram illustrating an example of extracting target information based on high-resolution fused BEV features according to one embodiment of the present disclosure.
[0021] FIG. 9 is a diagram showing an example of training a model used in the present method according to one embodiment of the present disclosure.
[0022] FIG. 10 is a diagram showing an example of target information according to one embodiment of the present disclosure.
[0023] FIG. 11 is a diagram illustrating an example of generating hybrid target information according to one embodiment of the present disclosure.
[0024] FIG. 12 is a flowchart illustrating an example of a method for extracting bird's eye view information according to one embodiment of the present disclosure.
[0025] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the attached drawings. However, in the following description, specific descriptions of widely known functions or configurations will be omitted if they may unnecessarily obscure the gist of the present disclosure.
[0026] In the attached drawings, identical or corresponding components are assigned the same reference numerals. Furthermore, in the description of the embodiments below, duplicate descriptions of identical or corresponding components may be omitted. However, even if a description of a component is omitted, it is not intended that such component is not included in any embodiment.
[0027] The advantages and features of the disclosed embodiments, and methods for achieving them, will become clearer with reference to the embodiments described below, along with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure the completeness of the disclosure and to fully inform those skilled in the art of the scope of the invention.
[0028] The terms used in this specification will be briefly explained, followed by a detailed description of the disclosed embodiments. The terms used in this specification have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of engineers working in the relevant field, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the relevant description of the invention. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on their meanings and the overall content of the present disclosure.
[0029] In this specification, singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, plural expressions include singular expressions unless the context clearly indicates otherwise. When a part of the specification is said to include a component, this does not exclude other components, but rather implies that other components may be included, unless otherwise specifically stated.
[0030] Also, the term 'module' or 'part' used in the specification means a software or hardware component, and the 'module' or 'part' performs certain roles. However, the 'module' or 'part' is not limited to software or hardware. The 'module' or 'part' may be configured to reside on an addressable storage medium and may be configured to execute one or more processors. Thus, as an example, the 'module' or 'part' may include at least one of components such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, or variables. The functionality provided within the components and 'modules' or 'parts' may be combined into a smaller number of components and 'modules' or 'parts', or further separated into additional components and 'modules' or 'parts'.
[0031] According to one embodiment of the present disclosure, a 'module' or 'unit' may be implemented as a processor and a memory. 'Processor' should be broadly construed to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, and the like. In some circumstances, a 'processor' may also refer to an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field-programmable gate array (FPGA), and the like. A 'processor' may also refer to a combination of processing devices, such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors in conjunction with a DSP core, or any other such combination of configurations. In addition, 'memory' should be broadly construed to include any electronic component capable of storing electronic information. 'Memory' may refer to various types of processor-readable media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage, registers, etc. Memory is said to be in electronic communication with the processor if the processor can read information from, and / or write information to, the memory. Memory integrated in a processor is in electronic communication with the processor.
[0032] In the present disclosure, the "system" may include, but is not limited to, at least one of a server device and a cloud device. For example, the system may be comprised of one or more server devices. As another example, the system may be comprised of one or more cloud devices. As yet another example, the system may be configured and operated by a combination of a server device and a cloud device.
[0033] In the present disclosure, 'display' may refer to any display device associated with a computing device, for example, any display device capable of displaying any information / data controlled by or provided from the computing device.
[0034] In the present disclosure, 'each of the plurality of As' or 'each of the plurality of As' may refer to each of all components included in the plurality of As, or may refer to each of some components included in the plurality of As.
[0035] In the present disclosure, a "machine learning model" may include any model used to infer an answer to a given input. In one embodiment, the machine learning model may include an artificial neural network model including an input layer, multiple hidden layers, and an output layer. Here, each layer may include multiple nodes. In the present disclosure, each of the multiple machine learning models is described as a separate machine learning model, but this is not limited thereto, and some or all of the multiple machine learning models may be implemented as a single machine learning model. Furthermore, a single machine learning model may include multiple machine learning models. In the present disclosure, the terms "machine learning model" and "artificial neural network model" may be used interchangeably to refer to the same or similar models. In the present disclosure, at least some of the models referred to as "models," "encoders," or "decoders" may be implemented as machine learning models.
[0036] In the present disclosure, a 'region of interest' may be a region to be analyzed in recognizing the surrounding environment for smooth driving of a driving device. According to one embodiment, the region of interest may be determined within a predefined range based on a point where the driving device is located. For example, when the forward direction of the driving device is defined as the +Y-axis direction, the region of interest may be defined as an area within a distance of -p meters to +p meters in the X-axis direction and within a distance of -q meters to +q meters in the Y-axis direction from the point where the driving device is located. In the present disclosure, when the region of interest is defined as an area within a distance of -r meters to +r meters in the X-axis direction and the Y-axis direction from the point where the driving device is located, the region of interest may be expressed as (-rm, +rm).
[0037] In the present disclosure, a 'Bird's Eye View (BEV)' may refer to a view of an area of interest (or an area including the area of interest) viewed from above. In the present disclosure, a Bird's Eye View may be denoted as BEV.
[0038] In this disclosure, 'resolution' may refer to the degree to which data represents a region of interest in detail. The shorter the length of the region corresponding to the height or width of one pixel of data (i.e., the narrower the area corresponding to one pixel), the higher the resolution may be considered. For example, if the height or width of one pixel of specific data corresponds to an area of 2.0 m in length (i.e., 2.0 m / px), one pixel can represent information for an area of 4 m2. As another example, if the height or width of one pixel of specific data corresponds to an area of 0.5 m in length (i.e., 0.5 m / px), one pixel can represent information for an area of 0.25 m2. In other words, since data of 0.5 m / px represents an area of interest in more detail than data of 2.0 m / px, data of 0.5 m / px can have a higher resolution than data of 2.0 m / px. High-resolution data for a specific region of interest may have a larger data size than low-resolution data for a specific region of interest of the same size.
[0039] FIG. 1 is a diagram illustrating an example of a method for extracting bird's-eye view information according to one embodiment of the present disclosure. According to one embodiment, a driving device may be equipped with multiple sensors to recognize the surrounding environment. For example, the driving device may be equipped with a first sensor and a second sensor. Specifically, the first sensor may include multiple cameras, and the second sensor may include one or more LiDARs. However, the scope of the present disclosure is not limited thereto, and the driving device may additionally or alternatively be equipped with other sensors (e.g., radar) in addition to the sensors described above. For convenience of explanation, the present disclosure focuses on a method for extracting BEV features based on data acquired by the first sensor and the second sensor; however, the scope of the present disclosure is not limited thereto, and the same or similar method may be performed based on data acquired by other sensors.
[0040] According to one embodiment, an information processing system (e.g., an information processing system included in a driving device or an information processing system capable of communicating with the driving device, etc.) can extract a single high-resolution fused BEV feature (132) from a plurality of BEV features extracted based on data acquired by a first sensor and a second sensor.
[0041] First, the information processing system can receive first sensor data (112) acquired by the first sensor at the location where the driving device is located. For example, the information processing system can receive multiple viewpoint images captured by multiple cameras mounted on the driving device when the driving device is located at a specific location. Then, the information processing system can extract first sensor BEV features (114) for a region of interest (e.g., (-50 m, 50 m)) based on the first sensor data (112) using a first BEV feature extraction model (110). Here, the extracted first sensor BEV features (114) may be low-resolution BEV features. For example, the first sensor BEV features (114) may have a first resolution (e.g., 2.0 m / px).
[0042] In addition, the information processing system can receive second sensor data (122) acquired by the second sensor at the location where the driving device is located. According to one embodiment, the second sensor may be a different type of sensor than the first sensor, and the second sensor data (122) may be data of a different modality than the first sensor data (112). For example, the information processing system can receive point cloud data acquired by the lidar mounted on the driving device at the same time or at the same location when the multiple viewpoint images are captured from the multiple cameras mounted on the driving device. The information processing system can extract the second sensor BEV feature (124) for the region of interest based on the second sensor data (122) using the second BEV feature extraction model (120). Here, the extracted second sensor BEV feature (124) may be a low-resolution BEV feature. For example, the second sensor BEV feature (124) may have the same first resolution as the first sensor BEV feature (114). The BEV feature extraction model used to extract the sensor BEV feature based on the sensor data may vary depending on the type of sensor from which the sensor data was acquired and / or the modality of the sensor data.
[0043] The information processing system can extract a single high-resolution fused BEV feature (132) based on the first sensor BEV feature (114) and the second sensor BEV feature (124) using the high-resolution BEV feature extraction model (130). For example, the information processing system can extract a high-resolution fused BEV feature (132) having a second resolution (e.g., 0.5 m / px) higher than the first resolution based on the first sensor BEV feature (114) having a first resolution and the second sensor BEV feature (124) having the first resolution. According to one embodiment, the information processing system can fuse the first sensor BEV feature (114) and the second sensor BEV feature (124) at a low resolution to generate a low-resolution fused BEV feature, and super-resolve the generated low-resolution fused BEV feature to generate a high-resolution fused BEV feature (132).
[0044] As described above, by performing BEV feature fusion at low resolution, which can cause a bottleneck when the data size is large, fused BEV features can be efficiently extracted using limited computing resources. Furthermore, by super-resolving the BEV features fused at low resolution, high-resolution fused BEV features (132) can be extracted, which provide more detailed information about the surrounding environment. In other words, both computing resource savings and detailed environmental recognition can be achieved simultaneously.
[0045] The information processing system can perform a target task (e.g., map segmentation, object detection, etc.) for recognizing the surrounding environment of the driving device based on the extracted high-resolution fused BEV features (132). For example, the information processing system can extract target information (142) based on the high-resolution fused BEV features (132) using a decoder (140) suitable for the target task. By performing the target task using the high-resolution fused BEV features (132) that represent detailed information about the surrounding environment, the surrounding environment can be accurately recognized, and thus the driving device can drive safely.
[0046] In one embodiment, the information processing system can perform a method for extracting bird's-eye view information in real time while the vehicle is driving. For example, as the vehicle drives, the location and area of interest of the vehicle may change, and new first sensor data and new second sensor data may be acquired. The information processing system can perform the above-described method in real time, such as extracting new high-resolution fused BEV features based on the new first sensor data and the new second sensor data, and performing targeting operations based on the extracted new high-resolution fused BEV features.
[0047] Although FIG. 1 and the above description illustrate and describe the extraction of a single high-resolution fused BEV feature (132) based on two sensor BEV features, the scope of the present disclosure is not limited thereto. For example, the information processing system may extract a high-resolution fused BEV feature based on one sensor BEV feature or three or more sensor BEV features.
[0048] FIG. 2 is a block diagram illustrating the internal configuration of an information processing system (200) according to one embodiment of the present disclosure. The information processing system (200) may include a memory (210), a processor (220), a communication module (230), and an input / output interface (240). The information processing system (200) may be configured to communicate information and / or data with an external system via a network using the communication module (230).
[0049] The memory (210) may include any non-transitory computer-readable recording medium. According to one embodiment, the memory (210) may include a permanent mass storage device such as a read only memory (ROM), a disk drive, a solid state drive (SSD), a flash memory, etc. As another example, a permanent mass storage device such as a ROM, an SSD, a flash memory, a disk drive, etc. may be included in the information processing system (200) as a separate permanent storage device distinct from the memory. In addition, the memory (210) may store an operating system and at least one program code (e.g., code for executing high-resolution fused BEV feature extraction, etc., which is installed and operated in the information processing system (200).
[0050] These software components may be loaded from a computer-readable recording medium separate from the memory (210). This separate computer-readable recording medium may include a recording medium directly connectable to the information processing system (200), for example, a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, a memory card, etc. As another example, the software components may be loaded into the memory (210) through a communication module (230) other than a computer-readable recording medium. For example, at least one program may be loaded into the memory (210) based on a computer program (e.g., a program for executing high-resolution fused BEV feature extraction, etc.) that is installed by files provided by developers or a file distribution system that distributes installation files of applications through the communication module (230).
[0051] The processor (220) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. The instructions may be provided to a user terminal (not shown) or another external system via a memory (210) or a communication module (230). For example, the processor (220) may extract a single high-resolution fused BEV feature from a plurality of BEV features extracted based on data acquired by the first sensor and the second sensor. In addition, the processor (220) of the information processing system (200) may be configured to manage, process, and / or store information and / or data received from a plurality of user terminals and / or a plurality of external systems.
[0052] The communication module (230) may provide a configuration or function for a user terminal (not shown) and an information processing system (200) to communicate with each other via a network, and may provide a configuration or function for the information processing system (200) to communicate with an external system (e.g., a separate cloud system, etc.). For example, control signals, commands, data, etc. provided under the control of the processor (220) of the information processing system (200) may be transmitted to the user terminal and / or the external system via the communication module (230) and the network via the communication module of the user terminal and / or the external system.
[0053] In addition, the input / output interface (240) of the information processing system (200) may be a means for interfacing with a device (not shown) for input or output that is connected to the information processing system (200) or that the information processing system (200) may include. In FIG. 2, the input / output interface (240) is illustrated as an element configured separately from the processor (220), but is not limited thereto, and the input / output interface (240) may be configured to be included in the processor (220). The information processing system (200) may include more components than those illustrated in FIG. 2. However, there is no need to explicitly illustrate most of the conventional technology components.
[0054] FIG. 3 is a diagram illustrating an internal configuration of a processor (220) of an information processing system according to one embodiment of the present disclosure. According to one embodiment, the processor (220) may include a sensor-specific feature extraction unit (310), a low-resolution feature extraction unit (320), a super-resolution unit (330), and a target information extraction unit (340). The internal configuration of the processor (220) of the information processing system illustrated in FIG. 3 is merely an example and may be implemented differently. For example, at least a part of the configuration of the processor (220) may be omitted, another configuration may be added, and at least a part of the operations or processes performed by the processor (220) may be performed by another configuration (e.g., a processor of another device that is communicatively connected to the information processing system, etc.).
[0055] The sensor-specific feature extraction unit (310) can extract multiple sensor BEV features from each of multiple sensor data acquired by multiple sensors mounted on the driving device. According to one embodiment, the sensor-specific feature extraction unit (310) can utilize an appropriate BEV feature extraction model depending on the type of sensor from which the sensor data was acquired and / or the modality of the sensor data.
[0056] For example, the sensor-specific feature extraction unit (310) may receive first sensor data acquired by a first sensor mounted on the driving device. Specifically, it may receive multiple viewpoint images captured by multiple cameras mounted on the driving device. The multiple viewpoint images may all be images captured at different viewpoints at the same time or at the same location. This will be described in more detail below with reference to FIG. 4.
[0057] The sensor-specific feature extraction unit (310) can extract the first sensor BEV feature for the region of interest based on the first sensor data using the first BEV feature extraction model. The first BEV feature extraction model may be a model for extracting the first sensor BEV feature based on the first sensor data. For example, when the first sensor is a camera, the first BEV feature extraction model may include a first sensor encoder and a view transformation model. This will be described in more detail below with reference to FIG. 6. According to one embodiment, the first sensor BEV feature extracted by the sensor-specific feature extraction unit (310) may be a low-resolution BEV feature. For example, the first sensor BEV feature may have a first resolution (e.g., 2.0 m / px).
[0058] In addition, the sensor-specific feature extraction unit (310) may receive second sensor data acquired by a second sensor mounted on the driving device at the same time or at the same location as the first sensor data. According to one embodiment, the second sensor may be a different type of sensor than the first sensor, and the second sensor data may be data of a different modality than the first sensor data. For example, the sensor-specific feature extraction unit (310) may receive point cloud data acquired by a lidar mounted on the driving device at the same time or at the same location as the location of the driving device at the corresponding time as the time at which the multiple viewpoint images were captured.
[0059] The sensor-specific feature extraction unit (310) may extract second sensor BEV features for an area of interest based on second sensor data using a second BEV feature extraction model. The second BEV feature extraction model may be a model for extracting second sensor BEV features based on second sensor data. For example, when the second sensor is a lidar, the second BEV feature extraction model may include a lidar encoder and a smoothing model. This will be described in more detail below with reference to FIG. 6. According to one embodiment, the second sensor BEV features extracted by the sensor-specific feature extraction unit (310) may be low-resolution BEV features. For example, the second sensor BEV features may have the same first resolution (e.g., 2.0 m / px) as the first sensor BEV features.
[0060] The low-resolution feature extraction unit (320) can generate a low-resolution fused BEV feature by fusing the first sensor BEV feature and the second sensor BEV feature extracted by the sensor-specific feature extraction unit (310) at a low resolution. For example, the low-resolution feature extraction unit (320) can generate a low-resolution fused BEV feature based on the first sensor BEV feature and the second sensor BEV feature extracted by the sensor-specific feature extraction unit (310) using a low-resolution BEV feature fusion model. For example, the low-resolution BEV feature fusion model can include a fusion model and an enhancing model, and the low-resolution feature extraction unit (320) can generate a low-resolution fused BEV feature by fusing the first sensor BEV feature and the second sensor BEV feature using the fusion model and enhancing the fused BEV feature using the enhancing model. According to one embodiment, the low-resolution fused BEV features generated by the low-resolution feature extraction unit (320) may have a first resolution (e.g., 2.0 m / px).
[0061] The super-resolution unit (330) can convert the low-resolution fused BEV features generated by the low-resolution feature extraction unit (320) into high-resolution fused BEV features using the super-resolution model. For example, the super-resolution unit (330) can convert the low-resolution fused BEV features of the first resolution into high-resolution fused BEV features of the second resolution (e.g., 0.5 m / px) higher than the first resolution. The process of extracting high-resolution fused BEV features by the low-resolution feature extraction unit (320) and the super-resolution unit (330) will be described in more detail below with reference to FIG. 7.
[0062] The target information extraction unit (340) can perform target tasks (e.g., map segmentation, object detection, etc.) for recognizing the surrounding environment of the driving device based on the high-resolution fused BEV features. For example, the information processing system can extract target information based on the high-resolution fused BEV features using a decoder suitable for the target task. For example, the target information extraction unit (340) can generate a semantic map based on the high-resolution fused BEV features using a first decoder that performs map segmentation. As another example, the target information extraction unit (340) can generate object detection information based on the high-resolution fused BEV features using a second decoder that performs object detection.
[0063] FIG. 4 is a diagram illustrating an example of first sensor data according to an embodiment of the present disclosure, and FIG. 5 is a diagram illustrating an example of second sensor data according to an embodiment of the present disclosure. According to an embodiment, the first sensor and the second sensor may be different types of sensors, and the first sensor data and the second sensor data may be data of different modalities. For example, the first sensor may include a camera (C1, C2, C3, C4, C5, C6), the second sensor may include a lidar, the first sensor data may include a plurality of viewpoint images (410, 420, 430, 440, 450, 460), and the second sensor data may include point cloud data (500).
[0064] FIG. 4 illustrates an example of first sensor data. As illustrated, a driving device (400) may be equipped with a plurality of cameras (C1, C2, C3, C4, C5, C6). The plurality of cameras (C1, C2, C3, C4, C5, C6) may capture images from different viewpoints. That is, the plurality of cameras (C1, C2, C3, C4, C5, C6) may be arranged to capture images from different directions so as to cover all 360° directions around the driving device (400). According to one embodiment, the first sensor data may include a plurality of viewpoint images (410, 420, 430, 440, 450, 460) captured by each of the plurality of cameras (C1, C2, C3, C4, C5, C6) mounted on the driving device (400). For example, the multiple viewpoint images (410, 420, 430, 440, 450, 460) may all be captured by each of the multiple cameras (C1, C2, C3, C4, C5, C6) at the same time, or may be images of different viewpoints captured by each of the multiple cameras (C1, C2, C3, C4, C5, C6) when the driving device (400) is at the same position.
[0065] FIG. 5 illustrates an example of second sensor data. The second sensor data may be data acquired by the second sensor at the same time as the first sensor data or data acquired by the second sensor at the same location as the location of the driving device (400) when the first sensor data was acquired. For example, the driving device (400) may further be equipped with a lidar, and the second sensor data may include point cloud data (500) acquired by the lidar at the same time as the times when the plurality of viewpoint images (410, 420, 430, 440, 450, 460) were captured or data acquired by the lidar at the same location as the location of the driving device (400) when the plurality of viewpoint images (410, 420, 430, 440, 450, 460) were acquired by the first sensor.
[0066] FIG. 6 is a diagram illustrating an example of extracting BEV features for each sensor according to one embodiment of the present disclosure. According to one embodiment, an information processing system may extract a plurality of sensor BEV features (114, 124) from each of a plurality of sensor data (112, 122) acquired by a plurality of sensors mounted on a driving device. In one embodiment, the information processing system may utilize an appropriate BEV feature extraction model depending on the type of sensor from which the sensor data was acquired and / or the modality of the sensor data (112, 122).
[0067] For example, the information processing system can extract first sensor BEV features (114) for a region of interest based on first sensor data (112) including a plurality of viewpoint images using a first BEV feature extraction model (110). According to one embodiment, the first BEV feature extraction model (110) can include a first sensor encoder (610) and a view transformation model (620). The information processing system can extract a plurality of camera features based on each of the plurality of viewpoint images using the first sensor encoder (610), and then perform view transformation on the plurality of camera features using the view transformation model (620), thereby generating the first sensor BEV features (114) of the first resolution.
[0068] For example, according to one embodiment of the present disclosure, a process of extracting a first sensor BEV feature (114) of a first resolution (e.g., 2.0 m / px) for a region of interest (e.g., (-50 m, 50 m)) based on first sensor data (112) including multiple viewpoint images may be expressed by the following mathematical expression 1.
[0069]
[0070] Here, is the first sensor BEV feature (114), s is the scale factor, R is information about the region of interest and resolution, I is the first sensor data (112), E ω,i represents the first sensor encoder (610), and T represents the view transformation model (620).
[0071] For example, the information processing system can extract second sensor BEV features (124) for an area of interest based on second sensor data (122) including point cloud data using a second BEV feature extraction model (120). According to one embodiment, the second BEV feature extraction model (120) can include a second sensor encoder (630) and a flattening model (640). The information processing system can generate second sensor BEV features (124) having a first resolution by voxelizing the point cloud data, extracting lidar features using the second sensor encoder (630), and performing flattening (e.g., Z-axis flattening) on the lidar features using the flattening model (640).
[0072] For example, according to one embodiment of the present disclosure, a process of extracting a second sensor BEV feature (124) of a first resolution (e.g., 2.0 m / px) for an area of interest (e.g., (-50 m, 50 m)) based on second sensor data (122) including point cloud data may be expressed by the following mathematical expression 2.
[0073]
[0074] Here, is the second sensor BEV feature (124), s is the scale factor, R is information about the region of interest and resolution, P is the second sensor data (122), E ω,p represents the second sensor encoder (630), and Flatten represents the flattening model (640).
[0075] FIG. 7 is a diagram illustrating an example of extracting a high-resolution fused BEV feature (132) according to one embodiment of the present disclosure. An information processing system can extract a single high-resolution fused BEV feature (132) based on a first sensor BEV feature (114) and a second sensor BEV feature (124) using a high-resolution BEV feature extraction model (130). For example, the information processing system can extract a high-resolution fused BEV feature (132) of a second resolution higher than the first resolution based on a first sensor BEV feature (114) of a first resolution and a second sensor BEV feature (124) of the first resolution.
[0076] In one embodiment, the high-resolution BEV feature extraction model (130) may include a low-resolution BEV feature fusion model (710) and a super-resolution model (730). The information processing system may use the low-resolution BEV feature fusion model (710) to fuse the first sensor BEV features (114) and the second sensor BEV features (124) at low resolution to generate a low-resolution fused BEV feature (720), and may use the super-resolution model (730) to convert the low-resolution fused BEV feature (720) into a high-resolution fused BEV feature (132).
[0077] For example, first, the information processing system can generate a low-resolution fused BEV feature (720) based on the first sensor BEV feature (114) and the second sensor BEV feature (124) using a low-resolution BEV feature fusion model (710). Specifically, the low-resolution BEV feature fusion model (710) can include a fusion model (712) and an enhancing model (714). The information processing system can fuse the first sensor BEV feature (114) and the second sensor BEV feature (124) using the fusion model (712). Any fusion model that fuses a plurality of features to output a single fused feature can be used as the fusion model (712). According to one embodiment, the fusion model (712) can include at least one of a convolutional neural network (CNN) or a transformer-based model. The information processing system can then generate a low-resolution fused BEV feature (720) for the region of interest by enhancing the fused BEV feature at a low resolution using an enhancement model (714). In one embodiment, the low-resolution fused BEV feature (720) can have a first resolution (e.g., 2.0 m / px).
[0078] For example, according to one embodiment of the present disclosure, a process of generating a low-resolution fused BEV feature (720) based on a first sensor BEV feature (114) and a second sensor BEV feature (124) can be expressed by the following mathematical expression 3.
[0079]
[0080] Here, is a low-resolution fusion BEV feature (720), is an enhancing model (714), z LR is a fused BEV feature (before passing the enhancing model (714)), is a fusion model (712), is the first sensor BEV feature (114), represents the second sensor BEV feature (124) and s represents the scale factor.
[0081] In one embodiment, the enhancement model (714) may include a feature pyramid network (FPN) and / or a multi-head self-attention layer. For example, the process of generating a low-resolution fused BEV feature (720) for a region of interest by enhancing the fused BEV feature using the enhancement model (714) including the FPN and the multi-head self-attention layer may be expressed by the following mathematical expression (4).
[0082]
[0083] Here, j is the index, The process of predicting can be repeated recursively M times. Q, K and V are the query, key and value of the multi-head self-attention layer, respectively, and the Mth predicted value ( ) is a low-resolution fusion BEV feature ( , 720) can be.
[0084] Then, the information processing system can convert the low-resolution fused BEV feature (720) into a high-resolution fused BEV feature (132) using the super-resolution model (730). For example, the information processing system can convert the low-resolution fused BEV feature (720) of a first resolution (e.g., 2.0 m / px) into a high-resolution fused BEV feature (132) of a second resolution (e.g., 0.5 m / px) higher than the first resolution.
[0085] When the scale factor is s, the length of the area corresponding to the height or width of one pixel of data can be 1 / s times. Accordingly, when the low-resolution fused BEV feature (720) for the region of interest is super-resolved with a scale factor s≥1, the height and width of the super-resolved high-resolution fused BEV feature (132) can each increase by s times. For example, when the low-resolution fused BEV feature (720) of 2.0 m / px is super-resolved with a scale factor s=4, it can be converted into a high-resolution fused BEV feature (132) with a resolution of 2.0 * 1 / 4 = 0.5 m / px. That is, if the area of interest is (-50m, 50m), the size of the low-resolution fused BEV feature (720) may be 50(H) X 50(W), and the size of the low-resolution fused BEV feature (720) may be 200(H) X 200(W).
[0086] Any model that super-resolves low-resolution features into high-resolution features can be used as the super-resolution model (730). For example, a generative adversarial network (GAN)-based super-resolution model, a peak signal-to-noise ratio (PSNR-oriented) super-resolution model, etc. can be used as the super-resolution model (730).
[0087] For example, the process of converting a low-resolution fused BEV feature (720) into a high-resolution fused BEV feature (132) using a PSNR-oriented super-resolution model can be expressed by the following mathematical expression 5.
[0088]
[0089] Here, is a high-resolution fusion BEV feature (132), is a low-resolution fusion BEV feature (720), is a super-resolution model (730), and is an artificial neural network ( ) operations and pixel shuffle (PS) can be performed sequentially.
[0090] As described above, by performing BEV feature fusion at low resolution, which can cause a bottleneck when the data size is large, fused BEV features can be efficiently extracted using limited computing resources. Furthermore, by super-resolving the BEV features fused at low resolution, high-resolution fused BEV features (132) can be extracted, which provide more detailed information about the surrounding environment. In other words, both computing resource savings and detailed environmental recognition can be achieved simultaneously.
[0091] FIG. 8 is a diagram illustrating an example of extracting target information (142) based on high-resolution fused BEV features (132) according to one embodiment of the present disclosure. According to one embodiment, an information processing system may perform a target task (e.g., map segmentation, object detection, etc.) for recognizing the surrounding environment of a driving device based on the extracted high-resolution fused BEV features (132). For example, the information processing system may extract target information (142) for a region of interest based on the high-resolution fused BEV features (132) using a decoder (140) suitable for the target task. According to one embodiment, the decoder (140) may include a convolutional artificial neural network (CNN).
[0092] As a specific example, the information processing system may generate a semantic map for a region of interest based on high-resolution fused BEV features (132) using a first decoder that performs map segmentation. As another example, the information processing system may generate object detection information for a region of interest based on high-resolution fused BEV features (132) using a second decoder that performs object detection.
[0093] For example, the process of extracting target information (142) based on high-resolution fused BEV features (132) using a decoder (140) can be expressed by the following mathematical expression 6.
[0094]
[0095] Here, The extracted target information (142), is a decoder (140), represents a high-resolution fusion BEV feature (132).
[0096] FIG. 9 is a diagram illustrating an example of training a model used in the present method according to one embodiment of the present disclosure. According to one embodiment, the information processing system can obtain reference target information (900) indicating correct answer data for a target task. The information processing system can calculate a loss based on the extracted target information (142) and the reference target information (900). For example, the information processing system can calculate a cross-entropy loss or a focal loss based on the extracted target information (142) and the reference target information (900). The information processing system can train the model by updating at least some of the weights of the models (130, 710, 712, 714, 730, 140) so that the calculated loss is minimized.
[0097] In one embodiment, the weights of most of the model (e.g., the low-resolution BEV feature fusion model (710)) can be frozen, and only at least some of the weights of the super-resolution model (730) and / or the decoder (140) can be updated to fine-tune the model for different target tasks, different data sizes, and / or different resolutions, in order to train the model for different target tasks, different region-of-interest sizes, different data sizes, and / or different resolutions.
[0098] FIG. 10 is a diagram illustrating examples of target information according to one embodiment of the present disclosure. FIG. 10 illustrates examples (1010, 1020, 1030, 1040) of semantic maps extracted based on fused BEV features generated using various methods and a corresponding correct semantic map (GT, 1000). In the semantic maps illustrated in FIG. 10, the first class represents a road, the second class represents a lane, the third class represents a stop line, and the fourth class represents a sidewalk. The first example (1010), the second example (1020), and the third example (1030) are comparative examples, and the fourth example (1040) is an example of a semantic map extracted based on high-resolution fused BEV features according to one embodiment of the present disclosure. Referring to FIG. 10, it can be seen that the fourth example (1040) according to one embodiment of the present disclosure provides a high-quality semantic map most similar to the correct semantic map. Thus, according to the present disclosure, by performing a target task using high-resolution fused BEV features that precisely represent information about the surrounding environment, the surrounding environment can be accurately recognized, thereby enabling the driving device to drive safely.
[0099] FIG. 11 is a diagram illustrating an example of generating hybrid target information (1150) according to one embodiment of the present disclosure. According to one embodiment, the information processing system can generate target information for a region of interest based on multiple fused BEV features having different resolutions.
[0100] For example, the information processing system may divide the area of interest into a first area including a point where the driving device is located (i.e., an area near the driving device) and a second area excluding the first area from the area of interest.
[0101] The information processing system can convert the low-resolution fused BEV feature (720) into a first resolution fused BEV feature (1112) using a first super-resolution model (1110) having a scale factor s1 (e.g., s1=1). Then, the information processing system can extract first resolution target information (1122) for the region of interest based on the first resolution fused BEV feature (1112) using a first decoder (1120). Furthermore, the information processing system can convert the low-resolution fused BEV feature (720) into a second resolution fused BEV feature (1132) using a second super-resolution model (1120) having a scale factor s2 (e.g., s2>s1). Then, the information processing system can extract second resolution target information (1142) for the first region based on the first resolution fused BEV feature (1112) using the second decoder (1140). The information processing system can generate hybrid target information (1150) based on information (1152) for the second region among the first resolution target information (1122) and information (1154) for the first region among the second resolution target information (1142). In one embodiment, when the scale factor s1 applied to the first super-resolution model (1110) is 1, the low-resolution fused BEV feature (720) may be used as the first resolution fused BEV feature (1112).
[0102] As described above, the information processing system can efficiently provide information about the surrounding environment under limited resources by providing high-quality target information based on high-resolution BEV features that provide detailed information about the surrounding environment for an area close to the driving device, and providing target information based on low-resolution BEV features for an area not close to the driving device.
[0103] FIG. 12 is a flowchart illustrating an example of a bird's-eye view information extraction method (1200) according to one embodiment of the present disclosure. The bird's-eye view information extraction method (1200) may be performed by at least one processor (e.g., at least one processor of an information processing system). According to one embodiment, the bird's-eye view information extraction method (1200) may be performed in real time while the driving device is driving.
[0104] First, the processor may receive first sensor data and second sensor data acquired by a first sensor and a second sensor mounted on the driving device at a specific point where the driving device is located (S1210). According to one embodiment, the first sensor and the second sensor may be different types of sensors, and the first sensor data and the second sensor data may be data of different modalities. For example, the first sensor may include a plurality of cameras mounted on the driving device, and the first sensor data may include a plurality of viewpoint images acquired by each of the plurality of cameras. Additionally or alternatively, the second sensor may include a lidar mounted on the driving device, and the second sensor data may include point cloud data acquired by the lidar.
[0105] Then, the processor can extract first sensor BEV features and second sensor BEV features, which are bird's-eye view information for an area of interest, based on the first sensor data and the second sensor data (S1220). Here, the area of interest can include an area within a predefined range based on a specific point.
[0106] For example, if the first sensor data includes multiple viewpoint images acquired by each of the multiple cameras, the processor may perform view transformation on the multiple viewpoint images to extract the first sensor BEV features. Additionally or alternatively, if the second sensor data includes point cloud data acquired by the lidar, the processor may perform flattening on the point cloud data to extract the second sensor BEV features.
[0107] Then, the processor can generate a first fused BEV feature of a first resolution based on the first sensor BEV feature and the second sensor BEV feature using a low-resolution BEV feature fusion model (S1230). According to one embodiment, the low-resolution BEV feature fusion model can include a multi-head self-attention layer. Additionally or alternatively, the low-resolution BEV feature fusion model can include a convolutional artificial neural network.
[0108] Thereafter, the processor can use the super-resolution model to convert the first fused BEV feature into a second fused BEV feature of a second resolution higher than the first resolution (S1240).
[0109] In one embodiment, the processor may extract target information based on the second fused BEV features using a decoder that performs a target task. For example, the target task may include map segmentation, and the target information may include a semantic map for the region of interest. In another example, the target task may include object detection, and the target information may include object detection information for the region of interest.
[0110] In one embodiment, the region of interest may include a first region including a specific point where the driving device is located, and a second region excluding the first region. The processor may generate hybrid target information including target information extracted based on BEV features of different resolutions for the first region and the second region.
[0111] For example, the processor may use a first decoder to extract first target information based on a first fused BEV feature of a first resolution, and may use a second decoder to extract second target information based on a second fused BEV feature of a second resolution. Then, the processor may generate hybrid target information based on information about a second region among the first target information and information about a first region among the second target information.
[0112] As another example, the processor may use the first super-resolution model to transform the first fused BEV feature into a second fused BEV feature at a second resolution. Furthermore, the processor may use the second super-resolution model to transform the first fused BEV feature into a third fused BEV feature at a third resolution for the region of interest, wherein the third resolution may be the same as or higher than the first resolution and lower than the second resolution. Then, the processor may use the first decoder to extract first target information based on the third fused BEV feature at the third resolution, and use the second decoder to extract second target information based on the second fused BEV feature at the second resolution. Thereafter, the processor may generate hybrid target information based on the information about the second region among the first target information and the information about the first region among the second target information.
[0113] According to one embodiment, the information processing system can train the model and / or decoder used in the bird's-eye view information extraction method (1200). For example, the information processing system can calculate a loss based on target information and reference target information for the region of interest. Then, the processor can train the model and / or decoder by updating at least some of the weights of the model and decoder based on the loss. In one embodiment, the information processing system can fine-tune the model and / or decoder by updating only at least some of the weights of the super-resolution model or the weights of the decoder while keeping the weights of the low-resolution BEV feature fusion model fixed.
[0114] The flowchart of FIG. 12 and the description above are merely examples, and the present disclosure is not limited thereto. For example, at least some of the steps of the method may be added / changed / deleted, or the order of at least some of the steps of the method may be changed.
[0115] The above-described method may be provided as a computer program stored on a computer-readable recording medium for execution on a computer. The medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording means or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program instructions, including ROM, RAM, and flash memory. In addition, examples of other media may include recording or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.
[0116] The methods, operations, or techniques of the present disclosure may be implemented by various means. For example, these techniques may be implemented in hardware, firmware, software, or a combination thereof. Those skilled in the art will appreciate that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software will depend on the particular application and the design requirements imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but such implementations should not be construed as departing from the scope of the present disclosure.
[0117] In a hardware implementation, the processing units used to perform the techniques may be implemented within one or more ASICs, DSPs, GPUs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described herein, a computer, or a combination thereof.
[0118] Accordingly, the various exemplary logical blocks, modules, and circuits described in connection with the present disclosure may be implemented or performed by any combination of a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or those designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0119] In a firmware and / or software implementation, the techniques may be implemented as instructions stored on a computer-readable medium, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, a compact disc (CD), a magnetic or optical data storage device, etc. The instructions may be executable by one or more processors and may cause the processor(s) to perform certain aspects of the functionality described herein.
[0120] When implemented in software, the techniques may be stored on or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media includes both computer storage media and communication media, including any medium that facilitates transfer of a computer program from one place to another. Storage media may be any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. In addition, any connection is suitably made to a computer-readable medium.
[0121] For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, digital subscriber line, or wireless technologies such as infrared, radio, and microwave are included within the definition of media. Disk and disc, as used herein, includes compact discs, laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks usually reproduce data magnetically, whereas discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0122] A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. Alternatively, the processor and the storage medium may reside as discrete components in the user terminal.
[0123] While the embodiments described above have been described as utilizing aspects of the presently disclosed subject matter in one or more standalone computer systems, the present disclosure is not limited thereto and may be implemented in conjunction with any computing environment, such as a network or distributed computing environment. Furthermore, aspects of the present disclosure may be implemented in multiple processing chips or devices, and storage may be similarly affected across multiple devices. Such devices may include personal computers, network servers, and portable devices.
[0124] While the present disclosure has been described in connection with certain embodiments herein, various modifications and variations may be made without departing from the scope of the present disclosure, which would be apparent to those skilled in the art. Furthermore, such modifications and variations are intended to fall within the scope of the claims appended to this specification.
Claims
1. A method for extracting bird's eye view information, performed by at least one processor, A step of receiving first sensor data and second sensor data acquired by a first sensor and a second sensor mounted on a driving device at a specific point where the driving device is located; A step of extracting a first sensor BEV feature and a second sensor BEV feature, which are bird's eye view (BEV) information for an area of interest, based on the first sensor data and the second sensor data; A step of generating a first fused BEV feature of a first resolution based on the first sensor BEV feature and the second sensor BEV feature using a low-resolution BEV feature fusion model; and A step of converting the first fused BEV feature into a second fused BEV feature of a second resolution higher than the first resolution using a super-resolution model. Including, A method for extracting bird's eye view information, wherein the above region of interest includes an area within a predefined range based on the specific point.
2. In paragraph 1, The first sensor and the second sensor are different types of sensors, A method for extracting bird's eye view information, wherein the first sensor data and the second sensor data are data of different modalities.
3. In paragraph 1, The first sensor includes a plurality of cameras mounted on the driving device, A method for extracting bird's eye view information, wherein the first sensor data includes a plurality of viewpoint images acquired by each of the plurality of cameras.
4. In paragraph 3, The step of extracting the first sensor BEV feature and the second sensor BEV feature is: A step of extracting a first sensor BEV feature by performing view transformation on the above multiple viewpoint images. A method for extracting bird's eye view information, including:
5. In paragraph 1, The second sensor includes a LiDAR mounted on the driving device, A method for extracting bird's eye view information, wherein the second sensor data includes point cloud data acquired by the lidar.
6. In paragraph 5, The step of extracting the first sensor BEV feature and the second sensor BEV feature is: A step of extracting second sensor BEV features by performing flattening on the above point cloud data. A method for extracting bird's eye view information, including:
7. In paragraph 1, The above low-resolution BEV feature fusion model is a method for extracting bird's eye view information, including a multi-head self-attention layer.
8. In paragraph 1, The above low-resolution BEV feature fusion model is a method for extracting bird's eye view information, including a convolutional neural network (CNN).
9. In paragraph 1, A step of extracting target information based on the second fused BEV feature using a decoder that performs a target task. A method for extracting bird's eye view information, which further includes:
10. In paragraph 9, A method for extracting bird's eye view information, wherein the target task includes map segmentation, and the target information includes a semantic map for the region of interest.
11. In paragraph 9, A method for extracting bird's eye view information, wherein the target task includes object detection, and the target information includes object detection information for the region of interest.
12. In paragraph 9, A step of calculating a loss based on the target information and reference target information for the region of interest; and Based on the above loss, a step of updating at least some of the weights of the super-resolution model or the weights of the decoder. Including more, A method for extracting bird's eye view information, wherein the weights of the above low-resolution BEV feature fusion model are fixed.
13. In paragraph 9, The region of interest includes a first region including the specific point and a second region excluding the first region in the region of interest, The step of extracting the above target information is: A step of extracting first target information based on the first fused BEV feature of the first resolution using a first decoder; A step of extracting second target information based on the second fused BEV feature of the second resolution using a second decoder; and A step of generating hybrid target information based on information about the second area among the first target information and information about the first area among the second target information. A method for extracting bird's eye view information, including:
14. In paragraph 9, The region of interest includes a first region including the specific point and a second region excluding the first region in the region of interest, The step of converting to the above second fusion BEV feature is: A step of converting the first fused BEV feature into the second fused BEV feature of the second resolution using the first super-resolution model, The above method, A step of converting the first fused BEV feature into a third fused BEV feature of a third resolution for the region of interest using a second super-resolution model, wherein the third resolution is greater than or equal to the first resolution and less than or equal to the second resolution. Including more, The step of extracting the above target information is: A step of extracting first target information based on the third fused BEV feature of the third resolution using a first decoder; A step of extracting second target information based on the second fused BEV feature of the second resolution using a second decoder; and A step of generating hybrid target information based on information about the second area among the first target information and information about the first area among the second target information. A method for extracting bird's eye view information, including:
15. In paragraph 1, A method for extracting bird's eye view information, wherein the method is performed in real time while the driving device is driving.
16. A computer program stored on a computer-readable recording medium for executing the method according to any one of Articles 1 to 15 on a computer.
17. As an information processing system, memory; and At least one processor connected to said memory and configured to execute at least one computer-readable program contained in said memory Including, At least one program above, Receive first sensor data and second sensor data acquired by a first sensor and a second sensor mounted on the driving device at a specific point where the driving device is located, Based on the first sensor data and the second sensor data, a first sensor BEV feature and a second sensor BEV feature, which are bird's eye view information for an area of interest, are extracted, Using a low-resolution BEV feature fusion model, a first fused BEV feature of a first resolution is generated based on the first sensor BEV feature and the second sensor BEV feature, Includes instructions for converting the first fused BEV feature into a second fused BEV feature of a second resolution higher than the first resolution using a super-resolution model, An information processing system, wherein the above area of interest includes an area within a predefined range based on the specific point.
Citation Information
Patent Citations
Moving image processing device, moving image processing method, and moving image processing program
JP2009253667A
Surround View Monitoring System and Image Signal Processing Method thereof
KR101809727B1
Electronic device for obtaining three-dimension object based on camera and radar sensor fusion, and operaing method thereof
KR102168753B1
Method for preparing iron phosphate and by-product fertilizer using ammonium phosphate
KR102850486B1
Autonomous vehicle environmental perception software architecture
US20220398851A1