Using hyperbolic projection to change the virtual view perspective
By combining the physical camera and depth sensor, the polar geometry is determined and the directional cost function is minimized to generate a high-quality virtual view. This solves the artifact problem when the viewpoint changes in the prior art and enables flexible viewpoint adjustment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-20
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies struggle to maintain the 3D structure of the captured scene and change the perspective of the virtual scene when generating virtual views, resulting in unwanted artifacts in the virtual images.
By capturing the input scene using a physical camera and a depth sensor, the controller determines the polar geometry between the actual pose of the capture device and the desired pose of the virtual camera. The controller then uses polar coordinates to resample the pixels of the input image and minimize the directional cost function to generate the output image of the virtual camera.
It enables flexible changes in the perspective of the virtual scene while maintaining the three-dimensional structure, reduces artifacts in the virtual image, and improves the quality of the virtual view.
Smart Images

Figure CN114998427B_ABST
Abstract
Description
Technical Field
[0001] The technical field generally relates to generating virtual views based on captured image data. In particular, the description relates to changes in the virtual view's perspective. More specifically, the description relates to systems and methods for generating a virtual view of a virtual camera based on an input scene captured by a capturing device, which includes a physical camera and depth sensors located in the same location or spatially separated. Background Technology
[0002] Modern vehicles are typically equipped with one or more optical cameras configured to provide image data to vehicle occupants. For example, the image data shows a predetermined viewpoint of the environment surrounding the vehicle.
[0003] Under certain conditions, it may be desirable to shift the perspective to image data provided by optical cameras. For this purpose, a so-called virtual camera is used, and image data captured by one or more physical cameras is modified to display the captured scene from a different, desired perspective; the modified image data can be referred to as a virtual scene or output image. The desired perspective on the virtual scene can be changed according to the occupants' wishes. A virtual scene can be generated based on multiple images captured from different perspectives. However, merging image data from image sources located in different locations can create unwanted artifacts in the virtual image.
[0004] Therefore, it is desirable to provide systems and methods for generating virtual views of scenes captured by physical cameras, which have improved quality of virtual scenes, maintain the three-dimensional structure of the captured scenes, and are able to change the viewing angle of the virtual scenes.
[0005] Furthermore, other desirable features and characteristics of the invention will become apparent from the following detailed description and appended claims, taking into account the accompanying drawings and the foregoing technical and background information. Summary of the Invention
[0006] A method is provided for generating a virtual view of a virtual camera based on an input scene. In one embodiment, the method includes: capturing the input scene by a capture device; determining the actual pose of the capture device by a controller; determining the desired pose of a virtual camera for displaying the virtual view by the controller; defining the polar geometry between the actual pose of the capture device and the desired pose of the virtual camera by the controller; and generating an output image for the virtual camera by the controller based on the polar relationship between the actual pose of the capture device, the input scene, and the desired pose of the virtual camera.
[0007] In one embodiment, the method includes capturing an input scene by a physical camera and a depth sensor located at the same position. Capturing the input scene by the capture device includes capturing an input image by the physical camera; assigning depth information to pixels of the input image by the depth sensor; determining the actual pose of the capture device by a controller includes determining the actual pose of the physical camera by the controller; defining the polar geometry between the actual pose of the capture device and the desired pose of the virtual camera by the controller includes defining the polar geometry between the actual pose of the physical camera and the desired pose of the virtual camera by the controller; and generating an output image for the virtual camera by the controller includes: resampling the depth information of the pixels of the input image in polar coordinates by the controller; identifying target pixels on the input epipolar lines of the physical camera by the controller; generating a disparity map for one or more output epipolar lines of the virtual camera by the controller; and generating an output image based on one or more output epipolar lines by the controller.
[0008] In one embodiment, identifying target pixels on the input epipolar line of a physical camera by the controller includes minimizing a directional cost function by the controller to identify the target pixels.
[0009] In another embodiment, for each output pixel of each of one or more output epipolar lines, a directional cost function minimized by the controller is performed to identify the target pixel.
[0010] In another embodiment, the method further includes the controller determining the disparity along the output epipolar line for each output pixel by minimizing the directional cost function.
[0011] In another embodiment, generating an output image by the controller based on one or more output epipolar lines includes generating the output image by the controller acquiring pixels by determining a parallax.
[0012] In another embodiment, the directional cost function is defined as:
[0013]
[0014] And the directional cost function is minimized according to the following:
[0015]
[0016] in:
[0017] DC is the directional cost function;
[0018] m is the pixel along the input epipolar line of the physical camera;
[0019] It is a unit vector pointing from the center of the virtual camera to the pixel direction on the output epipolar line of the virtual camera. For this unit vector, the corresponding physical pixel m along the input epipolar line of the physical camera will be identified.
[0020] It is the vector between the physical camera center and the virtual camera center;
[0021] It corresponds to the 3D point position of pixel m;
[0022] Is for a given The physical pixel on the input epipolar line corresponding to the DC minimum value.
[0023] In another embodiment, the assignment of depth information to pixels of the input image by the depth sensor includes the assignment of depth information by the depth sensor to each pixel captured by the physical camera.
[0024] In another embodiment, the assignment of depth information to pixels of the input image by the depth sensor includes the depth information of each pixel captured by the physical camera determined by the depth sensor based on the input image.
[0025] In another embodiment, the method includes capturing an input scene by a physical camera and a depth sensor spatially separated from the physical camera. Capturing the input scene by the capture device includes capturing an input image by the physical camera and capturing depth data associated with the input image by the depth sensor. Determining the actual pose of the capture device by the controller includes determining the actual pose of the physical camera and the actual pose of the depth sensor; defining the polar geometry between the actual pose of the capture device and the desired pose of the virtual camera by the controller includes defining the polar geometry between the actual pose of the depth sensor and the desired pose of the virtual camera; and generating an output image for the virtual camera by the controller includes: generating dense depth data for the desired pose of the virtual camera by the controller; projecting the dense depth data for the desired pose of the virtual camera onto the input image by the controller; and generating the output image of the virtual camera based on the dense depth data projected onto the input image by the controller.
[0026] In one embodiment, generating dense depth data for the desired pose of the virtual camera by the controller includes minimizing an directional cost function by the controller to estimate the depth of a pixel.
[0027] In another embodiment, for each output pixel of one or more output epipolar lines of the output image, the controller performs a directional cost function to estimate the pixel's depth.
[0028] In another embodiment, the method further includes, when the depth sensor is a dense depth sensor, the depth sensor assigns depth information to the pixels of the input image, and the controller resamples the depth information of the pixels of the input image in polar coordinates.
[0029] In another embodiment, the method further includes, when the depth sensor is a sparse depth sensor, establishing a lookup table and mapping the epipolar lines of the output image to a set of points near the corresponding epipolar lines of the depth sensor.
[0030] In another embodiment, the method further includes minimizing a directional cost function over a set of relevant points in the point cloud by a controller. Preferably, the point cloud is converted into a dense depth map via triangulation or Voronoi mosaicking.
[0031] In another embodiment, the directional cost function is defined as:
[0032]
[0033] And the directional cost function is minimized according to the following:
[0034]
[0035] in:
[0036] DC is the directional cost function;
[0037] m is the pixel along the input epipolar line of the physical camera;
[0038] It is a unit vector pointing from the center of the virtual camera to the pixel direction on the output epipolar line of the virtual camera. For this unit vector, the corresponding physical pixel m along the input epipolar line of the physical camera will be identified.
[0039] It is the vector between the physical camera center and the virtual camera center;
[0040] It corresponds to the 3D point position of pixel m;
[0041] Is for a given The physical pixel on the input epipolar line corresponding to the DC minimum value.
[0042] A system is provided for generating a virtual view of a virtual camera based on an input scene. The system includes a capture device having a physical camera and a depth sensor, and a controller. The capture device is configured to capture the input scene. The controller is configured to determine the actual pose of the capture device, determine the desired pose of the virtual camera for displaying the virtual view, define the polar geometry between the actual pose of the capture device and the desired pose of the virtual camera, and generate an output image for the virtual camera based on the polar relationship between the actual pose of the capture device, the input scene, and the desired pose of the virtual camera.
[0043] Preferably, the system is configured to implement the functions described above with reference to the embodiments of the method. In particular, the system is configured to implement the steps described in embodiments where the physical camera and depth sensor are located in the same location, and the steps described in embodiments where the physical camera and depth sensor are spatially separated.
[0044] In one embodiment, the physical camera and the depth sensor of the capture device are located in the same location.
[0045] In one embodiment, the physical camera and depth sensor are spatially separated, for example, isolated from each other.
[0046] A vehicle is provided, comprising a system for generating a virtual view of a virtual camera based on an input scene. The system includes a capture device having a physical camera and a depth sensor, and a controller. The capture device is configured to capture the input scene. The controller is configured to determine the actual pose of the capture device, determine the desired pose of the virtual camera for displaying the virtual view, define the polar geometry between the actual pose of the capture device and the desired pose of the virtual camera, and generate an output image for the virtual camera based on the polar relationship between the actual pose of the capture device, the input scene, and the desired pose of the virtual camera.
[0047] Preferably, the system included in the vehicle is the system described with reference to one of the above embodiments. Attached Figure Description
[0048] Exemplary embodiments will now be described in conjunction with the following accompanying drawings, wherein the same reference numerals denote the same elements, and wherein:
[0049] Figure 1 This is a schematic diagram of a vehicle with a controller that enables the generation of virtual views;
[0050] Figure 2 It is a schematic diagram based on the polar geometry principle of two cameras;
[0051] Figure 3 This is a schematic diagram of the extreme reprojection of pixels according to the first embodiment;
[0052] Figure 4 This is a schematic diagram of a method for generating a virtual view according to the first embodiment;
[0053] Figure 5 This is a schematic diagram of a method for generating a virtual view according to the second embodiment;
[0054] Figure 6 This is a schematic diagram of the extreme reprojection of pixels according to the second embodiment;
[0055] Figure 7 This is a schematic diagram of the extreme reprojection of pixels according to the second embodiment. Detailed Implementation
[0056] The following detailed description is exemplary in nature only and is not intended to limit application and use. Furthermore, it is not intended to be bound by any express or implied theory set forth in the foregoing technical fields, background art, summary of the invention, or the following detailed description. As used herein, the term "module" refers individually or in any combination of any hardware, software, firmware, electronic control components, processing logic, and / or processor devices, including but not limited to: application-specific integrated circuits (ASICs), electronic circuits, processors (shared, dedicated, or grouped), and memory executing one or more software or firmware programs, combinational logic circuits, and / or other suitable components providing the described functionality.
[0057] Embodiments of this disclosure can be described herein based on functional and / or logical block components and various processing steps. It should be understood that such block components can be implemented by any number of hardware, software, and / or firmware components configured to perform specified functions. For example, embodiments of this disclosure may employ various integrated circuit components, such as memory elements, digital signal processing elements, logic elements, lookup tables, etc., which can perform various functions under the control of one or more microprocessors or other control devices. Furthermore, those skilled in the art will understand that embodiments of this disclosure can be practiced in conjunction with any number of systems, and the systems described herein are merely exemplary embodiments of this disclosure.
[0058] For the sake of brevity, conventional techniques relating to signal processing, data transmission, signaling, control, and other functional aspects of the system (as well as the various operating components of the system) may not be described in detail herein. Furthermore, the connecting lines shown in the various figures included herein are intended to represent exemplary functional relationships and / or physical connections between the various elements. It should be noted that many alternative or additional functional relationships or physical connections may exist in the embodiments of this disclosure.
[0059] refer to Figure 1 The image illustrates a vehicle 10 according to various embodiments. The vehicle 10 generally includes a chassis 12, a body 14, front wheels 16, and rear wheels 18. The body 14 is disposed on the chassis 12 and substantially encloses the components of the vehicle 10. The body 14 and chassis 12 may together form a frame. The wheels 16 and 18 are each rotatably connected to the chassis 12 near a corresponding corner of the body 14.
[0060] In various embodiments, vehicle 10 is an autonomous vehicle. An autonomous vehicle is, for example, a vehicle automatically controlled to transport passengers from one location to another. Vehicle 10 is described as a passenger car in the illustrated embodiments, but it should be understood that any other means of transportation may be used, including motorcycles, trucks, sport utility vehicles (SUVs), recreational vehicles (RVs), boats, aircraft, etc. In exemplary embodiments, the autonomous vehicle is a Level 2 or higher level automation system. A Level 2 automation system signifies “partial automation.” However, in other embodiments, the autonomous vehicle may be a so-called Level 3, Level 4, or Level 5 automation system. A Level 3 automation system signifies conditional automation. A Level 4 system signifies “high automation,” referring to the driving mode-specific performance of the automated driving system for all aspects of a dynamic driving task, even if the human driver does not appropriately respond to intervention requests. A Level 5 system signifies “full automation,” referring to the full-time performance of the automated driving system for all aspects of a dynamic driving task under all road and environmental conditions manageable by a human driver.
[0061] However, it should be understood that vehicle 10 can also be a conventional vehicle without any autonomous driving capabilities. Vehicle 10 can implement the functions and methods described herein for generating virtual views and using hypergravity projection to change the virtual view perspective, in order to assist the driver of vehicle 10.
[0062] As shown, vehicle 10 typically includes a propulsion system 20, a transmission system 22, a steering system 24, a braking system 26, a sensor system 28, an actuator system 30, at least one data storage device 32, at least one controller 34, and a communication system 36. In various embodiments, the propulsion system 20 may include an internal combustion engine, an electric motor such as a traction motor, and / or a fuel cell propulsion system. The transmission system 22 is configured to transmit power from the propulsion system 20 to wheels 16 and 18 according to a selectable speed ratio. According to various embodiments, the transmission system 22 may include a stepped-ratio automatic transmission, a continuously variable transmission (CVT), or other suitable transmission. The braking system 26 is configured to provide braking torque to wheels 16 and 18. In various embodiments, the braking system 26 may include friction brakes, line brakes, regenerative braking systems such as electric motors, and / or other suitable braking systems. The steering system 24 affects the position of wheels 16 and 18. Although shown as including a steering wheel for illustrative purposes, in some embodiments contemplated within the scope of this disclosure, the steering system 24 may not include a steering wheel.
[0063] Sensor system 28 includes one or more sensing devices 40a-40n that sense observable conditions of the external and / or internal environments of vehicle 10. Sensing devices 40a-40n may include, but are not limited to, radar, lidar, GPS, optical cameras, thermal cameras, ultrasonic sensors, and / or other sensors. Actuator system 30 includes one or more actuator devices 42a-42n that control one or more vehicle features, such as, but not limited to, propulsion system 20, drivetrain 22, steering system 24, and braking system 26. In various embodiments, vehicle features may also include internal and / or external vehicle features, such as, but not limited to, doors, trunk, and cabin features, such as ventilation, music, lighting, etc. (not numbered).
[0064] Communication system 36 is configured to conduct wireless communication with other entities 48, such as, but not limited to, other vehicles (“V2V” communication), infrastructure (“V2I” communication), remote systems and / or personal devices (about Figure 2 (Description in more detail below). In an exemplary embodiment, communication system 36 is a wireless communication system configured to communicate using the IEEE 802.11 standard or via a wireless local area network (WLAN) using cellular data communication. However, additional or alternative communication methods, such as dedicated short-range communication (DSRC) channels, are also contemplated within the scope of this disclosure. A DSRC channel refers to a one-way or two-way short-to-medium-range wireless communication channel specifically designed for automotive use, along with a set of corresponding protocols and standards.
[0065] Data storage device 32 stores data for functions used to automatically control vehicle 10. In various embodiments, data storage device 32 stores a defined map of the navigable environment. In various embodiments, the defined map may be predefined by a remote system and obtained from the remote system (see reference). Figure 2 (To be described in more detail). For example, the defined map can be assembled by a remote system and transmitted to the autonomous vehicle 10 (wirelessly and / or via wire) and stored in a data storage device 32. It is understood that the data storage device 32 may be part of the controller 34, separate from the controller 34, or part of the controller 34 and a separate system.
[0066] The controller 34 includes at least one processor 44 and a computer-readable storage device or medium 46. The processor 44 can be any custom or commercially available processor, central processing unit (CPU), graphics processing unit (GPU), auxiliary processor among multiple processors associated with the controller 34, semiconductor-based microprocessor (in the form of a microchip or chipset), macroprocessor, any combination thereof, or any device typically used to execute instructions. For example, the computer-readable storage device or medium 46 can include volatile and non-volatile storage in read-only memory (ROM), random access memory (RAM), and persistent active memory (KAM). KAM is persistent or non-volatile memory that can be used to store various operational variables when the processor 44 is powered off. The computer-readable storage device or medium 46 can be implemented using any of many known storage devices, such as PROM (programmable read-only memory), EPROM (electrical PROM), EEPROM (electrically erasable PROM), flash memory, or any other electrical, magnetic, optical, or combined storage device capable of storing data, some of which represent executable instructions used by the controller 34 in controlling and performing the functions of the vehicle 10.
[0067] The instructions may include one or more separate programs, each including an ordered list of executable instructions for implementing logical functions. When executed by processor 44, the instructions receive and process signals from sensor system 28, execute logic, calculations, methods, and / or algorithms for automatically controlling components of vehicle 10, and generate control signals to actuator system 30 based on the logic, calculations, methods, and / or algorithms to automatically control components of vehicle 10. Figure 1 Only one controller 34 is shown, but embodiments of vehicle 10 may include any number of controllers 34 that communicate via any suitable communication medium or combination of communication media and cooperate to process sensor signals, perform logic, calculations, methods and / or algorithms, and generate control signals to automatically control the features of vehicle 10.
[0068] Typically, according to an embodiment, vehicle 10 includes a controller 34 that implements a method for generating a virtual view of a virtual camera based on an input scene captured by a capture device. The capture device includes, for example, a physical camera and a depth sensor. One of the sensing devices 40a to 40n is an optical camera, and another of these sensing devices 40a to 40n is a physical depth sensor (such as a lidar, radar, ultrasonic sensor, etc.).
[0069] Vehicle 10 is designed to perform a method for generating a virtual view of a scene captured by physical camera 40a and depth sensor 40b located in the same position or spatially separated.
[0070] In one embodiment, a method for generating a virtual view of a virtual camera based on an input scene includes the following steps: capturing the input scene by a capture device including a physical camera 40a and a depth sensor 40b; determining the actual pose of the capture device by a controller 34; determining the desired pose of the virtual camera for displaying the virtual view by the controller 34; defining the polar geometry between the actual pose of the capture device and the desired pose of the virtual camera by the controller 34; and generating an output image of the virtual camera by the controller 34 based on the polar relationship between the actual pose of the capture device, the input scene, and the desired pose of the virtual camera.
[0071] In one embodiment, the method includes the following steps: capturing an input scene by a physical camera 40a and a depth sensor 40b located at the same position; wherein capturing the input scene by the capture device includes capturing an input image by the physical camera 40a; allocating depth information to pixels of the input image by the depth sensor 40b; wherein determining the actual pose of the capture device by a controller 34 includes determining the actual pose of the physical camera 40a by the controller 34; wherein defining the polar geometry between the actual pose of the capture device and the desired pose of the virtual camera by the controller 34 includes defining the polar geometry between the actual pose of the physical camera 40a and the desired pose of the virtual camera by the controller 34; and wherein generating an output image for the virtual camera by the controller includes: resampling the depth information of the pixels of the input image in polar coordinates by the controller 34; identifying target pixels on the input polar lines of the physical camera by the controller 34; generating a disparity map for one or more output polar lines of the virtual camera by the controller 34; and generating an output image based on one or more output polar lines by the controller 34. The vehicle 10 includes a display 50 for displaying the output image to a user or occupant of the vehicle 10. Note that the sensor's attitude can be measured or estimated using a specific attitude measurement device or attitude estimation module. The controller 34 described herein obtains the attitude of the physical camera and / or depth sensor from these attitude measurement devices or attitude estimation modules, i.e., determines the attitude by reading or obtaining specific attitude values, and uses the determined attitude values in the steps of the method described herein.
[0072] The input image is captured by a physical camera 40a, such as an optical camera configured to capture a color image of the environment. The physical camera 40a is positioned at the vehicle 10 such that it can cover a specific field of view around the vehicle. Depth information is assigned to the pixels of the input image to obtain or estimate the distance between the physical camera 40a and objects represented by the pixels of the input image. The depth information can be assigned to each pixel of the input image by a dense or sparse depth sensor or a module configured to determine depth based on image information.
[0073] The desired viewing position and viewing direction of the virtual camera 40a can be referred to as the desired field of view of the virtual camera. In addition, inherent calibration parameters of the virtual camera can be provided to determine the virtual camera's field of view, resolution, and other optional or additional parameters. The desired field of view can be a field of view defined by the vehicle user. Therefore, the vehicle user or occupant can select the virtual camera's field of view of the vehicle's surrounding environment.
[0074] The desired pose of a virtual camera can include its viewing position and orientation relative to a reference point or frame, such as the virtual camera's viewing position and orientation relative to the vehicle. The desired pose is the virtual point where the user wants the virtual camera to be located, including the direction the virtual camera is pointing. Vehicle users can change the desired pose to generate virtual views of the vehicle and its environment from different viewing positions and orientations.
[0075] The actual pose of the physical camera is determined to have information about the viewpoint from which it captures the input image.
[0076] The depth sensor 40b can be a physical depth sensor or a module that assigns depth information to pixels or objects in the input image based on image information (which can be called a virtual depth sensor). Examples of physical depth sensors are ultrasonic sensors, radar sensors, lidar sensors, etc. These sensors are configured to determine the distance to a physical object. The distance information determined by the physical depth sensor is then assigned to the pixels of the input image. A virtual depth sensor determines or estimates depth information based on image information. If the depth information provided by the virtual depth sensor is consistent, it may be sufficient to generate an appropriate output image for the pose of the virtual camera. Absolutely accurate depth information is not necessarily required.
[0077] In another embodiment, the method includes capturing an input scene by a physical camera and a depth sensor spatially separated from the physical camera. Capturing the input scene by the capture device includes capturing an input image by a physical camera 40a and capturing depth data associated with the input image by a depth sensor 40b; determining the actual pose of the capture device by the controller includes determining the actual pose of the physical camera 40a by a controller 34 and determining the actual pose of the depth sensor 40b by a controller 34; defining the polar geometry between the actual pose of the capture device and the desired pose of the virtual camera by the controller 34 includes defining the polar geometry between the actual pose of the depth sensor 40b and the desired pose of the virtual camera by the controller 34; and generating an output image for the virtual camera by the controller includes: generating dense depth data for the desired pose of the virtual camera by the controller 34; projecting the dense depth data for the desired pose of the virtual camera onto the input image by the controller 34; and generating an output image of the virtual camera based on the dense depth data projected onto the input image by the controller 34.
[0078] In this embodiment, the input image is captured by a physical camera 40a, such as an optical camera configured to capture a color image of the environment. Depth data is captured by a physical depth sensor 40b spaced apart from the physical camera 40a; that is, the input image and depth data are captured by sensors located at different positions and having different perspectives of the scene. Dense depth data is generated by the controller for the pose of the virtual camera, that is, objects visible from the desired position of the virtual camera are assigned depth information to obtain the distance of objects captured by the physical camera from the virtual camera. The dense depth data is applied to the input image such that objects shown in the input image are assigned depth information with respect to the virtual camera.
[0079] Typically, the depth sensor 40b can be either a sparse depth sensor or a dense depth sensor. A sparse depth sensor provides depth information for some pixels and regions of the input image, but not all pixels. A sparse depth sensor cannot provide a continuous depth map. A dense depth sensor provides depth information for every pixel of the input image. When the depth sensor is a dense depth sensor, reprojection of the depth values onto the image from the virtual camera is performed, as in embodiments with depth sensors located in the same position.
[0080] Figure 1 The vehicle 10 is shown in general, including a system for generating a virtual view of a virtual camera based on an input scene. The system includes a physical camera 40a, a depth sensor 40b (either located in the same location as the physical camera or spatially separated from it), and a controller 34. The system is configured to perform the steps of the two methods described herein.
[0081] Figure 2The epipolar geometry principle is illustrated exemplarily with respect to a first camera 102 having a camera center C1 and a second camera 112 having a camera center C2. A first epipolar line 104 is defined within the first camera 102. A ray 106 defines the position of pixel P (denoted as 110) on the epipolar line 104. The position of the same pixel P 110 is also defined on the epipolar line 114 by a ray 116 extending from the camera center C2 of the second camera 112 to pixel P. Reference symbol 118 is a vector between the two camera centers C1 and C2. Given vector 118 and the known position of pixel P on the epipolar line 104, as well as the distance between the camera center C1 and pixel P, the position of pixel P on the epipolar line 114 can be determined. Using this principle, the scene captured by the first camera 102 can be used to calculate the scene as if it were being observed by the second camera 112. The virtual position of the second camera 112 can change. Therefore, when the position of the second camera 112 changes, the position of the pixel on the epipolar line 114 also changes. In the various embodiments described herein, virtual view perspective changes are enabled. This virtual view perspective change can be advantageously used to generate surround views and for trailer applications. Using polar geometry to generate the virtual view takes into account the three-dimensional characteristics of the vehicle environment, especially when generating the virtual view of the second camera 112, by taking into account the depth of pixel P (the distance between pixel P and the camera center C1 of the first camera 102).
[0082] In various embodiments, the method includes: defining the polar geometry between the desired viewing orientation and position (second camera 112) and the depth sensor pose of the first camera 102; using depth-assisted polar reprojection; finding the parallax along the epipolar line for each output pixel of the image to be generated by the virtual camera by minimizing the directional cost function; and generating the output image by acquiring pixels with the calculated parallax. Note that parallax involves the difference between pixel positions on cameras located at different locations. Parallax is related to the distance between objects and cameras at different locations. The greater the distance, the smaller the parallax of the object or the pixel representing the object.
[0083] Figure 3 The illustration schematically depicts a specific use case applied to a device with a virtual camera 122. Figure 2 The basic relationship is as follows: A first camera 102 with a first epipolar line 104 is used to capture the environment. Multiple pixels P and their positions along the epipolar line 104, as well as the distance of each pixel from the camera center C1, are captured by the first camera 102. A ray R 106 indicates the position of pixel P 110 along the epipolar line 104. Using the known distance between the camera center C1 and pixel P, and the vector T118 between the center C1 of the physical camera and the center C2 of the virtual camera 122, the position of pixel P 110 on the epipolar line 124 of the virtual camera 122 can be determined.
[0084] exist Figure 3 middle, It is a unit vector from the center C2 of the virtual camera 112, pointing in the direction of pixel 110 on the epipolar line 124. For this unit vector, the corresponding pixel m on the epipolar line 104 of the physical camera will be identified. It is the vector between the camera centers C1 and C2. It corresponds to the 3D point position of pixel m. Is for a given The physical pixel on epipolar line 104 corresponding to the minimum of the directional cost function, where the directional cost function is defined as:
[0085]
[0086] And its minimum value is determined as:
[0087]
[0088] Figure 4 A first embodiment of a method for generating a virtual view of a scene captured by a physical camera and a depth sensor located at the same position is schematically illustrated. At 141, viewpoint pose data (position and orientation of the virtual camera 122) is captured. This information can be selected by the user or occupant of vehicle 10 and corresponds to the desired position of the virtual camera used to view the surrounding scene. At 142, input camera 102 pose data is acquired. Note that at 141 and 142, the inherent calibration parameters of the viewpoint pose data and / or the input camera pose data can preferably be acquired. For example, the inherent parameters of a viewpoint camera typically include focal length, principal point, distortion model, etc. At 143, the polar geometry between the physical camera (i.e., the input camera) 102 and the virtual camera 122 is defined using the input camera 102 pose data and the desired pose data of the virtual camera 122. To define the polar geometry, the viewpoint pose data and / or the inherent calibration parameters of the input camera can be used. At 144, an input depth map including depth information from an input image (received from the physical camera) is received, and at 145, the input image is received. Based on the polar geometry generated in step 143, in step 146, the input depth map from step 144 and the input image from step 145 are resampled in polar coordinates. In step 147, for each output pixel and each output epipolar line (of the virtual camera's output image), the directional cost is minimized to find the target pixel on the corresponding input epipolar line. Then, in step 148, a disparity map is created, which forms the basis for generating the output image to be displayed by the virtual camera in step 149. Figure 4 The steps shown are preferably used to generate the output image of a virtual camera based on the input image of a physical camera and a depth sensor located at the same location or a depth-from-mono-module that provides depth information based on the input image.
[0089] Figure 3 The related descriptions indicate how to perform polar reprojection on a depth sensor located in the same position as the physical (input) camera. Polar geometry is defined between the input physical camera 102 and the virtual camera 122 with the desired position and viewing direction, and the directional cost is minimized for each output pixel to find the output pixel parallax.
[0090] Figure 5 A second embodiment of a method for generating a virtual view of a scene captured by a physical camera and a spatially separated depth sensor is schematically illustrated. At 153, polar geometry is defined based on viewpoint pose data of the virtual camera retrieved at 151 (as desired or selected by the user) and depth sensor pose data retrieved at 156 from a physical depth sensor spatially separated from the physical camera. At 151 and 156, inherent calibration parameters of the viewpoint pose data and / or depth sensor pose data may preferably be obtained. The inherent calibration parameters of the viewpoint pose data and / or depth sensor pose data can then be used to define the polar geometry. The polar geometry defined at 153 and the input depth data retrieved from the depth sensor at 154 are used at 157 to minimize the directional cost per output pixel for each output epipolar line to estimate the depth of the output pixels. Based on the minimized directional cost, dense depth data of the virtual camera's viewpoint is generated at 158. The dense depth data generated at 158, the input camera pose data (position and orientation of the physical camera) retrieved at 152, and the input image retrieved at 155 are used at 159 to project the dense depth data onto the input camera and find the target pixels used at 160 to generate the output image of the virtual camera. At 152, the inherent calibration parameters of the input camera pose data can preferably be obtained. For example, the inherent parameters of the viewpoint camera typically include focal length, principal point, distortion model, etc.
[0091] Figure 6 and Figure 7 The following are references to the input physical camera 102 (with center Cp), virtual camera 122 (with center Cv), and depth sensor 130 (with center CD) for description. Figure 5 During the process, the depth sensor 130 is spatially separated from the physical camera 102 and located in different positions and spaced apart from the physical camera 102.
[0092] Initially, polar geometry is defined between depth sensor 130 and virtual camera 122. Depth sensor 130 determines depth data 136 (for a set of point cloud points) at different points within its field of view. Polar geometry is defined as a reference... Figure 3 As stated above. However, in Figure 6In this context, the epipolar line 134 of the depth sensor, the epipolar line 124 of the virtual camera, and the vector 128 between the center CD of the depth sensor 130 and the center Cv of the virtual camera 122 are used to define the epipolar geometry. If the depth sensor 130 is dense, then as in the reference... Figure 3 and Figure 4 The process involves performing a re-projection of depth values onto the virtual camera. If the depth sensor is sparse, a lookup table is built to map the virtual camera epipolar lines to a set of point cloud points near the corresponding epipolar lines of the depth sensor. For each output pixel, the directional cost is minimized over a set of related point cloud points. The dense virtual camera depth map is projected onto the input image of the physical camera to find the corresponding input pixels, and depth information is assigned to the input pixels, as follows: Figure 7 As shown. To speed up the process, sparse point clouds can be converted into dense depth maps using triangulation or Voronoi tessellation. These dense depth maps can then be used for highly re-projected depth values onto a virtual camera, as described above.
[0093] It should be understood that the controller 34 of vehicle 10 implements reference Figures 3 to 7 The described function involves the sensors of vehicle 10 providing the necessary information: a physical camera provides the input image, and depth sensors, whether located in the same location or spatially separated, provide depth information. This depth information can be obtained from the input image using a technique known as single-depth.
[0094] Due to the realistic representation of the scene's 3D structure, object occlusion may occur after a change in viewpoint. Occlusion detection is typically not part of reprojection. However, a typical characteristic of occluded regions is a very high directional cost function value, making them detectable. In addition to the pure directional term, the cost function can be regularized by including a so-called data smoothing term, which comprises, for example, the sum of the color differences between adjacent pixels in the output image. Preferably, nearest-neighbor interpolation (equivalent to an integer value of the optimal disparity) is used to avoid erroneous slopes over large depth differences. However, for sparse data, more complex interpolation modules can be used depending on the depth differences.
[0095] While at least one exemplary embodiment has been presented in the foregoing detailed description, it should be understood that numerous variations exist. It should also be understood that the one or more exemplary embodiments are merely examples and are not intended to limit the scope, applicability, or configuration of this disclosure in any way. Rather, the foregoing detailed description will provide those skilled in the art with a convenient roadmap for implementing one or more exemplary embodiments. It should be understood that various changes can be made to the function and arrangement of the elements without departing from the scope of this disclosure as set forth in the appended claims and their legal equivalents.
Claims
1. A method for generating a virtual view of a virtual camera based on an input scene, the method comprising: capturing, by a capture device, input scene data, wherein the capture device comprises a physical camera and a depth sensor spatially separated from the physical camera, wherein the input scene data comprises an input image and depth data related to the input image; determining, by a controller, an actual pose of the physical camera and the depth sensor; determining, by the controller, a desired pose of the virtual camera for displaying the virtual view; defining, by the controller, epipolar geometry between the actual pose of the depth sensor and the desired pose of the virtual camera; generating, by the controller, an output image for the virtual camera based on the epipolar relationship between the actual pose of the depth sensor, the input scene data, and the desired pose of the virtual camera.
2. The method of claim 1, wherein generating, by the controller, the output image for the virtual camera comprises: generating, by the controller, dense depth data for the desired pose of the virtual camera; projecting, by the controller, the dense depth data for the desired pose of the virtual camera onto the input image; and generating, by the controller, the output image for the virtual camera based on the dense depth data projected onto the input image.
3. The method of claim 2, wherein generating, by the controller, the dense depth data for the desired pose of the virtual camera comprises: minimizing, by the controller, a directional cost function to estimate depth of a pixel.
4. A system for generating a virtual view of a virtual camera based on an input scene, the system comprising: a capture device having a physical camera and a depth sensor spatially separated from the physical camera configured to capture an input scene, wherein the input scene data comprises an input image and depth data related to the input image; a controller configured to determine an actual pose of the physical camera and an actual pose of the depth sensor; wherein the controller is configured to determine a desired pose of the virtual camera for displaying the virtual view; wherein the controller is configured to define epipolar geometry between the actual pose of the physical camera and the actual pose of the depth sensor and the desired pose of the virtual camera; wherein the controller is configured to generate an output image for the virtual camera based on the epipolar relationship between the actual pose of the physical camera and the actual pose of the depth sensor, the input scene data, and the desired pose of the virtual camera.
5. A vehicle comprising the system of claim 4.
Citation Information
Patent Citations
Image processing method and apparatus for calibrating depth of depth sensor
US20150279016A1
Image reconstruction for virtual 3D
US20200357128A1