Generating virtual images based on captured image data

By determining the epipolar geometry and the minimum matching angle index, the image data fusion artifact problem in the virtual view is solved, generating a clear virtual view suitable for autonomous vehicles.

CN115705716BActive Publication Date: 2026-08-04GM GLOBAL TECHNOLOGY OPERATIONS LLC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GM GLOBAL TECHNOLOGY OPERATIONS LLC
Filing Date
2022-05-25
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively fuse image data from different locations when generating virtual views, resulting in unwanted artifacts in the virtual images.

Method used

By determining the epipolar geometry between the actual pose of the capture device and the desired pose of the virtual camera, and using the depth information provided by the depth sensor, the pixels of the input image are resampled to generate a virtual image. Then, by comparing the minimum matching angle index and the adjacent index, pixels in the reverse occlusion area are removed to generate a clear virtual view.

Benefits of technology

It reduces artifacts in virtual views, provides clearer perspective changes and virtual scene display, and is suitable for virtual view generation of autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115705716B_ABST
    Figure CN115705716B_ABST
Patent Text Reader

Abstract

Systems and methods for generating virtual views of a virtual camera based on input images are described. A system for generating virtual views of a virtual camera based on input images can include a capture device comprising a physical camera and a depth sensor. The system can also include a controller configured to determine an actual pose of the capture device, determine a desired pose of a virtual camera for displaying a virtual view, define epipolar geometry between the actual pose of the capture device and the desired pose of the virtual camera, and generate a virtual image for the virtual camera depicting objects within the input image according to the desired pose of the virtual camera based on an epipolar relationship between the actual pose of the capture device, the input image, and the desired pose of the virtual camera.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This technical field generally relates to generating virtual views based on captured image data. In particular, this description relates to changes in the viewpoint of the virtual view. Background Technology

[0002] Modern vehicles are typically equipped with one or more optical cameras configured to provide image data to the vehicle's occupants. For example, the image data displays a predetermined viewpoint of the environment surrounding the vehicle.

[0003] In some situations, vehicle operators may want to change the perspective of image data provided by optical cameras. For this purpose, virtual cameras are used, modifying image data captured by one or more physical cameras to display the captured scene from a different, desired perspective. The modified image data can be called a virtual image or output image. The desired perspective on the virtual scene can be changed according to the occupants' wishes. Virtual scenes can be generated based on multiple images captured from different perspectives. However, fusing image data from image sources located in different locations may produce undesirable artifacts in the virtual image. Summary of the Invention

[0004] A method is provided for generating a virtual view of a virtual camera based on an input image. The method may include: determining the actual pose of a capture device by a controller; determining the desired pose of a virtual camera for displaying the virtual view by the controller; defining the epipolar geometry between the actual pose of the capture device and the desired pose of the virtual camera by the controller; and generating a virtual image depicting objects within the input image based on the epipolar geometry between the actual pose of the capture device, the input image, and the desired pose of the virtual camera, according to the desired pose of the virtual camera, wherein at least one pixel corresponding to the input image is selected based on a minimum matching angle index corresponding to the actual pose of the capture device and the desired pose of the virtual camera.

[0005] Among other features, the method includes: capturing an input image using a physical camera with a depth sensor at the same location; wherein capturing the input image by a capture device includes capturing the input image by a physical camera; assigning depth information to pixels of the input image by a depth sensor; wherein determining the actual pose of the capture device by a controller includes determining the actual pose of the physical camera by the controller; wherein defining the epipolar geometry between the actual pose of the capture device and the desired pose of the virtual camera by the controller includes defining the epipolar geometry between the actual pose of the physical camera and the desired pose of the virtual camera by the controller; and wherein generating an output image for the virtual camera by the controller includes: resampling the depth information of the pixels of the input image in epipolar coordinates by the controller; identifying target pixels on the input epipolar lines of the physical camera by the controller; generating a disparity map for one or more output epipolar lines of the virtual camera by the controller; and generating an output image based on one or more output epipolar lines by the controller.

[0006] Among other features, the method includes: monotonicizing the at least one minimum matching angle index by a controller based on a comparison of at least one minimum matching angle index with adjacent minimum matching angle indices.

[0007] Among other features, monotonicizing at least one minimum matching angle index further includes: comparing adjacent minimum matching angle indices with at least one minimum matching angle index; determining whether the difference between adjacent minimum matching angle indices and at least one minimum matching angle index is greater than a difference threshold; and removing the at least one minimum matching angle index when the difference is greater than the difference threshold.

[0008] Among other features, the minimum matching angle index corresponds to the pixel representing the object within the reverse occlusion region.

[0009] Among the other features, the minimum matching angle index is defined as

[0010] θv(θpi)=argmin θpi |θv(θpi)-θv|,

[0011] in:

[0012] θv(θpi) is the minimum matching angle exponent;

[0013] θv is the angle measurement between the axis extending from the center of the virtual camera and the vector T; and

[0014] θpi is the angular measurement of the i-th value, representing the angle between vectors perpendicular to the axis extending from the center of the physical camera.

[0015] Among other features, the method includes assigning depth information by a depth sensor to each pixel captured by a physical camera.

[0016] Among other features, the method includes: determining depth information for each pixel captured by a physical camera based on an input image using a depth sensor.

[0017] A system for generating a virtual view of a virtual camera based on an input image is disclosed. The system may include: a capture device including a physical camera and a depth sensor, the capture device configured to capture the input image; and a controller configured to: determine the actual pose of the capture device; determine the desired pose of the virtual camera for displaying the virtual view; define the epipolar geometry between the actual pose of the capture device and the desired pose of the virtual camera; and generate a virtual image depicting objects within the input image based on the epipolar relationship between the actual pose of the capture device, the input image, and the desired pose of the virtual camera, wherein at least one pixel corresponding to the input image is selected based on a minimum matching angle index corresponding to the actual pose of the capture device and the desired pose of the virtual camera.

[0018] Among other features, the controller is configured to monotonic the at least one minimum matching angle index based on a comparison of at least one minimum matching angle index with an adjacent minimum matching angle index.

[0019] Among other features, the controller is configured to: compare adjacent minimum matching angle indices with at least one minimum matching angle index; determine whether the difference between adjacent minimum matching angle indices and at least one minimum matching angle index is greater than a difference threshold; and remove the at least one minimum matching angle index when the difference is greater than the difference threshold.

[0020] Among the other features, at least one minimum matching angle index corresponds to a pixel representing an object within the reverse occlusion region.

[0021] Among the other features, the minimum matching angle index is defined as

[0022] θv(θpi)=argmin θpi |θv(θpi)-θv|,

[0023] in:

[0024] θv(θpi) is the minimum matching angle exponent;

[0025] θv is the angle measurement between the axis extending from the center of the virtual camera and the vector T; and

[0026] θpi is the angular measurement of the i-th value, representing the angle between vectors perpendicular to the axis extending from the center of the physical camera.

[0027] A vehicle is disclosed comprising a system for generating a virtual view of a virtual camera based on an input image. The system may include: a capture device including a physical camera and a depth sensor, the capture device configured to capture the input image; and a controller configured to: determine the actual pose of the capture device; determine the desired pose of the virtual camera for displaying the virtual view; define the epipolar geometry between the actual pose of the capture device and the desired pose of the virtual camera; and generate a virtual image depicting objects within the input image based on the epipolar relationship between the actual pose of the capture device, the input image, and the desired pose of the virtual camera, wherein at least one pixel corresponding to the input image is selected based on a minimum matching angle index corresponding to the actual pose of the capture device and the desired pose of the virtual camera.

[0028] Among other features, the controller is configured to monotonic the at least one minimum matching angle index based on a comparison of at least one minimum matching angle index with an adjacent minimum matching angle index.

[0029] Among other features, the controller is configured to: compare adjacent minimum matching angle indices with at least one minimum matching angle index; determine whether the difference between adjacent minimum matching angle indices and at least one minimum matching angle index is greater than a difference threshold; and remove the at least one minimum matching angle index when the difference is greater than the difference threshold.

[0030] Among the other features, at least one minimum matching angle index corresponds to a pixel representing an object within the reverse occlusion region.

[0031] Among the other features, the minimum matching angle index is defined as

[0032] θv(θpi)=argmin θpi |θv(θpi)-θv|,

[0033] in:

[0034] θv(θpi) is the minimum matching angle exponent;

[0035] θv is the angle measurement between the axis extending from the center of the virtual camera and the vector T; and

[0036] θpi is the angular measurement of the i-th value, representing the angle between vectors perpendicular to the axis extending from the center of the physical camera. Attached Figure Description

[0037] Exemplary embodiments will be described below with reference to the accompanying drawings, in which the same reference numerals denote the same elements, and wherein:

[0038] Figure 1 It is a schematic diagram of a vehicle including a controller configured to generate virtual views;

[0039] Figure 2 It is a schematic diagram of the epipolar geometry principle of two cameras;

[0040] Figure 3 This is a block diagram illustrating an exemplary method for generating a virtual view;

[0041] Figure 4 This is a schematic diagram of a capture device that captures an input image from both a first-person perspective and the perspective of a virtual image relative to the capture device; and

[0042] Figure 5 This is another schematic diagram of a capture device that captures an input image from a first-person perspective and from the perspective of a virtual image relative to the capture device. Detailed Implementation

[0043] As described herein, a system can generate a virtual view of a scene captured by a physical camera, such that the virtual view displays a different perspective than that captured by the physical camera. Since some pixels may be included within a reverse occlusion region, the system can determine one or more pixels to include within the virtual view. For example, the system can determine the minimum matching angle index corresponding to a pixel. The system also compares the minimum matching angle index with adjacent minimum matching angle indices to determine oscillatory behavior. When the difference between the minimum matching angle index and adjacent minimum matching angle indices is greater than a difference threshold, the system removes the minimum matching angle index with the larger value.

[0044] See Figure 1 A vehicle 10 is shown according to various embodiments. The vehicle 10 generally includes a chassis 12, a body 14, front wheels 16, and rear wheels 18. The body 14 is mounted on the chassis 12 and substantially surrounds the components of the vehicle 10. The body 14 and the chassis 12 may together form a frame. The front wheels 16 and the rear wheels 18 are each rotatably coupled to the chassis 12 near a respective corner of the body 14.

[0045] In various embodiments, vehicle 10 is an autonomous vehicle. For example, an autonomous vehicle is a vehicle that is automatically controlled to transport passengers from one location to another. Vehicle 10 is described as a passenger car in the illustrated embodiment, but it should be understood that any other means of transportation, including motorcycles, trucks, SUVs, RVs, ships, aircraft, etc., may also be used. In one exemplary embodiment, the autonomous vehicle is a Level 2 or higher automation system. A Level 2 automation system means “partial automation.” However, in other embodiments, the autonomous vehicle may be a so-called Level 3, Level 4, or Level 5 automation system. A Level 3 automation system means “conditional automation.” A Level 4 system means “high automation,” referring to the driving mode-specific performance of the autonomous driving system for all aspects of the dynamic driving task, even when the human driver does not respond appropriately to intervention requests. A Level 5 system means “full automation,” referring to the all-time performance of the autonomous driving system for all aspects of the dynamic driving task under all road and environmental conditions that a human driver can handle.

[0046] As shown in the figure, the vehicle 10 typically includes a propulsion system 20, a transmission system 22, a steering system 24, a braking system 26, a sensor system 28, an actuator system 30, at least one data storage device 32, at least one controller 34, and a communication system 36. In various embodiments, the propulsion system 20 may include an internal combustion engine, an electric motor such as a traction motor, and / or a fuel cell propulsion system. The transmission system 22 is configured to transmit power from the propulsion system 20 to the front wheels 16 and the rear wheels 18 according to a selectable speed ratio. According to various embodiments, the transmission system 22 may include a stepped automatic transmission, a continuously variable transmission (CVT), or other suitable transmission. The braking system 26 is configured to provide braking torque to the front wheels 16 and the rear wheels 18. In various embodiments, the braking system 26 may include friction brakes, brake-by-wire brakes, regenerative braking systems such as electric motors, and / or other suitable braking systems. The steering system 24 influences the position of the front wheels 16 and the rear wheels 18. Although depicted as including a steering wheel for illustrative purposes, in some embodiments contemplated within the scope of this disclosure, the steering system 24 may not include a steering wheel.

[0047] Sensor system 28 includes one or more sensing devices 40a-140n that sense observable conditions of the external and / or internal environment of vehicle 10. Sensing devices 40a-40n may include, but are not limited to, radar, lidar, GPS, optical cameras, thermal imaging cameras, ultrasonic sensors, and / or other sensors. Actuator system 30 includes one or more actuator devices 42a-42n that control one or more vehicle features, such as, but not limited to, propulsion system 20, transmission system 22, steering system 24, and braking system 26. In various embodiments, vehicle features may also include internal and / or external vehicle features, such as, but not limited to, doors, trunk, and passenger compartment features (such as air conditioning systems, in-vehicle entertainment systems, and lighting systems).

[0048] Communication system 36 is configured to wirelessly exchange information with other entities 48 (such as, but not limited to, other vehicles (“V2V” communication), infrastructure (“V2I” communication), remote systems and / or personal devices) (see reference). Figure 2 (A more detailed description has been provided). In one exemplary embodiment, communication system 36 is a wireless communication system configured to communicate via a wireless local area network (WLAN) employing the IEEE 802.11 standard or via cellular data communication. However, additional or alternative communication methods, such as Dedicated Short Range Communication (DSRC) channels, are also considered to be within the scope of this disclosure. DSRC channels refer to one-way or two-way short-to-medium range wireless communication channels specifically designed for automobiles, along with a corresponding set of protocols and standards.

[0049] Data storage device 32 stores data for the functions of the automated control vehicle 10. In various embodiments, data storage device 32 stores a defined map of the navigable environment. In various embodiments, the defined map may be predefined by a remote system and obtained from that remote system (see reference). Figure 2 (A more detailed description has been provided). For example, the defined map can be assembled by a remote system and transmitted to vehicle 10 (wirelessly and / or via wire) and stored in data storage device 32. It is understood that data storage device 32 can be part of controller 34, independent of controller 34, or part of controller 34 and an independent system.

[0050] The controller 34 includes at least one processor 44 and a computer-readable storage device or medium 46. The processor 44 can be any custom or commercially available processor, central processing unit (CPU), graphics processing unit (GPU), auxiliary processor among several processors associated with the controller 34, semiconductor-based microprocessor (in the form of a microchip or chipset), macroprocessor, any combination thereof, or any device generally used for executing instructions. For example, the computer-readable storage device or medium 46 can include volatile and non-volatile storage in read-only memory (ROM), random access memory (RAM), and persistent powered-on memory (KAM). KAM is persistent or non-volatile memory that can be used to store various operational variables when the processor 44 is powered off. The computer-readable storage device or medium 46 may be implemented using any of a variety of known storage devices, such as programmable read-only memory (PROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or any other electrical, magnetic, optical, or combined storage device capable of storing data (some of which represent executable instructions used by the controller 34 in controlling and performing the functions of the vehicle 10).

[0051] The instructions may include one or more separate programs, each comprising a sequence of executable instructions for implementing logical functions. When executed by processor 34, the instructions receive and process signals from sensor system 28, perform logic, calculations, methods, and / or algorithms to automatically control components of vehicle 10, and generate control signals to actuator system 30 to automatically control components of vehicle 10 based on logic, calculations, methods, and / or algorithms. Although in Figure 1 Only one controller 34 is shown, but embodiments of the vehicle 10 may include any number of controllers 34 that communicate via any suitable communication medium or combination of communication media and collaboratively process sensor signals, execute logic, calculations, methods and / or algorithms, and generate control signals to automatically control the features of the vehicle 10.

[0052] Typically, according to one embodiment, the vehicle 10 includes a controller 34 that generates a virtual view of a virtual camera based on input images captured by a capture device. The capture device includes, for example, a physical camera and a depth sensor. At least one of the sensing devices 40a to 40n is an optical camera, and at least one of the sensing devices 40a to 40n is a physical depth sensor (e.g., LiDAR, radar, ultrasonic sensor, etc.). In one or more embodiments, the vehicle 10 can generate a virtual view of a scene captured by a physical camera 40a with depth sensors 40b located at the same position or spatially separated.

[0053] The input image is captured by a physical camera 40a (e.g., an optical camera that captures a color image of the environment). The physical camera 40a is positioned at the vehicle 10 such that it can cover a specific field of view of the environment surrounding the vehicle 10. Depth information is assigned to the pixels of the input image to estimate the distance between the physical camera 40a and the objects represented by the pixels of the input image. The depth information may be assigned to each pixel of the input image by a dense or sparse depth sensor or by a module configured to determine depth based on image information.

[0054] The desired observation position and direction of a virtual camera can be referred to as the desired field of view (POV) of the virtual camera. Furthermore, the virtual camera's inherent calibration parameters can be provided to determine its field of view, resolution, and other optional or additional parameters. The desired POV can be defined by the vehicle user. Therefore, the vehicle user or occupants can select the virtual camera's POV view of the vehicle's surroundings.

[0055] The desired pose of the virtual camera can include the viewing position and viewing direction relative to a reference point or reference frame, such as the viewing position and viewing direction of the virtual camera relative to the vehicle. The desired pose is the virtual point where the user wants the virtual camera to be located, including the direction the virtual camera is pointing. The vehicle user can change the desired pose to generate virtual views of the vehicle 10 and its environment from different viewing positions and different viewing directions. The actual pose of the physical camera is determined to have information about the viewpoint used to capture the input image.

[0056] The depth sensor 40b can be a physical depth sensor or a module that assigns depth information to pixels or objects in the input image based on image information (e.g., a virtual depth sensor). Examples of physical depth sensors include ultrasonic sensors, radar sensors, lidar sensors, etc. These sensors are configured to determine the distance to a physical object shown within the input image. The distance information determined by the physical depth sensor is then assigned to the pixels of the input image. The virtual depth sensor determines or estimates depth information based on image information. If the depth information provided by the virtual depth sensor is consistent, it may be sufficient to generate an appropriate output image for the pose of the virtual camera.

[0057] Typically, the depth sensor 40b can be either a sparse depth sensor or a dense depth sensor. A sparse depth sensor provides depth information for some pixels and regions of the input image, but not for all pixels. A sparse depth sensor does not provide a continuous depth map. A dense depth sensor provides depth information for every pixel of the input image. When the depth sensor is a dense depth sensor, depth values ​​are reprojected onto the image from the virtual camera using depth sensors at the same location.

[0058] Figure 1 A vehicle 10 is shown, which includes a system for generating a virtual view of a virtual camera based on input images. The system includes a physical camera 40a, a depth sensor 40b, and a controller 34. In various embodiments, the depth sensor 40b may be located in the same position as the physical camera 40a. In another exemplary embodiment, the depth sensor 40b may be spatially separated from the physical camera 40a. The controller 34 may receive and analyze data from the physical camera 40a and the depth sensor 40b substantially continuously and / or periodically. For convenience, the first camera 102 may refer to the physical camera 40a, and the second camera 112 may refer to the virtual camera.

[0059] Figure 2 The epipolar geometry principle is illustrated for a first camera 102 with camera center C1 and a second camera 112 with camera center C2. A first epipolar line 104 is defined in the first camera 102. Ray 106 defines the position of pixel P (denoted as 110) on the epipolar line 104. The position of the same pixel P 110 is also defined on the epipolar line 114 by ray 116 extending from the camera center C2 of the second camera 112 to pixel P. Reference numeral 118 is a vector between the two camera centers C1 and C2. Given vector 118, the known position of pixel P on the epipolar line 104, and the distance between camera center C1 and pixel P, the position of pixel P on the epipolar line 114 can be determined. Using this principle, the scene captured by the first camera 102 can be used to calculate the scene observed by the second camera 112. The virtual position of the second camera 112 can change. Therefore, when the position of the second camera 112 changes, the pixel position on the epipolar line 114 also changes. In the various embodiments described herein, the virtual view perspective is allowed to change. This virtual view perspective change can be advantageously used for surround view generation and trailer-mounting applications. When epipolar geometry is used to generate the virtual view, the three-dimensional nature of the vehicle environment is taken into account, particularly by considering the depth of pixel P (the distance between pixel P and the camera center C1 of the first camera 102) when generating the virtual view of the second camera 112.

[0060] The controller 34 can generate a virtual image of the scene captured by a first camera 102 with a depth sensor (e.g., depth sensor 40b) at the same location. As used herein, the virtual image can be defined as a composite image generated by the controller 34 such that the virtual image depicts one or more objects within the scene captured by the first camera 102 from a desired viewing position and viewing direction. The controller 34 can receive input camera pose data and desired pose data from the first camera, i.e., from the perspective of the virtual camera. The controller 34 can also retrieve intrinsic parameters of the first camera 102, which may include focal length, principal point, distortion model, etc. The input camera pose data and desired pose data are used to define the epipolar geometry between the first camera and the second camera 122.

[0061] The controller 34 can define epipolar geometry using viewpoint pose data and / or the intrinsic calibration parameters of the input camera. The controller 34 can receive an input depth map including the input image and input depth information. Based on the epipolar geometry, the input depth map and the input image are resampled in epipolar coordinates. The controller 34 can employ a suitable directional cost function to minimize each output pixel, and for each output epipolar line, the controller 34 identifies target pixels along the corresponding input epipolar line. The controller 34 can generate a disparity map, which forms the basis for generating the virtual image.

[0062] However, one or more pixels within the input image may be in the reverse occlusion area. Figure 3 This is a block diagram illustrating a process 300 for mitigating back occlusion caused by viewpoint changes in a virtual image generated based on epipolar projection. The steps of process 300 can be executed by a controller 34. In step 302, an input depth map including depth information from an input image from a physical camera 102 is received, and in step 304, an input image is received. In step 306, intrinsic calibration and alignment parameters of viewpoint pose data and / or depth sensor pose are acquired. The intrinsic calibration parameters of the viewpoint pose data and / or depth sensor pose data can then be used to define epipolar geometry.

[0063] In step 308, epipolar geometry is generated based on the viewpoint pose data and depth sensor pose data of the virtual camera. For example, the inherent calibration and alignment parameters of the viewpoint pose data and / or depth sensor pose data can be used to generate the epipolar geometry. In step 308, the input image data and / or input depth map data can be resampled in epipolar coordinates. The input image data and / or input depth map data can be resampled in epipolar coordinates based on the generated epipolar geometry.

[0064] In step 310, the controller 34 calculates the matching angle index for each pixel. The matching angle index can be calculated based on Equation 1:

[0065] θv(θpi)=∠(T,T+Rpi(θpi)) Equation 1.

[0066] Controller 34 determines the minimum value of the matching angle index according to Equation 2:

[0067] θv(θpi)=arhmin θpi Equation 2: |θv(θpi)-θv|

[0068] See Figure 4 and Figure 5 θv is the angle measurement between the axis extending from the center 408 of the virtual camera (i.e., the epipolar line) and vector T; θpi is the angle measurement representing the i-th value of the angle between vector 402 and the axis extending perpendicularly from the center 404 of the physical camera; vector T is a known vector 406 between the center 404 of the physical camera and the center 408 of the virtual camera; and Rpi is the vector 412 representing the i-th value between the center 404 of the physical camera and the corresponding pixel P1, P2, or P3. Vector 412 represents the distance between the center of the physical camera and the corresponding portion of the object captured by the physical camera, represented by pixel P1, P2, or P3.

[0069] Controller 34 generates virtual images, or composite images, from the perspective of a virtual camera based on image data captured by a physical camera and a depth sensor. See also Figure 5 The first camera 102 captures an image of the displayed object 502 (e.g., a tree). The controller 34 generates a virtual image representing the captured object 502 from the desired viewpoint of the virtual camera. As shown, due to reverse occlusion, points P2 and P3 are within the field of view of the physical camera but not within the field of view of the virtual camera. Based on the reference frame of the virtual camera, the generated virtual image should include the pixels representing point P1, since point P1 is within the field of view of both the physical and virtual cameras.

[0070] Controller 34 calculates and classifies the minimum matching angle index for each pixel within the captured image. In step 312, controller 34 monotonicizes (i.e., removes) the minimum matching angle index value corresponding to a pixel within the reverse occlusion region. In one exemplary embodiment, controller 34 detects the oscillating behavior of the minimum matching angle index, i.e., the difference between adjacent minimum matching angle index values. For example, each minimum matching angle index can be compared with adjacent minimum matching angle indices to determine the difference between each value. If the difference between adjacent minimum matching angle indices is greater than a difference threshold, the minimum matching angle index with the larger value is removed (i.e., monotonicized).

[0071] In step 314, a virtual image is generated based on the input image from the physical camera. For example, the input depth map and the input image are resampled in polar coordinates based on epipolar geometry. A disparity map can then be generated based on appropriate epipolar calculations, and the virtual image is generated by the controller 34 based on this disparity map. For pixels corresponding to various portions of an object within the inverse occlusion region, pixels corresponding to the calculated minimum matching angle index are selected to be included in the virtual image. For example, pixels captured by the physical camera corresponding to the calculated minimum matching angle index (e.g., ...) Figure 5 Point P1 shown is selected to be included in the virtual image.

[0072] While at least one exemplary embodiment has been presented in the above detailed description, it should be understood that numerous variations exist. It should also be understood that the one or more exemplary embodiments are merely examples and are not intended to limit the scope, applicability, or configuration of this disclosure in any way. Rather, the above detailed description will provide those skilled in the art with a convenient roadmap for implementing one or more exemplary embodiments. It should be understood that various changes can be made to the function and arrangement of the elements without departing from the scope of this disclosure as set forth in the appended claims and their legal equivalents.

[0073] This detailed description is exemplary in nature and is not intended to limit the scope of this application and its uses. Furthermore, it is not intended to be limited by any explicit or implicit theory set forth in the foregoing technical field, background, summary of the invention, or the following detailed description. As used herein, the term "module" refers to one or any combination of any hardware, software, firmware, electronic control components, processing logic, and / or processor device, including but not limited to: application-specific integrated circuits (ASICs), electronic circuits, processors (shared processors, dedicated processors, or processor groups) and memories executing one or more software or firmware programs, combinational logic circuits, and / or other suitable components providing the said functionality.

[0074] This document describes embodiments of the present disclosure based on functional and / or logical block components and various processing steps. It should be understood that these block components can be implemented by any number of hardware, software, and / or firmware components configured to perform specified functions. For example, one embodiment of the present disclosure may employ various integrated circuit components (e.g., memory elements, digital signal processing elements, logic elements, or lookup tables, etc.) that can perform various functions under the control of one or more microprocessors or other control devices. Furthermore, those skilled in the art will understand that embodiments of the present disclosure can be practiced in conjunction with any number of systems, and the systems described herein are merely exemplary embodiments of the present disclosure.

Claims

1. A method for generating a virtual view of a virtual camera based on an input image, the method comprising: The controller determines the actual pose of the physical camera with a depth sensor at the same location; The input image is captured by the physical camera; Depth information is assigned to the pixels of the input image by the depth sensor; The controller determines the desired pose of the virtual camera for displaying the virtual view; The controller defines the epipolar geometry between the actual pose of the physical camera and the desired pose of the virtual camera; The controller resamples the depth information of the pixels of the input image in polar coordinates, identifies target pixels on the input epipolar line of the physical camera, generates a disparity map for one or more output epipolar lines of the virtual camera, and generates the output image based on the one or more output epipolar lines. Based on the epipolar relationship between the actual pose of the physical camera, the input image, and the desired pose of the virtual camera, the controller generates a virtual image for the virtual camera depicting objects within the input image, wherein at least one pixel corresponding to the input image is selected based on the minimum matching angle index corresponding to the actual pose of the physical camera and the desired pose of the virtual camera. The controller monotonicates at least one minimum matching angle index based on a comparison between at least one minimum matching angle index and adjacent minimum matching angle indices. Compare the adjacent minimum matching angle index with the at least one minimum matching angle index; Determine whether the difference between the adjacent minimum matching angle index and the at least one minimum matching angle index is greater than a difference threshold. as well as When the difference is greater than the difference threshold, the at least one minimum matching angle index is removed; The minimum matching angle index thereon corresponds to a pixel representing an object within the reverse occlusion region.

2. The method according to claim 1, The minimum matching angle index is defined as follows: in: is the minimum matching angle index; is a measure of the angle between the axis and vector T extending from the center of the virtual camera; and is an angle measurement value of the i-th value, representing the angle between the vectors normal to the axes extending from the center of the physical camera.

3. The method of claim 1, wherein assigning depth information from the depth sensor to the pixels of the input image comprises: Depth information is assigned to each pixel captured by the physical camera by the depth sensor.

4. The method of claim 1, wherein the allocation of depth information from the depth sensor to the pixels of the input image comprises: The depth sensor determines the depth information of each pixel captured by the physical camera based on the input image.