Sensor fusion-based perceptually enhanced surround view
The method and system address the challenge of incomplete surround views by integrating temporal mapping and history buffers to create a continuous surround view using sensor fusion, ensuring a reliable and comprehensive view around the vehicle.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- TEXAS INSTRUMENTS JAPAN LTD
- Filing Date
- 2020-09-08
- Publication Date
- 2026-06-01
AI Technical Summary
Existing vehicle surround view systems using multiple cameras face challenges in providing a comprehensive view of areas outside the field of view of any camera, and the failure of one or more cameras disrupts system operation, leading to gaps and inconsistencies in the surround view.
A method and system that utilize sensor fusion techniques, including temporal mapping and history buffers, to generate a perceptually enhanced surround view by integrating images from multiple cameras, even when certain cameras are disabled or out of view, using a virtual camera orientation based on physical camera orientations and vehicle movement history.
Enables a seamless and consistent surround view by filling gaps with historical camera images, maintaining a continuous view around the vehicle, even when cameras are disabled, thus enhancing the perception and reliability of the surround view system.
Smart Images

Figure 0007867964000001 
Figure 0007867964000002 
Figure 0007867964000003
Abstract
Description
[Technical Field]
[0001] Vehicles such as cars, aircraft, and robots are increasingly equipped with multiple external cameras to provide the driver with an external view of the area around the vehicle. These external views are typically used to assist in maneuvering the vehicle, such as when backing up or parking. Multiple camera views can be stitched together to form an external surround view of the vehicle. However, an external view of an area not within the field of view of any camera in such a system is unavailable. In addition, multiple cameras are required to generate these multi-camera views, and the failure of one or more cameras can disrupt the operation of such systems. Therefore, it is desirable to have improved techniques for perceptually enhanced surround views based on sensor fusion. [Overview of the project]
[0002] This disclosure relates to a method, the method comprising: obtaining a first position of a vehicle, wherein the vehicle has one or more cameras arranged around the vehicle, each camera being associated with a physical camera orientation indicating where each camera is located relative to the vehicle; capturing a first image of a first area with the first camera; associating the first image with a first position of the vehicle when the first image is captured; moving the vehicle in a certain direction such that the first area is no longer in the field of view of the first camera; obtaining a second position of the vehicle; determining a temporal camera orientation based on the physical camera orientation of the first camera and the second position of the vehicle; and rendering a view of the first area based on the temporal camera orientation and the first image.
[0003] Another aspect of the present disclosure relates to a non-temporary program storage device, the non-temporary program storage device includes instructions stored to cause one or more processors to acquire a first position of a vehicle, wherein the vehicle has one or more cameras arranged around the vehicle, each camera associated with a physical camera orientation indicating where each camera is located relative to the vehicle; to receive a first image of a first area from the first camera; to associate the first image with a first position of the vehicle when the first image is captured; to acquire a second position of the vehicle after the vehicle has moved in a certain direction such that the first area is no longer in the field of view of the first camera; to determine a temporal camera orientation based on the physical camera orientation of the first camera and the second position of the vehicle; and to represent a view of the first area based on the temporal camera orientation and the first image.
[0004] Another aspect of the present disclosure relates to a system for representing a view around a vehicle, the system comprising one or more cameras arranged around the vehicle, each camera associated with a physical camera orientation indicating where each camera is located relative to the vehicle; memory; and one or more processors operably coupled to the memory and one or more cameras. The one or more processors are configured to execute non-temporary instructions causing one or more processors to: acquire a first position of the vehicle; capture a first image of a first area with a first camera; associate the first image with the first position of the vehicle when the first image is captured; acquire a second position of the vehicle after the vehicle has moved in a certain direction such that the first area is no longer in the field of view of the first camera; determine a temporal camera orientation based on the physical camera orientation of the first camera and the second position of the vehicle; and represent a view of the first area based on the temporal camera orientation and the first image.
[0005] The techniques described herein are considered in relation to visible light cameras, and while a bowl shape is used to determine orientation relative to physical and virtual cameras, it should be understood that these techniques are not limited to such sensors or techniques for determining orientation. Rather, the techniques described herein are readily applicable to a wide range of sensor devices, including invisible light sensors or electromagnetic sensors, including infrared and near-infrared sensors, or cameras capable of capturing images across a wide range of electromagnetic frequencies. The techniques described herein are also applicable to other forms of determining orientation relative to physical and virtual cameras.
[0006] Next, refer to the attached drawings for a detailed explanation of various examples. [Brief explanation of the drawing]
[0007] [Figure 1A] This figure shows a technique for generating a 3D surround view according to an aspect of this disclosure. [Figure 1B] This figure shows a technique for generating a 3D surround view according to an aspect of this disclosure.
[0008] [Figure 2] This is a diagram of an exemplary three-dimensional (3D) bowl mesh for use in a surround-view system according to an aspect of this disclosure.
[0009] [Figure 3] This document describes a ray tracing process for mapping a virtual camera to a physical camera, in accordance with the aspects of this disclosure.
[0010] [Figure 4A] This document illustrates the exemplary effects of temporal mapping in accordance with the aspects of this disclosure. [Figure 4B] This document illustrates the exemplary effects of temporal mapping in accordance with the aspects of this disclosure.
[0011] [Figure 5]This flowchart shows a technique for an improved surround view according to the aspects of this disclosure.
[0012] [Figure 6] This document illustrates exemplary changes in the vehicle's attitude according to the aspects of this disclosure.
[0013] [Figure 7] This is a block diagram showing one embodiment of the system according to the aspects of this disclosure.
[0014] [Figure 8] This is a block diagram showing one embodiment of a computing device according to the aspects of this disclosure. [Modes for carrying out the invention]
[0015] Figure 1A shows a technique for generating a 3D surround view according to an aspect of the present disclosure. The process for generating the 3D surround view generates a composite image from a viewpoint that appears to be located directly above the vehicle when viewed straight down. Substantially, a virtual top view of the vicinity of the vehicle is provided.
[0016] A vehicle surround view system typically includes four to six fisheye cameras mounted around a vehicle 110. For example, the camera set may include one at the front of the vehicle 110, another at the rear of the vehicle 110, and one on each side of the vehicle 110. The images generated by each camera may be provided to an image signal processing system (ISP) which includes memory circuits for storing one or more frames of image data from each camera. The fisheye images 111-114 captured by each camera may be conceptually arranged around the vehicle 110, for example.
[0017] A general process for generating a surround view from multiple fisheye cameras is described in "Surround view camera system for ADAS on TI's TDAx SoCs" by Vikram Appia et al. in October 2015, which is incorporated herein by reference. The basic surround view camera solution typically includes two main algorithmic components: geometric alignment and synthetic view integration. Geometric alignment corrects the fisheye distortion for the input video frames and transforms them into a common bird's-eye perspective. The integration algorithm generates a synthetic surround view after geometric correction. Another main algorithm, called "photometric alignment," may be required to generate a seamlessly stitched surround view output. Photometric alignment corrects the luminance and color inconsistencies between adjacent views to achieve seamless stitching. Photometric correction is described in detail, for example, in U.S. Patent Application No. 14 / 642,510, filed on March 9, 2015, titled "Method, Apparatus and System for Processing a Display From a Surround View Camera Solution," which is incorporated herein by reference. [Non-Patent Document 1] “Surround view camera system for ADAS on TI's TDAx SoCs,” Vikram Appia etal, October 2015 [Patent Document 1] U.S. Patent Application No. 14 / 642,510 “Method, Apparatus and System for Processing a Display From a Surround View Camera Solution”
[0018] Camera system calibration may include both fisheye lens distortion correction (LDC) and perspective transformation. In the case of fisheye distortion correction, a radial distortion model can be used to remove the fisheye from the original input frame by applying the inverse transformation of the radial distortion function. After LDC, four extrinsic calibration matrices can be estimated for each camera so that all input views are properly registered within a single world coordinate system in order to transform the four input LDC corrected frames. A chart-based calibration technique can be used. The content of the chart is designed to facilitate accurate and reliable discovery and alignment of features by an algorithm. Chart-based calibration is discussed in detail, for example, in U.S. Patent Application No. 15 / 294,369, filed October 14, 2016, titled "Automatic Feature Point Detection for Calibration of Multi-Camera Systems", which is hereby incorporated by reference. [Patent Document 2] U.S. Patent Application No. 15 / 294,369 "Automatic Feature Point Detection for Calibration of Multi-Camera Systems"
[0019] Assuming that appropriate geometric alignment has already been applied to the input frames, the composite surround view 132 of FIG. 1B can be generated, for example, using a digital signal processor (DSP). The composite surround view uses data from all four input frames from a set of cameras. The overlapping regions are portions of the frames captured by two adjacent cameras that are from the same physical world, namely O{m,n}, where m = 1, 2, 3, 4 and n = (m + 1) mod 4. O{m,n} indicates the overlapping region between view m and view n, and n is the neighboring view of view m in clockwise order. At each position in O{m,n}, there are two available pixels, namely the image data from view m and its spatial counterpart from view n.
[0020] The calibrated camera system receives input video streams from four fisheye cameras and generates a surround view integration function that creates a composite 3D surround view 132. The LDC module can perform fisheye correction, perspective warp, alignment, and bilinear / bicubic interpolation on image frames from each of the four fisheye cameras. The LDC module may be a hardware accelerometer (HWA) module and may be incorporated as part of a DSP module or graphics processing unit (GPU), for example. The DSP module may also perform stitching and superimpose vehicle images, such as vehicle image 134, onto the final composite output image 132.
[0021] This integration produces a stitched output image using mappings encoded in a geometric LUT. In the overlapping region of the output frame where image data from two adjacent input frames is required, each output pixel is mapped to a pixel position in the two input images. In the overlapping region, the image data from the two adjacent images may be mixed, or a binary choice may be made to use data from one of the two images.
[0022] In areas where usable image data is unavailable, gaps may result in the stitched output image. For example, areas beneath a vehicle are generally not directly captured and may appear as blank or black areas in the stitched output image. Typically, these blank areas are filled with superimposed images of the vehicle, such as vehicle image 134. If a camera is disabled, the corresponding areas that would normally be captured by that camera may appear as blank or black areas in the stitched output image.
[0023] Figure 2 shows an exemplary three-dimensional (3D) bowl mesh 200 for use in a surround-view system according to an aspect of the present disclosure. In the case of a 3D image, the world around a vehicle may be represented by the shape of a bowl. Since there is no full depth to the scene, the bowl is a reasonable estimate of the shape of the world around the vehicle. This bowl can be any smooth, fluctuating surface. In this particular representation, the bowl 200 is used, which is a flat portion 201 in the area close to the vehicle and a curved portion shown as 202 and 203 at the front and rear as it moves away from the vehicle. In this example, both sides of the bowl may curve slightly upward as shown in 204. In other embodiments, other bowl shapes may be used.
[0024] Images such as the stitched-together output image may be superimposed onto the 3D bowl mesh 200 by, for example, a graphics processing unit (GPU) or image processor, and a set of virtual viewpoints or virtual cameras may be defined along with the mapping from the cameras used to create the stitched-together output image and the virtual viewpoint.
[0025] Figure 3 illustrates a ray tracing process 300 for mapping a virtual camera to a physical camera, according to an aspect of this disclosure. This example shows a cross-sectional view of a portion 302 of a bowl mesh, similar to the bowl mesh 200 in Figure 2. The bowl mesh 302 may include flat portions 304 and raised portions 306, similar to the flat portions 201 and raised portions 202 in Figure 2. As described in more detail above, a camera 308 equipped with a fisheye lens 310 may be mounted on the front of an actual vehicle. The virtual viewpoint 312 for the output image may be defined, for example, above the actual vehicle position.
[0026] When projected onto the bowl mesh 302, initial calibration of the camera may be used to provide a mapping of positions in the imaging area to pixels of the camera 308 equipped with a fisheye lens 310. This mapping may be prepared, for example, during the calibration stage and stored, for example, in a lookup table. As previously mentioned, the virtual viewpoint 312 may be defined at a location separate from the hardware camera 308. The mapping for the virtual viewpoint 312 may be defined by projecting rays from the virtual viewpoint 312 position on the virtual viewpoint image plane 314 and identifying the locations where the rays intersect the bowl mesh 302. Rays 316 and 318 are examples. For example, ray 316 intersects the flat portion 302 of the bowl mesh 302, and ray 318 intersects the raised portion 306 of the bowl mesh 302. The ray projection operation generates a mapping of any 2D point on the virtual image plane 314 having the corresponding coordinates of the bowl mesh 302. Subsequently, a mapping can be generated between the region visible to the virtual viewpoint 312 and the region visible to the camera 308, using the mapping between the camera 308 and the bowl mesh 302, as well as the mapping between the virtual viewpoint 312 and the bowl mesh 302.
[0027] According to the aspects of this study, the visible area for the virtual viewpoint 312 may include areas not visible to camera 308. In such cases, the mapping for the virtual viewpoint may be based on mapping between multiple cameras and the bowl mesh 302. Note that the virtual viewpoint can be arbitrarily positioned and is not limited to a standard view directly above the vehicle and surrounding area. For example, the virtual viewpoint may be defined to be above and slightly behind the vehicle to give the view a greater sense of 3D. In addition, in some cases, the viewpoint may be dynamically moved, for example by the user. In such cases, the mapping may be dynamically recalculated or based on a set of mappings recalculated for multiple defined positions. In some cases, areas not currently visible to any camera on the vehicle may have been previously imaged by one or more cameras on the vehicle. A temporal camera capable of providing images of the area may be used. The temporal camera can display images of the area even if cameras on the vehicle cannot directly image the area. These images of the area may have been captured at a previous point in time and may be used to provide images of the area and to provide a temporal dimension to the virtual camera viewpoint.
[0028] Figures 4A and 4B illustrate the exemplary effect of temporal mapping according to aspects of the present disclosure. Figure 4A shows a first example for representing a view below a vehicle, and Figure 4B shows a second example for representing a view when the camera is disabled. As shown in this example, in the case of a moving vehicle, an area that is not visible to the camera on the vehicle at a current time such as t1 may have been visible to the camera on the vehicle at an earlier time such as t0. In Figure 4A, at time t0, vehicle 402A, having a camera pointed in the direction of travel, in this case forward, is able to image an area 404 in front of vehicle 402A, including a reference area 406. At time t1, vehicle 402B has moved far enough forward that vehicle 402B is now over the previously imaged area 404 and the reference area 406. For clarity, it should be noted that the examples provided involve a forward-facing camera and a vehicle moving forward. However, those skilled in the art will understand that other cameras corresponding to the direction of travel may be used, such as a rear-facing camera for reversing.
[0029] In Figure 4B, again at time t0, vehicle 410A, which has a camera pointed in the direction of travel, here forward, is able to image the region 412 in front of vehicle 410A, which includes a reference region 414, corresponding to the field of view of the right-side camera of vehicle 410B. At time t1, vehicle 410B has traveled far enough forward that region 412 is now within the field of view of the right-side camera of vehicle 410B. In this case, the right camera 416 is inactive, and no image is received from the right-side camera. Although the reference region 414 cannot be directly imaged by the camera on vehicle 410B at time t1, the reference region 414 was previously imaged at an earlier time, such as t0. Therefore, a time dimension may be added to the mapping as between a virtual viewpoint and one or more camera views.
[0030] In accordance with aspects of this disclosure, one or more history buffers may be provided to store images captured by one or more cameras positioned on a vehicle. For example, a separate history buffer may be provided for each camera, or a central history buffer may be provided for some or all of the cameras. In some cases, the history buffer may be large enough to buffer images for a set time frame and / or distances for one or more cameras, which are supported by the history buffer. Using this history buffer, images for temporal cameras may be provided in a similar manner to live camera images for virtual cameras.
[0031] Figure 5 is a flowchart 500 illustrating a technique for an improved surround view according to an aspect of the present disclosure. In step 502, the method begins by acquiring a first position of a vehicle, the vehicle having one or more cameras arranged around the vehicle, each camera associated with a physical camera orientation indicating where each camera is located relative to the vehicle. Generally, the vehicle includes one or more cameras, such as cameras at the front, right, left, and rear of the vehicle, which are configured to capture images of an area around the vehicle. Each camera is associated with a physical orientation indicating the camera's position and orientation. In some cases, position information may be acquired by any known technique, such as using Global Positioning System (GPS) coordinates. In some cases, these GPS coordinates may be acquired by additional sensors, such as accelerometers or other inertial sensors. In step 504, the method includes capturing a first image of a first area with a first camera. For example, cameras arranged around a vehicle may capture a stream of images of an area around the vehicle. In some cases, one or more of these cameras may be fisheye cameras, and they are calibrated so that images from multiple cameras can be stitched together to generate a surround view of the vehicle. For example, images may be projected onto a bowl mesh and one or more virtual cameras used to generate a surround view of the vehicle.
[0032] In step 506, the method includes associating a first image with a first position of the vehicle when the first image is captured. For example, when an image is captured by a camera, the image is associated with the current position of the vehicle. These captured images and associated positions can be stored in an image or temporal buffer. The temporal buffer may be a single temporal buffer shared by one or more cameras, or multiple temporal buffers may be provided for each camera, such as a temporal buffer for each camera. Multiple temporal buffers may be interconnected. In some cases, images in the temporal buffer may be stored in chronological order, and the stored images may be based on one or more threshold distances between the position of the vehicle associated with an image and the position of the vehicle associated with another image already stored in the temporal buffer. In some cases, multiple images stored in the temporal buffer may be used to represent a portion of a single image for display. For example, the resolution of a fisheye camera may degrade relatively rapidly over a certain distance. To mitigate this degradation of resolution, multiple stored images can be combined to represent a single image. As a more concrete example, since images are stored in relation to location, when displaying a view of the area beneath a vehicle, a first portion of the area beneath the vehicle, such as the first third, may be displayed using a temporal image captured from a first distance up to that first third of the area. The second third of the area may be displayed using a second temporal image captured from a second distance up to the second third, the second distance being close but immediately following the first third of the area. Similarly, the third third of the area may be displayed using a third temporal image captured from a third distance up to the third third, the third distance being close but immediately following the second third of the area.
[0033] In some cases, images may be stored in a time buffer based on the vehicle's direction of travel. For example, if the vehicle is moving substantially forward, images from the forward-facing camera may be stored in the time buffer, but images from the rear-facing camera may not. Conversely, if the vehicle is moving substantially backward, images from the rear-facing camera may be stored in the time buffer, but images from the forward-facing camera may not. In some cases, images may be removed from the time buffer based on the maximum distance between the vehicle's position associated with the image in the time buffer and the vehicle's current position.
[0034] In step 508, the method includes moving the vehicle in a direction such that the first area is no longer within the field of view of the first camera. In the first example, the vehicle may be moved such that the first area is substantially below the vehicle. In the second example, the vehicle may be moved in a direction such that the first area is substantially not within the field of view of the first camera, but is within the expected field of view of the second camera. In this second example, the second camera is disabled or unavailable in other circumstances, and therefore the first area cannot be seen from the second camera. In some cases, the second camera can be replaced as a whole by a virtual camera. For example, the vehicle may include front and rear cameras for capturing views of the area, and temporal virtual cameras used to provide left and right views of the vehicle. In some cases, the viewing angle or area within the view of the temporal camera may be adjusted compared to the predicted field of view of the second camera, and the image quality of the other camera may be further limited at the edges of the area captured by the other camera, for example due to lens distortion, fisheye lens, etc., so the view provided by the temporal virtual camera may have reduced resolution, captured area, and / or area. Adjusting the viewing angle or area within the view of the temporal camera helps to reduce the impact of reduced image quality. In step 510, the second position of the vehicle is acquired.
[0035] In step 512, the method includes determining a temporal camera orientation based on the physical camera orientation of a first camera and a second position of the vehicle. For example, as will be further discussed below, the temporal camera orientation may be based on changes in the pre-calibrated physical camera orientation of the first camera and the orientation of the vehicle. Images stored in the temporal buffer may be selected based, for example, the current position of the vehicle, as well as a threshold distance between the current position of the vehicle and the position of the vehicle associated with the selected image.
[0036] In step 514, the method includes representing a view of a first area based on a virtual camera pose and a first image. In some cases, an image selected from a temporal buffer may be projected onto a bowl mesh. The view from the temporal camera may be determined as a temporal image, which may be displayed, for example, on a display in a vehicle. In some cases, images selected from the temporal buffer and projected onto the bowl mesh may be stitched together to form a composite temporal image. The view from the temporal camera may be based on the composite temporal image.
[0037] The temporal camera orientation can be determined to help generate an integrated historical view of areas previously captured by cameras on the vehicle. In some cases, information regarding changes in the vehicle's orientation can be obtained using a combination of GPS and an inertial measurement unit (IMU). For example, GPS position information may be provided by an augmented GPS and combined with rotation / translation information provided by an accelerometer or other inertial sensor to determine the vehicle's orientation at a given time. This orientation information can be associated with images stored in a history buffer.
[0038] FIG. 6 shows an exemplary change 600 in the vehicle's pose according to an aspect of the present disclosure. Generally, pose refers to the position and orientation of an actual or virtual object with respect to a coordinate system and is described in the form of an M matrix. In this example, a vehicle 602A having a first pose at time t0 moves to a second pose at time t1, where the vehicle 602B has a second pose that is multi-dimensionally different from the first pose. To handle the change in the vehicle's pose, the pose of the temporal camera can be based on the change in the vehicle's pose with respect to a previous time. This change in the vehicle's pose can be described by ΔM and can be based on changes in the vehicle's position, rotation, and / or translation. Thus, the pose of the temporal camera can be described by FC(t0) M W(t1) = FC(t0) M FC(t1) × FC(t1) M W(t1) where, in the above equation, FC represents the pose of the front camera, W represents the world coordinate system such as the coordinates of the bowl mesh at a specific point in time, FC(t0) M FC(t1) is provided by ΔM, FC(t1) M W(t1) is provided by camera calibration. Once the relative pose of the temporal camera is determined, the relative pose can be used as the pose for the virtual camera using the corresponding image stored in the history buffer.
[0039] Inconsistent selection of images from the history buffer can lead to problems with temporal consistency, flickering, or other artifacts. Distance thresholds may be used to help determine the correct image from the history buffer for use with a temporal camera. In some cases, a threshold distance from the camera may be defined for a temporal camera. For example, a threshold distance of 5 meters from the position of the vehicle's front camera may be defined for use with a temporal camera. In some cases, images in the history buffer may be arranged chronologically. When selecting an image from the history buffer, the translation component of the orientation of each image may be examined starting with the earliest one to determine whether the image was taken from a distance greater than the threshold distance. If an image was not taken from a distance greater than the threshold distance, the next image is examined until a first image with a distance greater than the threshold distance is found. The first image with a distance greater than the threshold distance may be selected as the image for use with the temporal camera.
[0040] To help maintain image consistency, in some cases, images stored in the time buffer may be removed based on changes in the vehicle's direction. For example, if a vehicle moving backward stops and then begins moving forward, the images stored in the time buffer for the rear-facing camera may be removed, and new images from the forward-facing camera may be stored in the time buffer. This helps keep the images stored in the time buffer current, as the position of objects may have shifted while the vehicle was moving in the other direction. Similarly, if the vehicle remains stationary for a certain threshold amount of time, the time buffer may be cleared because objects may have shifted. To communicate this to the vehicle driver, the transparency of the model, such as vehicle image 134 in Figure 1, may be reduced to make the model more opaque when the time buffer is invalidated. The transparency of the model may be increased to make the model less opaque when images are stored in the time buffer, in order to generate a view of the area captured by the time camera.
[0041] In accordance with aspects of this disclosure, buffer optimization techniques may be used to limit the number of images stored in the history buffer. For example, if a vehicle is moving slowly and the history buffer can be filled rapidly, it may not be necessary to store every possible image frame in the history buffer. To help reduce the number of images that need to be stored, images from a camera may be stored in the history at a discrete distance-frequency threshold. For example, when the distance-frequency threshold is set to 10 centimeters, images from a camera may only be stored in the history buffer if the translation associated with an image is greater than 10 centimeters from the most recently stored image.
[0042] In some cases, a maximum distance for storing images may also be set. For example, an image buffer may be configured to store images up to a maximum threshold distance, such as 5 meters, beyond a minimum threshold distance. The maximum number of images per camera supported by the image buffer can then be calculated, and the image buffer's size can be appropriately determined. For example, if the image buffer is configured to store images associated with a maximum threshold distance of 5 meters and a minimum threshold distance of 10 centimeters between images, the maximum number of images per camera that can be stored in the history buffer is 50 images.
[0043] Figure 7 is a block diagram showing one embodiment of system 700 according to an aspect of the present disclosure. This exemplary system 700 includes a number of cameras, such as cameras 702-708, arranged around a vehicle and coupled to a capture block 710. Block 712 may perform color correction operations (such as conversion from Bayer format to YUV420 format, color mapping, noise filtering, and gamma correction) using known or future-developed image processing methods, if necessary. Block 714 may perform automatic exposure control and white balance of the video sensors to achieve optimal image quality using known or future-developed techniques. Block 716 synchronizes all cameras 700-708 to ensure that each frame captured by the sensors is within the same time period. In some cases, position information provided by a position subsystem 726 may be associated with the synchronized frames captured by the cameras. The position subsystem may include, for example, a GPS sensor, as well as other sensors such as inertial or acceleration sensors. The synchronized frames may be stored in a temporal buffer 732. The buffer manager 736 can manage the images stored in the temporal buffer 732 by performing thresholding to determine whether to store a certain image, removing images based on the distance and direction of the vehicle's movement, the amount of time it is stationary, and managing which camera images are stored in the temporal buffer 732. In some cases, the temporal buffer may be optimized to store only images from the front and rear-facing cameras.
[0044] A mapping lookup table generated by the calibrator 724 can be used by the warp module 728 to warp input video frames directly provided by cameras 702-708, as well as images stored in the temporal buffer 732 based on virtual and temporal cameras. Thus, both fisheye distortion correction and viewpoint warping can be performed in a single operation using a predetermined viewpoint mapping.
[0045] The synthesizer module 730 is responsible for generating a composite video frame containing one frame from each video channel. The synthesis parameters can be modified depending on the virtual viewpoint. This module is similar to the integration block described above with respect to Figure 1. Instead of a fisheye input image, the synthesizer module 730 receives warp-corrected outputs for each camera image from the warp module 728.
[0046] The synthesizer module 730 can stitch together and mix images corresponding to the proximity camera and the temporal camera. The mixing position varies based on the position of the virtual view, and this information can also be used to offline encode the resulting world view mesh.
[0047] The display subsystem 734 may receive the video stream output from the synthesizer 730 and display the same stream on a connected display unit such as an LCD, monitor, or TV for viewing by the vehicle driver. The system may also be configured to display metadata such as detected objects, pedestrians, and warnings.
[0048] In the specific implementation described herein, four cameras are used. The same principle disclosed herein can be extended to up to N cameras in other embodiments, where N may be greater than or less than 4.
[0049] Camera calibration mapping data 718 may be generated by the calibration procedure in combination with the world view mesh and stored in the 3D bowl mesh table 720. As described in more detail above, the world view mesh 720 may be generated offline 722 and stored for future use by the calibrator module 724.
[0050] For each predefined virtual viewpoint, the calibrator module 724 reads the associated 3D bowl mesh table 720, compensates for the camera calibration parameters 718, and generates a 2D mesh lookup table for each of the four channels. This is typically a one-time operation, performed when the system is started, for example, when the system is placed inside a vehicle during the assembly process. This process can be repeated whenever a position change is detected for one of the cameras mounted on the vehicle. Thus, the 3D bowl mesh table 720 can be generated for each frame of the temporal camera, as the calibration of the temporal camera changes with each frame as the vehicle moves. In some embodiments, the calibration process can be repeated, for example, each time the vehicle is started.
[0051] In some cases, captured image data from a camera may not be suitable for use in conjunction with a temporal buffer. For example, when a vehicle such as a car is moving through congested traffic, the captured image from the camera may include images of other vehicles. Such images are unsuitable for use with a temporal camera that displays images of the area beneath the vehicle, for example. In such cases, the temporal camera may be disabled, for example, by making the vehicle model opaque when the captured image contains objects that invalidate its use for the temporal camera. Once the image is captured and stored in a temporal buffer that does not contain such objects, the transparency of the model may be increased to make the model less opaque. Objects in the captured image can be detected and identified using any known technique.
[0052] As shown in Figure 8, device 800 includes processing elements such as a processor 805, which includes one or more hardware processors, each of which may have one or more processor cores. Examples of processors include, but are not limited to, a central processing unit (CPU) or a microprocessor. Although not shown in Figure 8, the processing elements constituting processor 805 may also include one or more other types of hardware processing components, such as a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), and / or a digital signal processor (DSP). In some cases, processor 805 may be configured to perform tasks described in relation to modules 710-716, 724-730 in Figure 7.
[0053] Figure 8 shows a memory 810 that can be operationally and communicatively coupled to the processor 805. The memory 810 may be a non-temporary computer-readable storage medium configured to store various types of data. For example, the memory 810 may include one or more volatile devices, such as random-access memory (RAM). In some cases, the temporal buffer 730 in Figure 7 may be part of the memory 810. The non-volatile storage device 820 may include one or more disk drives, optical drives, solid-state drives (SSDs), tap drives, flash memory, electrically erasable programmable read-only memory (EEPROM), and / or any other type of memory designed to retain data for a certain duration after a power outage or shutdown operation. The non-volatile storage device 820 may also be used to store programs that are loaded into RAM when such programs are executed.
[0054] Those skilled in the art will be aware that software programs can be developed, encoded, and compiled in various computing languages for various software platforms and / or operating systems, and subsequently loaded and executed by the processor 805. In one embodiment, the software program compilation process may translate program code written in one programming language into another computer language so that the processor 805 can execute the programming code. For example, the software program compilation process may generate an executable program that provides encoded instructions (e.g., machine code instructions) for the processor 805 to achieve a specific, non-general, particular computing function.
[0055] After the compilation process, the encoded instructions may be loaded from storage 820, from memory 810, into processor 805 as computer executable instructions or process steps, and / or embedded within processor 805 (e.g., via cache or onboard ROM). Processor 805 may be configured to execute stored instructions or process steps in order to implement the instructions or process steps in order to translate the computing device into a non-general-purpose, specific, specially programmed machine or device. Stored data, for example, data stored by storage device 820, may be accessed by processor 805 to instruct one or more components within computing device 800 during the execution of computer executable instructions or process steps. Storage 820 may be divided or partitioned into multiple sections that can be accessed by different software programs. For example, storage 820 may include sections designated for a specific purpose, such as storing program instructions or data for updating the software of computing device 800. In one embodiment, the software to be updated includes ROM or firmware of the computing device. In some cases, computing device 800 may include multiple operating systems. For example, computing device 800 may include a general-purpose operating system used for normal operation. Computing device 800 may also include another operating system, such as a boot loader, to perform specific tasks such as upgrading and restoring the general-purpose operating system, and to enable access to computing device 800 at a level not generally available through the general-purpose operating system. Both the general-purpose operating system and the other operating system may access sections of storage 820 designated for specific purposes.
[0056] One or more communication interfaces may include wireless communication interfaces for interfacing with one or more wireless communication devices. In some cases, elements coupled to the processor may be contained on hardware shared with the processor. For example, the communication interface 825, storage 820, and memory 810 may be contained within a single chip or package, such as within a system-on-a-chip (SOC), along with other elements such as digital radio. Computing devices may also include input and / or output devices, not shown, examples of which include human input devices such as sensors and cameras, mice, keyboards, and touchscreens, monitors, display screens, tactile or motion generators, speakers, and light sources. For example, processed input from radar device 830 may be output from computing device 800 to one or more other devices via communication interface 825.
[0057] The above considerations are intended to be illustrative of the principles and various implementations of this disclosure. Those skilled in the art will recognize numerous variations and modifications once the above disclosure is fully understood. The following claims are intended to encompass all such variations and modifications.
[0058] While this specification has described conventional vehicles with an onboard driver, other embodiments may be implemented in vehicles where the "driver" is located far from the vehicle, such as autonomous vehicles that can be controlled from a remote site.
[0059] As used herein, the term “vehicle” may also apply to other types of devices, such as robots, industrial devices, and medical devices, where low-cost, low-power image processing from multiple cameras is advantageous for forming a virtual viewpoint in real time.
[0060] The techniques described herein may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the software may run on one or more processors, such as microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and digital signal processors (DSPs). Software performing such techniques may be initially stored in a computer-readable medium, such as a compact disc (CD), diskette, tape, file, memory, or any other computer-readable storage device, and then loaded onto a processor and executed. In some cases, the software may also be sold as a computer program product, including the computer-readable medium and packaging materials for the computer-readable medium. In some cases, software instructions may be distributed via a transmission path from a computer-readable medium on another digital system, such as via a removable computer-readable medium (e.g., a floppy disk, optical disc, flash memory, or USB key).
[0061] In this specification, the term “to connect” means either an indirect or direct wired or wireless connection. Therefore, when a first device connects to a second device, the connection may be via a direct connection or via an indirect connection through other devices and connections.
[0062] Modifications to the described embodiments are permitted within the claims, and other embodiments are also possible.
Claims
1. It is a method, To obtain the first position of the vehicle, A first camera positioned on the vehicle captures a first image of a first region in the first field of view of the first camera at a first time interval. When the first image is captured, the first image is associated with the first position of the vehicle, The first image of the first region is stored in an image buffer, To obtain a second position of the vehicle, wherein the first region is not in the first field of view of the first camera at the second position, The first camera captures a second image of a second region that at least partially overlaps the first region at a second time after the first time, When the second image is captured, the second image is associated with the second position of the vehicle, The second image of the second region is stored in the image buffer, Obtaining a third position of the vehicle after the vehicle has moved in a certain direction such that the first region is within the expected second field of view of a second camera positioned on the vehicle, wherein the third position of the vehicle is not captured by the second camera. When the first image is captured, the first physical camera orientation of the first camera is determined based on the first position, When the second image is captured, the second physical camera orientation of the first camera is determined based on the second position, Representing a view of the first region based on the first physical camera orientation, the second physical camera orientation, the third position, the first image stored in the image buffer, and the second image stored in the image buffer, To generate a 3D bowl mesh, Representing a 3D surround view related to the vehicle, which includes projecting at least the first image and the second image onto the 3D bowl mesh, Methods that include...
2. The method according to claim 1, A method such that the first field of view of the first camera does not overlap in time with the second field of view of the second camera.
3. The method according to claim 1, A method wherein the first field of view of the first camera is in a different direction from the second field of view of the second camera.
4. The method according to claim 1, Representing a view of the first region is Selecting candidate images from the aforementioned image buffer, Determining the distance between the current position of the vehicle and the position of the vehicle associated with the selected candidate image, The selected candidate image is chosen when the distance is greater than or equal to the minimum distance threshold. Representing the view based on the selected candidate image, A method that includes repeating the process.
5. The method according to claim 4, A method further comprising deleting an expired image from the image buffer based on the determination that the distance between the vehicle's position associated with the candidate image and the vehicle's current position is greater than a maximum distance threshold.
6. The method according to claim 1, The method further includes capturing a third image of the first region with a third camera. A method for representing a view of the first region, further based on a third image of the first region.
7. A non-temporary program storage device, One or more processors, To obtain the first position of the vehicle, Receiving a first image of a first region in the first field of view of the first camera from a first camera positioned in the vehicle at a first time interval, When the first image is captured, the first image is associated with the first position of the vehicle, The first image of the first region is stored in an image buffer, To obtain a second position of the vehicle, wherein the first region is not in the first field of view of the first camera at the second position, The first camera captures a second image of a second region that at least partially overlaps the first region at a second time after the first time, When the second image is captured, the second image is associated with the second position of the vehicle, The second image of the second region is stored in the image buffer, Obtaining a third position of the vehicle after the vehicle has moved in a certain direction such that the first region is within the expected second field of view of a second camera positioned on the vehicle, wherein the third position of the vehicle is not captured by the second camera. When the first image is captured, the first physical camera orientation of the first camera is determined based on the first position, When the second image is captured, the second physical camera orientation of the first camera is determined based on the second position, Representing a view of the first region based on the first physical camera orientation, the second physical camera orientation, the third position, the first image stored in the image buffer, and the second image stored in the image buffer, To generate a 3D bowl mesh, Representing a 3D surround view related to the vehicle, which includes projecting at least the first image and the second image onto the 3D bowl mesh, A non-temporary program storage device that contains stored instructions for performing certain actions.
8. A non-temporary program storage device according to claim 7, A non-temporary program storage device in which the first field of view of the first camera does not temporally overlap with the second field of view of the second camera.
9. A non-temporary program storage device according to claim 7, A non-temporary program storage device in which the first field of view of the first camera is in a different direction from the second field of view of the second camera.
10. A non-temporary program storage device according to claim 7, The stored instructions for representing a view of the first region are, in the one or more processors, Selecting candidate images from the aforementioned image buffer, Determining the distance between the current position of the vehicle and the position of the vehicle associated with the selected candidate image, The selected candidate image is chosen when the distance is greater than or equal to the minimum distance threshold. Representing the view based on the selected candidate image, A non-temporary program storage device that repeatedly performs the following process.
11. A non-temporary program storage device according to claim 10, The stored instructions are sent to one or more processors. A non-temporary program storage device that further causes the device to delete an expired image from the image buffer based on the determination that the distance between the vehicle's position associated with the candidate image and the vehicle's current position is greater than a maximum distance threshold.
12. A non-temporary program storage device according to claim 7, The stored instructions are sent to one or more processors. Further, the system captures a third image of the first region using a third camera. A non-temporary program storage device that represents a view of the first region, further based on a third image of the first region.
13. A system that represents the view around a vehicle, A first camera positioned on the aforementioned vehicle, A second camera positioned on the aforementioned vehicle, Memory and One or more processors operably coupled to the memory and the first and second cameras, wherein the one or more processors To obtain the first position of the vehicle, The first camera captures a first image of a first region in the first field of view of the first camera at a first time, When the first image is captured, the first image is associated with the first position of the vehicle, The first image of the first region is stored in the memory, To obtain a second position of the vehicle, wherein the first region is not in the first field of view of the first camera at the second position, The first camera captures a second image of a second region that at least partially overlaps the first region at a second time after the first time, When the second image is captured, the second image is associated with the second position of the vehicle, The second image of the second region is stored in the memory, Obtaining a third position of the vehicle after the vehicle has moved in a certain direction such that the first region is within the expected second field of view of a second camera positioned on the vehicle, wherein the third position of the vehicle is not captured by the second camera. When the first image is captured, the first physical camera orientation of the first camera is determined based on the first position, When the second image is captured, the second physical camera orientation of the first camera is determined based on the second position, Representing a view of the first region based on the first physical camera orientation, the second physical camera orientation, the third position, the first image stored in the memory, and the second image stored in the memory, To generate a 3D bowl mesh, Representing a 3D surround view related to the vehicle, which includes projecting at least the first image and the second image onto the 3D bowl mesh, One or more processors configured to execute non-temporary instructions, which cause the following to occur: A system that includes this.
14. The system according to claim 13, A system in which the first field of view of the first camera is in a different direction from the second field of view of the second camera.
15. The system according to claim 13, A system in which the first field of view of the first camera is in a different direction from the second field of view of the second camera.
16. The system according to claim 13, When the non-transient instruction is executed by the one or more processors, Selecting candidate images from the aforementioned memory, Determining the distance between the current position of the vehicle and the position of the vehicle associated with the selected candidate image, The selected candidate image is chosen when the distance is greater than or equal to the minimum distance threshold. Representing the view based on the selected candidate image, A system that causes the view to be represented by repeating the process at least once.
17. The system according to claim 16, A system in which, when executed by the one or more processors, the non-transient instruction further causes the one or more processors to delete an expired image from memory based on a determination that the distance between the vehicle's position associated with the candidate image and the vehicle's current position is greater than a maximum distance threshold.
18. A system that represents the view around a vehicle, A first camera positioned on the aforementioned vehicle, A second camera positioned on the aforementioned vehicle, Memory and One or more processors operably coupled to the memory and the first and second cameras, wherein the one or more processors To obtain the first position of the vehicle, The first camera captures a first image of a first region in the first field of view of the first camera, When the first image is captured, the first image is associated with the first position of the vehicle, The first image of the first region is stored in the memory, To obtain the second position of the aforementioned vehicle, The first camera captures a second image of a second region that at least partially overlaps the first region, When the second image is captured, the second image is associated with the second position of the vehicle, The second image of the second region is stored in the memory, Obtaining a third position of the vehicle after the vehicle has moved in a certain direction such that the first region is within the expected second field of view of a second camera positioned on the vehicle, wherein the third position of the vehicle is not captured by the second camera. When the first image is captured, the first physical camera orientation of the first camera is determined based on the first position, When the second image is captured, the second physical camera orientation of the first camera is determined based on the second position, Representing a view of the first region based on the first physical camera orientation, the second physical camera orientation, the third position, the first image stored in the memory, and the second image stored in the memory, To generate a 3D bowl mesh, Representing a 3D surround view related to the vehicle, which includes projecting at least the first image and the second image onto the 3D bowl mesh, One or more processors configured to execute non-temporary instructions, which cause the following to occur: Includes, When executed by the one or more processors, the non-transient instruction further causes the one or more processors to capture a third image of the first region with a third camera. A system in which, when executed by one or more processors, the non-transient instruction causes one or more processors to represent a view by at least representing a view of the first region based on a third image of the first region.