Sensor fusion-based perception-enhanced surround view
By acquiring the location and camera pose on the vehicle, capturing images with the camera, and moving the vehicle, combined with history buffers and virtual camera technology, the problems of incomplete field of view and the impact of multiple camera failures in the external view system of the vehicle are solved, and a stable panoramic view is achieved.
Patent Information
- Application Number
- CN202080062849.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-21
- Filing Date
- 2020-09-08
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2040-09-08
AI Technical Summary
Existing external view systems for vehicles suffer from incomplete field of view, and multi-camera systems are susceptible to camera malfunctions, leading to operational instability.
By obtaining the location of the vehicle and the camera pose, images are captured using one or more cameras, and a panoramic view is presented based on the time-based camera pose after the vehicle has moved. This is combined with history buffers and virtual camera technology to fill in gaps in the field of view.
It enables panoramic views of the surroundings of vehicles, solves the problem of incomplete field of view, improves the stability and reliability of the system, and reduces the reliance on multiple cameras.
Smart Images

Figure CN114423646B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle perception systems. Background Technology
[0002] Increasingly, vehicles such as cars, airplanes, and robots are equipped with multiple external cameras to provide the operator with an external view of the area surrounding the vehicle. These external views are often used to assist in maneuvering the vehicle, for example, when reversing or parking. Multiple camera views can be stitched together to form an external surround view of the vehicle. However, external views of areas not within the field of view of any camera in such systems may be unavailable. Furthermore, generating these multi-camera views requires multiple cameras, and the failure of one or more cameras can impede the operation of such systems. Therefore, an improved technique for sensor fusion-based perception-enhanced surround views is desired. Summary of the Invention
[0003] This disclosure relates to a method comprising: obtaining a first position of a vehicle having one or more cameras disposed around the vehicle, wherein each camera is associated with a physical camera pose indicating where each camera is positioned relative to the vehicle; capturing a first image of a first region by the first camera; as the first image is captured, associating the first image with the first position of the vehicle; moving the vehicle in one direction such that the first region is no longer within the field of view of the first camera; obtaining a second position of the vehicle; determining a temporal camera pose based on the physical camera pose of the first camera and the second position of the vehicle; and presenting a view of the first region based on the temporal camera pose and the first image.
[0004] Another aspect of this disclosure relates to a non-transitory program storage device containing instructions stored thereon that cause one or more processors to: obtain a first position of a vehicle having one or more cameras disposed around the vehicle, wherein each camera is associated with a physical camera pose indicating where each camera is positioned relative to the vehicle; receive a first image of a first region from a first camera; associate the first image with the first position of the vehicle when the first image is captured; obtain a second position of the vehicle after the vehicle has been moved in one direction such that the first region is no longer within the field of view of the first camera; determine a temporal camera pose based on the physical camera pose of the first camera and the second position of the vehicle; and present a view of the first region based on the temporal camera pose and the first image.
[0005] Another aspect of this disclosure relates to a system for presenting a view around a vehicle, the system comprising: one or more cameras disposed around the vehicle, wherein each camera is associated with a physical camera pose indicating where each camera is positioned relative to the vehicle; a memory; and one or more processors operatively coupled to the memory and the one or more cameras, wherein the one or more processors are configured to execute non-transitory instructions causing the one or more processors to: obtain a first position of the vehicle; capture a first image of a first region by a first camera; associate the first image with the first position of the vehicle when the first image is captured; obtain a second position of the vehicle after the vehicle has been moved in one direction such that the first region is no longer within the field of view of the first camera; determine a temporal camera pose based on the physical camera pose of the first camera and the second position of the vehicle; and present a view of the first region based on the temporal camera pose and the first image.
[0006] It should be understood that while the techniques described herein are discussed in the context of visible light cameras and the use of a bowl shape to determine the pose of physical and virtual cameras, nothing in this disclosure is intended to limit these techniques to such sensors and techniques used for pose determination. Rather, the techniques discussed herein are readily applicable to a wide range of sensor devices, including invisible light or electromagnetic sensors, including infrared, near-infrared, or cameras capable of capturing images across a wide range of electromagnetic frequencies. The techniques discussed herein are also further applicable to other methods of determining the pose of physical and virtual cameras. Attached Figure Description
[0007] For a detailed description of each example, please refer to the accompanying drawings, in which:
[0008] Figure 1A and 1B This is a diagram illustrating a technique for generating a 2D surround view according to aspects of this disclosure.
[0009] Figure 2 This is a description of an example three-dimensional (3D) bowl-shaped mesh for a surround view system according to aspects of this disclosure.
[0010] Figure 3 This describes a ray tracing process for mapping a virtual camera to a physical camera, according to aspects of this disclosure.
[0011] Figure 4A and 4B This describes an example of the effect of time mapping based on aspects of this disclosure.
[0012] Figure 5This is a flowchart illustrating a technique for enhancing a surround view according to aspects of this disclosure.
[0013] Figure 6 This describes examples of variations in the orientation of a vehicle according to aspects of this disclosure.
[0014] Figure 7 This is a block diagram of an embodiment of a system according to aspects of this disclosure.
[0015] Figure 8 This is a block diagram of an embodiment of a computing device according to aspects of the present disclosure. Detailed Implementation
[0016] Figure 1A This diagram illustrates a technique for generating a 3D surround view according to aspects of this disclosure. The process for generating a 3D surround view produces a composite image from a viewpoint that appears positioned directly above when the vehicle is looking downwards. Essentially, it provides a virtual top view of the vicinity around the vehicle.
[0017] A vehicle surround view system typically includes four to six fisheye cameras mounted around the vehicle 110. For example, the camera group includes one camera at the front of the vehicle 110, another camera at the rear of the vehicle 110, and one camera on each side of the vehicle 110. Images generated by each camera are provided to an image signal processing system (ISP) that includes memory circuitry for storing one or more frames of image data from each camera. For example, fisheye images 111 to 114 captured by each camera may conceptually be arranged around the vehicle 110.
[0018] The general process for generating surround views from multiple fisheye lens cameras is described in "Surround view camera system for ADAS on TI's TDAx SoCs" by Vikram Appia et al., October 2015, which is incorporated herein by reference. A basic surround view camera solution typically comprises two key algorithmic components: geometric alignment and composite view synthesis. Geometric alignment corrects fisheye distortion in the input video frame and transforms it into a common bird's-eye view. The synthesis algorithm generates a composite surround view after geometric correction. To produce a seamlessly stitched surround view output, another key algorithm, known as "photometric alignment," may be required. Photometric alignment corrects brightness and color mismatches between adjacent views to achieve seamless stitching. Photometric correction is described in detail in, for example, U.S. Patent Application 14 / 642,510, filed March 09, 2015, entitled “Method, Apparatus and System for Processing a Display From a Surround View Camera Solution,” which is incorporated herein by reference.
[0019] Camera system calibration can include both fisheye lens distortion correction (LDC) and viewpoint transformation. For fisheye distortion correction, a radial distortion model can be used to remove the fisheye from the original input frame by applying an inverse transformation of the radial distortion function. After LDC, four external calibration matrices (one for each camera) can be estimated to transform the four input LDC-corrected frames so that all input views are correctly registered in a single world coordinate system. A graph-based calibration method can be used. The graph is designed to facilitate the algorithm to accurately and reliably find and match features. Graph-based calibration is discussed in detail in, for example, U.S. Patent Application No. 15 / 294,369, filed October 14, 2016, entitled "Automatic FeaturePoint Detection for Calibration of Multi-Camera Systems," which is incorporated herein by reference.
[0020] Assuming correct geometric alignment has been applied to the input frame, then Figure 1BThe composite surround view 132 can be generated using, for example, a digital signal processor (DSP). The composite surround view uses data from all four input frames from a set of cameras. The overlap region is a portion of frames from the same physical world but captured by two adjacent cameras, i.e., O{m,n}, where m = 1, 2, 3, 4, and n = (m+1) mod 4. O{m,n} refers to the overlap region between view m and view n, where n is the adjacent view of view m in clockwise order. At each position in O{m,n}, there are two available pixels, i.e., the image data from view m and its spatial counterpart from view n.
[0021] A calibrated camera system generates a surround view synthesis function that receives input video streams from four fisheye cameras and creates a composite 3D surround view 132. The LDC module performs fisheye correction, perspective distortion, alignment, and bilinear / bicubic interpolation on image frames from each of the four fisheye cameras. For example, the LDC module can be a hardware accelerator (HWA) module and can be incorporated as part of a DSP module or graphics processing unit (GPU). The DSP module can also perform stitching on the final composite output image 132 and can overlay images of vehicles (e.g., vehicle image 134) onto the final composite output image 132.
[0022] This synthesis uses mappings encoded in a geometric LUT to create a stitched output image. In the overlapping region of the output frames, each output pixel is mapped to a pixel location in one of the two input images, where image data from two neighboring input frames is needed. In the overlapping region, image data from two neighboring images can be blended, or a binary decision can be performed to use data from one of the two images.
[0023] Areas lacking usable image data can result in holes in the stitched output image. For example, the area beneath a vehicle is typically not directly imaged and may appear as a blank or black area in the stitched output image. This blank area is usually filled by an overlay of the vehicle's image, such as vehicle image 134. In cases where the camera is deactivated, the corresponding area typically imaged by that camera may appear as a blank or black area in the stitched output image.
[0024] Figure 2This is a description of an example three-dimensional (3D) bowl-shaped mesh 200 for a surround view system according to aspects of this disclosure. For a 3D image, the world around a vehicle can be represented by the shape of a bowl. Due to the lack of full scene depth, a bowl is a reasonable assumption about the shape of the world around the vehicle. This bowl can be any smoothly varying surface. In this particular representation, a bowl 200 is used, which is flat 201 in the area near the vehicle and curved away from the vehicle, as indicated at 202 and 203 respectively at the front and rear. In this example, the bowl may only curve slightly upwards on each side, as indicated at 204. Other bowl shapes may be used for other embodiments.
[0025] For example, the image of the stitched output image can be overlaid onto the 3D bowl-shaped mesh 200, for example, via a graphics processing unit (GPU) or image processor, and a set of virtual viewpoints or virtual cameras can be defined together with mappings from the cameras used to create the stitched output image and the virtual viewpoints.
[0026] Figure 3 This describes a ray tracing process 300 for mapping a virtual camera to a physical camera, according to aspects of this disclosure. This example illustrates a process similar to... Figure 2 A cross-sectional view of portion 302 of the bowl-shaped mesh 200. The bowl-shaped mesh 302 may contain elements similar to... Figure 2 The flat portion 201 and the raised portion 202 have a flat portion 304 and a raised portion 306. A camera 308 with a fisheye lens 310 can be mounted in front of an actual vehicle, as described in more detail above. The virtual viewpoint 312 of the output image can be defined, for example, above the position of the actual vehicle.
[0027] The initial calibration of the camera can be used to provide a positional mapping (e.g., projected onto the bowl-shaped grid 302) to the pixels of the camera 308 with the fisheye lens 310 in the imaging area. This mapping can be prepared, for example, during the calibration phase and stored, for example, in a lookup table. As discussed above, a virtual viewpoint 312 can be defined at a location separate from the hardware camera 308. The mapping of the virtual viewpoint 312 can be defined by projecting rays from the virtual viewpoint 312 location in the virtual viewpoint image plane 314 and identifying the locations where the rays intersect the bowl-shaped grid 302. Rays 316 and 318 are examples. For example, ray 316 intersects the flat portion 302 of the bowl-shaped grid 302, and ray 318 intersects the convex portion 306 of the bowl-shaped grid 302. The ray projection operation produces a mapping of each 2D point on the virtual viewpoint image plane 314 to the corresponding coordinates of the bowl-shaped grid 302. Next, the mapping between the area visible from the virtual viewpoint 312 and the area visible from the camera 308 can be generated together with the mapping between the camera 308 and the bowl-shaped mesh 302, along with the mapping between the virtual viewpoint 312 and the bowl-shaped mesh 302.
[0028] According to this discussion, the visible area of virtual viewpoint 312 may include areas invisible to camera 308. In such cases, the mapping of the virtual viewpoint may be based on mappings between multiple cameras and the bowl-shaped grid 302. It should be noted that the virtual viewpoint can be placed arbitrarily and is not limited to a standard directly overhead view of the vehicle and its surrounding area. For example, the virtual viewpoint may be defined above and slightly behind the vehicle to provide a more 3D feel to the view. Additionally, in some cases, the viewpoint may be dynamically moved by the user, for example. In such cases, the mapping may be dynamically recalculated or based on a set of recalculated mappings from multiple defined locations. In some cases, areas not currently visible to any camera on the vehicle may have previously been imaged by one or more cameras on the vehicle. A time camera capable of providing images of the area can be used. The time camera can display images of the area even if the cameras on the vehicle cannot directly image the area. These images of the area may have been captured at a previous point in time and can be used to provide images of the area, thus providing a temporal dimension to the virtual camera viewpoint.
[0029] Figure 4A and 4B This describes an example of the effect of time mapping based on aspects of this disclosure. Figure 4A This describes the first instance used to present the view below the vehicle, and Figure 4B This illustrates a second example of a view where the camera is deactivated. As shown in this example, for a moving vehicle, an area where the camera on the vehicle is not visible at the current time point (e.g., t1) might have been visible to the camera on the vehicle at a previous time point (e.g., t0). Figure 4A In the example, at time t0, a vehicle 402A with a camera pointing in the direction of travel (forward in this case) is able to image area 404 (including reference area 406) in front of the vehicle 402A. At time t1, the vehicle 402B has traveled far enough forward that the vehicle 402B is now above the previously imaged area 404 and reference area 406. It should be noted that, for clarity, the provided example relates to a vehicle with a forward-facing camera that is moving forward. However, those skilled in the art will understand that other cameras corresponding to the direction of travel can be used, such as a rear-facing camera for reversing.
[0030] exist Figure 4BIn this scenario, at time t0, car 410A, with a camera pointing in the direction of travel (forward in this case), can image region 412 (including reference region 414) in front of car 410A. Region 412 corresponds to the field of view of the right-side camera of car 410B. At time t1, car 410B has traveled far enough that region 412 should now be within the field of view of the right-side camera of car 410B. In this case, the right-side camera 416 is deactivated and no image is received from it. Although reference region 414 cannot be directly imaged by the camera on car 410B at time t1, reference region 414 was previously imaged at a previous time point (e.g., t0). Therefore, a time dimension can be added to the mapping between the virtual viewpoint and one or more camera views.
[0031] According to aspects of this disclosure, one or more history buffers may be provided to store images captured by one or more cameras positioned around a vehicle. For example, a separate history buffer may be provided for each camera, or a central history buffer may be provided for some or all cameras. In some cases, the history buffer may be large enough to buffer images from the setup timeframes and / or distances of one or more cameras supported by the history buffer. This history buffer can be used to provide images from time-lapse cameras in a manner similar to live camera images from a virtual camera.
[0032] Figure 5 This is a flowchart 500 illustrating a technique for enhancing a surround view according to aspects of this disclosure. In step 502, the method begins by obtaining a first position of a vehicle having one or more cameras positioned around the vehicle, wherein each camera is associated with a physical camera pose indicating where each camera is positioned relative to the vehicle. Generally, the vehicle includes one or more cameras configured to capture images of an area around the vehicle, such as cameras at the front, right, left, and rear of the vehicle. Each camera is associated with a physical pose indicating the camera's position and orientation. In some cases, the position information can be obtained using any known technique, such as by using Global Positioning System (GPS) coordinates. In some cases, these GPS coordinates may be supplemented by additional sensors, such as accelerometers or other inertial sensors. In step 504, the method includes capturing a first image of the first area by the first camera. For example, cameras positioned around the vehicle can capture an image stream of the area around the vehicle. In some cases, one or more of these cameras may be fisheye cameras and have been calibrated such that images from multiple cameras can be stitched together to produce a surround view of the vehicle. For example, images can be projected onto a bowl-shaped grid, and one or more virtual cameras can be used to generate a surround view of a vehicle.
[0033] In step 506, the method includes associating the first image with a first location of the vehicle when the first image is captured. For example, because the image was captured by a camera, it is associated with the current location of the vehicle. These captured images and associated locations may be stored in an image or time buffer. The time buffer may be a single time buffer shared by one or more cameras, or multiple time buffers may be provided for each camera, such as a time buffer for each camera. Multiple time buffers may be interconnected. In some cases, the images in the time buffer may be stored chronologically, and the stored images may be based on one or more threshold distances between the location of the vehicle associated with the image and the location of the vehicle associated with another image already stored in the time buffer. In some cases, multiple images stored in the time buffer may be used to present a portion of a single image for display. For example, the resolution of a fisheye camera may decrease relatively quickly over a certain distance. To mitigate this reduced resolution, multiple stored images may be combined to present a single image. As a more specific example, because the images are stored along with their associated locations, when displaying a view beneath a vehicle, for example, the first portion of the first third of a zone beneath the vehicle can be displayed using a time image captured from a first distance to that first third of the zone. The second third of the zone can be displayed using a second time image captured from a second distance to that second third of the zone, which is adjacent to but immediately after the first third of the zone. Similarly, the third third of the zone can be displayed using a third time image captured from a third distance to that third third of the zone, which is adjacent to but immediately after the second third of the zone.
[0034] In some cases, images may be stored in a time buffer based on the vehicle's direction of travel. For example, if the vehicle is traveling primarily forward, images from a forward-facing camera may be stored in the time buffer, but images from a backward-facing camera may not. Conversely, if the vehicle is traveling primarily backward, images from a backward-facing camera may be stored in the time buffer, but images from a forward-facing camera may not. In some cases, images may be removed from the time buffer based on the maximum distance between the vehicle's position associated with the image in the time buffer and the vehicle's current position.
[0035] In step 508, the method includes moving the vehicle in one direction such that the first area is no longer within the field of view of the first camera. In a first instance, the vehicle is moved such that the first area is substantially below the vehicle. In a second instance, the vehicle is moved in one direction such that the first area is substantially outside the field of view of the first camera, but within the intended field of view of the second camera. In this second instance, the second camera is deactivated or otherwise unavailable, and therefore, the first area cannot be seen by the second camera. In some cases, it is possible that the second camera can be completely replaced by a virtual camera. For example, the vehicle may include a forward camera and a rear camera for capturing a view of an area, and a temporal virtual camera for providing views of the left and right sides of the vehicle. In some cases, the viewpoint or area in the temporal camera's view may be adjusted relative to the intended field of view of the second camera because the image quality of the other camera may be more limited at the edges of the area imaged by the other camera, for example due to lens distortion, fisheye lenses, etc., and the view provided by the temporal virtual camera may have reduced resolution, imaged area, and / or range. Adjusting the viewpoint or area in the time-lapse camera's view helps reduce the impact of reduced image quality. In step 510, the second position of the vehicle is obtained.
[0036] In step 512, the method includes determining a temporal camera pose based on the physical camera pose of the first camera and the second position of the vehicle. For example, as further discussed below, the temporal camera pose may be based on a pre-calibrated physical camera pose of the first camera and changes in the vehicle pose. Images stored in the temporal buffer may be selected, for example, based on the current position of the vehicle and a threshold distance between the current position of the vehicle and the position of the vehicle associated with the selected image.
[0037] In step 514, the method includes presenting a view of a first region based on the virtual camera pose and a first image. In some cases, an image selected from a time buffer may be projected onto a bowl-shaped grid. The view from the time camera may be determined as a time image, and this time image may be presented on a display, for example, within a vehicle. In some cases, images selected from the time buffer and projected onto the bowl-shaped grid may be stitched together to form a composite time image. The view from the time camera may be based on the composite time image.
[0038] To aid in generating a synthetic historical view of an area previously imaged by a camera on a vehicle, the pose of the time-lapse camera can be determined. In some cases, information related to changes in the vehicle's pose can be obtained using a combination of GPS and an inertial measurement unit (IMU). For example, GPS location information can be provided by augmented GPS and combined with rotation / translation information provided by an accelerometer or another inertial sensor to determine the vehicle's pose at a specific time. This pose information can then be correlated with images stored in a history buffer.
[0039] Figure 6 This describes an example of a change in the pose 600 of a vehicle according to aspects of this disclosure. Generally, pose refers to the position and orientation of a real or virtual object with respect to a coordinate system and is described in the form of an M-matrix. In this example, vehicle 602A, which has a first pose at time t0, moves to a second pose at time t1, while vehicle 602B has a second pose that differs from the first pose in multiple dimensions. To handle the change in the pose of the vehicle, the pose of the time camera can be based on the change in the vehicle's pose with respect to previous times. This change in the vehicle's pose can be described by ΔM and can be based on changes in the vehicle's position, rotation, and / or translation. Therefore, the pose of the time camera can be expressed by the formula... x Description, where FC represents the pose of the forward-facing camera at a specific time point, and W represents the world coordinate system at that specific time point, such as the coordinates of a bowl-shaped mesh, and where Provided by ΔM, and Provided via camera calibration. Once the relative pose of the time-lapse camera is determined, the relative pose can be used as the pose of the virtual camera using the corresponding image stored in the history buffer.
[0040] Inconsistent image selection from the history buffer can lead to problems related to temporal consistency, flickering, or other artifacts. To help determine the correct image from the history buffer for use with the time camera, a distance threshold can be used. In some cases, a threshold distance from the camera used for the time camera can be defined. For example, a threshold distance of five meters from the position of a forward-facing camera on a vehicle can be defined for use with the time camera. In some cases, the images in the history buffer can be arranged chronologically. When selecting an image from the history buffer, the pose translation component of each image can be checked from the earliest to determine if the image was taken from a distance greater than the threshold distance. If the image was not taken from a distance greater than the threshold distance, then the next image is checked until the first image with a distance greater than the threshold distance is found. The first image with a distance greater than the threshold distance can be selected as the image to be used with the time camera.
[0041] To help maintain image consistency, in some cases, images stored in the time buffer can be removed based on changes in the vehicle's direction. For example, if a vehicle traveling in the opposite direction stops and then begins to move forward, the image stored in the time buffer of the rearward camera can be removed, and a new image from the forward camera can be stored in the time buffer. This helps preserve images stored in the current time buffer, as objects may have shifted while the vehicle moves in the other direction. Similarly, if the vehicle has remained stationary for a certain threshold amount of time, the time buffer can be cleared, as objects may have shifted. To communicate this to the vehicle operator, the transparency of the model (e.g., vehicle image 134 in Figure 1) can be reduced to make the model more opaque when the time buffer is invalidated. The transparency of the model can be increased when an image is stored in the time buffer to make the model more transparent and produce a view of the area imaged by the time camera.
[0042] According to aspects of this disclosure, buffer optimization schemes can be used to limit the number of images stored in a history buffer. It may be unnecessary to store every possible image frame in the history buffer (e.g., if a vehicle is moving slowly), and the history buffer can be filled quickly. To help reduce the number of images that need to be stored, images from the camera can be stored in the history buffer discretely based on a distance frequency threshold. For example, if the distance frequency threshold is set to ten centimeters, if the translation associated with an image is more than ten centimeters away from the most recently stored image, then the image from the camera may only be stored in the history buffer.
[0043] In some cases, a maximum distance for storing images can also be set. For example, the image buffer can be configured to store images at a maximum threshold distance (e.g., five meters) that exceeds a minimum threshold distance. Then, the maximum number of images per camera supported by the image buffer can be calculated, and the size of the image buffer can be set appropriately. For example, if the image buffer is configured to store images associated with a maximum threshold distance of 5 meters (where the minimum threshold distance between images is ten centimeters), then the maximum number of images that can be stored in the history buffer per camera is 50 images.
[0044] Figure 7This is a block diagram of an embodiment of system 700 according to aspects of this disclosure. This example system 700 includes multiple cameras, such as cameras 700 to 708 positioned around the perimeter of a vehicle and coupled to capture block 710. If needed, block 712 performs color correction operations (e.g., conversion from Bayer format to YUV420 format, tone mapping, noise filtering, gamma correction, etc.) using known or later-developed image processing methods. Block 714 performs automatic exposure control of the video sensor and white balance to achieve optimal image quality using known or later-developed techniques. Block 716 synchronizes all cameras 700 to 708 to ensure that each frame captured from the sensor is within the same time period. In some cases, location information provided by location subsystem 726 may be associated with the synchronized frames captured by the cameras. The location subsystem may include, for example, a GPS sensor and other sensors, such as inertia or accelerometer sensors. The synchronized frames may be stored in time buffer 732. The buffer manager 736 manages images stored in the time buffer 732, for example, by performing thresholding to determine whether to store a specific image, removing images based on the distance and direction of vehicle travel, the amount of time at rest, etc., and managing from which camera images are stored in the time buffer 732. In some cases, the time buffer can be optimized to store only images from the forward and backward cameras.
[0045] The mapping lookup table generated by calibrator 724 can be used by warping module 728 to warp the input video frames directly provided by cameras 702 to 708 together with the images stored in time buffer 732 based on virtual and time cameras. Therefore, both fisheye distortion correction and viewpoint warping can be performed in a single operation using a predetermined viewpoint mapping.
[0046] The compositor module 730 is responsible for generating a composite video frame containing one frame from each video channel. The composite parameters can vary depending on the virtual viewpoint. This module is similar to the compositor block described above with respect to Figure 1. Instead of a fisheye input image, the compositor module 730 receives a distorted, modified output of each camera image from the distortion module 728.
[0047] The synthesizer block 730 can stitch together and blend images corresponding to neighboring cameras and time-lapse cameras. The blending position will change based on the position of the virtual view, and this information can also be encoded in the offline generated world-to-view grid.
[0048] The display subsystem 734 can receive video stream output from the synthesizer 730 and display the same content on a connected display unit for viewing by the driver of the vehicle, such as an LCD, monitor, TV, etc. The system can be configured to also display metadata of detected objects, pedestrians, alarms, etc.
[0049] In the specific implementation described herein, four cameras are used. In other embodiments, the same principles disclosed herein can be extended to N cameras, where N may be greater than or less than four.
[0050] Camera calibration mapping data 718 can be generated and stored in a 3D bowl-shaped mesh table 720 through a calibration procedure combined with the world-to-view mesh. As described in more detail above, the world-view mesh 720 can be generated offline 722 and stored for later use by the calibrator module 724.
[0051] For each predefined virtual viewpoint, the calibrator module 724 reads the associated 3D bowl-shaped mesh table 720, interprets the camera calibration parameters 718, and generates a 2D mesh lookup table for each of the four channels. This is typically a one-time operation and is completed at system startup, such as when the system is placed in a vehicle during assembly. This process can be repeated whenever a change in the position of one of the cameras mounted on the vehicle is sensed. Therefore, the 3D bowl-shaped mesh table 720 can be generated for each frame of the time-lapse camera, as the time-lapse camera calibration changes with each frame as the vehicle moves. In some embodiments, the calibration process can be repeated, for example, whenever the vehicle is started.
[0052] In some situations, image data captured by a camera may be ineffective when used in conjunction with a time buffer. For example, in a scenario where a vehicle, such as a car, is traveling in congested traffic, the image captured by the camera may contain images of other vehicles. As an example, such images would be unsuitable for use with a time camera that displays an image of the area beneath the vehicles. In such cases, the time camera can be deactivated, for example, by making the vehicle model opaque if the captured image contains objects that would render it unusable by the time camera. Once an image without such objects is captured and stored in the time buffer, the model's transparency can be increased to make it more transparent. Objects in the captured image can be detected and identified using any known techniques.
[0053] like Figure 8 The description states that device 800 includes processing elements, such as processor 805 containing one or more hardware processors, each of which may have one or more processor cores. Examples of processors include (but are not limited to) a central processing unit (CPU) or a microprocessor. Although Figure 8 Not specified, but the processing elements comprising processor 805 may also include one or more other types of hardware processing components, such as a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), and / or a digital signal processor (DSP). In some cases, processor 805 may be configured to perform combined... Figure 7The tasks described in modules 710 to 716 and 724 to 730.
[0054] Figure 8 The memory 810 is operatively and communicatively coupled to the processor 805. The memory 810 may be a non-transitory computer-readable storage medium configured to store various types of data. For example, the memory 810 may include one or more volatile devices, such as random access memory (RAM). In some cases, Figure 7 The time buffer 730 may be part of the memory 810. The non-volatile storage device 820 may include one or more disk drives, optical drives, solid-state drives (SSDs), tape drives, flash memory, electrically programmable read-only memory (EEPROM), and / or any other type of memory designed to retain data for a duration after a power-off or shutdown operation. The non-volatile storage device 820 may also be used for such programs stored in RAM during program execution.
[0055] Those skilled in the art will recognize that software programs can be developed, coded, and compiled in various computing languages for use on various software platforms and / or operating systems, and subsequently loaded and executed by processor 805. In one embodiment, the compilation process of a software program can transform program code written in a programming language into another computer language, enabling processor 805 to execute the programming code. For example, the compilation process of a software program can produce an executable program that provides processor 805 with encoded instructions (e.g., machine code instructions) to implement specific, non-general-purpose computational functions.
[0056] Following the compilation process, the encoded instructions are then loaded as computer-executable instructions or process steps from storage device 820, from memory 810, into processor 805, and / or embedded within processor 805 (e.g., via cache or onboard ROM). Processor 805 may be configured to execute the stored instructions or process steps to transform the computing device into a non-general-purpose, specific, specially programmed machine or device. For example, the stored data from storage device 820 may be accessed by processor 805 during the execution of the computer-executable instructions or process steps to indicate one or more components within computing device 800. Storage device 820 may be partitioned or divided into multiple segments accessible by different software programs. For example, storage device 820 may contain segments marked for a specific purpose, such as storing program instructions or data for updating software of computing device 800. In one embodiment, the software to be updated includes ROM or firmware of the computing device. In some cases, computing device 800 may contain multiple operating systems. For example, computing device 800 may include a general-purpose operating system for normal operation. Computing device 800 may also include another operating system for performing specific tasks, such as a bootloader, which may include updating and restoring the general-purpose operating system and allowing access to computing device 800 at levels normally unavailable through the general-purpose operating system. Both the general-purpose operating system and the other operating system may access segments of the identified storage device 820 for specific purposes.
[0057] One or more communication interfaces may include radio communication interfaces for interfacing with one or more radio communication devices. In some cases, components coupled to the processor may be included on hardware shared with the processor. For example, communication interface 825, storage device 820, and memory 810 may be included in a single chip or package, such as in a system-on-a-chip (SoC), along with other components, such as digital radio. The computing device may also include input and / or output devices not shown, examples of which include sensors, cameras, human input devices such as mice, keyboards, touchscreens, monitors, display screens, haptic or motion generators, speakers, lights, etc. For example, processed input from radar device 830 may be output from computing device 800 to one or more other devices via communication interface 825.
[0058] The foregoing discussion is intended to illustrate the principles and various embodiments of this disclosure. Once fully understanding the foregoing disclosure, those skilled in the art will recognize numerous variations and modifications. It is intended that the appended claims be interpreted as encompassing all such variations and modifications.
[0059] While conventional vehicles with onboard drives have been described herein, other embodiments may be implemented in vehicles in which the “drive” is remote from the vehicle, such as autonomous vehicles that can be controlled from a remote location.
[0060] As used in this article, the term "vehicle" can also be applied to other types of devices, such as robots, industrial devices, medical devices, etc., where low-cost, low-power processing of images from multiple cameras used to form a virtual viewpoint in real time is beneficial.
[0061] The techniques described in this disclosure can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the software can execute on one or more processors, such as microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), etc. The software executing the technology may initially be stored on a computer-readable medium, such as an optical disc (CD), magnetic disk, magnetic tape, file, memory, or any other computer-readable storage device, and then executed in a processor. In some cases, the software may also be sold as a computer program product, which includes a computer-readable medium and packaging material for the computer-readable medium. In some cases, the software instructions may be distributed via removable computer-readable media (e.g., floppy disks, optical discs, flash memory, USB keys), distributed from a computer-readable medium on another digital system via a transmission path, etc.
[0062] In this description, the term "couple" means an indirect or direct wired or wireless connection. Therefore, if a first device is coupled to a second device, that connection can be made either directly or indirectly via other devices and connections.
[0063] Modifications are possible in the described embodiments, and other embodiments are possible within the scope of the claims.
Claims
1. A method for presenting a view around a vehicle, comprising: A first position of a vehicle is obtained, the vehicle having one or more cameras positioned around the vehicle, and each camera being associated with a physical camera pose indicating where each camera is positioned relative to the vehicle; The first image of the first region is captured by the first camera; When the first image is captured, the first image is associated with the first location of the vehicle; After moving the vehicle in one direction so that the first area is no longer within the field of view of the first camera, the second position of the vehicle is obtained; The time camera pose is determined based on the physical camera pose of the first camera and the second position of the vehicle. A view of the first region is presented based on the time camera pose and the first image; The first image of the first region is stored in an image buffer, wherein the images in the image buffer are stored in chronological order, and the view is presented based on the images stored in the image buffer; and Based on the determination that the location of the vehicle associated with the candidate image and the current location of the vehicle is higher than the maximum distance threshold, expired images are deleted from the image buffer.
2. The method of claim 1, wherein moving the vehicle includes moving the vehicle in the direction such that the first region is below the vehicle, and wherein presenting the view includes presenting the first region below the vehicle.
3. The method according to claim 1, further comprising: A second image of a second region that at least partially overlaps with the first region is captured by the first camera; and The second image is stored in the image buffer.
4. The method of claim 3, wherein the third location of the vehicle associated with the second image is at least a minimum threshold distance from the first region; and The second image is stored based on the determination that the third position is at least at the minimum threshold distance from the first region.
5. The method of claim 1, wherein presenting the view comprises repeating the following steps: Select the candidate image from the image buffer; and Determine the distance between the current location of the vehicle and the location of the vehicle associated with the selected candidate image; If the distance is greater than or equal to the minimum threshold distance, then the selected candidate image is selected; and The view is presented based on the selected candidate images.
6. The method of claim 1, further comprising: A second image of the first region is captured by a second camera; and the view presenting the first region is further based on the second image of the first region.
7. A non-transitory program storage device comprising instructions stored thereon, the instructions causing one or more processors to perform the following operations: A first position of a vehicle is obtained, the vehicle having one or more cameras positioned around the vehicle, and each camera being associated with a physical camera pose indicating where each camera is positioned relative to the vehicle; Receive a first image of a first region from a first camera; When the first image is captured, the first image is associated with the first location of the vehicle; The second position of the vehicle is obtained after the vehicle has been moved in one direction so that the first area is no longer within the field of view of the first camera; The time camera pose is determined based on the physical camera pose of the first camera and the second position of the vehicle. A view of the first region is presented based on the time camera pose and the first image; The first image of the first region is stored in an image buffer, wherein the images in the image buffer are stored in chronological order, and the view is presented based on the images stored in the image buffer; and Based on the determination that the location of the vehicle associated with the candidate image and the current location of the vehicle is higher than the maximum distance threshold, expired images are deleted from the image buffer.
8. The non-transitory program storage device of claim 7, wherein when the vehicle moves in the direction such that the first region is below the vehicle, presenting the view includes presenting the first region below the vehicle.
9. The non-transitory program storage device of claim 7, wherein the stored instructions further cause one or more processors to perform the following operations: Receive a second image of a second region that at least partially overlaps with the first region from the first camera; and The second image is stored in the image buffer.
10. The non-transitory program storage device of claim 9, wherein the stored instructions further cause one or more processors to perform the following operations: Obtain the third location of the vehicle associated with the second image; Determine that the third location of the vehicle is at least a minimum threshold distance from the first area; and The second image is stored based on the determination that the third position is at least at the minimum threshold distance from the first region.
11. The non-transitory program storage device of claim 7, wherein the stored instructions for presenting the view further cause one or more processors to repeat the following steps: Select the candidate image from the image buffer; and Determine the distance between the current location of the vehicle and the location of the vehicle associated with the selected candidate image; If the distance is greater than or equal to the minimum threshold distance, then the selected candidate image is selected; and The view is presented based on the selected candidate images.
12. The non-transitory program storage device of claim 7, wherein the stored instructions further cause one or more processors to perform the following operations: A second image of the first region is captured by a second camera; and the view presenting the first region is further based on the second image of the first region.
13. A system for presenting a view around a vehicle, the system comprising: One or more cameras are positioned around the vehicle, wherein each camera is associated with a physical camera pose indicating where each camera is positioned relative to the vehicle. Memory; and One or more processors operatively coupled to the memory and the one or more cameras, wherein the one or more processors are configured to execute non-transitory instructions that cause the one or more processors to perform the following operations: Obtain the first position of the vehicle; The first image of the first region is captured by the first camera; When the first image is captured, the first image is associated with the first location of the vehicle; The second position of the vehicle is obtained after the vehicle has been moved in one direction so that the first area is no longer within the field of view of the first camera; The time camera pose is determined based on the physical camera pose of the first camera and the second position of the vehicle. A view of the first region is presented based on the time camera pose and the first image; The first image of the first region is stored in an image buffer, wherein the images in the image buffer are stored in chronological order, and the view is presented based on the images stored in the image buffer; and Based on the determination that the location of the vehicle associated with the candidate image and the current location of the vehicle is higher than the maximum distance threshold, expired images are deleted from the image buffer.
14. The system of claim 13, wherein presenting the view includes presenting the first region below the vehicle when the vehicle moves in the direction such that the first region is below the vehicle.
15. The system of claim 13, wherein the non-transitory instruction further causes the one or more processors to perform the following operations: Receive a second image of a second region that at least partially overlaps with the first region from the first camera; and The second image is stored in the image buffer.
Citation Information
Patent Citations
Automatic feature point detection for calibration of multi-camera systems
US10438081B2
Method, apparatus and system for processing a display from a surround view camera solution
US20150254825A1
Apparatus and method for generating peripheral image of vehicle
US20170148136A1
Periphery monitoring device
US20190149774A1