Image synthesis in multi-view automotive and robotics systems.
A closed-loop synthesis quality optimization module addresses distortions in automotive visualization by optimizing geometric and photometric scores, resulting in high-quality panoramic images that improve safety and accuracy in automotive and robotic systems.
Patent Information
- Application Number
- JP2021100390
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-31
- Filing Date
- 2021-06-16
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2041-06-16
AI Technical Summary
Conventional automotive visualization systems suffer from distortions and artifacts due to limitations in camera technology, leading to inaccurate and incomplete information, which compromises driver safety and the effectiveness of image processing.
A closed-loop synthesis quality optimization module is used to assess and optimize image compositing by calculating geometric and photometric scores, applying transformations to minimize distortions and ensure accurate alignment and color matching between multiple camera inputs.
The solution provides high-quality, artifact-free panoramic images that enhance safety by ensuring accurate and complete visualization, suitable for automotive and robotic systems.
Smart Images

Figure 0007725249000004 
Figure 0007725249000005 
Figure 0007725249000006
Abstract
Description
[Technical Field]
[0001] Automotive systems are increasingly using visualization solutions to assist drivers, provide information and enhance safety. [Background technology]
[0002] Various techniques for providing visualization in automotive applications suffer from various disadvantages. For example, cameras often cannot capture all areas where safety issues may arise or all areas that may be of interest to the viewer. Attempts to provide a wider field of view may result in images that are distorted and / or have artifacts due to inherent camera limitations or characteristics, such as location, orientation, and intrinsic parameters. Any distorted visualization or artifacts in automotive visualization make it more difficult for drivers, other vehicle users, safety monitors, and traditional image processing techniques to process the information provided by the automotive visualization system, because the artifacts and / or distorted visualization do not reflect the real world. These distortions and / or artifacts may reduce driver safety by providing inaccurate or incomplete information. Summary of the Invention [Means for solving the problem]
[0003] Embodiments of the present disclosure relate to improved image compositing in automotive or robotic systems using quality assessment feedback. Systems and methods are disclosed for constructing improved panoramic images generated in automotive or robotic systems or platforms to promote user safety and diagnostic capabilities. Image compositing involves combining image data from two or more sources, such as cameras, for display in a single visualization engine. That is, image data captured from each individual camera is combined or stitched to present a view of the entire scene. Traditional compositing processes often produce artifacts and anomalies that obscure the overall view of the scene and cause improper alignment between individual images, which is particularly problematic in vehicular applications where artifacts can appear in safety-critical areas.
[0004] Conventional automotive visualization systems may use naive image synthesis techniques such as homography estimation, which are open-loop, limited to pairwise synthesis, and require the estimation of correspondences between feature points identified between the image data of two individual images. However, because feature points arise only in the presence of distinct objects in the scene, this approach is ineffective in some scenes and produces geometric distortions, where the stitched images contain misaligned edges between overlapping portions. In addition, homography estimation produces geometric distortions in scenes with depth.
[0005] In contrast to conventional systems such as those described above, improved image synthesis in automotive systems uses quality assessment feedback in a closed-loop synthesis quality optimization module to guide various adjustments to each image to minimize artifacts and other distortions. This synthesis quality optimization module receives as input calibration parameters and image data from two or more cameras, and can optionally receive data from other vehicle sensors as well. The synthesis quality optimization module can calculate the overlap area of each image pair and calculate one or more synthesis quality scores based on the overlap area. If any synthesis quality score is below a threshold, the synthesis quality optimization module calculates and applies one or more geometric or photometric transformations to perform on a subset of the input image data. Based on the transformed image data, the synthesis quality optimization module loops until one or more synthesis quality scores are all equal to or greater than the threshold. During each loop, the synthesis quality optimization module calculates one or more synthesis quality scores and applies one or more geometric or photometric transformations.
[0006] The present system and method for improved image synthesis in automotive systems using quality assessment feedback is described in detail below with reference to the accompanying drawings. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 2 is a block diagram illustrating a compositing pipeline, in accordance with some embodiments of the present disclosure. [Figure 2A] FIG. 10 is an illustration of compositing using two images, according to some embodiments of the present disclosure. [Figure 2B] 1A-1C are diagrams of bowl projection image synthesis before and after optimization, according to some embodiments. [Figure 3] FIG. 1 is a block diagram illustrating a composition engine pipeline, according to some embodiments of the present disclosure. [Figure 4] FIG. 2 is a block diagram illustrating a compositing module for performing improved image compositing, according to some embodiments of the present disclosure. [Figure 5] FIG. 1 is a block diagram illustrating a synthesis quality optimization module, according to some embodiments of the present disclosure. [Figure 6] FIG. 10 is a block diagram illustrating a quality assessment module according to some embodiments of the present disclosure. [Figure 7] FIG. 1 illustrates a process for performing improved compositing using compositing quality optimization, according to some embodiments of the present disclosure. [Figure 8A] 1 is an illustration of an exemplary autonomous vehicle, according to some embodiments of the present disclosure. [Figure 8B] 8B is an illustration of camera positions and fields of view for the example autonomous vehicle of FIG. 8A, according to some embodiments of the present disclosure. [Figure 8C] FIG. 8B is a block diagram of an example system architecture of the example autonomous vehicle of FIG. 8A, in accordance with some embodiments of the present disclosure. [Figure 8D] FIG. 8B is a system diagram of communication between a cloud-based server and the example autonomous vehicle of FIG. 8A, according to some embodiments of the present disclosure. [Figure 9] FIG. 1 is a block diagram of an exemplary computing device suitable for use in implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0008] Systems and methods are disclosed for an improved image synthesis pipeline using quality assessment feedback for use by automotive systems and platforms or robotics systems and platforms featuring surround view capabilities.
[0009] In a multi-view system, such as those used in a surround-view vehicle system, each camera captures an image corresponding to a portion of the scene. The captured image data from each camera is combined, or stitched, to present a view of the entire scene in a single image through the composition of two or more images. Composition is performed using a composition engine or composition pipeline as a component of the surround-view vehicle system or any other system in which multiple overlapping images must be combined.
[0010] When a surround view vehicle system, as described further herein, stitches two or more images captured by vehicle cameras, one or more of the resulting images are often improperly aligned, containing artifacts and anomalies that obscure or confuse the overall view of the scene. An example of this is described further below in conjunction with FIG. 2A. Conventional approaches for image synthesis focus on homography estimation to align unstitched images.
[0011] Using homography estimation, conventional open-loop surround-view automotive systems merge overlapping images by estimating correspondences between feature points identified in the individual image data. That is, systems that implement homography estimation to perform image fusion identify feature points in each of the two images to be merged and estimate how the feature points in each image correspond to each other. Using the correspondences, overlapping areas are identified and the images are merged. However, because feature points only occur in the presence of distinct objects in the scene, conventional techniques, such as homography estimation, are ineffective for some scenes and can produce geometric distortions that appear when the stitched images contain misaligned edges between overlapping portions. Additionally, homography estimation can produce geometric distortions of the overall view in scenes with depth.
[0012] Conventional methods also have limited applicability and are only useful for pairwise synthesis. These conventional methods cannot be used for multi-view systems where three or more cameras are used. In addition to the limitations of homograph estimation, multi-view systems introduce photometric distortions when individual cameras have different exposure levels, which results in inconsistent color values across the entire stitched scene, as discussed further below in conjunction with Figure 2A.
[0013] In contrast, closed-loop feedback-based techniques for performing image synthesis in automotive systems and platforms or robotics systems and platforms featuring surround-view capabilities can be used to synthesize two or more images in settings where homograph estimation is unavailable or results in significant artifacts. In addition, feedback-based techniques are robust in the face of photometric distortions and can perform color correction to match images used as input to automotive systems and platforms or robotics systems and platforms featuring surround-view capabilities that implement the feedback-based techniques.
[0014] A feedback-based approach for image synthesis in automotive systems and platforms, or robotics systems and platforms featuring surround view capabilities, and other systems requiring image synthesis, uses a closed-loop synthesis quality optimization module to guide the geometric adjustment and photometric placement between two or more stitched images. As described below in conjunction with Figures 4-6, this synthesis quality optimization module receives as input image data and intrinsic and extrinsic calibration parameters from two or more cameras in the automotive systems and platforms, or robotics systems and platforms featuring surround view capabilities, and may optionally receive data from other vehicle sensors, as further described herein.
[0015] Each image in a pair of images is first transformed from fisheye space to a synthesis space, as captured from each camera, as further described herein in conjunction with FIG. 8B. As described below in conjunction with FIGS. 1 and 4, a synthesis engine or module then stitches two or more individual images in synthesis space by projecting each image into the image coordinates of a virtual camera, where the virtual camera represents the stitched or combined pair of images, including an overlap region. This overlap region is calculated for each image pair and indicates the portion of each image from each camera that overlaps in the virtual camera's image space. Based on this overlap region, a synthesis quality optimization module performs one or more quality assessment operations to calculate one or more synthesis quality scores and determine one or more transformations to be applied to the individual input images of each image pair.
[0016] For example, in one embodiment, the composition quality optimization module performs one or more quality assessment operations to calculate a geometric score and a photometric score for each stitched image, as described further below in conjunction with FIG. 6. The geometric score is a numerical value that represents the amount or degree of structural errors resulting from the composition of two or more images into a single image by the composition engine, as described further below in conjunction with FIGS. 5 and 6. In one embodiment, a higher geometric score indicates that the stitched image contains fewer structural errors and more closely represents the composition engine's ideal combination of the two images. A lower geometric score indicates that the stitched image contains more structural errors that represent an improper composition by the composition engine. The photometric score is a numerical value that represents the degree of color difference between the two images to be stitched, as described further below in conjunction with FIGS. 2 and 6. In one embodiment, a higher score indicates that each image to be stitched is similar to the other in terms of color properties. A lower score indicates that each image to be stitched contains significant color property differences.
[0017] As described further below in conjunction with FIG. 5, the compositing quality optimization module analyzes both the geometric score and the photometric score. In one embodiment, if the geometric score is below a threshold (indicating low compositing quality), the compositing quality optimization module applies a geometry alignment. During the geometry alignment, the compositing quality optimization module calculates a 3D transformation to be applied to one or more of the stitched images, as described further below in conjunction with FIG. 5. In one embodiment, the 3D transformation is a geometric transformation that modifies one or more of the stitched images and adjusts the extrinsic calibration of one camera while leaving the extrinsic calibration of the other camera unchanged, such that the geometric quality matrix is maximized. After the compositing quality optimization module calculates the 3D transformation, the compositing quality optimization module applies the transformation to produce maximally aligned stitched or virtual camera images.
[0018] In one embodiment, if the photometric score is below a threshold (indicating poor color matching between the two input images), the composition quality optimization module applies a photometric adjustment, as described further below in conjunction with FIG. 5. During photometric adjustment, the composition quality optimization module improves photometric quality by matching color intensities and color angles between the input images from the two cameras. The composition quality optimization module calculates and applies a transformation to match the intensity and angle of the destination input image to the source input image for each RGB channel, where the transformation is based on applying scaling and offsets to the image's color channels. To find the optimal scale and offset, the composition quality optimization module directly searches for values that maximize the photometric score and, consequently, composition quality. To improve performance and reduce the time required to determine the optimal scale and offset, the composition quality optimization module, in one embodiment, reduces the search space by calculating an initial scale and offset as the average of both the source and destination images.
[0019] Because automotive systems require real-time performance, as described further below in conjunction with FIG. 5, in one embodiment, the composition quality optimization module uses optical flow vectors to reduce the search space required to compute 3D transformations when performing geometric adjustments. The reduced search space facilitates real-time determination of 3D transformations. To reduce the search space, the optical flow vectors facilitate the composition quality optimization module's identification of directions for rotating the image to improve the geometric quality score, thereby removing any transformations that do not rotate the image in that particular direction.
[0020] While the synthesis of two images from two cameras in an automotive surround view system will be used extensively for illustrative purposes, it should be noted that the techniques described herein can be adapted for other uses. For example, improved image synthesis using the various techniques described herein can be used in one embodiment to facilitate an improved virtual reality system. In another embodiment, improved image synthesis using the various techniques described herein can be used to facilitate an autonomous or semi-autonomous vehicle or other automotive application, where stitched images are input to a neural network or other machine learning technique used by one of the vehicle's systems (e.g., control system, emergency braking). Improved image synthesis using the various techniques described herein can be used in another embodiment to improve security camera systems. In another embodiment, improved image synthesis using the various techniques described herein can be used for medical applications. For example, improved image synthesis can be used to stitch images and / or video from multiple orthogonal cameras to provide a specialist or robotic surgeon with an improved view of the entire area around an organ where surgery is being performed. In another example, improved image synthesis using the various techniques described herein can be used to align projectors in movies or other visualization systems that require the combination of multiple camera views (e.g., IMAX). In another example, improved image synthesis using the various techniques described herein can be used by construction equipment, for example, by showing an excavator operator the area being excavated, perhaps when the view is obstructed by a bucket or other construction equipment component.
[0021] The compositing techniques described herein, in one embodiment, can be used to facilitate an automotive surround view system with three or more cameras generating three or more images to be combined or stitched together. In an automotive surround view system with three or more cameras generating images to be stitched together, the improved image compositing pipeline using the quality assessment feedback described herein, in one embodiment, is applied in two different ways. First, in one embodiment, a compositing quality optimization module may evaluate and improve the quality of each overlap region in a series between three or more images captured from the three or more cameras to be combined or stitched together. The adjusted image is used as a reference image when adjusting the next overlap region. Second, in one embodiment, the compositing quality optimization module optimizes the compositing quality of all overlap regions between three or more images captured from the three or more cameras to be combined or stitched together. To optimize the geometric quality, the compositing quality optimization module summarizes the geometric scores of all overlap regions between the images to form an overall geometric score for the stitched image as a whole. Based on this score, the composition quality optimization module searches the potential transformation space for each image to optimize the overall geometric score, as described below in conjunction with FIG. 5. To optimize photometric quality, the composition quality optimization module summarizes the photometric scores from all overlapping regions among three or more input images to form an overall photometric score for the stitched image. Based on this score, the composition quality optimization module searches the scale and offset space for each input image to maximize the overall photometric score.
[0022] In the foregoing and following descriptions, various techniques are described. For purposes of explanation, specific configurations and details are set forth in order to provide a thorough understanding of possible ways of implementing the techniques. However, it will also be apparent that the techniques described below may be practiced in different configurations without the specific details. Additionally, well-known features may be omitted or simplified to avoid obscuring the techniques being described.
[0023] Referring to FIG. 1, FIG. 1 is an exemplary architecture for performing video and / or image composition according to some embodiments of the present disclosure. It should be understood that this and other configurations described herein are provided by way of example only. Other configurations and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those illustrated, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as individual or distributed components or in conjunction with other components, and in any suitable combination and location. Various functions described herein as being performed by entities may be implemented by hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in a memory.
[0024] FIG. 1 is a block diagram illustrating a compositing pipeline or compositing block according to some embodiments of the present disclosure. In the compositing pipeline, a compositing engine 110 receives a multiview input 104. In one embodiment, the multiview input is data such as images from a camera 102, data from a sensor 106, and / or calibration parameters 108. In one embodiment, the multiview input 104 includes a camera 102, such as a surround camera 874, a wide-view camera 870, a mid-range camera 898, a long-range camera 898, or a stereo camera 868, as described further below in conjunction with FIG. 8B. In one embodiment, the camera 102 is any hardware device for capturing image and / or video data. For example, in one embodiment, the camera 102 is a fisheye camera or is equipped with a fisheye lens to generate images having a fisheye projection or perspective. In another embodiment, the camera 102 is any other type of camera for facilitating visualization 118 in an automotive surround view system. In another embodiment, camera 102 is any type of camera for facilitating visualization in a system that performs image synthesis, as described further herein.
[0025] Two or more cameras 102 generate images for the multiview input 104. In one embodiment, the images are data generated by one or more cameras 102 representing locations within a field of view, such as an area around a vehicle in an automotive surround view system. In one embodiment, the multiview input 104 includes calibration parameters 108. In one embodiment, the calibration parameters 108 are a set of data values that can be used to configure one or more cameras 102. For example, in one embodiment, the calibration parameters 108 include color, contrast, and brightness levels. In another embodiment, the calibration parameters 108 are a set of data values associated with the vehicle or other components of the image capture system, such as a sensor 106. In one embodiment, the sensor is a software and / or hardware component for collecting data related to the environment, such as a RADAR sensor 860 or a LiDAR sensor 864, as described further herein. The calibration parameters 108 include intrinsic and extrinsic parameters. In one embodiment, the intrinsic parameters are focal length, optical center or principal point, skew factor, and / or distortion parameters. In one embodiment, the extrinsic parameters are camera rotation and translation. For example, in one embodiment, the extrinsic calibration parameters 108 include location data representing rotation and other geometric adjustments to one or more images from one or more cameras 102. In one embodiment, the multi-view input includes two or more images from two or more cameras 102 and calibration parameters 108 associated with the two or more images. In one embodiment, each image is associated with its own individual set of calibration parameters 108. In another embodiment, two or more cameras 102 share one or more sets of calibration parameters 108. In one embodiment, the calibration parameters 108 may be unavailable. If the calibration parameters 108 are unavailable, feature matching techniques, such as homography estimation, are used to estimate a transformation that can be used to transform two or more multi-view input 104 images into a common stitching space or a single image.
[0026] The composition engine 110 receives two or more images and associated calibration parameters 108 from the cameras 102, as well as optional sensor 106 data, as multiview input 104. In one embodiment, the composition engine 110 is data values and software instructions that, when executed, combine and position two or more images from two or more cameras 102 into a single image having a common coordinate system. In one embodiment, the composition engine 110 receives multiview input 104 from an ingest block that includes two or more cameras 102. In another embodiment, the composition engine 110 receives multiview input 104 from any group or implementation of two or more cameras 102.
[0027] Multi-view input 104, including images from cameras 102, optional sensor 106 data, and calibration parameters 108, is used by the synthesis engine to generate output 112, including one or more stitched images 114 and updated calibration parameters 116. The one or more stitched images 114 are image data generated by synthesis engine 110 based at least in part on two or more input images from cameras 102. In one embodiment, stitched image 114 is a 2D image or texture, including any data format supported by automotive surround view systems described further herein. In another embodiment, stitched image 114 is a 3D image or texture.
[0028] The composition engine 110 combines or stitches two or more multiview input 104 images by aligning each multiview input 104 image in a common coordinate system as a single image, or stitched image 114, as described further below in conjunction with FIG. 4. In one embodiment, the composition engine 110 aligns two or more multiview input 104 images into one or more stitched images 114 by aligning the images based on camera calibration parameters 108. In another embodiment, the composition engine 110 aligns two or more multiview input 104 images into one or more stitched images 114 by identifying one or more features or landmarks and adjusting the calibration parameters 108 to align each of the two or more multiview input 104 images. In another embodiment, the composition engine 110 aligns two or more multiview input 104 images based at least in part on sensor 106 data. In another embodiment, the composition engine 110 uses any other technique to align each multiview input 104 image into the output 112 stitched image 114.
[0029] In one embodiment, adjustments made to the multiview input 104 image results in updated calibration parameters 116. The updated calibration parameters 116 output 112 by the composition engine 110 include data representing new calibration values that can be used by two or more cameras 102 for further capture of image and / or video data or to adjust additional components, such as sensors 106. In one embodiment, the updated calibration parameters 116 include adjusted calibration data corresponding to structural or geometric adjustments, e.g., rotation and translation. In another embodiment, the updated calibration parameters 116 include adjusted calibration data corresponding to color or photometric adjustments, e.g., brightness and / or contrast. In one embodiment, the updated calibration parameters 116 include both geometric and photometric adjustments.
[0030] The synthesis engine 110 determines the quality of one or more stitched images, as described further below in conjunction with Figures 5 and 6. Based on the quality analysis, the synthesis engine determines one or more transformations to apply to the multiview input 104 images to generate an optimal output 112 stitched image 114. In one embodiment, a single-pass analysis is performed to determine the photometric and geometric quality of the stitched image 114 generated by the synthesis engine 110. In another embodiment, the synthesis engine 110 performs multiple passes using a feedback mechanism to improve the quality of the output 112 stitched image 114, as described further below in conjunction with Figure 5. Any adjustments made to the multiview input 104 images are reflected by updated calibration parameters 116.
[0031] In one embodiment, the output 112, including one or more stitched images 114 and any updated calibration parameters 116, is used by a visualization 118 module, such as a visualization engine or block in an automotive surround view system. In another embodiment, the output 112 from the composition engine 110 is used by any other visualization 118 system to display or otherwise use one or more stitched images 114. In one embodiment, one or more additional transformations are applied by a composition pipeline prior to visualization 118, as further described below in conjunction with FIG. 3.
[0032] FIG. 2A is an illustration of compositing using two images, according to some embodiments of the present disclosure. During compositing, two or more cameras capture a left image 202 and a right image 204. In one embodiment, the left image 202 is data including a representation of visual information corresponding to a left location within a scene, and the right image 204 is data including a representation of visual information corresponding to a right location within the same scene. As discussed above in conjunction with FIG. 1 and further described herein, in an automotive surround-view camera system or other multi-view system, the left image 202 and the right image 204 include an overlap region 206. In one embodiment, the overlap region 206 is data representing a shared projection space within a scene captured by two or more cameras. In one embodiment, the overlap region 206 includes one or more landmarks or features that can be used to position the left image 202 and the right image 204. In one embodiment, both the left image 202 and the right image 204 include one or more landmarks or features within the overlap region 206.
[0033] The compositing engine generates a stitched image 208 based at least in part on the left image 202 and the right image 204, for example, as described above in conjunction with Figure 1 and further below in conjunction with Figures 4-6. In one embodiment, the stitched image 208 includes a stitched region 210 that includes projection information about a scene from the overlap region 206 of the left image 202 and the right image 204. The stitched region 210 is data that represents a projection within a scene that includes the overlap region 206 of the left image 202 combined or aligned with the overlap region 206 of the right image 204.
[0034] In one embodiment, stitched region 216 of stitched image 208 includes a geometric error, where elements from left image 212 of stitched region 216 do not properly align with elements from right image 214 of stitched region 216. The geometric error is quantified by the synthesis engine using a geometric quality score, as described below in conjunction with FIG. 5 . In one embodiment, stitched region 222 of stitched image 208 includes a photometric error, where portions from left image 218 within stitched region 222 have different color levels compared to portions from right image 220 within stitched region 222. The different color levels are quantified by the synthesis engine using a photometric quality score, as described below in conjunction with FIG. 5 . The synthesis engine calculates or otherwise determines one or more transformations or other adjustments to be applied to left image 202 and / or right image 204 based at least in part on the geometric and photometric quality scores, as described further below in conjunction with FIGs. 5 and 6 .
[0035] FIG. 2B is a diagram of a bowl projection image synthesis before 222 and after 224 optimization, according to some embodiments. The bowl projection image synthesis before optimization 222 is a directly synthesized image including several images stitched together by an automotive surround view system. Before optimization, visual artifacts are visible, such as a misaligned building edge in the upper right quadrant of the bowl projection image synthesis 222. The bowl projection image synthesis after optimization 224 is a mixed and / or stitched image that corrects for geometric and photometric distortions using various techniques described further herein. After optimization 224, the visual artifacts visible in the bowl projection image synthesis before optimization 222 are mitigated or unobservable.
[0036] 3 is a block diagram illustrating a composition engine 302 pipeline according to some embodiments of the present disclosure. The composition engine 302 pipeline, such as that shown in FIG. 3, is a series of software and / or hardware modules that, when executed, combines multiview input 304 images into a single image containing elements of the multiview input 304 to produce a scene. In one embodiment, the composition engine 302 pipeline is a component of an automotive surround view system, as described further herein. In another embodiment, the composition engine is a component of any other visualization system, such as virtual reality or any other application that requires combining multiview inputs 304 into a single output 314 for visualization.
[0037] In one embodiment, the composition engine 302 receives as input a multi-view input 304 from a data ingest block comprising two or more cameras and calibration parameters, as described above in conjunction with Figure 1 and further herein. In another embodiment, the composition engine 302 receives as input any other data that can be used to facilitate image composition.
[0038] In one embodiment, the composition engine 302 includes a dewarp 306 module or dewarping module. In one embodiment, the dewarp 306 module is data values and software instructions that, when executed, transform the multiview input 304 images from image space to composition space. In one embodiment, the image space is a fisheye projection of the image data, or image data including fisheye coordinates. In one embodiment, the composition space is an equirectangular or equirectangular view projection space. In another embodiment, the composition space is a rectilinear projection space. In one embodiment, the composition space is any projection space usable by the composition engine 302 to project, combine, or otherwise use the multiview input 302 images to generate output 314 usable for visualization.
[0039] In one embodiment, the dewarp 306 module uses forward mapping to project the multiview input 304 from 2D image space to 3D composition space or other projection space. In forward mapping, the dewarp 306 module scans each multiview input 304 image pixel by pixel and copies them to the appropriate location in the composition space image. In another embodiment, the dewarp 306 module uses backward mapping to project the multiview input 304 from 2D image space to 3D composition space or other projection space. In backward mapping, the dewarp 306 module traverses each pixel in the destination composition space image and fetches or samples the correct pixel from one or more multiview input 304 images. A single point in an image from one input camera maps to an epipolar line in another input camera. Each multiview input 304 image is projected from the 2D input space to a 3D composition space, where the depth of each pixel is adjustable based on one or more transformations, as described below in conjunction with FIG. 5. In one embodiment, each 2D multiview input 304 image is mapped to an initial depth of zero for regions closer to the multiview camera sensor, as described further herein. In another embodiment, if the displacement between multiview input 304 images is small or negligible, or if the scene depth is far away, each 2D multiview input 304 image is mapped to an initial depth of infinity. For example, if each multiview input 304 camera captures images that contain no overlap or represent a scene at a far distance, no depth adjustment will result in overlapping regions of the multiview input 304 images, and the initial depth is mapped as infinity. In another embodiment, each 2D multiview input 304 image is mapped to any depth necessary to represent the viewpoint associated with the multiview input 304 image. In another embodiment, the dewarp 306 module maps each 2D multiview input 304 image to any depth estimated by any sensor or additional method, such as LiDAR-based depth estimation, monocular camera-based depth estimation, or multiview camera-based depth estimation.In one embodiment, the dewarp 306 module maps each 2D multiview input 304 image to an arbitrary depth estimated using any technique for depth estimation available in a surround-view system. In one embodiment, the dewarp 306 module uses one or more calibration parameters received by the composition engine 302 as the multiview input 304. In another embodiment, the dewarp 306 module does not use calibration parameters received by the composition engine 302 as the multiview input 304.
[0040] In one embodiment, the composition engine 302 includes a composition 308 module. In one embodiment, the composition 308 module is data values and software instructions that, when executed, combine and align input images, e.g., camera frames received as part of the multiview input, into one or more output 314 images having a common coordinate system. In one embodiment, the composition 308 module performs alignment of two or more multiview input 304 images, e.g., video frames or camera frames, based on one or more multiview input 304 calibration parameters. The composition 308 module aligns two or more multiview input 304 images, and in one embodiment, uses feature information from the two or more images. In another embodiment, the composition 308 module aligns two or more multiview input 304 images using disparity-based fine-tuning. In another embodiment, the composition 308 module detects structural artifacts or misalignments and refines calibration parameters to adjust the stitched image. To detect structural artifacts or inconsistencies, the synthesis 308 module uses quality analysis and feedback to improve synthesis quality, as further described below in conjunction with FIGS.
[0041] In one embodiment, the Compositing 308 module combines two or more Multiview Input 304 images during 2D to 3D dewarping 306 using backward mapping. That is, a stitched image representing a 3D virtual camera projection space receives information for each 3D pixel location from each of the two or more Multiview Input 304 images. Individual pixels corresponding to the multiple Multiview Input 304 images are blended using a Blending 310 module. In one embodiment, the Blending 310 module is a set of data values and software instructions that, when executed, create a smooth pixel-level transition between the two or more Multiview Input 304 images combined or projected into a shared 3D projection space by the Compositing 308 module. In one embodiment, the Blending 310 module combines pixel data from two or more Multiview Input 304 images to generate an output pixel. In one embodiment, the Blending 310 module combines individual pixel data from the 3D stitched space into an Output 314 2D image using either forward mapping or backward mapping, as described above. In another embodiment, the Blending 310 module combines data from multiple pixels in each of two or more Multiview Input 304 images to generate individual or groups of output pixels. In one embodiment, the Blending 310 module performs various blending techniques, such as alpha or weighted blending, multi-band blending, gradient blending, or optimal cut. In another embodiment, the Blending 310 module performs any other blending technique to combine or otherwise blend pixel data from the multiview input images to achieve a smooth transition between the multiview input images in the stitched output 314 image.
[0042] In one embodiment, the composition engine 302 includes a projection 312 module. In one embodiment, the projection 312 module is data values and software instructions that, when executed, project stitched image data from the composition space into a fixed bowl or bowl-view projection space, as described above. In one embodiment, the projection 312 module generates output 314 usable by a visualization system, as described above in conjunction with FIG. 1 . This output 314 includes data structures, such as image data and calibration parameters, usable to visualize a scene, e.g., the area around a vehicle in an automotive surround-view system. For example, in one embodiment, the output 314 includes data structures that can be consumed by an embedded visualization system or a non-embedded visualization block. In one embodiment, these output 314 data structures include 2D stitched images and / or textures. In another embodiment, the output 314 data structures include 3D stitched images and / or textures. The output 314 data structures, in one embodiment, are rectified camera images. In one embodiment, the output 314 data structure includes data representing a bowl mesh or texture mesh mapping.
[0043] 3, the composition engine pipeline 302 includes multiple software modules for performing image and / or video composition. In one embodiment, the composition engine 302 includes additional modules for facilitating image and / or video composition specific to the system in which the composition engine 302 is integrated or otherwise used. For example, in one embodiment, the composition engine 302 includes one or more reconstruction modules for reconstructing undercarriage texture data based on vehicle odometer data in an automotive system. In other systems, various software and / or hardware modules may be used to perform one or more additional steps to facilitate composition engine 302 operation.
[0044] 4 is a block diagram illustrating a Compositing 402 module for performing improved image composition according to some embodiments of the present disclosure. In one embodiment, the Compositing 402 module receives as input 404 two or more 2D images received from two or more cameras. In another embodiment, the Compositing 402 module receives as input 404 a 3D virtual camera image that includes combined or stitched image data resulting from a dewarp module performing an inverse mapping between the 2D image data from two or more cameras and a 3D virtual camera or projection space.
[0045] In one embodiment, the Compositing 402 module, when executed, generates the stitched image 406 by executing software instructions that transform two or more 2D input images into a 3D projection space, or composition space, using inverse mapping, where each pixel in the 3D projection space is fetched from one or more 2D input images. In another embodiment, the Dewarping module generates the stitched image, as described above in conjunction with FIG. 3, and the Compositing 402 module only performs composition quality optimization 408.
[0046] The Compositing 402 module includes a Compositing Quality Optimization 408 module. In one embodiment, the Compositing Quality Optimization 408 module is data values and software instructions that, when executed, calculate one or more quality scores based at least in part on the 3D stitched image generated 406 by the Compositing 402 module, as described further below in conjunction with FIG. 4. In another embodiment, the Compositing Quality Optimization 408 module calculates one or more quality scores based at least in part on the 3D stitched image generated by the Dewarp module and provided as input to the Compositing 402 module, as described above in conjunction with FIG. 3. In another embodiment, the Compositing Quality Optimization 408 module calculates one or more quality scores based on the 2D stitched image input 404 to or generated by the Compositing 402 module.
[0047] The composition quality optimization 408 module analyzes the 2D or 3D stitched image to generate one or more scores, and if any of the one or more scores are below a threshold, the composition quality optimization 408 module calculates and applies one or more transformations to the input 404 images used to generate the 2D or 3D stitched image. In one embodiment, the threshold used to analyze the scores calculated by the composition quality optimization module is a constant value. In another embodiment, the threshold used to analyze the scores calculated by the composition quality optimization module is variable. In one embodiment, the threshold is determined using one or more neural network operations to infer an optimal or acceptable threshold.
[0048] In one embodiment, the Compositing Quality Optimization 408 module provides feedback to any module responsible for generating the stitched image 406, such as those described above. In one embodiment, this feedback is one or more transformations to be applied to one or more 2D input 404 images. In another embodiment, the Compositing Quality Optimization 408 module feeds back one or more updated calibration parameters that can be used by two or more cameras in conjunction with two or more 2D input images.
[0049] In one embodiment, the Compositing Quality Optimization 408 module outputs a 3D virtual camera projection or stitched image that includes one or more transformations applied to one or more 2D input 404 images. In another embodiment, the Compositing Quality Optimization 408 module projects the 3D virtual camera projection or stitched image into 2D space using inverse mapping and outputs 416 the resulting 2D stitched image.
[0050] 5 is a block diagram illustrating a Composition Quality Optimization module according to some embodiments of the present disclosure. In one embodiment, the Composition Quality Optimization 502 module receives as input 504 a 3D stitched image generated 506 using a dewarp module or a compositing module, as described above in conjunction with FIGS. 3 and 4. In another embodiment, the Composition Quality Optimization 502 module receives as input 504 a 2D stitched image. The Composition Quality Optimization 502 module includes a Quality Assessment 508 block. In one embodiment, the Quality Assessment 508 block is data values and software instructions that, when executed, generate one or more scores as a result of one or more 2D or 3D input images.
[0051] The quality assessment 508 block, in one embodiment, generates a geometric quality score. As further described below in conjunction with FIG. 6, the geometric quality score is a data value that indicates how well two or more 2D input 504 images were stitched 506 into a 2D or 3D stitched image. That is, in one embodiment, the geometric quality score score geometric indicates whether the stitched image contains errors. The quality assessment 508 block determines a geometric quality score by applying one or more mathematical analyses of the 2D and / or 3D input 504 images in conjunction with the 2D or 3D stitched 506 image, as further described below in conjunction with FIG. geometric The quality assessment 508 block, in one embodiment, dynamically calculates a geometric quality score score based on the image content. geometric The image content, in one embodiment, includes objects detected by an object detector, a lane marking detector, a traffic light detector, or any other object detection essential to safety and visual quality. For example, when a mismatch occurs in a human-sensitive area, e.g., a vehicle, a pedestrian, a lane marking, or any other visual safety indicator, the quality assessment 508 block calculates a geometric quality score score geometric Lower.
[0052] The quality assessment 508 block, in one embodiment, includes a photometric quality score photometric As will be further described in conjunction with FIG. 6, a photometric quality score photometric is a data value that indicates how well each color space associated with each input image from two or more cameras matches in the 2D or 3D stitched 506 image. In one embodiment, the photometric quality score, photometric indicates whether the color and / or brightness and contrast levels between the input 504 images are significantly different, resulting in a distorted 2D or 3D stitched 506 image. In one embodiment, the Quality Assessment 508 block calculates a photometric score based on the color intensity score and the color angle score. photometric Given two cameras c1 and c2, the color intensity score is calculated as follows:
number
number
[0053] The composite quality optimization module 502 includes a quality score verification block 510. In one embodiment, the quality score verification 510 block is a data value and software instructions that, when executed, compares one or more quality scores output by the quality assessment 508 block to one or more threshold number data values. In another embodiment, the quality score verification 510 block compares one or more quality scores output by the quality assessment 508 block to one or more numerical values calculated by a classifier using one or more machine learning techniques, e.g., neural networks. In one embodiment, the one or more machine learning techniques learn a mapping between one or more objective composite quality scores and one or more subjective composite quality scores. In one embodiment, the objective quality score is a score that reflects the overall mismatch within a scene captured by one or more images. The subjective quality score is a score that reflects the mismatch between objects within a scene captured by one or more images. For example, in certain scenarios, human vision may be more focused on vehicles, pedestrians, lane markings, or other obstacles when driving and / or parking. In one embodiment, an automotive surround view system using a synthesis quality optimization module 502 to perform quality assessment 508 focuses more on inconsistencies in these focused areas of human vision and uses one or more machine learning techniques to map the objective quality scores to lower subjective quality scores. Various examples of objective and subjective scores are further described below in conjunction with FIG. 6.
[0054] The Quality Score Verification 510 block calculates the geometric quality score calculated by the Quality Assessment 508 block. geometric If it determines that the composite quality score is less than the threshold, then the composite quality optimization 502 performs geometry placement 512. In one embodiment, geometry placement 512, when performed, optimizes the composite quality score score geometricThe geometry 512 is data values and software instructions that align the unstitched input 504 images to improve alignment. In one embodiment, the technique for performing alignment is any technique for improving synthesis alignment, e.g., pixel / feature / depth / parallax-based alignment, improved depth estimation, or any other technique for performing alignment. In one embodiment, the geometry 512 calculates one or more transformations to adjust portions of the 2D input 504 or 3D stitched 506 images. In another embodiment, the geometry 512 calculates one or more transformations to adjust calibration parameters associated with the source camera c1 and the target camera c2. In another embodiment, the geometry 512 optimizes depth and / or parallax-based alignment to improve alignment. The quality of the depth / parallax-based alignment and depth estimation is measured using a synthesis quality score (score). geometric and the arrangement or transformation or method is evaluated by a composite quality score score geometric The parameters that are adjusted to optimize the
[0055] The composite quality optimization 502 performs the geometry 512 by computing the 3D optimal transformations for adjusting the extrinsic parameters {A2, [R2|T2]} of c2 without changing the extrinsic parameters {A1, [R1|T1]} of c1, where A is the intrinsic parameter and [R|T] is the extrinsic rotation and translation. In one embodiment, the composite quality optimization 502 in the geometry 512 computes the 3D optimal transformations for performing the geometry 512 by directly searching the neighborhood of the extrinsic space [R2|T2] of c2. That is, the composite quality optimization 502 maximizes the score geometric When applied to the extrinsic parameters of c2, the resulting rotation and translation adjustment [R optimal |T optimal ] * For example, the 3D optimal transformation is calculated as follows: [R optimal |T optimal ] * =argmax score geometric (Search the neighborhood of [R2,T2]) Here, searching the neighborhood of [R2, T2] is an adjustment that results in c2 being rotated and translated around c1. The synthesis quality optimization 502 in the geometry arrangement 512 allows for different scores. geometric Calculating the individual transformations that result in a value, and the maximum score geometric The 3D optimal transformation [R optimal |T optimal ] * Determine.
[0056] In another embodiment, the synthesis quality optimization 502 during geometry 512 calculates the 3D optimal transformation to perform geometry 512 by reducing the search space via optical flow vectors. The relationship between the 3D synthesis space coordinate M and the projected 2D image space coordinate m is defined as follows: m=sA([RT]M) where s is the depth scale factor. The geometry 512 block estimates optical flow vectors between the 2D input 504 images from source camera c1 and the 2D input 504 images from destination camera c2 using any estimation method, such as phase correlation, block-based methods, differential methods such as Lucas-Kanade or Horn-Schunck, individual optimization methods, or any other optical flow estimation technique. After the geometry 512 block estimates optical flow vectors between the 2D input 504 images from source camera c1 and the 2D input 504 images from destination camera c2 in the region where the 2D input 504 images from source camera c1 and the 2D input 504 images from destination camera c2 overlap, the 3D transformation [R0|T0] of the destination 2D image from c2 is estimated using the relationship between M and m, as described above. The 3D deformation [R0|T0] is the optimal score based on the optical flow vector estimation error. geometricAs a result, the geometry 512 block applies a confidence value to the optical flow vectors and evaluates a neighborhood of the 3D transformation [R0|T0], where the neighborhood of the 3D transformation [R0|T0] represents a reduced search neighborhood of [R2|T2]. The use of the estimated optical flow vectors by the geometry 512 block of the synthesis quality optimization 502 results in a guided search direction and a reduced search space. In one embodiment, the 3D optimal transformation is then calculated as follows: [R optimal |T optimal ] * =argmax score geometric (Search the neighborhood of [R0,T0]). In one embodiment, we score temporal consistency during optimization. geometric The 3D optimal transformation that we integrate is a frame-time 3D volume optimization rather than an individual frame optimization.
[0057] The Quality Score Verification 510 block returns the photometric quality score calculated by the Quality Assessment 508 block. photometric If it determines that the photometric quality score is less than the threshold, then the composition quality optimization 502 performs a photometric adjustment 514. In one embodiment, the photometric adjustment 514 adjusts the photometric quality score score photometric These are data values and software instructions that, when executed, adjust color values between unstitched input 504 images to improve color matching.
[0058] In one embodiment, the photometric adjustment 514 attempts to match color intensities and color angles, as described above, between the 2D input 504 image from source camera c1 and the 2D input 504 image from target camera c2 in the overlapping regions of the 2D input 504 image from source camera c1 and the 2D input 504 image from target camera c2. The photometric adjustment 514 performs optimal transformations on the individual R, G, and B channels of the 2D image from target camera c2, such as: scale * optimal offset * ) to calculate: Target image = optimal scale *Source image +optimal offset
[0059] In one embodiment, the photometry adjustment 514 directly retrieves the offset and scale values for various values of scale and offset as follows: (optimal scale * optimal offset * )=argmax score photometric (scale,offset). The photometric quality score SCO is maximized when the color values of the 2D input 504 image from the source camera c1 are as close as possible to the color values of the 2D input 504 image from the destination camera c2.
[0060] In another embodiment, photometric adjustment 514 reduces the search space and improves performance by calculating an initial scale and an initial offset to begin the search using the equations above. The initial scale is calculated during photometric adjustment 514 as follows:
number
[0061] In one embodiment, if the quality score verification 510 block within the composition quality optimization 502 module determines that all quality scores, as described above and further below in conjunction with FIG. 6, exceed a threshold, the composition quality optimization 502 module outputs 516 an optimized 3D stitched image to the projection block, as described above in conjunction with FIG. 3. In another embodiment, the quality score verification 510 block determines that one or more quality scores, e.g., score geometric or score photometric , or score geometric and score photometricIf it determines that both of the thresholds are exceeded, the Compositing Quality Optimization 502 module outputs 516 the optimized 3D stitched image to the Projection block.
[0062] 6 is a block diagram illustrating a quality assessment 602 block according to some embodiments of the present disclosure. The quality assessment 602 block calculates one or more quality scores 614, 618 associated with the input stitched image 606, as described above in conjunction with FIG. 5. In one embodiment, the quality assessment 602 block calculates the one or more quality scores 614, 618 based at least in part on a comparison between the left image 604 and the stitched image 606 or the right image 608 and the stitched image 606. In another embodiment, the quality assessment 602 block calculates the one or more quality scores 614, 618 based at least in part on a comparison between the left image 604 and the right image 608. In one embodiment, the quality assessment 602 block calculates the one or more quality scores 614, 618 using the left image 604, the right image 608, and the stitched image 606. In one embodiment, the left image 604 is captured by a left-facing camera c, as described further herein. L In one embodiment, the right image 608 is data representing a 2D input image from a right-facing camera c R is data representing a 2D input image from
[0063] The quality assessment 602 block calculates the overlap 610 between two or more input images 604, 608 using the stitched image 606. In one embodiment, the overlap 610 is data containing information about areas where the left image 604 and the right image 608 share a coordinate space in the stitched image 606. In one embodiment, the coordinate space is a 3D projection space or a virtual camera space. In another embodiment, the coordinate space is a 2D projection space. In one embodiment, the coordinate space includes overlap 610 information from both the left image 604 and the right image 608. In another embodiment, the coordinate space includes overlap 610 information from either the left image 604 or the right image 608.
[0064] In one embodiment, the quality assessment 602 block includes a geometric score 614, or a geometric quality score as described above in conjunction with FIG. geometric The high-pass 612 operation is performed to calculate the geometric score 614. In one embodiment, the high-pass 612 operation is data values and software instructions that, when executed, calculate the geometric score 614 using one or more mathematical analysis techniques, including objective quality metrics and / or subjective quality metrics. In one embodiment, the high-pass 612 operation determines the objective quality metric and calculates the geometric score 614 by averaging the high-frequency structural similarity index measure (SSIM) between the left image 604 or right image 608 and the stitched image 606. The SSIM is a perceptual-based model that considers image degradation as a perceived change in structural information while also incorporating terms for important perceptual phenomena, such as luminance masking and contrast masking. In another embodiment, the high-pass 612 operation determines the objective quality metric and calculates the geometric score 614 by generating a perceptual quality significance map (PQSM) between the left image 604 or right image 608 and the stitched image 606. The PQSM is an array whose elements represent the relative perceptual quality importance levels of corresponding areas and / or regions between images. In one embodiment, the high-pass 612 operation calculates the geometric score 614 using root mean square error (RMSE), peak signal to noise ratio (PSNR), or any other objective quality metric that can be used to quantify the difference between the left image 604 or right image 608 and the stitched image 606. In another embodiment, the high-pass 612 operation calculates the geometric score 614 using any subjective quality metric, such as mean opinion score.
[0065] In one embodiment, the quality assessment 602 block includes a photometric score 618, or a photometric quality score, as described above in conjunction with FIG. 5. photometricThe low pass 616 operation is performed to calculate a photometric score 618. In one embodiment, the low pass 616 operation is data values and software instructions that, when executed, calculate a photometric score 618. In one embodiment, the low pass 616 operation calculates the photometric score 618 as described above in conjunction with FIG. 5 . In another embodiment, the low pass 616 operation calculates the photometric score 618 based at least in part on a spectral angle mapper (SAM) to determine photometric color quality and / or an intensity magnitude ratio (IMR) to determine photometric intensity quality. In one embodiment, the low pass 616 operation calculates the photometric score 618 using any other technique that can be used to quantify the color difference between the left image 604 or right image 608 and the stitched image 606. In one embodiment, the low pass 616 operation, as described further herein, calculates a photometric score 618 based on a color space, such as RGB, CIELAB, CMYK, YIQ, YCbCr, YUV, HSV, HSL, or any other color space usable by the left image 604 and the right image 608 captured by two or more cameras.
[0066] Referring now to FIG. 7 , each block of method 700 described herein includes computational processes that may be implemented using any combination of hardware, firmware, and / or software. For example, various functions may be implemented by a processor executing instructions stored in a memory. The method may also be implemented as computer-usable instructions stored on a computer storage medium. The method may be provided by a standalone application, a service, or a hosted service (standalone or in combination with another hosted service), or a plug-in to another product, to name a few. Additionally, method 700 is described with respect to the video and / or image composition pipeline of FIG. 1 as performed by a surround view system in an automotive application, by way of example. However, these methods may additionally or alternatively be performed by any one system or any combination of systems, including, but not limited to, those described herein.
[0067] 7 is a flow diagram illustrating a method 700 for performing improved image compositing using compositing quality optimization, according to some embodiments of the present disclosure. In method 700, at block 702, a compositing module or compositing pipeline receives two or more images from two or more cameras, as well as calibration parameters associated with the two or more cameras used to generate the two or more images, as described above in conjunction with FIG.
[0068] In block 704, the compositing module or compositing pipeline transforms each input image into a composite space, as described above in conjunction with FIG. 3, and generates an initial stitched image in block 706. In one embodiment, blocks 704 and 706 are combined, and a stitched image is generated 706 as a result of transforming two or more images into composite space 704. In another embodiment, blocks 704 and 706 are performed independently, and a stitched image is generated 706 by the compositing module or compositing pipeline after two or more 2D input images are transformed into composite space 704. In one embodiment, the compositing module or compositing pipeline transforms two or more 2D input images into a single 3D virtual camera or stitch space using forward or inverse mapping. In another embodiment, the compositing module or compositing pipeline transforms two or more 2D input images into a single 2D image.
[0069] In block 708, the quality assessment block of the compositing module calculates geometric scores using various techniques described above in conjunction with FIG. 6. In block 710, the quality assessment block of the compositing module calculates photometric scores using various techniques described above in conjunction with FIGs. 5 and 6. Based on the geometric and / or photometric scores calculated in blocks 708 and 710, the quality score verification block determines whether any of the geometric and / or photometric scores are below one or more thresholds. If neither the geometric nor photometric scores are below one or more thresholds 712, the compositing module or compositing pipeline transforms the resulting stitched image to projection space in block 718, as described above in conjunction with FIG. 3.
[0070] If the geometric score is less than the threshold 712, the compositing module reduces the geometric distortion in block 714 by calculating an optimal 3D geometric transformation, as described above in conjunction with Figure 5. If the photometric score is less than the threshold 712, the compositing module reduces the photometric distortion in block 716 by applying one or more optimal photometric transformations, as described above in conjunction with Figure 5.
[0071] Exemplary Autonomous Vehicle 8 is a diagram of an example autonomous vehicle 800 according to some embodiments of the present disclosure. Autonomous vehicle 800 (alternatively referred to herein as “vehicle 800”) may include, but is not limited to, a passenger vehicle, such as a car, a truck, a bus, a first responder vehicle, a shuttle, an electric or moped, a motorcycle, a fire engine, a police vehicle, an ambulance, a boat, a construction vehicle, a submarine, a drone, and / or another type of vehicle (e.g., unmanned and / or carrying one or more passengers). Autonomous vehicles are generally described in terms of levels of automation as defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201806, published June 15, 2018; Standard No. J3016-201609, published September 30, 2016; and previous and future versions of this standard). Mobile vehicle 800 may be capable of functionality according to one or more of levels 3 through 5 of autonomous driving. For example, mobile vehicle 800 may be capable of conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on the embodiment.
[0072] Mobile vehicle 800 may include components such as a vehicle chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components. Mobile vehicle 800 may include a propulsion system 850, such as an internal combustion engine, a hybrid power plant, a fully electric engine, and / or another propulsion system type. Propulsion system 850 may be connected to a drive train of mobile vehicle 800, which may include a transmission, to enable propulsion of mobile vehicle 800. Propulsion system 850 may be controlled in response to receiving a signal from a throttle / accelerator 852.
[0073] A steering system 854, which may include a steering wheel, may be used to steer the vehicle 800 (e.g., along a desired course or route) when the propulsion system 850 is operating (e.g., when the vehicle is moving). The steering system 854 may receive signals from a steering actuator 856. A steering wheel may be optional for fully automated (Level 5) functionality.
[0074] Brake sensor system 846 may be used to operate vehicle brakes in response to receiving signals from brake actuators 848 and / or brake sensors.
[0075] A controller 836, which may include one or more CPUs, system on chip (SoC) 804 (FIG. 8C), and / or GPUs, can provide signals (e.g., representations of commands) to one or more components and / or systems of the vehicle 800. For example, the controller can send signals to operate vehicle brakes via one or more brake actuators 848, to operate a steering system 854 via one or more steering actuators 856, and / or to operate a propulsion system 850 via one or more throttle / accelerators 852. The controller 836 may include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operational commands (e.g., signals representing commands) to enable rhythmic driving and / or assist a driver in operating the vehicle 800. The controllers 836 may include a first controller 836 for autonomous driving functions, a second controller 836 for functional safety functions, a third controller 836 for artificial intelligence functions (e.g., computer vision), a fourth controller 836 for infotainment functions, a fifth controller 836 for redundancy in emergency situations, and / or other controllers. In some instances, a single controller 836 may handle two or more of the foregoing functions, and two or more controllers 836 may handle a single function and / or any combination thereof.
[0076] Controller 836 may provide signals to control one or more components and / or systems of vehicle 800 in response to sensor data (e.g., sensor inputs) received from one or more sensors. Sensor data may be received from, for example, and without limitation, global navigation satellite system sensors 858 (e.g., global positioning system sensors), RADAR sensors 860, ultrasonic sensors 862, LIDAR sensors 864, inertial measurement unit (IMU) sensors 866 (e.g., accelerometers, gyroscopes, magnetic compasses, magnetometers, etc.), microphones 896, stereo cameras 868, wide-view cameras 870 (e.g., fisheye cameras), infrared cameras 872, surround cameras 874 (e.g., 360-degree cameras), long-range and / or medium-range cameras 898, speed sensors 844 (e.g., for measuring the speed of the moving vehicle 800), vibration sensors 842, steering sensors 840, brake sensors (e.g., as part of a brake sensor system 846), and / or other sensor types.
[0077] One or more of the controllers 836 may receive input (e.g., represented by input data) from the instrument cluster 832 of the vehicle 800 and provide output (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 834, an audible annunciator, a loudspeaker, and / or other components of the vehicle 800. The output may include information such as vehicle velocity, speed, time, map data (e.g., HD map 822 of FIG. 8C ), position data (e.g., the location of the vehicle 800 on a map, etc.), direction, the locations of other vehicles (e.g., an occupancy grid), information about objects and object situations as known by the controller 836, etc. For example, the HMI display 834 may display information regarding the presence of one or more objects (e.g., road signs, warning signs, traffic light changes, etc.) and / or a driving maneuver that the moving vehicle has performed, is performing, or will perform (e.g., changing lanes now, taking exit 34B in 3.22 km (2 miles), etc.).
[0078] The mobile vehicle 800 further includes a network interface 824 that can communicate over one or more networks using one or more wireless antennas 826 and / or a modem. For example, the network interface 824 can be capable of communication over LTE, WCDMA, UMTS, GSM, CDMA2000, etc. The wireless antenna 826 can also enable communication between objects in the environment (e.g., mobile vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth LE, Z-Wave, ZigBee, etc., and / or low power wide-area networks (LPWANs) such as LoRaWAN, SigFox, etc.
[0079] 8B is an illustration of camera positions and fields of view of the exemplary autonomous vehicle 800 of FIG. 8A, according to some embodiments of the present disclosure. The cameras and their respective fields of view are one illustrative example and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be located at different positions on the vehicle 800.
[0080] The camera type may include, but is not limited to, a digital camera adapted for use with components and / or systems of the mobile vehicle 800. The camera may be capable of operating at Automotive Safety Integrity Level (ASIL) B and / or at another ASIL. The camera type may be capable of any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The camera may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some instances, the color filter array may include a red clear clear clear (RCCC) color filter array, a red clear clear blue (RCCB) color filter array, a red blue green clear (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras, such as cameras with RCCC, RCCB, and / or RBGC color filter arrays, may be used in an effort to increase light sensitivity.
[0081] In some instances, one or more of the cameras may be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-function mono camera may be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlamp control. One or more of the cameras (e.g., all cameras) may simultaneously record and provide image data (e.g., video).
[0082] One or more of the cameras may be mounted in a mounting part, such as a custom-designed (e.g., 3D printed) part, to filter out stray light and reflections from within the vehicle (e.g., reflections from the dashboard reflected in the windshield mirror) that may interfere with the camera's image data capture ability. Referring to a side mirror mounting part, the side mirror part may be custom 3D printed so that the camera mounting plate fits the shape of the side mirror. In some instances, the camera may be integrated into the side mirror. For side view cameras, the camera may also be integrated into four posts at each corner of the cabin.
[0083] A camera (e.g., a forward-facing camera) with a field of view that includes a portion of the environment in front of the vehicle 800 may be used for surround view to aid in identifying a forward path and obstacles and, with the assistance of one or more controllers 836 and / or control SoCs, to provide information essential for generating an occupancy grid and / or determining a preferred vehicle path. Forward-facing cameras may be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. Forward-facing cameras may also be used for ADAS functions and systems, including other functions such as lane departure warning (LDW), autonomous cruise control (ACC), and / or traffic sign recognition.
[0084] Various cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform including a complementary metal oxide semiconductor (CMOS) color imager. Another example can be a wide-view camera 870 that can be used to understand objects that come into view from the periphery (e.g., pedestrians, crossing traffic, or bicycles). While only one wide-view camera is shown in FIG. 8B, any number of wide-view cameras 870 can be present in the vehicle 800. Additionally, a long-range camera 898 (e.g., a long-view stereo camera pair) can be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. The long-range camera 898 can also be used for object detection and classification, as well as basic object tracking.
[0085] One or more stereo cameras 868 may also be included in the forward-facing configuration. The stereo camera 868 may include an integrated control unit with an extensible processing unit, which may provide programmable logic (e.g., FPGA) and a multi-core microprocessor with a CAN or Ethernet interface integrated on a single chip. Such a unit may be used to generate a 3D map of the vehicle's environment, including distance estimates for all points in the image. An alternative stereo camera 868 may include a compact stereo vision sensor, which may include two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from the vehicle to objects of interest and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 868 may be used in addition to or instead of those described herein.
[0086] Cameras having a field of view that includes portions of the environment to the sides of the mobile vehicle 800 (e.g., side-view cameras) may be used for surround view, providing information used to create and update the occupancy grid and generate side-impact collision warnings. For example, surround cameras 874 (e.g., four surround cameras 874 as shown in FIG. 8B ) may be positioned around the mobile vehicle 800. The surround cameras 874 may include wide-view cameras 870, fisheye cameras, 360-degree cameras, and / or the like. For example, four fisheye cameras may be positioned at the front, rear, and sides of the mobile vehicle. In an alternative arrangement, the mobile vehicle may use three surround cameras 874 (e.g., left, right, and rear) and utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround view camera.
[0087] A camera having a field of view that includes the portion of the environment behind the moving vehicle 800 (e.g., a rearview camera) may be used for parking assistance, surround view, rear collision warning, and creating and updating an occupancy grid. As described herein, a wide variety of cameras may be used, including, but not limited to, cameras that are also suitable as forward-facing cameras (e.g., long-range and / or medium-range camera 898, stereo camera 868, infrared camera 872, etc.).
[0088] FIG. 8C is a block diagram of an example system architecture for the example autonomous vehicle 800 of FIG. 8A , in accordance with some embodiments of the present disclosure. It should be understood that this and other arrangements described herein are merely illustrative. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted altogether. Furthermore, many of the elements described herein are functional entities that may be implemented as separate or distributed components or in conjunction with other components, and in any suitable combination and location. Various functions described herein as being performed by entities may be implemented by hardware, firmware, and / or software. For example, various functions may be implemented by a processor executing instructions stored in a memory.
[0089] Each of the components, features, and systems of the mobile vehicle 800 in FIG. 8C is shown connected via a bus 802. The bus 802 may include a controller area network (CAN) data interface (alternatively referred to as a "CAN bus"). The CAN may be a network within the mobile vehicle 800 used to help control various features and functions of the mobile vehicle 800, such as braking, acceleration, braking, steering, windshield wiper operation, etc. The CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus may be read to determine steering angle, ground speed, engine revolutions per minute (RPM), button position, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.
[0090] Although the bus 802 is described herein as being a CAN bus, this is not intended to be limiting. For example, FlexRay and / or Ethernet may be used in addition to or as an alternative to a CAN bus. Additionally, although a single line is used to represent the bus 802, this is not intended to be limiting. There may be any number of buses 802, which may include, for example, one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some instances, two or more buses 802 may be used to perform different functions and / or for redundancy. For example, a first bus 802 may be used for collision avoidance functions, and a second bus 802 may be used for actuation control. In any instance, each bus 802 may communicate with any of the components of the vehicle 800, and two or more buses 802 may communicate with the same component. In some instances, each SoC 804, each controller 836, and / or each computer in the vehicle may have access to the same input data (e.g., input from sensors in the vehicle 800) and may be connected to a common bus, such as a CAN bus.
[0091] Mobile vehicle 800 may include one or more controllers 836, such as those described herein with respect to FIG. 8A. Controller 836 may be used for a variety of functions. Controller 836 may be coupled to any of various other components and systems of mobile vehicle 800 and may be used for control of mobile vehicle 800, artificial intelligence of mobile vehicle 800, infotainment for mobile vehicle 800, and / or the like.
[0092] Mobile vehicle 800 may include a system-on-chip (SoC) 804. SoC 804 may include a CPU 806, a GPU 808, a processor 810, a cache 812, an accelerator 814, a data store 816, and / or other components and features not shown. SoC 804 may be used to control mobile vehicle 800 in a variety of platforms and systems. For example, SoC 804 may be coupled in a system (e.g., that of mobile vehicle 800) with an HD map 822 that can obtain map refreshes and / or updates via a network interface 824 from one or more servers (e.g., server 878 of FIG. 8D ).
[0093] The CPU 806 may include a CPU cluster or CPU complex (alternatively referred to as a "CCPLEX"). The CPU 806 may include multiple cores and / or L2 caches. For example, in some embodiments, the CPU 806 may include eight cores in a coherent multiprocessor configuration. In some embodiments, the CPU 806 may include four dual-core clusters, each with its own dedicated L2 cache (e.g., a 2M L2 cache). The CPU 806 (e.g., a CCPLEX) may be configured to support simultaneous cluster operation, allowing any combination of clusters of CPUs 806 to be active at any given time.
[0094] The CPU 806 may implement power management capabilities including one or more of the following features: individual hardware blocks may be automatically clock gated when idle to conserve dynamic power; each core clock may be gated when the core is not actively executing instructions by executing a WFI / WFE instruction; each core may be independently power gated; each core cluster may be independently clock gated when all cores are clock gated or power gated; and / or each core cluster may be independently power gated when all cores are power gated. The CPU 806 may further implement an enhanced algorithm for managing power states, where allowable power states and expected wake-up times are specified and hardware / microcode determines the best power state for entering the cores, clusters, and CCPLEX. The processing cores may support simplified power state entry sequences in software with work offloaded to microcode.
[0095] The GPU 808 may include an integrated GPU (alternatively referred to herein as an "iGPU"). The GPU 808 may be programmable and efficient for parallel workloads. In some instances, the GPU 808 may use an enhanced tensor instruction set. The GPU 808 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache having at least 96 KB of storage capacity) and two or more of the streaming microprocessors may share a cache (e.g., an L2 cache having 512 KB of storage capacity). In some embodiments, the GPU 808 may include at least eight streaming microprocessors. The GPU 808 may use a computer-based application programming interface (API). Additionally, the GPU 808 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0096] The GPU 808 may be power-optimized for best performance in automotive and embedded use cases. For example, the GPU 808 may be fabricated on FinFET (Fin field-effect transistor) chips. However, this is not intended to be limiting, and the GPU 808 may be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor may incorporate several mixed-precision processing cores partitioned into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores may be partitioned into four processing blocks. In such an example, each processing block may be assigned 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA tensor cores for deep learning matrix operations, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Additionally, the streaming microprocessor may include independent parallel integer and floating-point data paths to provide efficient execution of workloads with a mix of computational and addressing operations. Streaming microprocessors may include independent thread scheduling capabilities to allow finer-grained synchronization and coordination among concurrent threads. Streaming microprocessors may include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.
[0097] The GPU 808 may, in some instances, include a high bandwidth memory (HBM) and / or 16 GB HBM2 memory subsystem to provide up to 900 GB / s of peak memory bandwidth. In some instances, synchronous graphics random-access memory (SGRAM), such as graphics double data rate type five synchronous random-access memory (GDDR5), may be used in addition to or in place of the HBM memory.
[0098] The GPU 808 may include unified memory technology that includes access counters to enable more accurate movement of memory pages to the processors that access them most frequently, thereby improving the efficiency of storage areas shared between processors. In some instances, address translation service (ATS) support may be used to enable the GPU 808 to directly access the CPU 806 page tables. In such instances, when the GPU 808 memory management unit (MMU) experiences a miss, an address translation request may be sent to the CPU 806. In response, the CPU 806 may consult its page table for a virtual-to-real mapping of addresses and send the translation back to the GPU 808. As such, unified memory technology may enable a single unified virtual address space for both CPU 806 and GPU 808 memory, thereby simplifying GPU 808 programming and porting of applications to the GPU 808.
[0099] Additionally, GPU 808 may include access counters that can record the frequency of GPU 808's accesses to the memory of other processors. The access counters can help ensure that memory pages are moved to the physical memory of the processors that are accessing the pages most frequently.
[0100] The SoC 804 may include any number of caches 812, including those described herein. For example, the cache 812 may include an L3 cache available to both the CPU 806 and the GPU 808 (e.g., connected to both the CPU 806 and the GPU 808). The cache 812 may include a write-back cache that can record line state, such as by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache may include 4 MB or more, depending on the implementation, although smaller cache sizes may also be used.
[0101] SoC 804 may include an arithmetic logic unit (ALU) that may be utilized in performing processing for any of various tasks or operations (e.g., processing DNNs) of vehicle 800. In addition, SoC 804 may include a floating point unit (FPU) (or other math coprocessor or math coprocessor type) for performing mathematical operations within the system. For example, SoC 804 may include one or more FPUs integrated as execution units within CPU 806 and / or GPU 808.
[0102] The SoC 804 may include one or more accelerators 814 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the SoC 804 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. The large on-chip memory (e.g., 4 MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other operations. The hardware acceleration cluster may be used to complement the GPU 808 and to offload some of the GPU 808's tasks (e.g., to free up more cycles for the GPU 808 to perform other tasks). As an example, the accelerator 814 may be used for target workloads that are sufficiently stable to be suitable for acceleration (e.g., perception, convolutional neural networks (CNNs), etc.). As used herein, the term "CNN" may include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and Faster RCNNs (e.g., as used for object detection).
[0103] The accelerator 814 (e.g., a hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA may include one or more tensor processing units (TPUs), which can be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured and optimized to perform image processing functions (e.g., CNN, RCNN, etc.). The DLA may also be optimized for a specific set of neural network types and floating-point operations, as well as inference. The DLA design can provide more performance per millimeter than a general-purpose GPU, significantly exceeding the performance of a CPU. The TPU can perform several functions, including, for example, single-instance convolution functions, supporting INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.
[0104] The DLA can quickly and efficiently run neural networks, particularly CNNs, on processed or unprocessed data for any of a variety of functions, including, but not limited to: CNNs for object identification and detection using data from camera sensors, CNNs for distance estimation using data from camera sensors, CNNs for emergency vehicle detection and identification using data from microphones, CNNs for face recognition and moving vehicle owner identification using data from camera sensors, and / or CNNs for security and / or safety related events.
[0105] The DLA can perform any function of the GPU 808, and by using an inference accelerator, for example, a designer can target either the DLA or the GPU 808 for any function. For example, a designer can focus on processing CNNs and floating-point operations on the DLA, and offload other functions to the GPU 808 and / or other accelerators 814.
[0106] The accelerator 814 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA may provide a balance between performance and flexibility. For example, each PVA may include, but is not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.
[0107] The RISC cores may interact with an image sensor (e.g., an image sensor in any of the cameras described herein), an image signal processor, and / or the like. Each RISC core may include any amount of memory. The RISC cores may use any of several protocols, depending on the embodiment. In some instances, the RISC cores may execute a real-time operating system (RTOS). The RISC cores may be implemented using one or more integrated circuit devices, application specific integrated circuits (ASICs), and / or memory devices. For example, the RISC cores may include an instruction cache and / or tightly coupled RAM.
[0108] The DMA may allow components of the PVA to access system memory independent of the CPU 806. The DMA may support any number of features used to provide optimizations to the PVA, including, but not limited to, supporting multi-dimensional addressing and / or circular addressing. In some instances, the DMA may support up to six or more dimensions of addressing, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.
[0109] A vector processor may be a programmable processor that can be designed to efficiently and flexibly execute computer vision algorithm programming and provide signal processing capabilities. In some instances, a PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may act as the PVA's primary processing engine and may include a vector processing unit (VPU), an instruction cache, and / or a vector memory (e.g., VMEM). The VPU core may include a digital signal processor, such as a single instruction, multiple data (SIMD), or very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can increase throughput and speed.
[0110] Each vector processor may include an instruction cache and may be coupled to dedicated memory. As a result, in some instances, each vector processor may be configured to execute independently of other vector processors. In other instances, the vector processors included in a particular PVA may be configured to employ data parallelism. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other instances, the vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms on sequential images or portions of an image. In particular, any number of PVAs may be included in a hardware-accelerated cluster, and any number of vector processors may be included in each PVA. Additionally, the PVA may include additional error correcting code (ECC) memory to enhance overall system security.
[0111] The accelerator 814 (e.g., a hardware acceleration cluster) may include a computer vision network-on-chip and SRAM to provide high-bandwidth, low-latency SRAM for the accelerator 814. In some instances, the on-chip memory may include, for example, and without limitation, at least 4 MB of SRAM consisting of eight field-configurable memory blocks that may be accessible by both the PVA and DLA. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and DLA can access the memory through a backbone that provides the PVA and DLA with high-speed access to the memory. The backbone may include a computer vision network-on-chip that interconnects the PVA and DLA to the memory (e.g., using the APB).
[0112] The computer vision network-on-chip may include an interface that determines, prior to the transmission of any control signals, addresses, or data, that both the PVA and DLA provide ready and valid signals. Such an interface may provide separate phases and separate channels for transmitting control signals, addresses, and data, as well as burst-type communication for continuous data transfer. This type of interface may conform to the ISO 26262 or IEC 61508 standards, although other standards and protocols may also be used.
[0113] In some instances, SoC 804 may include a real-time ray tracing hardware accelerator, such as that described in U.S. Patent Application No. 16 / 101,232, filed August 10, 2018. The real-time ray tracing hardware accelerator may be used to quickly and efficiently determine the location and scale of objects (e.g., within a world model) to generate real-time visualization simulations for RADAR signal interpretation, for acoustic propagation synthesis and / or analysis, for SONAR system simulation, for general wave propagation simulation, for comparison to LIDAR data for localization and / or other functions, and / or other uses. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing-related operations.
[0114] The accelerator 814 (e.g., a hardware accelerator cluster) has diverse applications for autonomous driving. The PVA may be a programmable vision accelerator that can be used for critical processing stages in ADAS and autonomous vehicles. The capabilities of the PVA make it well suited to algorithmic domains that require predictable processing at low power and low latency. In other words, the PVA works well for semi-dense or dense regular computations on small data sets that require predictable execution times along with low latency and low power. Therefore, because the PVA is efficient at object detection and operating on integer computations, in the context of a platform for autonomous vehicles, the PVA is designed to run classic computer vision algorithms.
[0115] For example, according to one embodiment of the present technology, PVA is used to perform computer stereo vision. A semi-global matching-based algorithm may be used in some instances, but this is not intended to be limiting. Many applications for Level 3-5 autonomous driving require motion estimation / stereo matching on the fly (e.g., structure from motion, pedestrian recognition, lane detection, etc.). PVA can perform computer stereo vision functions with input from two monocular cameras.
[0116] In some instances, PVAs may be used to perform dense optical flow. For example, PVAs may be used to process raw RADAR data (e.g., using a 4D Fast Fourier Transform) to provide a processed RADAR signal before emitting the next RADAR pulse. In other instances, PVAs are used for time of flight depth processing, for example, by processing raw time of flight data to provide processed time of flight data.
[0117] DLA can be used to implement any type of network to enhance control and driving safety, including, for example, a neural network that outputs a confidence measure for each object detection. Such a confidence value can be interpreted as a probability or as providing the relative "weight" of each detection compared to other detections. This confidence value allows the system to make further decisions regarding which detections should be considered true positives rather than false positives. For example, the system can set a confidence threshold and consider only detections above the threshold as true positives. In an automatic emergency braking (AEB) system, a false positive detection would cause a moving vehicle to automatically apply emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered to trigger AEB. DLA can implement a neural network that regresses the confidence value. The neural network may receive as its input at least some subset of parameters, such as bounding box dimensions, a ground plane estimate obtained (e.g., from another subsystem), an inertial measurement unit (IMU) sensor 866 output that correlates with the vehicle 800 orientation, range, and 3D position estimate of the object obtained from the neural network and / or other sensors (e.g., a LIDAR sensor 864 or a RADAR sensor 860), and others.
[0118] The SoC 804 may include a data store 816 (e.g., memory). The data store 816 may be on-chip memory of the SoC 804 and may store neural networks to be executed by the GPU and / or DLA. In some instances, the data store 816 may have a capacity large enough to store multiple instances of the neural network for redundancy and safety. The data store 816 may comprise an L2 or L3 cache 812. References to the data store 816 may include references to memory associated with the GPU, DLA, and / or other accelerators 814, as described herein.
[0119] The SoC 804 may include one or more processors 810 (e.g., embedded processors). The processors 810 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management capabilities and related security enforcement. The boot and power management processor may be part of the SoC 804 boot sequence and may provide run-time power management services. The boot power and management processor may provide clock and voltage programming, assist with system low-power state transitions, manage the SoC 804 thermal and temperature sensors, and / or manage the SoC 804 power state. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 804 may use the ring oscillator to detect the temperature of the CPU 806, GPU 808, and / or accelerator 814. If the temperature is determined to exceed a threshold, the boot and power management processor may enter a temperature fault routine, place the SoC 804 in a lower power state, and / or place the vehicle 800 in a Chauffeur safe stop mode (e.g., bring the vehicle 800 to a safe stop).
[0120] The processor 810 may further include a set of embedded processors that can perform the functions of an audio processing engine. The audio processing engine may be an audio subsystem that allows full hardware support for multi-channel audio through multiple interfaces and a wide and flexible range of audio I / O interfaces. In some instances, the audio processing engine is a dedicated processor core that includes a digital signal processor with dedicated RAM.
[0121] The processor 810 may further include an always-on processor engine that can provide the necessary hardware features to support low-power sensor management and wake use cases. The always-on processor engine may include a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0122] The processor 810 may further include a safety cluster engine that includes a processor subsystem dedicated to handling safety management for automotive applications. The safety cluster engine may include two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In safety mode, the two or more cores may operate in lockstep mode and function as a single core with comparison logic to detect any differences between their operations.
[0123] The processor 810 may further include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.
[0124] The processor 810 may further include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.
[0125] The processor 810 may include a video image compositor, which may be a processing block (e.g., implemented in a microprocessor) that implements video post-processing functions required by the video playback application to produce the final image for the player window. The video image compositor may perform lens distortion correction on the wide-view camera 870, the surround camera 874, and / or the in-cabin surveillance camera sensor. The in-cabin surveillance camera sensor is preferably monitored by a neural network running on a separate instance of the advanced SoC, configured to identify in-cabin events and respond appropriately. The in-cabin system may perform lip reading to activate cellular service and make phone calls, dictate emails, change the vehicle's destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Certain features are available to the driver only when operating in autonomous mode and are disabled otherwise.
[0126] The video image combiner may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, when motion occurs in the video, the noise reduction reduces the weight of information provided by adjacent frames and appropriately weights spatial information. When an image or portion of an image does not contain motion, the temporal noise reduction performed by the video image combiner can use information from previous images to reduce noise in the current image.
[0127] The video image compositor may also be configured to perform stereo rectification on the input stereo lens frames. The video image compositor may further be used for user interface compositing when the operating system desktop is in use, and the GPU 808 is not required to continuously render new surfaces. Even when the GPU 808 is powered on and actively performing 3D rendering, the video image compositor may be used to offload the GPU 808 to improve performance and responsiveness.
[0128] The SoC 804 may further include a mobile industry processor interface (MIPI) camera serial interface, a high-speed interface for receiving video and input from a camera, and / or a video input block that may be used for camera and related pixel input functions. The SoC 804 may further include an input / output controller that may be controlled by software and that may be used to receive I / O signals that are not committed to a specific role.
[0129] The SoC 804 may further include a wide range of peripheral interfaces to enable communication with peripherals, audio codecs, power management, and / or other devices. The SoC 804 may be used to process data from cameras (e.g., connected via gigabit multimedia serial links and Ethernet), sensors (e.g., LIDAR sensors 864, RADAR sensors 860, etc., which may be connected via Ethernet), data from the bus 802 (e.g., vehicle 800 speed, steering wheel position, etc.), and GNSS sensors 858 (e.g., connected via Ethernet or CAN bus). The SoC 804 may further include a dedicated high-performance mass storage controller, which may include its own DMA engine and may be used to offload routine data management tasks from the CPU 806.
[0130] The SoC 804 may be an end-to-end platform with a flexible architecture spanning levels 3-5 of automation, thereby providing a comprehensive functional safety architecture that leverages and efficiently uses computer vision and ADAS techniques for diversity and redundancy, and provides a platform for a flexible, reliable driving software stack along with deep learning tools. The SoC 804 may be faster, more reliable, and more energy- and space-efficient than conventional systems. For example, when the accelerator 814 is combined with the CPU 806, GPU 808, and data store 816, it can provide a fast and efficient platform for levels 3-5 of autonomous vehicles.
[0131] This technology therefore offers capabilities and functionality not achievable by conventional systems. For example, computer vision algorithms can be implemented on a central processing unit (CPU), which can be configured using a high-level programming language, such as the C programming language, to execute a wide variety of processing algorithms across a wide variety of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, including those related to execution time and power consumption. Specifically, many CPUs cannot execute complex object detection algorithms in real time, a requirement for in-vehicle ADAS applications and practical Level 3-5 autonomous vehicles.
[0132] In contrast to conventional systems, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, the technology described herein allows multiple neural networks to run simultaneously and / or serially and the results to be combined to enable Level 3-5 autonomous driving capabilities. For example, a CNN running on the DLA or dGPU (e.g., GPU820) can include text and word recognition, enabling the supercomputer to read and understand traffic signs, including signs for which the neural network was not specifically trained. The DLA can further include a neural network that can identify, interpret, and provide a semantic understanding of the signs and pass the semantic understanding to a route planning module running on the CPU complex.
[0133] As another example, multiple neural networks may be run simultaneously, as required for Level 3, 4, or 5 operation. For example, a warning sign consisting of "Caution: Flashing lights indicate icy conditions" along with a lightning flash may be interpreted independently or collectively by several neural networks. The sign itself may be identified as a traffic sign by a first deployed neural network (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" may be interpreted by a second deployed neural network that notifies the vehicle's route planning software (preferably running on a CPU complex) that icy conditions exist when the flashing light is detected. The flashing light may be identified by running a third deployed neural network over multiple frames, informing the vehicle's route planning software of the presence (or absence) of the flashing light. All three neural networks may run simultaneously, such as within the DLA and / or on the GPU 808.
[0134] In some instances, a CNN for facial recognition and vehicle owner identification can use data from the camera sensor to identify the presence of a legitimate driver and / or owner of the vehicle 800. An always-on sensor processing engine can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's side door, and in security mode, to disable operation of the vehicle when the owner leaves the vehicle. In this manner, the SoC 804 provides security against theft and / or vehicle hijacking.
[0135] In another example, a CNN for emergency vehicle detection and identification can detect and identify emergency vehicle sirens using data from microphone 896. In contrast to conventional systems that use general classifiers to detect sirens and manually extract features, SoC 804 uses CNNs for environmental and urban sound classification, as well as visual data classification. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative terminal velocity of emergency vehicles (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area in which the mobile vehicle is operating, as identified by GNSS sensor 858. Thus, for example, when operating in Europe, the CNN would attempt to detect European sirens, and when in the United States, the CNN would attempt to identify only North American sirens. After an emergency vehicle is detected, a control program can be used to perform emergency vehicle safety routines, such as slowing down the mobile vehicle, stopping it at the side of the road, parking it, and / or idling it, with the assistance of ultrasonic sensor 862, until the emergency vehicle has passed.
[0136] The vehicle may include a CPU 818 (e.g., a discrete CPU or dCPU) that may be coupled to the SoC 804 via a high-speed interconnect (e.g., PCIe). The CPU 818 may include, for example, an X86 processor. The CPU 818 may be used to perform any of a variety of functions, including, for example, reconciling potentially inconsistent results between the ADAS sensors and the SoC 804 and / or monitoring the status and health of the controller 836 and / or the infotainment SoC 830.
[0137] Vehicle 800 may include a GPU 820 (e.g., a discrete GPU or dGPU) that may be coupled to SoC 804 via a high-speed interconnect (e.g., NVIDIA's NVLINK). GPU 820 may provide additional artificial intelligence functionality, such as by running redundant and / or different neural networks, and may be used to train and / or update neural networks based on input (e.g., sensor data) from sensors in vehicle 800.
[0138] The mobile vehicle 800 may further include a network interface 824, which may include one or more wireless antennas 826 (e.g., one or more wireless antennas for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). The network interface 824 may be used to enable wireless connections with the cloud via the Internet (e.g., with the server 878 and / or other network devices), with other mobile vehicles, and / or with computing devices (e.g., passenger client devices). To communicate with other mobile vehicles, a direct link may be established between the two mobile vehicles and / or an indirect link may be established (e.g., through a network and via the Internet). A direct link may be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link may provide the mobile vehicle 800 with information about mobile vehicles in its vicinity (e.g., vehicles in front of, beside, and / or behind the mobile vehicle 800). This functionality may be part of the collaborative adaptive cruise control functionality of the mobile vehicle 800.
[0139] The network interface 824 may include an SoC that provides modulation and demodulation functions and enables the controller 836 to communicate over a wireless network. The network interface 824 may include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. The frequency conversion may be performed through well-known processes and / or may be performed using a superheterodyne process. In some instances, the radio frequency front end functionality may be provided by a separate chip. The network interface may include wireless functionality for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0140] Mobile vehicle 800 may further include a data store 828, which may include off-chip (e.g., off-SoC 804) storage. Data store 828 may include one or more memory elements, including RAM, SRAM, DRAM, VRAM, flash, hard disk, and / or other components and / or devices capable of storing at least one bit of data.
[0141] Vehicle 800 may further include GNSS sensors 858 (e.g., GPS and / or assisted GPS sensors) that assist with mapping, perception, occupancy grid generation, and / or route planning functions. Any number of GNSS sensors 858 may be used, including, for example, but not limited to, a GPS that uses a USB connector with an Ethernet to serial (RS-232) bridge.
[0142] The mobile vehicle 800 may further include a RADAR sensor 860. The RADAR sensor 860 may be used by the mobile vehicle 800 for long-range mobile vehicle detection, even in darkness and / or severe weather conditions. The RADAR functional safety level may be ASIL B. In some instances, the RADAR sensor 860 may use the CAN and / or bus 802 for control and to access object tracking data (e.g., to transmit data generated by the RADAR sensor 860), with access to Ethernet for accessing raw data. A wide variety of RADAR sensor types may be used. For example, and without limitation, the RADAR sensor 860 may be suitable for front, rear, and side RADAR use. In some instances, a pulse-Doppler RADAR sensor is used.
[0143] The RADAR sensor 860 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, and short-range side coverage. In some instances, long-range RADAR may be used for adaptive cruise control functions. Long-range RADAR systems may provide a wide field of view achieved by two or more independent scans, such as within a 250-meter range. The RADAR sensor 860 may help distinguish between static and moving objects and may be used by ADAS systems for emergency brake assist and forward collision warning. Long-range RADAR sensors may include monostatic multimodal RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In one example with six antennas, the center four antennas may create a focused beam pattern designed to record the surroundings of the moving vehicle 800 at high speeds with minimal interference from traffic in adjacent lanes. The other two antennas may widen the field of view, allowing for rapid detection of moving vehicles entering or leaving the moving vehicle's lane.
[0144] As an example, a medium-range RADAR system may include a range of up to 860 meters (front) or 80 meters (rear) and a field of view of up to 42 degrees (front) or 850 degrees (rear). A short-range RADAR system may include, but is not limited to, a RADAR sensor designed to be mounted on either end of a rear bumper. When mounted on either end of a rear bumper, such a RADAR sensor system can create two beams that constantly monitor the blind spots behind and adjacent to a moving vehicle.
[0145] Short-range RADAR systems may be used in ADAS systems for blind spot detection and / or lane change assist.
[0146] The mobile vehicle 800 may further include ultrasonic sensors 862. The ultrasonic sensors 862, which may be positioned on the front, rear, and / or sides of the mobile vehicle 800, may be used for parking assistance and / or for creating and updating an occupancy grid. A variety of ultrasonic sensors 862 may be used, and different ultrasonic sensors 862 may be used for different ranges of detection (e.g., 2.5 m, 4 m). The ultrasonic sensors 862 may operate at an ASIL B functional safety level.
[0147] The mobile vehicle 800 may include a LIDAR sensor 864. The LIDAR sensor 864 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 864 may be functional safety level ASIL B. In some instances, the mobile vehicle 800 may include multiple (e.g., two, four, six, etc.) LIDAR sensors 864 that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).
[0148] In some instances, the LIDAR sensor 864 may be capable of providing a list of objects and their distances in a 360-degree field of view. Commercially available LIDAR sensors 864 may have an advertised range of approximately 100 m, with an accuracy of 2 cm to 3 cm, and support for a 100 Mbps Ethernet connection, for example. In some instances, one or more non-protruding LIDAR sensors 864 may be used. In such instances, the LIDAR sensor 864 may be implemented as a small device that may be integrated into the front, rear, sides, and / or corners of the vehicle 800. In such instances, the LIDAR sensor 864 may have a range of 200 m, even for low-reflecting objects, and provide up to a 120-degree horizontal and 35-degree vertical field of view. A front-mounted LIDAR sensor 864 may be configured for a horizontal field of view between 45 and 135 degrees.
[0149] In some instances, LIDAR technology such as 3D flash LIDAR may also be used. 3D flash LIDAR uses a laser flash as a transmitter to illuminate the surroundings of the vehicle up to approximately 200 meters. The flash LIDAR unit includes a receptor that records the laser pulse transit time and the reflected light at each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LIDAR may enable a highly accurate and distortion-free image of the surroundings to be generated with every laser flash. In some instances, four flash LIDAR sensors may be deployed, one on each side of the vehicle 800. Available 3D flash LIDAR systems include solid-state 3D steering array LIDAR cameras (e.g., non-scanning LIDAR devices) with no moving parts other than the blower. Flash LIDAR devices may use 5 nanosecond Class I (eye-safe) laser pulses per frame and may capture reflected laser light in the form of a 3D range point cloud and coregistered intensity data. By using flash LIDAR, and because flash LIDAR is a solid-state device with no moving parts, the LIDAR sensor 864 may be less susceptible to motion blur, vibration, and / or shock.
[0150] The mobile vehicle may further include an IMU sensor 866. In some instances, the IMU sensor 866 may be positioned at the center of the rear axle of the mobile vehicle 800. The IMU sensor 866 may include, for example, but not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some instances, such as in a six-axis application, the IMU sensor 866 may include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 866 may include an accelerometer, a gyroscope, and a magnetometer.
[0151] In some embodiments, the IMU sensor 866 may be implemented as a miniature, high-performance GPS-Aided Inertial Navigation System (GPS / INS) that combines micro-electro-mechanical system (MEMS) inertial sensors, a highly sensitive GPS receiver, and advanced Kalman filtering algorithms to provide estimates of position, velocity, and attitude. As such, in some instances, the IMU sensor 866 may enable the vehicle 800 to estimate heading without requiring input from a magnetic sensor by directly observing and correlating changes in velocity from the GPS to the IMU sensor 866. In some instances, the IMU sensor 866 and the GNSS sensor 858 may be combined in a single integrated unit.
[0152] The mobile vehicle may include a microphone 896 placed in and / or around the mobile vehicle 800. The microphone 896 may be used for emergency vehicle detection and identification, among other things.
[0153] The vehicle may further include any number of camera types, including stereo cameras 868, wide-view cameras 870, infrared cameras 872, surround cameras 874, long-range and / or mid-range cameras 898, and / or other camera types. The cameras may be used to capture image data around the entire exterior of the vehicle 800. The types of cameras used depend on the implementation and requirements of the vehicle 800, and any combination of camera types may be used to achieve the desired coverage around the vehicle 800. Additionally, the number of cameras may vary depending on the implementation. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. The cameras may support, by way of example only, Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each camera is described in further detail herein with reference to FIGS. 8A and 8B.
[0154] The vehicle 800 may further include a vibration sensor 842. The vibration sensor 842 may measure vibrations of vehicle components, such as an axle. For example, a change in vibration may indicate a change in the road surface. In another example, when two or more vibration sensors 842 are used, the difference in vibration may be used to determine friction or slippage of the road surface (e.g., when the difference in vibration is between a powered axle and a free-spinning axle).
[0155] The mobile vehicle 800 may include an ADAS system 838. In some instances, the ADAS system 838 may include an SoC. The ADAS system 838 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.
[0156] The ACC system may use a RADAR sensor 860, a LIDAR sensor 864, and / or a camera. The ACC system may include longitudinal ACC and / or lateral ACC. The longitudinal ACC monitors and controls the distance to the vehicle directly ahead of the vehicle 800 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. The lateral ACC performs distance maintenance and advises the vehicle 800 to change lanes when necessary. The lateral ACC is related to other ADAS applications such as LCA and CWS.
[0157] CACC uses information from other moving vehicles, which may be received from other moving vehicles via a wireless link via the network interface 824 and / or wireless antenna 826, or indirectly via a network connection (e.g., via the Internet). A direct link may be provided by a vehicle-to-vehicle (V2V) communication link, while an indirect link may be an infrastructure-to-vehicle (I2V) communication link. Generally, V2V communication concepts provide information about the immediately preceding moving vehicle (e.g., the moving vehicle directly ahead of the moving vehicle 800 that is in the same lane as the moving vehicle 800), while I2V communication concepts provide information about traffic further ahead. A CACC system may include either or both I2V and V2V information sources. Given information about moving vehicles ahead of the moving vehicle 800, CACC may be more reliable, potentially allowing for smoother traffic flow and reducing road congestion.
[0158] The FCW system is designed to warn the driver of hazards so that the driver can take corrective action. The FCW system uses a forward-facing camera and / or RADAR sensor 860 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to driver feedback such as a display, speaker, and / or vibration components. The FCW system can provide warnings in the form of an audio or visual alarm, vibration, and / or a quick brake pulse.
[0159] An AEB system can detect an imminent forward collision with another moving vehicle or other object and automatically apply the brakes if the driver does not take corrective action within specified time or distance parameters. The AEB system can use a forward-facing camera and / or RADAR sensor 860 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid the collision; if the driver does not take corrective action, the AEB system can automatically apply the brakes as part of an effort to prevent, or at least mitigate, the effects of the predicted collision. The AEB system may include techniques such as dynamic brake support and / or collision imminent braking.
[0160] The LDW system provides visual, audible, and / or tactile warnings, such as vibration of the steering wheel or seat, to alert the driver when the mobile vehicle 800 crosses a lane marking. The LDW system does not activate when the driver indicates an intentional lane departure by activating a turn signal. The LDW system may use a forward-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to driver feedback, such as a display, speaker, and / or vibration components.
[0161] The LKA system is a modification of the LDW system, which provides steering input or braking to correct the vehicle 800 if it begins to drift out of its lane.
[0162] The BSW system detects and warns the driver of a moving vehicle in the vehicle's blind spot. The BSW system can provide visual, audible, and / or tactile warnings to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses a turn signal. The BSW system can use a rear-facing camera and / or RADAR sensor 860 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to driver feedback, e.g., a display, speaker, and / or vibration component.
[0163] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the range of the rear camera when the vehicle 800 is backing up. Some RCTW systems include AEB to ensure vehicle brakes are applied to avoid a collision. The RCTW system can use one or more rear-facing RADAR sensors 860 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to driver feedback, e.g., a display, speaker, and / or vibration components.
[0164] Because conventional ADAS systems alert the driver and allow the driver to determine whether a safety condition truly exists and act accordingly, conventional ADAS systems can be prone to producing false positives that, while not usually catastrophic, can be annoying and distracting to the driver. However, in an autonomous vehicle 800, when results conflict, the vehicle 800 itself must decide whether to heed results from a primary computer or a secondary computer (e.g., the first controller 836 or the second controller 836). For example, in some embodiments, the ADAS system 838 may be a backup and / or secondary computer that provides perception information to a backup computer rationality module. The backup computer rationality monitor can run redundant software on hardware components to detect failures in perception and dynamic driving tasks. Output from the ADAS system 838 may be provided to a supervisory MCU. When outputs from the primary and secondary computers conflict, the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.
[0165] In some instances, the primary computer may be configured to provide a reliability score to the supervising MCU indicating the reliability of the primary computer in a selected outcome. If the reliability score exceeds a threshold, the supervising MCU may follow the primary computer's instructions regardless of whether the secondary computers provide conflicting or inconsistent results. If the reliability score does not meet the threshold, and the primary and secondary computers provide different (e.g., conflicting) results, the supervising MCU may arbitrate between the computers to determine the appropriate outcome.
[0166] The supervisory MCU may be configured to execute a neural network trained and configured to determine, based on outputs from the primary and secondary computers, conditions under which the secondary computer will provide a false alarm. Thus, the neural network in the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot be trusted. For example, when the secondary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW identifies a metal object that is not actually dangerous, such as a sewer grate or manhole cover, which triggers an alarm. Similarly, when the secondary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to ignore the LDW when a bicyclist or pedestrian is present and lane departure is, in fact, the safest maneuver. In embodiments including a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or a GPU suitable for executing the neural network with associated memory. In a preferred embodiment, the supervising MCU may comprise and / or be included as a component of the SoC 804 .
[0167] In other instances, the ADAS system 838 may include a secondary computer that performs ADAS functions using traditional rules of computer vision. As such, the secondary computer may use classical computer vision rules (if-then), and the presence of a neural network in the supervisory MCU may improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity may make the overall system more fault-tolerant, particularly to failures caused by software (or software-hardware interface) functions. For example, if a software bug or error exists in software running on the primary computer and non-identical software code running on the secondary computer provides the same overall result, the supervisory MCU may have greater confidence that the overall result is correct and that a bug in the software or hardware used by the primary computer has not caused a critical error.
[0168] In some instances, the output of the ADAS system 838 can be fed to the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, if the ADAS system 838 indicates a forward collision warning due to an object directly ahead, the perception block can use this information when identifying the object. In other instances, the secondary computer can have its own neural network that is trained as described herein, thus reducing the risk of false positives.
[0169] The mobile vehicle 800 may further include an infotainment SoC 830 (e.g., an in-vehicle infotainment system (IVI)). Although shown and described as an SoC, the infotainment system need not be an SoC and may include two or more separate components. The infotainment SoC 830 may include a combination of hardware and software that may be used to provide audio (e.g., music, personal digital assistants, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephony (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation system, reverse parking assist, wireless data system, vehicle-related information such as fuel level, total distance traveled, brake fuel level, oil level, door opening / closing, air filter information, etc.) to the mobile vehicle 800. For example, the infotainment SoC 830 may include a radio, a disc player, a navigation system, a video player, USB and Bluetooth connectivity, a car computer, in-car entertainment, Wi-Fi, steering wheel audio controls, hands-free voice control, a heads-up display (HUD), an HMI display 834, telematics devices, a control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. The infotainment SoC 830 may further be used to provide information (e.g., visual and / or audible) to a user of the vehicle, such as information from an ADAS system 838, autonomous driving information such as planned vehicle maneuvers, trajectory, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0170] The infotainment SoC 830 may include GPU functionality. The infotainment SoC 830 may communicate with other devices, systems, and / or components of the mobile vehicle 800 via the bus 802 (e.g., CAN bus, Ethernet, etc.). In some instances, the infotainment SoC 830 may be coupled to the supervisory MCU so that the infotainment system's GPU can perform some self-driving functions in the event of a failure of the primary controller 836 (e.g., the primary and / or backup computer of the mobile vehicle 800). In such instances, the infotainment SoC 830 may place the mobile vehicle 800 in a Chauffeur safe stop mode, as described herein.
[0171] The mobile vehicle 800 may further include an instrument cluster 832 (e.g., a digital dash, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 832 may include a controller and / or a supercomputer (e.g., a separate controller or supercomputer). The instrument cluster 832 may include a set of instruments such as a speedometer, fuel level, oil pressure, a tachometer, an odometer, turn signals, a gear shift position indicator, a seat belt warning light, a parking brake warning light, an engine malfunction light, airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some instances, information may be displayed and / or shared between the infotainment SoC 830 and the instrument cluster 832. In other words, the instrument cluster 832 may be included as part of the infotainment SoC 830, or vice versa.
[0172] 8D is a system diagram of communication between the cloud-based server and example autonomous vehicle 800 of FIG. 8A in accordance with some embodiments of the present disclosure. System 876 may include server 878, network 890, and vehicle 800. Server 878 may include multiple GPUs 884(A)-884(H) (collectively referred to herein as GPUs 884), PCIe switches 882(A)-882(H) (collectively referred to herein as PCIe switches 882), and / or CPUs 880(A)-880(B) (collectively referred to herein as CPUs 880). GPUs 884, CPUs 880, and PCIe switches may be interconnected with a high-speed interconnect, such as, but not limited to, an NVLink interface 888 developed by NVIDIA and / or a PCIe connection 886. In some instances, the GPUs 884 are connected via an NVLink and / or NVSwitch SoC, and the GPUs 884 and PCIe switch 882 are connected via a PCIe interconnect. While eight GPUs 884, two CPUs 880, and two PCIe switches are illustrated, this is not intended to be limiting. Depending on the embodiment, each server 878 may include any number of GPUs 884, CPUs 880, and / or PCIe switches. For example, the servers 878 may each include 8, 16, 32, and / or more GPUs 884.
[0173] Server 878 can receive image data from the mobile vehicles via network 890, representing images showing unexpected or changed road conditions, such as recently started road construction. Server 878 can transmit neural network 892, updated neural network 892, and / or map information 894, including information about traffic and road conditions, to the mobile vehicles via network 890. Updates to map information 894 can include updates to HD map 822, such as information about construction sites, potholes, detours, flooding, and / or other obstacles. In some instances, neural network 892, updated neural network 892, and / or map information 894 may result from new training and / or experience represented in data received from any number of mobile vehicles in the environment and / or based on training performed at a data center (e.g., using server 878 and / or other servers).
[0174] The server 878 may be used to train a machine learning model (e.g., a neural network) based on training data. The training data may be generated by a mobile vehicle and / or generated in a simulation (e.g., using a game engine). In some instances, the training data is tagged (e.g., if the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other instances, the training data is not tagged and / or preprocessed (e.g., if the neural network does not require supervised learning). The training may be performed according to any one or more classes of machine learning techniques, including, but not limited to, the following classes: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including preliminary dictionary learning), rule-based machine learning, anomaly detection, and variations or combinations thereof. After the machine-learned model is traced, it may be used by the vehicle (e.g., transmitted to the vehicle via network 890) and / or it may be used by server 878 to remotely monitor the vehicle.
[0175] In some instances, server 878 can receive data from mobile vehicles and apply the data to state-of-the-art real-time neural networks for real-time intelligent inference. Server 878 can include deep learning supercomputers and / or dedicated AI computers powered by GPUs 884, such as the DGX and DGX Station machines developed by NVIDIA. However, in some instances, server 878 can include deep learning infrastructure that uses only CPU-powered data centers.
[0176] The deep learning infrastructure of server 878 may be capable of rapid real-time inference and may use that capability to evaluate and verify the health of the processor, software, and / or associated hardware within mobile vehicle 800. For example, the deep learning infrastructure may receive periodic updates from mobile vehicle 800 (e.g., via computer vision and / or other machine learning object classification techniques), such as a sequence of images and / or objects where mobile vehicle 800 was located within the sequence of images. The deep learning infrastructure may run its own neural network to identify objects and compare them to objects identified by mobile vehicle 800; if the results are inconsistent and the infrastructure concludes that the AI within mobile vehicle 800 is not functioning properly, server 878 may send a signal to mobile vehicle 800 instructing its failsafe computer to take control, notify passengers, and complete a safe parking maneuver.
[0177] For inference, server 878 may include a GPU 884 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of a GPU-powered server and inference acceleration can enable real-time responsiveness. In other instances, such as when less performance is required, servers powered by CPUs, FPGAs, and other processors may be used for inference.
[0178] Exemplary Computing Device 9 is a block diagram of an example computing device 900 suitable for use in implementing some embodiments of the present disclosure. The computing device 900 may include an interconnection system 902 that indirectly or directly couples the following devices: memory 904, one or more central processing units (CPUs) 906, one or more graphics processing units (GPUs) 908, a communication interface 910, I / O ports 912, input / output components 914, a power supply 916, one or more presentation components 918 (e.g., displays), and one or more logic units 920.
[0179] While the various blocks in FIG. 9 are depicted as connected via interconnection system 902 with lines, this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component 918, such as a display device, may be considered an I / O component 914 (e.g., if the display is a touch screen). As another example, CPU 906 and / or GPU 908 may include memory (e.g., memory 904 may represent a storage device in addition to the memory of GPU 908, CPU 906, and / or other components). In other words, the computing devices in FIG. 9 are merely exemplary. Categories such as “workstation,” “server,” “laptop,” “desktop,” “tablet,” “client device,” “mobile device,” “handheld device,” “gaming console,” “electronic control unit (ECU),” “virtual reality system,” “augmented reality system,” and / or other device or system types are all intended to be within the scope of the computing devices in FIG. 9 and therefore will not be distinguished from one another.
[0180] Interconnect system 902 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. Interconnect system 902 may include one or more bus or link types, such as an industry standard architecture (ISA) bus, an extended industry standard architecture (EISA) bus, a video electronics standards association (VESA) bus, a peripheral component interconnect (PCI) bus, a peripheral component interconnect express (PCIe) bus, and / or another type of bus or link. In some embodiments, direct connections exist between components. As an example, CPU 906 may be directly connected to memory 904. Further, CPU 906 may be directly connected to GPU 908. When direct or point-to-point connections exist between components, interconnect system 902 may include a PCIe link to implement the connections. In these examples, a PCI bus need not be included in computing device 900.
[0181] Memory 904 may include any of a variety of computer-readable media. Computer-readable media may be any available media that can be accessed by computing device 900. Computer-readable media may include both volatile and nonvolatile media, and removable and non-removable media. By way of example, and not limitation, computer-readable media may comprise computer storage media and communication media.
[0182] Computer storage media may include both volatile and nonvolatile media, and / or removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 904 may store computer-readable instructions (e.g., representing programs and / or program elements), such as an operating system. Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by computing device 900. As used herein, computer storage media does not include the signals themselves.
[0183] Computer storage media may embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal such as a carrier wave or other transport mechanism and include any information delivery media. The term "modulated data signal" may refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, computer storage media may include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared and other wireless media. Combinations of any of the above should also be included within the scope of computer-readable media.
[0184] The CPU 906 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 900 to perform one or more of the methods and / or processes described herein. The CPU 906 may include one or more (e.g., 1, 2, 4, 8, 28, 72, etc.) cores, each capable of simultaneously processing multiple software threads. The CPU 906 may include any type of processor, and may include different types of processors depending on the type of computing device 900 implemented (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 900, the processor may be an Advanced RISC Machines (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). The computing device 900 may include one or more CPUs 906 within one or more microprocessors or auxiliary coprocessors, such as computational coprocessors.
[0185] In addition to or instead of CPU 906, GPU 908 may be configured to execute at least some of the computer-readable instructions to control one or more components of computing device 900 to perform one or more of the methods and / or processes described herein. One or more of GPUs 908 may be integrated GPUs (e.g., with one or more of CPUs 906) and / or one or more of GPUs 908 may be discrete GPUs. In an embodiment, one or more of GPUs 908 may be coprocessors of one or more of CPUs 906. GPU 908 may be used by computing device 900 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, GPU 908 may be used with GPGPU (General-Purpose Computing on a GPU) The GPU 908 may be used for graphics processing (GPU). The GPU 908 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU 908 may generate pixel data for an output image in response to rendering commands (e.g., rendering commands from the CPU 906 received via a host interface). The GPU 908 may include graphics memory, e.g., display memory, for storing pixel data or any other suitable data, e.g., GPGPU data. The display memory may be included as part of the memory 904. GPU 908 may include two or more GPUs operating in parallel (e.g., via a link). The link may connect the GPUs directly (e.g., using NVLINK) or may connect the GPUs via a switch (e.g., using NVSwitch). When coupled together, each GPU 908 may generate pixel data or GPGPU data for a different portion of the output or for a different output (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU may include its own memory or may share memory with other GPUs.
[0186] In addition to or instead of CPU 906 and / or GPU 908, logic unit 920 may be configured to execute at least some of the computer-readable instructions to control one or more of computing devices 900 to perform one or more of the methods and / or processes described herein. In an embodiment, CPU 906, GPU 908, and / or logic unit 920 may discretely or jointly execute any combination of methods, processes, and / or portions thereof. One or more of logic units 920 may be part of and / or integrated with one or more of CPU 906 and / or GPU 908, and / or one or more of logic units 920 may be discrete components to or otherwise external to CPU 906 and / or GPU 908. In an embodiment, one or more of logic units 920 may be a coprocessor of one or more of CPU 906 and / or GPU 908.
[0187] Examples of logic unit 920 include one or more processing cores and / or components thereof, such as a tensor core (TC), a tensor processing unit (TPU), a pixel visual core (PVC), a vision processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multiprocessor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application specific integrated circuit (ASIC), a floating point unit (FPU), an I / O element, a peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) element, and / or the like.
[0188] The communications interface 910 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 900 to communicate with other computing devices over electronic communications networks, including wired and / or wireless communications. The communications interface 910 may include components and functionality to enable communication over any of several different networks, such as a wireless network (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), a wired network (e.g., communicating over Ethernet or InfiniBand), a low-power wide area network (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.
[0189] The I / O ports 912 may enable the computing device 900 to be logically coupled to other devices, including I / O components 914, presentation components 918, and / or other components, some of which may be built into (e.g., integrated with) the computing device 900. Exemplary I / O components 914 include a microphone, mouse, keyboard, joystick, game pad, game controller, satellite dish, scanner, printer, wireless device, etc. The I / O components 914 may provide a natural user interface (NUI) that processes air gestures, voice, or other physiological input generated by the user. In some cases, the input may be sent to an appropriate network element for further processing. The NUI may implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, on-screen and adjacent-screen gesture recognition, air gestures, head and eye tracking, and touch recognition in connection with the display of the computing device 900 (as described in more detail below). The computing device 900 may include a depth camera, such as a stereoscopic camera system, an infrared camera system, an RGB camera system, touch screen technology, and combinations thereof, for gesture detection and recognition. Additionally, the computing device 900 may include an accelerometer or gyroscope (e.g., as part of an inertia measurement unit (IMU)) to enable detection of movement. In some instances, the output of the accelerometer or gyroscope may be used by the computing device 900 to render immersive augmented or virtual reality.
[0190] The power supply 916 may include a hardwired power supply, a battery power supply, or a combination thereof. The power supply 916 may provide power to the computing device 900 to enable the components of the computing device 900 to operate.
[0191] The presentation component 918 may include a display (e.g., a monitor, a touch screen, a television screen, a heads-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component 918 can receive data from other components (e.g., GPU 908, CPU 906, etc.) and output data (e.g., as images, video, sound, etc.).
[0192] The present disclosure may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions, such as program modules, being executed by a computer or other machine, such as a personal digital assistant or other handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. The present disclosure may be implemented in a variety of configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The present disclosure may also be implemented in distributed computing environments where tasks are performed by remote processing devices linked through a communications network.
[0193] As used herein, the term "and / or" in reference to two or more elements should be interpreted to mean one element only or a combination of elements. For example, "element A, element B, and / or element C" may include element A only, element B only, element C only, elements A and B, elements A and C, elements B and C, or elements A, B, and C. Additionally, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Furthermore, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0194] The subject matter of the present disclosure has been described with specificity to meet statutory requirements. However, that description itself is not intended to limit the scope of the disclosure. Rather, the inventors contemplate that the claimed subject matter may be implemented in other ways, including different steps or combinations of steps similar to those described herein, in conjunction with other current or future technologies. Furthermore, although the terms "step" and / or "block" may be used herein to connote different elements of the method used, these terms should not be construed as implying any particular order among the various steps disclosed herein unless and when the order of individual steps is explicitly described.
Claims
1. determining an overlap region corresponding to an overlap between the first image and the second image; determining a difference between the first image and the second image within the overlap region, the difference comprising calculating at least one score indicative of a difference between the first image and the second image within the overlap region; transforming at least the first image based at least in part on the score; Including, inferring a second score from said score using one or more neural networks; transforming at least the first image based at least in part on the second score; The method further comprises:
2. calculating one or more deformations as a result of the difference being less than a threshold; applying the one or more transformations to one of the first image or the second image; The method of claim 1 further comprising:
3. The method of claim 1 , further comprising combining the first image and the second image into a third image based at least in part on the overlap region.
4. the at least one score calculated is based at least in part on one or more geometric differences between the first image and the second image within the overlap region; determining the difference includes applying one or more geometric transformations to the first image or the second image as a result of the one or more geometric differences; The method of claim 1 , comprising:
5. the at least one score calculated is based at least in part on one or more photometric differences between the first image and the second image within the overlap region; determining the difference includes applying one or more photometric transformations to the first image or the second image as a result of the one or more photometric differences; The method of claim 1 , comprising:
6. determining the difference includes calculating at least one first score indicative of a geometric difference between the first image and the second image within the overlap region; calculating at least one second score indicative of a photometric difference between the first image and the second image within the overlap region as a result of transforming at least the first image; calculating one or more transformations to be applied to at least the first image based at least in part on the at least one second score; The method of claim 1 , comprising:
7. determining the difference calculating at least one score indicative of a difference between the first and second images within the overlap region; determining whether the at least one score exceeds a threshold; combining the first and second images into a third image based at least in part on the overlap region as a result of the at least one score exceeding a threshold; The method of claim 1 , comprising:
8. A vehicle or robot implementing the method of claim 1, a first image capture device for capturing the first image; a second image capture device for capturing the second image; a display device for displaying a stitched image generated based at least in part on at least the transformed first image; one or more processors; a memory containing instructions executable by said one or more processors to cause said one or more processors to perform the method of claim 1; A vehicle or robot comprising:
9. 1. A vehicle system, comprising: one or more processors; When executed by the one or more processors, determining an overlap region corresponding to an overlap between the first image and the second image; calculating a score indicative of a difference between the first image and the second image within the overlap region; and transforming at least the first image based at least in part on the score; a memory containing instructions for causing the vehicle system to Equipped with The memory further comprises: inferring a second score from said score using one or more neural networks; transforming at least the first image based at least in part on the second score; and and further comprising instructions that, in response to being executed by the one or more processors, cause the vehicle system to:
10. The memory further comprises: calculating one or more transformations resulting from said score being less than a threshold; applying the one or more transformations to one of the first image or the second image; 10. The vehicle system of claim 9, comprising instructions that, in response to being executed by the one or more processors, cause the vehicle system to:
11. the first image is associated with a first time and the third image is associated with a second time; the score indicating a difference between the first image and the third image; The vehicle system of claim 10.
12. The memory calculating the score based at least in part on optical flow vectors usable to facilitate identification of one or more differences between the first and second images within the overlap region; applying one or more geometric transformations to at least the first image as a result of the one or more object differences; The vehicle system of claim 9 , further comprising instructions that, in response to being executed by the one or more processors, cause the vehicle system to:
13. The memory calculating the score based at least in part on one or more color differences between each color channel of the first image and the second image within the overlap region; applying one or more photometric transformations to the first image or the second image as a result of the one or more color differences; The vehicle system of claim 9 , further comprising instructions that, in response to being executed by the one or more processors, cause the vehicle system to:
14. 10. The vehicle system of claim 9, wherein the memory further comprises instructions that, in response to being executed by the one or more processors, cause the vehicle system to combine the first image and the second image into a third image based at least in part on the overlap region.
15. The memory calculating a second score as a result of transforming at least the first image; calculating one or more transformations to be applied to at least the first image based at least in part on the second score; applying the one or more transformations to at least the first image; The vehicle system of claim 9 , further comprising instructions that, in response to being executed by the one or more processors, cause the vehicle system to:
16. 10. The vehicle system of claim 9, wherein the memory further comprises instructions that, in response to being executed by the one or more processors, cause the vehicle system to combine the first image and the second image into a third image based at least in part on the overlap region as a result of the score exceeding a threshold.
17. one or more controllers; a network interface; The display and a propulsion system; two or more surround cameras operable to capture the first image and the second image; The vehicle system of claim 9 further comprising:
18. 1. A robotic system comprising: one or more processors coupled to a computer-readable medium; or determining an overlap region corresponding to an overlap between the first image and the second image; calculating a score indicative of a difference between the first image and the second image within the overlap region; and transforming at least the first image based at least in part on the score; said computer-readable medium storing executable instructions to cause Equipped with the computer-readable medium comprising: inferring a second score from said score using one or more neural networks; transforming at least the first image based at least in part on the second score; and and further comprising instructions that, when executed by the one or more processors, cause the one or more processors to:
19. the computer-readable medium comprising: calculating a second score as a result of transforming at least the first image; calculating one or more transformations to be applied to at least the first image based at least in part on the second score; applying the one or more transformations to at least the first image; 20. The robotic system of claim 18, further comprising instructions that, when executed by the one or more processors, cause the one or more processors to:
20. the computer-readable medium comprising: calculating the score based at least in part on one or more geometric differences between the first image and the second image within the overlap region; applying one or more three-dimensional transformations to at least the first image as a result of the one or more geometric differences; 20. The robotic system of claim 18, further comprising instructions that, when executed by the one or more processors, cause the one or more processors to:
21. the computer-readable medium comprising: calculating the score based at least in part on one or more color differences between the first image and the second image; adjusting one or more color parameters of at least the first image as a result of the one or more color differences; 20. The robotic system of claim 18, further comprising instructions that, when executed by the one or more processors, cause the one or more processors to:
22. the computer-readable medium comprising: Transforming at least the first image into a third image; determining a second overlap region corresponding to a second overlap of the fourth image and the fifth image; calculating a second score indicative of a difference between the fourth image and the fifth image within the second overlap region; Transforming at least the fourth image into a sixth image based at least in part on the second score; combining the third image and the sixth image to create video data; 20. The robotic system of claim 18, further comprising instructions that, when executed by the one or more processors, cause the one or more processors to:
23. the computer-readable medium comprising: calculating one or more transformations resulting from said score being less than a threshold; applying the one or more transformations to one of the first image or the second image; 20. The robotic system of claim 18, further comprising instructions that, when executed by the one or more processors, cause the one or more processors to:
24. the first image associated with a first time and the third image associated with a second time in the time interval; the score indicating a temporal difference between the first image and the third image during the time interval; 20. The robotic system of claim 18.
25. one or more controllers; one or more sensors; a communication interface; two or more cameras operable to capture at least the first image and the second image; 20. The robotic system of claim 18, further comprising:
Citation Information
Patent Citations
Image processor and method
JP2008077666A
Panoramic image synthesizer, panoramic image synthesis method, and program
JP2011119974A
3D Rendering for Surround View Using Predefined Viewpoint Lookup Table
JP2019503007A
Seamless image stitching
US20200020075A1
Image Quality Assessment
US20200236280A1