Image coordination for image stitching systems and applications

Through the image signal processing (ISP) parameter coordination function, the rendering parameters between match and mix images are solved, and the problem of discontinuity of rendering parameters in image stitching under different lighting conditions is achieved, and smooth rendering of image stitching and high-quality image output are achieved.

CN119991426APending Publication Date: 2025-05-13NVIDIA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411618847.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-13
Filing Date
2024-11-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing image stitching technology is prone to discontinuity of rendering parameters between images captured under different lighting conditions, resulting in the appearance of artifacts and affecting the image quality and the accuracy of downstream processes.

Method used

Smooth rendering of images is achieved by using the Image Signal Processing (ISP) parameter coordination function, input image processing metadata parameters associated with the image, match the rendering parameters on the overlapping boundary between the image and mix the image.

Benefits of technology

During the image stitching process, the discontinuity of rendering parameters is avoided, and the artifacts appear in the stitched image are basically avoided, which improves the image quality and the accuracy of downstream processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991426A_ABST
    Figure CN119991426A_ABST
Patent Text Reader

Abstract

The invention relates to image coordination for image stitching systems and applications. In various examples, metadata-based image coordination for image stitching systems and applications is disclosed. Systems and methods are disclosed for pre-processing images for rendering parameters with the effect of mixing these parameters at the boundaries between the images to facilitate smooth rendering when these images are stitched together. An image signal processing (ISP) parameter coordination function may input metadata parameters associated with a set of images to match and mix one or more of rendering parameters on overlapping boundaries between the images prior to applying the images to a stitching algorithm. Scaling of metadata parameters may be performed using a parametric gain function. Pixels in two images located on the boundary are adjusted to the same boundary metadata parameter value and smoothed based on a parametric gain function. Discontinuity of rendering parameters is avoided, thereby substantially avoiding corresponding artifacts in the resulting stitched image.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Image stitching is a technique used in digital photography and computer graphics to create panoramic or 360-degree surround images by combining multiple overlapping images into a single seamless composite image. Stitching algorithms can be used to align common features or reference points that appear in two or more photos and blend pixel values ​​at the edges of the images to create smooth transitions. The result is a single image with a larger field of view than the original images. In automotive applications, the stitched images can be used to provide a visual display image of the vehicle's surroundings to the driver of the vehicle, giving the driver a better understanding of the surrounding environment and supplementing and / or replacing the view previously provided by the vehicle's rearview mirrors. Summary of the invention

[0002] Embodiments of the present disclosure relate to image harmonization for image stitching systems and applications. Systems and methods are disclosed for preprocessing images using metadata about one or more rendering parameters, the effect of which is to blend these parameters at the boundaries between images to achieve smooth rendering when the images are stitched together.

[0003] Compared to conventional systems associated with stitching techniques, embodiments of the present disclosure at least in part provide an image signal processing (ISP) parameter coordination function that inputs image processing metadata parameters associated with a set of images to match and blend one or more rendering parameters on overlapping boundaries between images before applying the images to a stitching algorithm (a process referred to herein as coordination). Image processing metadata can be used as parameters to control the rendering of image data captured by a sensor within an image frame. The rendering parameters applied to adjust the image captured by the sensor can be defined by the image metadata, such as but not limited to white balance temperature, global tone mapping (GTM), exposure time, brightness, contrast, gamma, chroma, noise reduction, saturation and / or sharpness. In some embodiments, the rendering parameters may include lens distortion and / or lens shading correction parameters applied to compensate for the optical characteristics of the sensor hardware.

[0004] The scaling may be performed using a parameter gain function that inputs an initial value of the metadata parameter, as applied by the camera that captured the first image and the signal image. For example, the ISP parameter coordination function may calculate a boundary metadata parameter value based on an average of a first metadata parameter value for the first image and a second metadata parameter value for the second image. The parameter gain function may then be used for each image to gradually adjust the metadata parameter from a location near or at the center of the image, where the applied initial metadata parameter value remains unchanged, to a boundary edge between the images, where the applied metadata parameter is the boundary metadata parameter value. Thus, the same boundary metadata parameter value will be used to adjust pixels on either side of the boundary edge where the images are to be stitched, where the metadata parameter value smoothly follows the parameter gain function. During the stitching process, discontinuities in the rendering parameters are avoided at the boundary region where the original images are blended together, thereby substantially avoiding corresponding artifacts in the resulting stitched image. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] The present system and method for metadata-based image coordination for image stitching will be described in detail below with reference to the accompanying drawings, wherein:

[0006] Figure 1 is a data flow diagram illustrating a metadata coordinated image stitching system according to some embodiments of the present disclosure;

[0007] Figure 2 is a diagram illustrating metadata coordination between two images according to some embodiments of the present disclosure;

[0008] Figure 3 is a diagram illustrating coordination of metadata parameters of images to render a stitched image according to some embodiments of the present disclosure;

[0009] Figure 4 and Figure 5 is a diagram illustrating attenuation of metadata artifacts provided by coordination of metadata parameters of images to render a stitched image according to some embodiments of the present disclosure;

[0010] Figure 6 is a flow chart of a method for metadata coordinated image stitching according to some embodiments of the present disclosure;

[0011] Fig. 7A is an illustration of an example autonomous vehicle according to some embodiments of the present disclosure;

[0012] Figure 7B According to some embodiments of the present disclosure Fig. 7A Examples of autonomous vehicle camera positions and fields of view;

[0013] Figure 7CAccording to some embodiments of the present disclosure Fig. 7A a block diagram of an example system architecture for an example autonomous vehicle;

[0014] Fig.7D is a cloud-based server and Fig. 7A System diagram of communication between example autonomous vehicles;

[0015] Figure 8 is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and

[0016] Fig. 9 is a block diagram of an example data center suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0017] Systems and methods related to image coordination for image stitching are disclosed. Although the present disclosure may be related to an example autonomous or semi-autonomous vehicle or machine 700 (referred to herein alternatively as "vehicle 700" or "self-machine 700"), the example Figures 7A-7D For example, the systems and methods described herein may be used, without limitation, by non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more advanced driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, flying boats, ships, space shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, engineering vehicles, trains, underwater vehicles, remotely controlled vehicles (e.g., drones), and / or other vehicle types. Furthermore, while the present disclosure may be described with respect to image stitching of vehicle surrounding images, this is not limiting, and the systems and methods described herein may be used for augmented reality, virtual reality, mixed reality, robotics, security and surveillance, autonomous or semi-autonomous machine applications, and / or any other technical field in which stitched images may be used.

[0018] The present disclosure relates to techniques for stitching images together to form a composite image of a scene. As described herein, systems and methods are provided for preprocessing images according to one or more rendering parameters, the effect of which is to blend these parameters at the boundaries between images to achieve smooth rendering when these images are stitched together. Image stitching is a technique used in digital photography and computer graphics to create panoramic or 360-degree surround view images by combining multiple overlapping images into a single seamless composite image.

[0019] However, problems may arise when the cameras that captured the original images applied different camera parameters to capture the scene. For example, two cameras that simultaneously and independently captured overlapping images of a scene may generate images with different white balance (color temperature), tone mapping, exposure, lens shading correction, color correction, noise reduction, sharpening, and / or other rendering parameters. Although traditional image stitching can produce a coherent image, the resulting image may contain visible artifacts regarding image brightness, color, noise, and / or clarity due to the discontinuity of rendering parameters at the boundary areas where the original images are blended together.

[0020] Artifacts that appear between images captured by different cameras often become more severe under different lighting conditions. For example, when an ego vehicle drives into a garage at night, the lighting conditions inside the garage may be very different from the lighting conditions outside the garage. As a result, individual cameras in a set of surround view cameras may apply exposure, white balance, and / or other rendering parameters that are significantly different from other cameras in the same set. When images captured from a well-lit garage using bright light condition parameters are stitched with overlapping images captured from a darker driveway using dim light condition parameters, artifacts caused by the different lighting condition parameters often appear clearly at the seam where the images are stitched. For example, pixels with green tint artifacts may appear at the seam between two images with significantly different white balance parameter settings. Such artifacts can adversely affect image quality assessment, visualization quality (e.g., user experience), and / or the accuracy of downstream processes that may use the stitched images. For example, a stitched image depicting a rear-view scene displayed to a driver should ideally present the driver with the same general appearance as what they would see behind the car. Similarly, the side view images providing the blind spot view should display the same general appearance as the driver would see if he were able to perceive the blind spot area. Rendering artifacts in the stitching process can distract or cause confusion to the driver (e.g., mistaking the artifacts for unexpected objects). Furthermore, in some implementations, rendering artifacts in the stitching process can cause downstream processes (e.g., artificial intelligence (AI) based processes) that use the stitched images to misclassify elements in the scene or have a lower confidence level in the prediction than it would have if the artifacts were not present in the stitched image.

[0021] In an attempt to address the problem that arises when a group of different cameras use different rendering parameters when capturing images, some image stitching techniques select one camera from a group of cameras as a reference camera and select an image from the reference camera as the reference image. Using this stitching technique (often referred to as a global color transfer technique), the global color statistics (e.g., color channel means and / or standard deviations) of the images from each of the other cameras in the group are transformed to match the global color statistics of the reference image. However, there are several disadvantages to techniques based on global color transfer. For example, techniques based on color transfer are sensitive to the decision of which camera is selected as the reference camera from a group of cameras. That is, depending on the selection of the reference camera, the resulting stitched image may differ in overall appearance. For example, if the forward-looking camera of a car is selected as a reference, it may capture a brighter scene than the rear-view camera. Global color transfer can be applied between a left-view camera and a right-view camera that produce images that share an overlapping area with the forward-looking camera. However, the rear-view camera may not share an overlapping view with the forward-looking camera, so either the left-view camera image or the right-view camera image will be selected as a reference to perform global color transfer of the rear-view camera. Since the left and right cameras may observe very different scenes under different lighting conditions, the left and right cameras may apply different rendering parameters to their captured images (e.g., apply different white balance color temperatures). Their global color statistics may be correspondingly different, so the appearance of the adjusted image from the rear view camera may be very different depending on whether the left or right camera is selected as a reference. Therefore, when the front view image, the left view image, the rear view image, and the right view image are stitched together to form a 360-degree surround view, a noticeable discontinuity in the rendering parameters may appear at the seam between the rear view image and the right view image or the left view image, thereby producing visible artifacts along the seam in the stitched image.

[0022] Global color transfer techniques are also challenged when adjacent cameras observe different scenes due to spatial differences, or have objects that significantly fill the overlapping area. For example, if a car is parked between the fields of view of two adjacent cameras, the color statistics transferred to the reference image may not be representative of the color statistics of the entire image. In some global color transfers, the two camera images can be segmented to find matching blocks that appear in both camera images, and such blocks are selected to define the color statistics to be transferred between the images. However, the segmentation and matching algorithms can be computationally intensive, making it impossible to render real-time visualizations for real-time surround view systems.

[0023] Compared to these existing stitching techniques, embodiments of the present disclosure at least in part provide an image signal processing (ISP) parameter coordination function that inputs image processing metadata parameters associated with a set of images to match and blend one or more rendering parameters on overlapping boundaries between images before applying the images to a stitching algorithm (a process referred to herein as coordination). The image processing metadata parameters control the rendering of image data captured by a sensor within an image frame. The image processing metadata parameters can control the rendering relative to rendering parameters that, for example, adjust the captured image for specific lighting conditions and / or camera characteristics. The rendering parameters applied to adjust the image captured by the sensor can be defined by image metadata, such as, but not limited to, white balance temperature, global tone mapping (GTM), exposure time, brightness, contrast, gamma, chroma, noise reduction, saturation, and / or sharpness. In some embodiments, the rendering parameters may include lens distortion and / or lens shading correction parameters applied to compensate for the optical characteristics of the sensor hardware.

[0024] Coordination may be applied on a per image frame basis, wherein metadata parameters applied to pixels within a first image may be adjusted and scaled in a direction toward an image boundary with a second image, starting from a location at (or near) the center of the first image. Scaling may be performed using a parameter gain function, which inputs the initial values ​​of metadata parameters applied by the camera that captured the first image and the signal image. For example, for the first image, the ISP parameter coordination function may calculate a boundary metadata parameter value based on an average of a first metadata parameter value for the first image and a second metadata parameter value for the second image. The parameter gain function may then be used to gradually adjust the metadata parameters applied to the first image from a selected location near or at the center of the image (where the initial metadata parameter value applied remains unchanged) to a boundary edge between images (where the metadata parameter applied is the boundary metadata parameter value). For the second image, the ISP parameter coordination function uses the same boundary metadata parameter value (e.g., based on an average of the first metadata parameter value for the first image and the second metadata parameter value for the second image). The parameter gain function may be used to gradually adjust the metadata parameters applied to the second image from a location near or at the center of the image to a boundary edge between images. In this way, pixels located on either side of the boundary edge where the images are to be stitched are adjusted using the same boundary metadata parameter value, wherein the metadata parameter value smoothly follows the parameter gain function. Thus, during the stitching process, discontinuities in the rendering parameters are avoided at the boundary regions where the original images are blended together, thereby substantially avoiding corresponding artifacts in the resulting stitched image. The coordination between adjacent images before stitching can be applied to a single metadata parameter or to multiple metadata parameters.

[0025] In some embodiments, coordination can be applied between an image and a plurality of other images to be stitched therewith. For example, a 360-degree surround image can be formed by stitching together four partially overlapping images captured by four cameras (e.g., a front camera, a left camera, a rear camera, and a right camera). In this case, the left side of each image can be coordinated with the image to be stitched to its left side, and the right side of each image can be coordinated with the image to be stitched to its right side. When the images are stitched together, each image is coordinated across the seams formed with its adjacent images-thereby avoiding rendering parameter discontinuities that amplify artifacts in the stitched image. Coordination can be performed for each boundary of the image and the overlapping adjacent images. For example, coordination can be performed with adjacent images on the left and / or right side of the image, adjacent images above and / or below the image, or any combination thereof. In some embodiments, a set of metadata parameter coordination of the image can be applied sequentially, and / or simultaneously applied by mathematically combining parameter gain functions and calculating the resulting combined parameter gain adjustment.

[0026] For example, parameter coordination may be formed for a white balance metadata parameter. Different light sources may produce light with different color distributions, so the color of objects in a scene may differ, depending on the light source that illuminates the object. An image captured under light with a higher color temperature may appear more blue, while an image captured under light with a lower color temperature may appear more yellow. White balance, also known as color temperature and measured in degrees Kelvin (K), is a metadata parameter that can be adjusted when an image is captured so that white objects appear white in the captured image. Auto white balance is a camera rendering parameter setting (usually on by default) in which the camera senses the ambient lighting conditions and automatically adjusts the white balance so that the colors in the captured image appear natural. Therefore, the auto white balance of two different cameras capturing different but overlapping areas of a scene will select and adjust the images they capture using different white balance color temperatures, because factors such as the lighting angle, shadows, glare, and / or the presence or absence of reflective surfaces may cause each camera to evaluate the ambient lighting conditions differently. For example, although two cameras may simultaneously capture images containing the same red object, the red object may appear different in one image than in the other due to the different white balance adjustments applied by the two cameras to their respective images.

[0027] White balance coordination compensates for this difference in white balance applied between the two images so that when the two images share an image boundary at an overlap boundary, the applied white balance metadata parameters are the same and a smooth transition of white balance is achieved from the center of one image to the center of the other image. The gradient of the white balance metadata parameters from the center of the image to the image boundary is a function of a parameter gain function. The parameter gain function can be implemented using a simple linear ramp function or a more complex curve, such as an s-curve (e.g., a standard logistic function curve). The parameter gain function can be calibrated so that it has a unity gain value at or near the center of the image. For image boundaries, the parameter gain function is calibrated based on a function of the corresponding white balance parameters selected by each camera so that pixels at the image boundary in each image will be rendered using the same white balance color temperature. Therefore, the gain at the image boundary between the first image and the second image can be expressed as:

[0028]

[0029] Where WB1 is the white balance temperature applied by camera 1 to image 1, WB2 is the white balance temperature applied by camera 2 to image 2, W' is the position of the image boundary relative to the frame of the first image, and 0 is the position of the image boundary relative to the frame of the second image. The white balance gain at the image center of the first image and the second image is unity gain, which can be expressed as:

[0030] g1WB(Width / 2)=1, g2WB(Width / 2)=1,

[0031] Where "Width / 2" represents the location of the center of the image for each image. For the first image, the white balance applied to the pixel may be adjusted according to g1WB(x) x WB1 = parametric gain function 1(x) x WB1, where x = Width / 2 to W'. For the second image, the white balance applied to the pixel will be adjusted according to g2WB(x) x WB2 = parametric gain function 2(x), where x = 0 to Width / 2. Note that the specific parametric gain function used in each image may apply a different shape of curve.

[0032] As another example, a parameter coordination can be formed for a global tone mapping (GTM) metadata parameter. Tone mapping can be used to map one set of colors to another set of colors, for example, to provide the appearance of a high dynamic range (HDR) image from a source image with a more limited dynamic range. As with white balance coordination, GTM coordination compensates for differences in the GTM applied between two images so that the GTM metadata parameters are the same, where the two images share image boundaries at overlapping boundaries, and a smooth transition of the GTM is achieved from the center of one image to the center of the other image. The gradient of the GTM metadata parameter from the image center to the image boundary is also adjusted according to a function of the parameter gain function. To coordinate the GTM, the gain at the image boundary between the first image located to the right of the second image can be expressed as:

[0033]

[0034] Where LUT1GTM and LUT2GTM are coordinated GTM 256-point lookup tables (LUTs). LUT1GTM is obtained by interpolation between the original GTM1origLUT applied to image 1 by camera 1 and the average GTM LUT = f(GTM1origLUT, GTM2origLUT). LUT2GTM is obtained by interpolation between the original GTM2origLUT applied to image 2 by camera 2 and the average GTM LUT, which is derived as a function of GTM1origLUT and GTM2origLUT. The GTM gain at the image center of the first image and the second image may be unity gain, which may be expressed as:

[0035] g1GTM(Width / 2)=1, g2GTM(Width / 2)=1 。

[0036] In this way, coordination for different metadata parameters can be applied based on calculating common metadata parameter values ​​between images at image boundaries and applying a parameter gain function within each image to smooth the application of the metadata parameter from a location within the image (e.g., the center of the respective image) to the shared boundary. Once the desired set of metadata parameters are coordinated, the images can be applied to the image stitching algorithm. As described above, in addition to white balance and GTM, other metadata parameters that can be coordinated using this process include, but are not limited to, parameters for exposure, lens shading correction, color correction, noise reduction, sharpening, color filter array (CFA) pattern, focal length correction, lens distortion correction (e.g., barrel distortion correction and pincushion distortion), and / or other rendering parameters that affect the way the image is rendered. In some embodiments, a series of coordinations can be applied to a set of images. While the order in which the coordination is performed may affect the final overall appearance of the stitched images, the choice of which metadata parameters to coordinate and in what order to coordinate can be based on the specific use case for generating the stitched images.

[0037] Metadata parameters associated with each image may be communicated to the ISP parameter coordination function in a variety of ways. For example, metadata parameters including rendering parameters selected by the camera may be communicated by the camera to the ISP parameter coordination function as metadata embedded and / or carried with the image data of each captured image frame. In some embodiments, the metadata parameters may be communicated by the camera to the ISP parameter coordination function via a channel separate from a channel carrying image data (e.g., pixel color data). In some cases, one or more metadata parameters may represent static characteristics or settings associated with the camera that do not change with each captured image (e.g., lens distortion correction and / or sensor or pixel size). Rather than being communicated by the camera to the ISP parameter coordination function with each captured image, these metadata parameters may be stored in memory and recalled by the ISP parameter coordination function as needed.

[0038] As described above, the parameter gain function can be implemented using any function that produces the desired gradient of the metadata parameter blending. Possible parameter gain functions can include using a simple linear ramp function or a function that produces a more complex and / or non-linear curve, such as a logistic function curve. The parameter gain function can be selected and / or adjusted on an image-by-image basis based on the metadata parameters to be coordinated and / or the individual lighting conditions affecting each image. For example, if two adjacent images have significantly different brightness levels, the steepness or rate of change of the parameter gain function (e.g., the logistic growth rate of the logistic function) can be adjusted to change more quickly to promote more seamless parameter blending. In low light conditions, a more gradual smoothing rate of white balance on an image can better resolve parameter discontinuities at image boundaries than a steeper smoothing rate. In some embodiments, the parameter gain function curve can be selected from a library of predefined curves by the ISP parameter coordination function. In some embodiments, the parameter gain function curve can be dynamically calculated by the ISP parameter coordination function.

[0039] Although examples of coordination parameters are discussed in the context of stitching together images from a set of cameras to produce a 360 degree surround image, coordination can be applied to other camera array configurations for other image stitching use cases. For example, multiple cameras can be arranged as a hemispherical camera array. Such a camera array can be used by an aerial drone to view the ground below or the sky above, and stitch the captured images together to generate a continuous wide-angle fisheye image for visualization, navigation, and / or other purposes. It should be understood that the embodiments described herein can be used in the context of vehicles, such as cars, trucks, trains, airplanes, spacecraft, and / or ships, but can be extended to other machinery, such as remote-controlled devices (e.g., robots and drones), industrial and / or construction machinery (e.g., lifts and cranes), and / or any other application where images can be stitched together to generate a composite image.

[0040] refer to Figure 1 , Figure 1 is an example data flow diagram for a metadata coordinated image stitching system 100 according to some embodiments of the present disclosure. It should be understood that this arrangement and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used in addition to or in place of the arrangements and elements shown, and some elements may be omitted entirely. In addition, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in combination with other components, and may be implemented in any suitable combination and location. The various functions performed by the entities described herein may be performed by hardware, firmware, and / or software. For example, the various functions may be performed by a processor executing instructions stored in a memory. In some embodiments, the systems, methods, and processes described herein may be used with Figures 7A-7D Example of autonomous vehicle 700, Figure 8 The example computing device 800 and / or Fig. 9 The example data center 900 may be implemented with similar components, features and / or functions.

[0041] like Figure 1 As shown, the metadata coordinated image stitching system 100 may include a parameter coordination function 120 that generates stitched image data 152 (e.g., a stitched image) of a three-dimensional (3D) environment (e.g., surrounding an object, such as a vehicle) based on image data 110 captured by one or more image sensors 105 and image metadata parameters 112 associated with the image data 110. The image data 110 may include image frames of a plurality of images. The image sensors 105 may include, for example, RGB, infrared (IR) and / or RGB-IR cameras, and / or other cameras, such as for FIG. 7A to FIG. 7D The camera depicted in vehicle 700. Image data 110 is not limited to any particular color space. For example, image data 110 may include images represented using a color space such as, but not limited to, YUV, RGB, CIELAB (L*a*b*), or other color spaces. Image sensor 105 may include one or more cameras from an object or from a participant, such as FIG. 7A to FIG. 7D The image sensors 105 may be configured to include a stereo camera 768, a wide angle camera 770 (e.g., a fisheye camera), an infrared camera 772, a surround camera 774 (e.g., a 360° camera), and / or a long-range and / or medium-range camera 798 of the autonomous vehicle 700. The image sensors 105 may be configured to generate one or more of the image data 110 of the 3D environment surrounding the self-object or self-participant and the image metadata parameters 112. In embodiments where multiple image sensors 105 are used, the multiple image sensors 105 may view a common area of ​​the 3D environment and overlapping portions of their respective fields of view, such that the image data 110 (e.g., images) from different image sensors 105 represent at least partially overlapping images of the 3D environment.

[0042] The image metadata parameters 112 are parameters that can be used to control image rendering within an image frame of the image data 110. The image processing metadata parameters of the image metadata parameters 112 can control rendering relative to rendering parameters that, for example, adjust an image captured by the image sensor 105 for a particular lighting condition and / or camera characteristics. The rendering parameters contained within the image metadata parameters 112 associated with a set of image data 110 may include parameters such as, but not limited to, white balance temperature, global tone mapping (GTM), exposure time, brightness, contrast, gamma, chromaticity tint, noise reduction, saturation, and / or sharpness. In some embodiments, the rendering parameters may include lens distortion and / or lens shading correction parameters to compensate for optical characteristics of the image sensor 105 hardware.

[0043] The image metadata parameters 112 associated with each image represented by the image data 110 may be communicated to the parameter coordination function 120 in a variety of ways. For example, metadata parameters including rendering parameters selected by the image sensor 105 may be communicated by the image sensor 105 to the parameter coordination function 120 as metadata embedded in and / or carried with the image data 110 for each captured image frame. In some embodiments, one or more of the image metadata parameters 112 may be communicated by the image sensor 105 to the parameter coordination function 120 via a channel separate from the channel carrying the image data 110. In some cases, one or more of the metadata parameters may represent static characteristics or settings associated with the image sensor 105 that do not change with each captured image (e.g., lens distortion correction and / or sensor or pixel size). Such metadata parameters may be stored in the memory 114 and recalled by the parameter coordination function 120 as needed, rather than being communicated by the image sensor 105 to the parameter coordination function 120 with or in addition to the image data 110.

[0044] In the metadata coordinated image stitching system 100, the parameter coordination function 120 can process the image data 110 and the corresponding image metadata parameters 112 (as described herein) to generate a metadata parameter coordinated version of the image data 110 (shown as coordinated metadata image data 128). In some embodiments, the coordinated metadata image data 128 includes a plurality of image frames from the image data 110 that have been coordinated for at least one metadata parameter by the parameter coordination function 120 (as described herein). The stitching module 150 can stitch the frames of the coordinated metadata image data 128 into stitched image data 152 (e.g., a 360° surround image, a wide angle image, a fisheye image, and / or a panoramic image) using any stitching technique. In some embodiments, the stitched image data 152 can be further processed using image processing 155 based on one or more image processing techniques. For example, in some embodiments, the image processing 155 may coordinate colors, apply image sharpening, noise reduction, gamma correction, exposure compensation, and / or apply other filters and / or adjustments to portions of the stitched image data 152. In some embodiments, the image processing 155 may perform color coordination and / or other image processing, such as described with respect to the color coordinator disclosed in U.S. patent application Ser. No. 17 / 959,934, filed on Oct. 4, 2022, entitled “IMAGE STITCHING WITH COLOR HARMONIZATION FOR SURROUND VIEW SYSTEMS AND APPLICATIONS,” the entire contents of which are incorporated herein by reference.

[0045] Rendering module 160 may cause visualization 165 to be rendered (e.g., on a monitor visible from an occupant or operator of the subject or from a participant) of at least a portion of stitched image data 152. In some embodiments, rendering module 160 projects stitched image data 152, or a portion thereof, onto a 3D representation of a 3D environment (e.g., a 3D bowl that models the 3D environment), renders a view of the projected stitched image data 152 from the perspective of a virtual camera, and / or renders the rendered view as visualization 165.

[0046] In some embodiments, the stitched image data 152 may be used by one or more downstream navigation components 170 of the self-machine, such as the controller 736 discussed below. For example, the downstream navigation component 170 may implement functions such as object avoidance navigation functions and / or a world model manager, a path planner, a control component, a positioning component, an obstacle avoidance component, an actuation component, etc. to perform operations for controlling the self-machine in the environment. In some embodiments, the downstream navigation component 170 may include one or more deep neural networks (DNNs) that generate one or more predictions and / or inferences about the 3D environment based at least on the stitched image data 152.

[0047] For some embodiments, the downstream navigation component 170 may include at least one or more path planning functions 172 (e.g., path planning functions for the self-machine 700) and / or actuation and control 174 (e.g., steering or braking actuators or another controller discussed herein with respect to the self-machine 700). For example, the path planning function 172 may include a configuration space manager, a free space manager, a reachability manager, and a path evaluator. The configuration space manager may manage a posture configuration space that represents a posture including a position and orientation of the self-machine in its environment. The free space manager and the reachability manager may process the posture configuration space to determine one or more paths for maneuvering from a current posture to a target posture in the posture configuration space based at least in part on the stitched image data 152. The path evaluator may determine one or more suggested paths or potential paths for the vehicle based at least on an evaluation by the reachability manager.

[0048] To generate the coordinated metadata image data 128 from the image data 110, the parameter coordination function 120 includes parameter blending 122 to blend the application of one or more of the image metadata parameters 112 on the images represented by the image data 110. The parameter blending 122 may apply a blend of one or more of the image metadata parameters 112 within each image frame of the image data 110 based on one or more metadata parameter blending definitions 124 that define, for example, algorithms for calculating boundary metadata parameter values ​​and / or selection and / or characteristics of parameter gain functions applied to perform parameter coordination between images represented in the image data 110 in preparation for stitching the images by the stitching module 150. For example, in some embodiments, the metadata parameter blending definitions 124 may include definitions for performing parameter coordination for one or more of white balance temperature, global tone mapping (GTM), exposure time, brightness, contrast, gamma, chroma, noise reduction, saturation, sharpness, lens distortion, and / or lens shading correction parameters, but are not limited thereto. In some embodiments, the parameter coordination function 120 may select (e.g., based on user input and / or input from another system of the vehicle 700 ) which metadata parameters from the image metadata parameters 112 are to be processed for parameter coordination, and control the parameter blending 122 to perform parameter coordination of the image data 110 using corresponding definitions from the metadata parameter blending definition 124 .

[0049] In some embodiments, parameter coordination function 120 performs parameter blending 122 to apply parameter coordination to frames represented by image data 110 on a per-image basis. Figure 2 , Figure 2 An example of parameter coordination between a first image 210 and a second image 220 represented by the image data 110 is shown. The two images share an overlap region 205 that defines an overlapping image boundary 206 between the first image 210 and the second image 220. In this example, parameter coordination between the first image 210 and the second image 220 is applied at the image boundary 206 between the center of the first image 210 (shown at 212) and the center of the second image 220 (shown at 222).

[0050] Assume that the width of the first image 210 is W1 pixels, and the width of the second image 220 is W2 pixels. Then, the center of the first image 210 can be defined as pixel column W1 / 2, and the center of the second image 220 can be defined as pixel column W2 / 2. In some embodiments, the image boundary 206 with respect to the first image 210 can be defined at pixel column W', where W' < W1, and this quantity takes into account the overlap between the first image 210 and the second image 220. The image boundary 206 with respect to the second image 220 can be defined at pixel column 0 of the second image 220, which can indicate that the columns of the second image 220 will be aligned with the pixel column W' at the image boundary 206.

[0051] To achieve parameter coordination between the first image 210 and the second image 220 with respect to the selected metadata parameter, the parameter mixer 122 can calculate the boundary metadata parameter value based on the algorithm defined by the metadata parameter mixing definition 124 for the selected metadata parameter. The boundary metadata parameter value can be calculated based on the average of the first metadata parameter value of the first image 210 and the second metadata parameter value of the second image 220. For example, for an implementation where the selected metadata parameter for parameter coordination is white balance, the parameter mixer 122 can input the first white balance parameter WB1 from the image metadata parameter 112 and input the second white balance parameter WB2 from the image metadata parameter 112. The first white balance parameter WB1 indicates the white balance temperature applied to the first image 210 when the first image 210 is captured, and the second white balance parameter WB2 indicates the white balance temperature applied to the second image 220 when the second image 220 is captured. Based on the seam or overlap region location, the boundary metadata parameter can be calculated using a weighted value of the first metadata parameter value of the first image 210 and the second metadata parameter value of the second image 220. For example, if the seam or overlap region location is closer to the first image, the weight of the metadata from the first image for the boundary metadata parameter is greater than the weight of the metadata from the second image. The seam and overlap regions can be of any shape, not limited to a straight line. In this case, different optimized mixing curves can be applied to each row. In some embodiments, different optimized mixing curves g1(x) and g2(x) can be applied to each row to account for camera lens radial distortion or parameter mixing in the radial direction. Although Figure 2 An example showing a case of horizontal overlap between images is presented, but this example is not intended to be limiting. For example, the overlap between two images can be horizontal, vertical, or in both horizontal and vertical directions.

[0052] The parameter gain function (shown at 230) of the first image 210 can be calibrated to have unity gain (gain of 1) at or near the center of the image (e.g., at pixel column W1 / 2), so that the white balance at that location in the first image is WB1. The parameter gain function (shown at 240) of the second image 220 can be calibrated to have unity gain (gain of 1) at or near the center of the image (e.g., at pixel column W2 / 2), so that the white balance at that location in the first image is WB2. The boundary metadata parameter value of the first image 210 can be calculated based on the average of WB1 and WB2, for example Where g1WB is the gain of the parameter gain function of the first image 210 applied to the white balance WB1 at column W' of the first image 210. The boundary metadata parameter value of the second image 220 can be calculated based on the average of WB1 and WB2, such as Where g2WB is the gain of the parametric gain function of the second image 220 applied to the white balance WB2 at the zero column of the second image 220. The resulting g1WB(W') and g2WB(0) boundary metadata parameter values ​​applied to the first image 210 and the second image 220 at the overlapping image boundary 206, respectively, adjust the corresponding images to be rendered using the same white balance color temperature within the overlapping region 205.

[0053] In addition, the parametric gain function 230 of the first image 210 produces a gradient in white balance color temperature by applying a gain of 1 to the white balance at the image center 212, applying a gain g1WB(W′) at the image boundary 206, and applying a smooth gradient in the white balance metadata parameters between the image center 212 and the image boundary 206 that follows the parametric gain function 230. Similarly, the parametric gain function 240 of the second image 220 produces a gradient in white balance color temperature by applying a gain of 1 to the white balance at the image center 222, applying a gain g2WB(0) at the image boundary 206, and applying a smooth gradient in the white balance metadata parameters between the image center 222 and the image boundary 206 that follows the parametric gain function 240.

[0054] In this way, pixels located on either side of the image boundary 206 where the first image 210 and the second image 220 are to be stitched are adjusted using the same boundary metadata parameter value, wherein the metadata parameter value smoothly follows the parameter gain function. During the stitching process, discontinuities in rendering parameters are avoided at the image boundary 206 where the images are blended together, thereby substantially avoiding corresponding artifacts in the resulting stitched image data 152.

[0055] The specific parameter gain function and / or algorithm used to calculate the boundary metadata parameter value may vary depending on the metadata parameter to which parameter coordination is applied, and / or may be specified by the corresponding metadata parameter blend definition selected by the parameter coordination function 120 from the metadata parameter blend definition 124. For example, for parameter coordination applied to a global tone mapping (GTM) metadata parameter, the metadata parameter blend definition 124 may provide an algorithm for calculating the boundary metadata parameter value based on a GTM 256-point lookup table (LUT) coordinated with LUT1GTM and LUT2GTM, as described previously herein.

[0056] The parameter gain function specified by the metadata parameter blend definition 124 and / or applied by the parameter blend 122 may be implemented using any function that produces the desired gradient of the metadata parameter blend. Possible parameter gain functions may include, for example, linear ramp functions and / or functions that produce more complex and / or non-linear curves, such as logistic function curves. In some embodiments, the parameter gain function applied by the parameter blend 122 may be selected and / or adjusted on an image-by-image basis based on the metadata parameters to be coordinated and / or the individual lighting conditions affecting each image. For example, in some embodiments, the parameter coordination function 120 may perform an ambient lighting assessment 126 on the images represented by the image data 110. If two adjacent images have significantly different brightness levels, the parameter blend 122 may adjust the steepness or rate of change of the parameter gain function (e.g., the logistic growth rate of a logistic function) to change the applied gain more quickly, thereby facilitating a more seamless parameter blend. In low light conditions, the parameter blend 122 may adjust the steepness or rate of change of the parameter gain function to achieve a more gradual rate smoothing. In some embodiments, the parameter gain function curve may be selected by the parameter coordination function 120 from a parameter gain function library 125 that includes predefined curves for coordinating different metadata parameters and / or coordinating based on different ambient lighting conditions.

[0057] Reference now Figure 3 , Figure 3 Parameter coordination applied to a set of images to render a stitched image (e.g., a 360-degree surround image) is shown. Figure 3 As shown, a 360-degree surround image may be formed from image data 110 representing four partially overlapping images captured by four cameras: a front image 312 captured by a front camera, a left image 314 captured by a left camera, a rear image 316 captured by a rear camera, and a right image 318 captured by a right camera. In this example, a parameter coordination function 120 may apply a parameter blend 122 to coordinate the right side of each image with the image to be stitched to its right, e.g., for Figure 2Similarly, the parameter coordination function 120 can apply parameter blending 122 to coordinate the left side of each image with the image to be stitched to its left, as for Figure 2 220 in the figure. The parameter coordination function 120 may output the resulting coordinated metadata image data 128 to the stitching module 150. When the coordinated images from the coordinated metadata image data 128 are stitched together, each image is coordinated across the seams formed with its adjacent images - thereby avoiding rendering parameter discontinuities that amplify artifacts in the stitched images. This metadata parameter coordination may be performed for each boundary of the image with overlapping adjacent images. For example, the coordination may be performed using adjacent images to the left and / or right of the image, adjacent images above and / or below the image, or any combination thereof. Figure 3 Metadata parameter coordination shown. In some embodiments, a set of metadata parameter coordination of an image may be applied sequentially and / or simultaneously by mathematically combining parameter gain functions and calculating the resulting combined parameter gain adjustment.

[0058] Figure 4 and Figure 5 Schematic diagrams are respectively shown of example metadata parameter coordination applied to image data 110 representing four overlapping images to produce stitched image data 152 representing a 360 degree surround view image.

[0059] exist Figure 4 , image 410 shows a stitching of four overlapping images where no metadata parameter coordination was applied prior to stitching to produce a 360-degree surround view image. In contrast, image 420 shows a stitching of the same four overlapping images where metadata parameter coordination (particularly white balance coordination and GTM coordination) was applied prior to stitching to produce a 360-degree surround view image. In image 410, visual artifacts 412a, 412b, 412c, and 412d are observable at image boundaries between the stitched images. For example, at visual artifact 412c, discontinuities at image boundaries can be observed in both the rendering of parking lot 414 and sky 416. In image 420, where white balance coordination and GTM coordination were applied to image data 110 prior to stitching, visual artifact 412c that would otherwise appear at 422 is greatly attenuated, and the images of parking lot 414 and sky 416 regions in image 420 appear relatively smooth in appearance.

[0060] exist Figure 5, image 510 shows a stitching of four overlapping images where no metadata parameter coordination was applied prior to stitching to produce a 360 degree surround view image. In contrast, image 520 shows a stitching of the same four overlapping images where metadata parameter coordination (particularly white balance coordination and GTM coordination) was applied prior to stitching to produce a 360 degree surround view image. In image 510, visual artifacts 512a, 512b, 512c, and 512d are observable at image boundaries between the stitched images. For example, at visual artifacts 512a, 512b, and 512c, discontinuities at image boundaries are observable in the rendering of parking lot 515 that are difficult to distinguish from the image of the legitimately rendered shadow in image 510. In image 520, where white balance coordination and GTM coordination were applied to image data 110 prior to stitching, visual artifacts 512a, 512b, 512c, and 512d are all significantly attenuated relative to the rendering of parking lot 515.

[0061] Figure 4 and Figure 5 In the two examples of metadata parameter coordination in FIG. 1 , during the stitching process, discontinuities in rendering parameters are avoided at the boundary areas where the original images are blended together, thereby substantially avoiding corresponding artifacts in the resulting stitched image.

[0062] Reference now Figure 6 , Figure 6 6 is a flowchart illustrating a method 600 for image stitching based on metadata coordination according to some embodiments of the present disclosure. Figure 6 The features and elements described in method 600 may be used in conjunction with, combined with, or substituted for elements of any other embodiment discussed herein, and vice versa. In addition, it should be understood that Figure 6 The functional, structural, and other descriptions of the elements of the embodiments described in the accompanying drawings may apply to the same or similarly named or described elements in any figures and / or embodiments described herein, and vice versa.

[0063] Each block of the method 600 described herein includes a computing process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in a memory. The methods can also be embodied as computer usable instructions stored on a computer storage medium. The methods can be provided by a stand-alone application, a service or a hosted service (standalone or in combination with another hosted service), or a plug-in to another product, to name a few. In addition, the method 600 is directed to Figure 1 The metadata coordinated image stitching system 100 is described by way of example. However, the methods may additionally or alternatively be performed by any one system or any combination of systems, including but not limited to the systems described herein.

[0064] As discussed in more detail herein, the method may include stitching a first image with a second image that at least partially overlaps the first image at an image boundary, the stitching being based at least on adjusting a metadata parameter of the first image by blending from a first value of the metadata parameter at a first location within the first image to a second value of the metadata parameter at a second location at the image boundary, wherein the second value is calculated based at least on the first value of the metadata parameter and a third value of the metadata parameter associated with the second image.

[0065] The method 600 includes, at block B602, determining a first metadata parameter value associated with rendering a first image and a second metadata parameter value associated with rendering a second image, wherein the first image at least partially overlaps the second image at an image boundary. The first metadata parameter value and the second metadata parameter value may be associated with rendering parameters, such as, but not limited to, white balance, tone mapping, exposure time, brightness, contrast, gamma, chroma, noise reduction, saturation, sharpness, color filter array (CFA) pattern, lens shading correction, lens distortion correction, focal length correction, barrel distortion correction, or pincushion distortion. In some embodiments, the first metadata parameter value and / or the second metadata parameter may be based on metadata communicated with the first image. For example, metadata parameters including rendering parameters selected by an image sensor may be communicated by the image sensor to the parameter coordination function as metadata embedded in and / or carried with the image data for each captured image frame of the image data. In some embodiments, one or more of the image metadata parameters may be communicated by the image sensor to the parameter coordination function via a channel separate from a channel carrying the image data.

[0066] The first image may be captured by a first camera, and the second image may be captured by a second camera having a field of view that overlaps with the first camera. In some embodiments, the first metadata parameter value and / or the second metadata parameter may be based on metadata associated with characteristics of the respective first camera or second camera that captured the first image or the second image. Thus, in some cases, one or more metadata parameters may represent static characteristics or settings associated with the image sensor that do not change with each captured image (e.g., lens distortion correction and / or sensor or pixel size). Rather than being transmitted to the parameter coordination function by the image sensor along with the image data or as a supplement thereto, these metadata parameters may be stored in a memory and recalled as needed by the parameter coordination function.

[0067] The method 600 includes, at block B604, calculating a boundary metadata parameter value based on at least the first metadata parameter value and the second metadata parameter value. In some embodiments, the boundary metadata parameter value may be calculated based on an average value calculated using the first metadata parameter value and the second metadata parameter value. To achieve parameter coordination between the first image and the second image with respect to a particular metadata parameter, the parameter blending function may calculate the boundary metadata parameter value based on an algorithm defined by the metadata parameter blending definition for the metadata parameter. The boundary metadata parameter value may be calculated based on an average value of the first metadata parameter value of the first image and the second metadata parameter value of the second image, as described herein with respect to Figure 1 and Figure 2 For example, for an implementation in which the metadata parameter selected by parameter coordination is white balance, parameter blending may input a first white balance parameter WB1 from the image metadata parameters and a second white balance parameter WB2 from the image metadata parameters, the first white balance parameter WB1 indicating a white balance temperature applied to the first image when the first image was captured, and the second white balance parameter WB2 indicating a white balance temperature applied to the second image when the second image was captured.

[0068] The method 600 includes, at block B606, adjusting at least a portion of the first image using a first parameter gain function based at least on a first metadata parameter value and a boundary metadata parameter value. The first parameter gain function may include, for example, a linear function, a non-linear curve function, an S-curve function, or a logistic curve function. The method 600 includes, at block B608, adjusting at least a portion of the second image using a second parameter gain function based at least on a second metadata parameter value and a boundary metadata parameter value. The second parameter gain function may include, for example, a linear function, a non-linear curve function, an S-curve function, or a logistic curve function. The parameter gain function may be used to gradually adjust the metadata parameter applied to the image from a location near or at the center of the image (where the initial metadata parameter value applied remains unchanged - e.g., a gain of 1) to a boundary edge between the images (where the metadata parameter applied is a boundary metadata parameter value). In this way, pixels located on either side of the boundary edge where the images are to be stitched will be adjusted using the same boundary metadata parameter value, where the metadata parameter value smoothly follows the parameter gain function. Thus, during the stitching process, discontinuities in rendering parameters are avoided at the boundary regions where the original images are blended together, thereby substantially avoiding corresponding artifacts in the resulting stitched image. The coordination between adjacent images before stitching may be applied to a single metadata parameter or to multiple metadata parameters.

[0069] In some embodiments, at least one of the first parameter gain function and the second parameter gain function can be adjusted based on ambient lighting condition data from a sensor. For example, in some embodiments, the parameter coordination function can perform an ambient lighting assessment on an image represented by the image data. If two adjacent images have significantly different brightness levels, the parameter blending can adjust the steepness or rate of change of the parameter gain function (e.g., the logistic growth rate of the logistic function) to change the applied gain more quickly, thereby facilitating a more seamless parameter blending. In low light conditions, the parameter blending can adjust the steepness or rate of change of the parameter gain function to achieve a more gradual rate smoothing. In some embodiments, the parameter coordination function can select a parameter gain function curve from a parameter gain function library, which includes predefined curves for coordinating different metadata parameters and / or coordinating based on different ambient lighting conditions.

[0070] Method 600 includes at box B610: stitching the first image with the second image at the image boundary. The second image may include, for example, a surround stitched image based at least on the first image and the second image, a fisheye stitched image based at least on the first image and the second image, and / or a panoramic view stitched image based at least on the first image and the second image. In some embodiments, stitching may include stitching the first image with the second image based on an overlapping image boundary between the first image and the second image, and / or stitching the first image with the third image based on an overlapping image boundary between the first image and the third image. In some embodiments, the stitching module may stitch frames of the coordinated metadata image data into stitched image data (e.g., a 360° surround image, a wide-angle image, a fisheye image, and / or a panoramic image) using any stitching technique. In some embodiments, the stitched image data may be further processed using image processing based on one or more image processing techniques.

[0071] The rendering module 160 may render a visualization 165 of at least a portion of the stitched image data 152 (e.g., on a monitor visible to an occupant or operator of the self-object or self-participant). In some embodiments, the rendering module 160 projects the stitched image data 152 or a portion thereof onto a 3D representation of a 3D environment (e.g., a 3D bowl that models the 3D environment), renders a view of the projected stitched image data 152 from the perspective of a virtual camera, and / or causes the rendered view to be presented as the visualization 165. The stitched image data 152 may also (or alternatively) be used by one or more downstream navigation components of the self-machine, such as the controller 736 discussed below. For example, the downstream navigation component may implement functionality such as object avoidance navigation functionality and / or a world model manager, a path planner, a control component, a localization component, an obstacle avoidance component, an actuation component, etc., to perform operations for controlling the self-machine in the environment. In some embodiments, the downstream navigation component may include one or more deep neural networks (DNNs) that generate one or more predictions and / or inferences about the 3D environment based at least on the stitched image data.

[0072] The systems and methods described herein may be used by, but are not limited to, non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more adaptive driver assistance systems (ADAS)), manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, aircraft, boats, space shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, engineering vehicles, trains, underwater vehicles, remotely controlled vehicles (e.g., drones), and / or other vehicle types. In addition, the systems and methods described herein may be used for a variety of purposes, such as, but not limited to, for machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or participant simulation and / or digital twins, data center processing, conversational AI (artificial intelligence), light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of 3D assets, cloud computing, generative AI, and / or any other suitable application.

[0073] The disclosed embodiments can be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, navigation systems, smart area monitoring systems, systems performing deep learning operations, systems performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems including one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least in part in a data center, systems performing conversational AI operations, systems for implementing one or more language models (e.g., one or more large language models (LLMs), systems performing light transport simulations, systems performing collaborative content creation of 3D assets, systems implemented at least in part using cloud computing resources, and / or other types of systems.

[0074] Example Autonomous Vehicle

[0075] Fig. 7A7 is an illustration of an example autonomous vehicle 700 according to some embodiments of the present disclosure. Autonomous vehicle 700 (alternatively referred to herein as "vehicle 700") may include, but is not limited to, a passenger vehicle such as a car, a truck, a bus, an emergency vehicle, a shuttle, an electric or motorized bicycle, a motorcycle, a fire truck, a police car, an ambulance, a boat, an engineering vehicle, an underwater vessel, a robotic vehicle, a drone, an airplane, a vehicle coupled to a trailer (e.g., a semi-tractor semi-trailer truck for hauling freight), and / or other types of vehicles (e.g., unmanned and / or capable of accommodating one or more passengers). Autonomous vehicles are generally described in terms of automation levels defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (Standard No. J3016-201806, issued on June 15, 2018, Standard No. J3016-201609, issued on September 30, 2016, and previous and future versions of the standard). The vehicle 700 is capable of implementing one or more of the functions of Level 3-5 that meet the autonomous driving level. The vehicle 700 is capable of implementing one or more of the functions of Level 1-5 that meet the autonomous driving level. For example, depending on the embodiment, the vehicle 700 is capable of driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). The term "autonomy" as used herein may include any and / or all types of autonomy of 700 or other machines, such as full autonomy, high autonomy, conditional autonomy, partial autonomy, assisted autonomy, semi-autonomy, primary autonomy, or other names.

[0076] The vehicle 700 may include components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. The vehicle 700 may include a propulsion system 750, such as an internal combustion engine, a hybrid power plant, an all-electric engine, and / or another type of propulsion system. The propulsion system 750 may be connected to a drive train of the vehicle 700, which may include a transmission, to achieve propulsion of the vehicle 700. The propulsion system 750 may be controlled in response to receiving a signal from a throttle / accelerator 752.

[0077] A steering system 754, which may include a steering wheel, may be used to steer the vehicle 700 (e.g., along a desired path or route) when the propulsion system 750 is operating (e.g., when the vehicle is in motion). The steering system 754 may receive signals from a steering actuator 756. For fully automated (Level 5) functionality, a steering wheel may be optional.

[0078] Brake sensor system 746 may be used to operate vehicle brakes in response to receiving signals from brake actuator 748 and / or brake sensors.

[0079] May include one or more system on chip (SoC) 704 ( Figure 7C ) and / or one or more controllers 736 of one or more GPUs can provide signals (e.g., representing commands) to one or more components and / or systems of the vehicle 700. For example, the one or more controllers can send signals to operate vehicle brakes via one or more brake actuators 748, operate the steering system 754 via one or more steering actuators 756, and operate the propulsion system 750 via one or more throttles / accelerators 752. The one or more controllers 736 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operating commands (e.g., signals representing commands) to achieve autonomous driving and / or assist a human driver in driving the vehicle 700. The one or more controllers 736 may include a first controller 736 for autonomous driving functions, a second controller 736 for functional safety functions, a third controller 736 for artificial intelligence functions (e.g., computer vision), a fourth controller 736 for infotainment functions, a fifth controller 736 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 736 may handle two or more of the above functions, two or more controllers 736 may handle a single function, and / or any combination thereof. In some embodiments, the controller 736 may control one or more operations of one or more components and / or systems of the vehicle 700 described herein based on the stitched image data 152.

[0080] The one or more controllers 736 may provide signals for controlling one or more components and / or systems of the vehicle 700 in response to sensor data (eg, sensor input) received from one or more sensors. Sensor data may be received from, for example and without limitation, a global navigation satellite system ("GNSS") sensor 758 (e.g., a global positioning system sensor), a RADAR sensor 760, an ultrasonic sensor 762, a LIDAR sensor 764, an inertial measurement unit (IMU) sensor 766 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 796, a stereo camera 768, a wide angle camera 770 (e.g., a fisheye camera), an infrared camera 772, a surround camera 774 (e.g., a 360 degree camera), a long-range and / or mid-range camera 798, a speed sensor 744 (e.g., for measuring the velocity of the vehicle 700), a vibration sensor 742, a steering sensor 740, a brake sensor (e.g., as part of a brake sensor system 546), one or more occupant monitoring system (OMS) sensors 701 (e.g., one or more interior cameras), and / or other sensor types. In some embodiments, one or more aspects of the parameter coordination function 120 and / or the stitching module 150 may be performed, at least in part, by one or more controllers 736.

[0081] One or more of the controllers 736 may receive inputs (e.g., represented by input data) from the instrument cluster 732 of the vehicle 700 and provide outputs (e.g., represented by output data, display data, etc.) via a human machine interface (HMI) display 734, an audible annunciator, a speaker, and / or via other components of the vehicle 700. These outputs may include information such as vehicle speed, velocity, time, map data (e.g., Figure 7C The HMI display 734 may include information such as a high-definition (HD) map 722 of the vehicle 700, location data (e.g., the location of the vehicle 700 on the map), directions, locations of other vehicles (e.g., an occupancy grid), information about objects and object states as sensed by the controller 736, and the like. For example, the HMI display 734 may display information about the presence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, leaving 34B in two miles, etc.). In some embodiments, the visualization 165 may be presented on the HMI display 734.

[0082] The vehicle 700 further includes a network interface 724 that can communicate via one or more networks using one or more wireless antennas 726 and / or a modem. For example, the network interface 724 may be capable of communicating via Long Term Evolution (LTE), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Global System for Mobile Communications (GSM), IMT-CDMA Multi-Carrier (CDMA2000), etc. The one or more wireless antennas 726 may also enable communication between objects (e.g., vehicles, mobile devices, etc.) in the environment using one or more local area networks such as Bluetooth, Bluetooth Low Energy (LE), Z-wave, ZigBee, etc. and / or one or more low power wide area networks (LPWAN) such as LoRaWAN, SigFox, etc.

[0083] Figure 7B For use according to some embodiments of the present disclosure Fig. 7A 700. The cameras and respective fields of view are an example embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included, and / or the cameras may be located at different locations on the vehicle 700. In some embodiments, the image sensor 105 may use information about Figure 7B The discussed one or more cameras are implemented.

[0084] The camera type for the camera may include, but is not limited to, a digital camera that may be suitable for use with components and / or systems of the vehicle 700. The camera may operate at an automotive safety integrity level (ASIL) B and / or at another ASIL. The camera type may have any image capture rate, such as 60 frames per second (fps), 120fps, 240fps, etc., depending on the embodiment. The camera may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a red-white-white-white (RCCC) color filter array, a red-white-white-blue (RCCB) color filter array, a red-blue-green-white (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, a clear pixel camera such as a camera with an RCCC, RCCB, and / or RBGC color filter array may be used in an effort to improve light sensitivity.

[0085] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-function monocular camera can be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. One or more of the cameras (e.g., all of the cameras) can simultaneously record and provide image data (e.g., video).

[0086] One or more of the cameras may be mounted in a mounting assembly such as a custom designed (three-dimensional (3D) printed) assembly to cut off stray light and reflections from within the car (e.g., reflections from the dashboard reflected in the windshield mirror) that may interfere with the camera's image data capture capabilities. With respect to the wing mirror mounting assembly, the wing mirror assembly may be custom 3D printed so that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras may be integrated into the wing mirror. For side view cameras, one or more cameras may also be integrated into the four pillars at each corner of the cab.

[0087] A camera (e.g., a front-facing camera) having a field of view that includes a portion of the environment in front of the vehicle 700 can be used for surround view to help identify the forward path and obstacles, as well as assist in providing information critical to generating an occupancy grid and / or determining a preferred vehicle path with the help of one or more controllers 736 and / or control SoCs. The front-facing camera can be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. The front-facing camera can also be used for ADAS functions and systems, including lane departure warning (LDW), autonomous cruise control (ACC), and / or other functions such as traffic sign recognition.

[0088] A variety of cameras may be used in the front-facing configuration, including, for example, a monocular camera platform including a complementary metal oxide semiconductor (CMOS) color imager. Another example may be a wide-angle camera 770, which may be used to sense objects (e.g., pedestrians, intersection traffic, or bicycles) entering the field of view from the periphery. Although Figure 7B Only one wide-angle camera is shown in the figure, but there may be any number (including zero) of wide-angle cameras 770 on the vehicle 700. In addition, any number of long-range cameras 798 (e.g., a long-view stereo camera pair) may be used for depth-based object detection, especially for objects for which a neural network has not been trained. Long-range cameras 798 may also be used for object detection and classification and basic object tracking.

[0089] Any number of stereo cameras 768 may also be included in the front configuration. In at least one embodiment, one or more stereo cameras 768 may include an integrated control unit including an expandable processing unit that may provide a multi-core microprocessor and programmable logic (FPGA) with an integrated controller area network (CAN) or Ethernet interface on a single chip. Such a unit can be used to generate a 3D map of the vehicle environment, including distance estimates for all points in the image. Alternative stereo cameras 768 may include a compact stereo vision sensor that may include two camera lenses (one on each side) and an image processing chip that can measure the distance from the vehicle to the target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. In addition to those described herein or alternatively, other types of stereo cameras 768 may be used.

[0090] A camera having a field of view of a portion of the environment including the sides of the vehicle 700 (e.g., a side view camera) can be used for surround viewing, providing information used to create and update the occupancy grid and generate side impact collision warnings. Figure 7B The four surround cameras 774 shown in FIG. 7 can be placed on the vehicle 700. The surround cameras 774 can include a wide-angle camera 770, a fisheye camera, a 360-degree camera, and / or the like. For example, the four fisheye cameras can be placed in front, behind, and on the sides of the vehicle. In an alternative arrangement, the vehicle can use three surround cameras 774 (e.g., left, right, and rear), and can utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround camera.

[0091] A camera having a field of view that includes a portion of the environment behind the vehicle 700 (e.g., a rear view camera) can be used to assist with parking, surround view, rear collision warning, and creating and updating occupancy grids. A variety of cameras can be used, including but not limited to cameras that are also suitable as front cameras as described herein (e.g., long-range and / or mid-range cameras 798, stereo cameras 768, infrared cameras 772, etc.).

[0092] A camera (e.g., one or more OMS sensors 701) whose field of view includes a portion of the interior environment in the cabin of the vehicle 700 may be used as part of an occupant monitoring system (OMS), such as, but not limited to, a driver monitoring system (DMS). For example, an OMS sensor (e.g., OMS sensor 701) may be used (e.g., by controller 736) to track the gaze direction, head posture, and / or blinking of an occupant and / or driver. The gaze information may be used to determine the attention level of an occupant or driver (e.g., to detect drowsiness, fatigue, and / or distraction), and / or to take responsive actions to prevent harm to an occupant or operator. In some embodiments, data from the OMS sensor may be used to implement gaze control operations triggered by a driver and / or non-driver passenger, such as, but not limited to, adjusting cabin temperature and / or airflow, opening and closing windows, controlling cabin lighting, controlling an entertainment system, adjusting rearview mirrors, adjusting seat positions, and / or other operations. In some embodiments, the OMS may be used for applications such as determining when an object and / or passenger remains in the cabin (e.g., by detecting the presence of a passenger after the driver leaves the vehicle).

[0093] Figure 7C For use according to some embodiments of the present disclosure Fig. 7A Block diagram of an example system architecture for an example autonomous vehicle 700. It should be understood that this arrangement and other arrangements described herein are set forth merely as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used in addition to or in place of those shown, and some elements may be omitted entirely. Further, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in combination with other components, and in any appropriate combination and location. The various functions described herein as being performed by an entity may be implemented by hardware, firmware, and / or software. For example, the various functions may be implemented by a processor executing instructions stored in a memory.

[0094] Figure 7C Each of the components, features, and systems of the vehicle 700 is illustrated as being connected via a bus 702. The bus 702 may include a controller area network (CAN) data interface (alternatively, referred to herein as a "CAN bus"). The CAN may be a network inside the vehicle 700 that assists in controlling various features and functions of the vehicle 700, such as the actuation of brakes, acceleration, braking, steering, windshield wipers, and the like. The CAN bus may be configured to have tens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus may be read to find steering wheel angle, ground speed, engine speed per minute (RPM), button position, and / or other vehicle status indicators. The CAN bus may be ASIL B compatible.

[0095] Although bus 702 is described as a CAN bus here, this is not intended to be limiting. For example, in addition to or alternatively to the CAN bus, FlexRay and / or Ethernet can be used. In addition, although bus 702 is represented by a single line, this is not intended to be limiting. For example, there can be any number of buses 702, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 702 can be used to perform different functions, and / or can be used for redundancy. For example, a first bus 702 can be used for a collision avoidance function, and a second bus 702 can be used for drive control. In any example, each bus 702 can communicate with any component of the vehicle 700, and two or more buses 702 can communicate with the same component. In some examples, each SoC 704, each controller 736, and / or each computer in the vehicle can have access to the same input data (e.g., input from a sensor of the vehicle 700), and can be connected to a common bus such as a CAN bus.

[0096] The vehicle 700 may include one or more controllers 736, such as those described herein. Fig. 7A Those controllers described. Controller 736 can be used for a variety of functions. Controller 736 can be coupled to any other different components and systems of vehicle 700, and can be used for control of vehicle 700, artificial intelligence of vehicle 700, infotainment for vehicle 700, and / or the like.

[0097] The vehicle 700 may include one or more system on chip (SoC) 704. The SoC 704 may include a CPU 706, a GPU 708, a processor 710, a cache 712, an accelerator 714, a data store 716, and / or other components and features not shown. The SoC 704 may be used to control the vehicle 700 in a variety of platforms and systems. For example, one or more SoCs 704 may be combined with an HD map 722 in a system (e.g., a system of the vehicle 700), and the HD map may be downloaded from one or more servers (e.g., a server) via a network interface 724. Fig.7D One or more servers 778) obtain map refreshes and / or updates. In some embodiments, one or more aspects of the parameter coordination function 120 and / or the stitching module 150 can be at least partially executed by one or more SoCs 704.

[0098] The CPU 706 may include a CPU cluster or a CPU complex (alternatively, referred to herein as a "CCPLEX"). The CPU 706 may include multiple cores and / or L2 caches. For example, in some embodiments, the CPU 706 may include eight cores in a coherent multiprocessor configuration. In some embodiments, the CPU 706 may include four dual-core clusters, each of which has a dedicated L2 cache (e.g., a 2MB L2 cache). The CPU 706 (e.g., CCPLEX) may be configured to support simultaneous cluster operations, such that any combination of clusters of the CPU 706 can be active at any given time.

[0099] CPU 706 may implement power management capabilities including one or more of the following features: each hardware block may be automatically clock gated when idle to save dynamic power; each core clock may be gated when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core may be independently power gated; each core cluster may be independently clock gated when all cores are clock gated or power gated; and / or each core cluster may be independently power gated when all cores are power gated. CPU 706 may further implement an enhanced algorithm for managing power states, in which allowed power states and expected wake-up times are specified, and hardware / microcode determines the optimal power state to enter for the core, cluster, and CCPLEX. The processing core may support a simplified power state entry sequence in software, with the work being offloaded to the microcode.

[0100] The GPU 708 may include an integrated GPU (alternatively, referred to herein as an "iGPU"). The GPU 708 may be programmable and efficient for parallel workloads. In some examples, the GPU 708 may use an enhanced tensor instruction set. The GPU 708 may include one or more streaming microprocessors, each of which may include an L1 cache (e.g., an L1 cache with at least 96KB storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB storage capacity). In some embodiments, the GPU 708 may include at least eight streaming microprocessors. The GPU 708 may use a computing application programming interface (API). In addition, the GPU 708 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0101] In the case of automotive and embedded use, GPU 708 can be power optimized to achieve optimal performance. For example, GPU 708 can be manufactured on fin field effect transistors (FinFETs). However, this is not intended to be limiting, and GPU 708 can be manufactured using other semiconductor manufacturing processes. Each streaming microprocessor can merge several mixed precision processing cores divided into multiple blocks. For example and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed precision NVIDIA tensor cores for deep learning matrix arithmetic, L0 instruction cache, thread beam (warp) scheduler, dispatch unit and / or 64KB register file. In addition, the streaming microprocessor can include independent parallel integer and floating point data paths to provide efficient execution of workloads using a mix of calculations and addressing calculations. The streaming microprocessor may include independent thread scheduling capabilities to allow for finer-grained synchronization and collaboration between parallel threads. A streaming microprocessor may include a combined L1 data cache and shared memory unit to increase performance while simplifying programming.

[0102] GPU 708 can include high bandwidth memory (HBM) and / or a 16GB HBM2 memory subsystem that provides a peak memory bandwidth of approximately 900 GB / s in some examples. In some examples, synchronous graphics random access memory (SGRAM), such as fifth generation graphics double data rate synchronous random access memory (GDDR5), can be used in addition to or in lieu of HBM memory.

[0103] The GPU 708 may include unified memory technology that includes access counters to allow memory pages to be more accurately migrated to the processor that accesses them most frequently, thereby improving the efficiency of memory ranges shared between processors. In some examples, address translation service (ATS) support can be used to allow the GPU 708 to directly access the CPU 706 page table. In such an example, when the GPU 708 memory management unit (MMU) experiences a miss, an address translation request can be transmitted to the CPU 706. In response, the CPU 706 can look up the virtual-physical mapping for the address in its page table and transmit the translation back to the GPU 708. In this way, unified memory technology can allow a single unified virtual address space to be used for the memory of both the CPU 706 and the GPU 708, thereby simplifying GPU 708 programming and porting applications to the GPU 708.

[0104] In addition, GPU 708 may include access counters that can track how often GPU 708 accesses the memory of other processors. Access counters can help ensure that memory pages are moved to the physical memory of the processor that accesses them most frequently.

[0105] SoC 704 may include any number of caches 712, including those described herein. For example, cache 712 may include an L3 cache available to both CPU 706 and GPU 708 (e.g., connected to both CPU 706 and GPU 708). Cache 712 may include a write-back cache that may track the state of a line, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, although smaller cache sizes may also be used.

[0106] The SoC 704 may include one or more arithmetic logic units (ALUs) that may be used to perform processing for any of a variety of tasks or operations with respect to the vehicle 700, such as processing a DNN. Additionally, the SoC 704 may include a floating point unit (FPU) or other math coprocessor or digital coprocessor type for performing mathematical operations within the system. For example, the SoC 704 may include one or more FPUs integrated as execution units within the CPU 706 and / or GPU 708.

[0107] SoC 704 may include one or more accelerators 714 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, SoC 704 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. The large on-chip memory (e.g., 4MB SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other calculations. The hardware acceleration cluster may be used to supplement GPU 708 and offload some tasks of GPU 708 (e.g., free up more cycles of GPU 708 for performing other tasks). As an example, accelerator 714 may be used for targeted workloads (e.g., perception, convolutional neural networks (CNNs), etc.) that are sufficiently stable to easily control acceleration. When used herein, the term "CNN" may include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).

[0108] The accelerator 714 (e.g., a hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA may include one or more tensor processing units (TPUs) that may be configured to provide an additional 10 trillion operations per second for deep learning applications and reasoning. The TPU may be an accelerator configured to perform image processing functions (e.g., for CNN, RCNN, etc.) and optimized for performing image processing functions. The DLA may be further optimized for a specific set of neural network types and floating point operations and reasoning. The design of the DLA may provide higher performance per millimeter than a general-purpose GPU, and far exceeds the performance of the CPU. The TPU may perform several functions, including a single instance convolution function, support, for example, INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.

[0109] The DLA can quickly and efficiently execute neural networks, particularly CNNs, on processed or unprocessed data for any of a wide variety of functions, such as, but not limited to: a CNN for object recognition and detection using data from a camera sensor; a CNN for distance estimation using data from a camera sensor; a CNN for emergency vehicle detection and identification and detection using data from a microphone; a CNN for facial recognition and vehicle owner identification using data from a camera sensor; and / or a CNN for safety and / or security-related events.

[0110] The DLA can perform any function of the GPU 708, and by using an inference accelerator, for example, the designer can target any function to either the DLA or the GPU 708. For example, the designer can focus the processing of CNNs and floating point operations on the DLA, and leave other functions to the GPU 708 and / or other accelerators 714.

[0111] The accelerator 714 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (PVA), which may be alternatively referred to herein as a computer vision accelerator. The PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA may provide a balance between performance and flexibility. For example, each PVA may include, for example and without limitation, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.

[0112] The RISC core can interact with an image sensor (e.g., an image sensor of any camera described herein), an image signal processor, and / or the like. Each of these RISC cores can include any amount of memory. Depending on the embodiment, the RISC core can use any of a number of protocols. In some examples, the RISC core can execute a real-time operating system (RTOS). The RISC core can be implemented using one or more integrated circuit devices, application specific integrated circuits (ASICs), and / or storage devices. For example, the RISC core can include an instruction cache and / or a tightly coupled RAM.

[0113] The DMA may enable components of the PVA to access system memory independently of the CPU 706. The DMA may support any number of features used to provide optimizations for the PVA, including, but not limited to, support for multi-dimensional addressing and / or circular addressing. In some examples, the DMA may support addressing in up to six or more dimensions, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.

[0114] The vector processor can be a programmable processor that can be designed to efficiently and flexibly perform programming for computer vision algorithms and provide signal processing capabilities. In some examples, the PVA can include a PVA core and two vector processing subsystem partitions. The PVA core can include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem can operate as the main processing engine of the PVA, and can include a vector processing unit (VPU), an instruction cache, and / or a vector memory (e.g., VMEM). The VPU core can include a digital signal processor, such as, for example, a single instruction multiple data (SIMD), a very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can enhance throughput and rate.

[0115] Each of the vector processors may include an instruction cache and may be coupled to a dedicated memory. As a result, in some examples, each of the vector processors may be configured to execute independently of the other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelism. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms on sequence images or portions of an image. Among other things, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each of these PVAs. In addition, the PVA may include additional error correction code (ECC) memory to enhance overall system security.

[0116] The accelerator 714 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and SRAM to provide high bandwidth, low latency SRAM for the accelerator 714. In some examples, the on-chip memory may include at least 4MB of SRAM consisting of, for example and not limited to, eight field-configurable memory blocks, which may be accessed by both the PVA and the DLA. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, a configuration circuit system, a controller, and a multiplexer. Any type of memory may be used. The PVA and DLA may access the memory via a backbone that provides high-speed memory access to the PVA and DLA. The backbone may include an on-chip computer vision network that interconnects the PVA and DLA to the memory (e.g., using APB).

[0117] The on-chip computer vision network may include an interface that determines that both the PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. Such an interface may provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface may comply with ISO 26262 or IEC 61508 standards, but other standards and protocols may also be used.

[0118] In some examples, SoC 704 may include a real-time ray tracing hardware accelerator such as described in U.S. Patent Application No. 16 / 101,232 filed on August 10, 2018. The real-time ray tracing hardware accelerator may be used to quickly and efficiently determine the position and range of objects (e.g., within a world model) in order to generate real-time visualization simulations for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for SONAR system simulations, for general wave propagation simulations, for comparison with LIDAR data for the purpose of positioning and / or other functions, and / or for other purposes. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing related operations.

[0119] The accelerator 714 (e.g., a hardware accelerator cluster) has a wide range of uses in autonomous driving. The PVA can be a programmable visual accelerator that can be used in key processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are a good match for algorithmic domains that require predictable processing, low power, and low latency. In other words, the PVA performs well on semi-dense or dense rule computations, and even on small data sets that require predictable runtimes with low latency and low power. Therefore, in the context of a platform for autonomous vehicles, the PVA is designed to run classic computer vision algorithms because they are effective at object detection and integer math.

[0120] For example, according to one embodiment of the technology, PVA is used to perform computer stereo vision. In some examples, semi-global matching based algorithms can be used, but this is not intended to be limiting. Many applications for Level 3-5 autonomous driving require instant motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.). PVA can perform computer stereo vision functions on inputs from two monocular cameras.

[0121] In some examples, PVA can be used to perform dense optical flow. According to the process raw RADAR data (e.g., using 4D fast Fourier transform) to provide processed RADAR. In other examples, PVA is used for time-of-flight depth processing, which is for example by processing raw time-of-flight data to provide processed time-of-flight data.

[0122] DLA can be used to run any type of network to enhance control and driving safety, including, for example, a neural network that outputs a confidence measure for each object detection. Such a confidence value can be interpreted as a probability, or as providing a relative "weight" of each detection compared to other detections. This confidence value enables the system to make further decisions about which detections should be considered true positive detections rather than false positive detections. For example, the system can set a threshold for confidence and only consider detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, a false positive detection will cause the vehicle to automatically perform emergency braking, which is obviously undesirable. Therefore, only the most confident detection should be considered a trigger for AEB. DLA can run a neural network for regressing confidence values. The neural network can take at least some subset of parameters as its input, such as bounding box dimensions, ground plane estimates obtained (e.g., from another subsystem), inertial measurement unit (IMU) sensor 766 output related to vehicle 700 orientation and distance, 3D position estimates of objects obtained from neural networks and / or other sensors (e.g., LIDAR sensor 764 or RADAR sensor 760), etc.

[0123] SoC 704 may include one or more data stores 716 (e.g., memory). Data store 716 may be on-chip memory of SoC 704 that may store neural networks to be executed on the GPU and / or DLA. In some examples, data store 716 may be large enough to store multiple instances of a neural network for redundancy and safety. Data store 716 may include an L2 or L3 cache 712. References to data store 716 may include references to memory associated with a PVA, DLA, and / or other accelerator 714 as described herein.

[0124] SoC 704 may include one or more processors 710 (e.g., embedded processors). Processor 710 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions and related safety implementations. The boot and power management processor may be part of the SoC 704 boot sequence and may provide runtime power management services. The boot power and management processor may provide clock and voltage programming, auxiliary system low power state transitions, SoC 704 thermal and temperature sensor management, and / or SoC 704 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to the temperature, and SoC 704 may use the ring oscillator to detect the temperature of CPU 706, GPU 708, and / or accelerator 714. If it is determined that the temperature exceeds a threshold, the boot and power management processor may enter a temperature fault routine and place SoC 704 in a lower power state and / or place the vehicle 700 in a driver safety parking mode (e.g., parking the vehicle 700 safely).

[0125] Processor 710 may further include a set of embedded processors that may be used as an audio processing engine. The audio processing engine may be an audio subsystem that allows for full hardware support for multi-channel audio through multiple interfaces and a wide range of flexible audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core having a digital signal processor with dedicated RAM.

[0126] The processor 710 may further include an always-on processor engine that may provide the necessary hardware features to support low-power sensor management and wake-up use cases. The always-on processor engine may include a processor core, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0127] The processor 710 may further include a safety cluster engine, which includes a dedicated processor subsystem that handles safety management of automotive applications. The safety cluster engine may include two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In safety mode, the two or more cores may operate in lockstep mode and act as a single core with comparison logic to detect any differences between their operations.

[0128] Processor 710 may further include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.

[0129] Processor 710 may further include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.

[0130] The processor 710 may include a video image compositer, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required for a video playback application to produce a final image for a player window. The video image compositer may perform lens distortion correction for the wide-angle camera 770, the surround camera 774, and / or for an in-cab surveillance camera sensor. The in-cab surveillance camera sensor is preferably monitored by a neural network running on another instance of the advanced SoC, configured to recognize in-cab events and respond accordingly. The in-cab system may perform lip reading to activate mobile phone service and place calls, dictate emails, change vehicle destinations, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Certain functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled in other circumstances.

[0131] The video image compositer may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, in the case of motion in the video, the noise reduction appropriately weights the spatial information and reduces the weight of information provided by adjacent frames. In the case where an image or portion of an image does not include motion, the temporal noise reduction performed by the video image compositer may use information from previous images to reduce noise in the current image.

[0132] The video image compositor may also be configured to perform stereoscopic rectification on the input stereoscopic footage frames. The video image compositor may further be used for user interface composition when the operating system desktop is in use and the GPU 708 does not need to continuously render new surfaces. Even when the GPU 708 is powered on and active for 3D rendering, the video image compositor may be used to offload the GPU 708 to improve performance and responsiveness.

[0133] The SoC 704 may further include a mobile industry processor interface (MIPI) camera serial interface for receiving video and input from a camera, a high-speed interface, and / or a video input block that may be used for camera and related pixel input functions. The SoC 704 may further include an input / output controller that may be controlled by software and may be used to receive I / O signals that are not committed to a specific role.

[0134] SoC 704 may further include a wide range of peripheral interfaces to enable communication with peripherals, audio codecs, power management and / or other devices. SoC 704 may be used to process data from cameras (connected via Gigabit multimedia serial links and Ethernet), sensors (e.g., LIDAR sensor 764, RADAR sensor 760, etc., which may be connected via Ethernet), data from bus 702 (e.g., speed of vehicle 700, steering wheel position, etc.), data from GNSS sensor 758 (connected via Ethernet or CAN bus). SoC 704 may further include dedicated high-performance mass storage controllers, which may include their own DMA engines, and which may be used to free up CPU 706 from routine data management tasks.

[0135] SoC 704 can be an end-to-end platform with a flexible architecture that spans automation levels 3-5, providing a comprehensive functional safety architecture that leverages and efficiently uses computer vision and ADAS technologies to achieve diversity and redundancy, together with deep learning tools to provide a platform for a flexible and reliable driving software stack. SoC 704 can be faster, more reliable, and even more energy efficient and space efficient than conventional systems. For example, when combined with CPU 706, GPU 708, and data storage 716, accelerator 714 can provide a fast and efficient platform for level 3-5 autonomous vehicles.

[0136] The technology thus provides capabilities and functionality that cannot be achieved by conventional systems. For example, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages ​​such as the C programming language to execute a wide variety of processing algorithms across a wide variety of visual data. However, CPUs often fail to meet the performance requirements of many computer vision applications, such as those related to, for example, execution time and power consumption. In particular, many CPUs are unable to execute complex object detection algorithms in real time, which is a requirement for in-vehicle ADAS applications and a requirement for practical Level 3-5 autonomous vehicles.

[0137] In contrast to conventional systems, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, the techniques described herein allow multiple neural networks to be executed simultaneously and / or sequentially, and the results to be combined together to achieve Level 3-5 autonomous driving functions. For example, a CNN executed on a DLA or dGPU (e.g., GPU 720) may include text and word recognition, allowing the supercomputer to read and understand traffic signs, including signs for which a neural network has not been specifically trained. The DLA may further include a neural network capable of recognizing, interpreting, and providing semantic understanding of the signs, and passing that semantic understanding to a path planning module running on the CPU complex.

[0138] As another example, as required for Level 3, 4, or 5 driving, multiple neural networks may be running simultaneously. For example, a warning sign consisting of "Caution: Flashing Lights Indicate Icing Conditions" along with electric lights may be interpreted by several neural networks independently or collectively. The sign itself may be recognized as a traffic sign by a deployed first neural network (e.g., a trained neural network), and the text "Flashing Lights Indicate Icing Conditions" may be interpreted by a deployed second neural network that informs the vehicle's path planning software (preferably executing on a CPU complex) that icing conditions exist when the flashing lights are detected. The flashing lights may be identified by operating a deployed third neural network over multiple frames that informs the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks may be running simultaneously, for example, within the DLA and / or on the GPU 708.

[0139] In some examples, a CNN for facial recognition and owner recognition can use data from the camera sensor to identify the presence of an authorized driver and / or owner of the vehicle 700. The always-on sensor processing engine can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in security mode, disable the vehicle when the owner leaves the vehicle. In this way, the SoC 704 provides security against theft and / or carjacking.

[0140] In another example, a CNN for emergency vehicle detection and identification can use data from microphone 796 to detect and identify emergency vehicle sirens. In contrast to conventional systems that use general classifiers to detect sirens and manually extract features, SoC 704 uses CNN to classify environmental and urban sounds and to classify visual data. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative closing rate of emergency vehicles (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area in which the vehicle is operating, as identified by the GNSS sensor 758. Thus, for example, when operating in Europe, the CNN will seek to detect European sirens, and when in the United States, the CNN will seek to identify only North American sirens. Once an emergency vehicle is detected, with the assistance of the ultrasonic sensor 762, the control program can be used to execute emergency vehicle safety routines to slow the vehicle, drive to the side of the road, stop the vehicle, and / or idle the vehicle until the emergency vehicle passes.

[0141] The vehicle may include a CPU 718 (e.g., a discrete CPU or dCPU) that may be coupled to the SoC 704 via a high-speed interconnect (e.g., PCIe). The CPU 718 may include, for example, an X86 processor. The CPU 718 may be used to perform any of a variety of functions, including, for example, arbitrating potentially inconsistent results between ADAS sensors and the SoC 704, and / or monitoring the status and health of the controller 736 and / or the infotainment SoC 730.

[0142] The vehicle 700 may include a GPU 720 (e.g., a discrete GPU or dGPU) that may be coupled to the SoC 704 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU 720 may provide additional artificial intelligence functionality, for example, by executing redundant and / or different neural networks, and may be used to train and / or update the neural network based on inputs (e.g., sensor data) from sensors of the vehicle 700.

[0143] The vehicle 700 may further include a network interface 724, which may include one or more wireless antennas 726 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). The network interface 724 can be used to enable wireless connections with the cloud (e.g., with a server 778 and / or other network devices), with other vehicles, and / or with computing devices (e.g., a passenger's client device) via the Internet. In order to communicate with other vehicles, a direct link can be established between the two vehicles, and / or an indirect link can be established (e.g., across a network and through the Internet). The direct link can be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide the vehicle 700 with information about vehicles approaching the vehicle 700 (e.g., vehicles in front of, to the side of, and / or behind the vehicle 700). This functionality can be part of the collaborative adaptive cruise control functionality of the vehicle 700.

[0144] The network interface 724 may include a SoC that provides modulation and demodulation functions and enables the controller 736 to communicate over a wireless network. The network interface 724 may include a radio frequency front end for up-conversion from baseband to radio frequency and down-conversion from radio frequency to baseband. The frequency conversion may be performed by a known process and / or may be performed using a super-heterodyne process. In some examples, the radio frequency front end function may be provided by a separate chip. The network interface may include a wireless function for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0145] The vehicle 700 may further include a data store 728, which may include off-chip storage (e.g., outside the SoC 704). The data store 728 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, a hard disk, and / or other components and / or devices that can store at least one bit of data.

[0146] The vehicle 700 may further include a GNSS sensor 758. The GNSS sensor 758 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) is used to assist in mapping, perception, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 758 may be used, including, for example and without limitation, GPS using a USB connector with an Ethernet to serial (RS-232) bridge.

[0147] The vehicle 700 may further include a RADAR sensor 760. The RADAR sensor 760 may be used by the vehicle 700 for remote vehicle detection even in darkness and / or inclement weather conditions. The RADAR functional safety level may be ASILB. The RADAR sensor 760 may use CAN and / or bus 702 (e.g., to transmit data generated by the RADAR sensor 760) for control and access to object tracking data, accessing Ethernet in some examples to access raw data. A variety of RADAR sensor types may be used. For example and without limitation, the RADAR sensor 760 may be suitable for front, rear, and side RADAR use. In some examples, a pulse Doppler RADAR sensor is used.

[0148] The RADAR sensor 760 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, and the like. In some examples, the long-range RADAR may be used for adaptive cruise control functions. The long-range RADAR system may provide a wide field of view (e.g., within a range of 250m) achieved through two or more independent browsing. The RADAR sensor 760 may help distinguish between static objects and moving objects, and may be used by the ADAS system for emergency braking assistance and forward collision warnings. The long-range RADAR sensor may include a single-station multimode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the central four antennas may create a focused beam pattern designed to record the surroundings of the vehicle 700 at a higher rate with minimal traffic interference from adjacent lanes. The other two antennas may expand the field of view, making it possible to quickly detect vehicles entering or leaving the lane of the vehicle 700.

[0149] As an example, a medium-range RADAR system may include a range of up to 760m (front) or 80m (rear) and a field of view of up to 42 degrees (front) or 750 degrees (rear). A short-range RADAR system may include, but is not limited to, a RADAR sensor designed to be mounted on both ends of a rear bumper. When mounted on both ends of a rear bumper, such a RADAR sensor system may create two beams that continuously monitor the blind spots behind and beside the vehicle.

[0150] Short-range RADAR systems can be used in ADAS systems for blind spot detection and / or lane change assistance.

[0151] The vehicle 700 may further include an ultrasonic sensor 762. The ultrasonic sensor 762, which may be placed on the front, rear, and / or sides of the vehicle 700, may be used for parking assistance and / or creating and updating an occupancy grid. A variety of ultrasonic sensors 762 may be used, and different ultrasonic sensors 762 may be used for different detection ranges (e.g., 2.5 m, 4 m). The ultrasonic sensor 762 may operate at a functional safety level of ASIL B.

[0152] The vehicle 700 may include a LIDAR sensor 764. The LIDAR sensor 764 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 764 may be ASIL B for functional safety level. In some examples, the vehicle 700 may include multiple LIDAR sensors 764 (e.g., two, four, six, etc.) that may use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).

[0153] In some examples, the LIDAR sensor 764 may be able to provide a list of objects and their distances for a 360-degree field of view. Commercially available LIDAR sensors 764 may have, for example, an advertised range of approximately 700m, an accuracy of 2cm-3cm, and support for 700Mbps Ethernet connections. In some examples, one or more non-protruding LIDAR sensors 764 may be used. In such examples, the LIDAR sensor 764 may be implemented as a small device that can be embedded in the front, back, side, and / or corner of the vehicle 700. In such examples, the LIDAR sensor 764 may provide a field of view of up to 120 degrees horizontally and 35 degrees vertically, with a range of 200m, even for low reflectivity objects. The front-mounted LIDAR sensor 764 may be configured for a horizontal field of view between 45 and 135 degrees.

[0154] In some examples, LIDAR technologies such as 3D flash LIDAR may also be used. 3D flash LIDAR uses a flash of laser as an emission source to illuminate the vehicle's surroundings up to about 200m. The flash LIDAR unit includes a receiver that records the laser pulse transmission time and the reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LIDAR can allow highly accurate and distortion-free images of the surrounding environment to be generated with each laser flash. In some examples, four flash LIDAR sensors can be deployed, one on each side of the vehicle 700. Available 3D flash LIDAR systems include solid-state 3D staring array LIDAR cameras (e.g., non-browsing LIDAR devices) with no moving parts other than fans. The flash LIDAR device can use 5 nanosecond Class I (eye-safe) laser pulses per frame, and can capture the reflected laser light in the form of a 3D range point cloud and co-registered intensity data. By using flash LIDAR, and because flash LIDAR is a solid-state device with no moving parts, the LIDAR sensor 764 may be less susceptible to motion blur, vibration, and / or shock.

[0155] The vehicle may further include an IMU sensor 766. In some examples, the IMU sensor 766 may be located at the center of the rear axle of the vehicle 700. The IMU sensor 766 may include, for example and without limitation, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, for example, in a six-axis application, the IMU sensor 766 may include an accelerometer and a gyroscope, and in a nine-axis application, the IMU sensor 766 may include an accelerometer, a gyroscope, and a magnetometer.

[0156] In some embodiments, the IMU sensor 766 can be implemented as a miniature high-performance GPS-assisted inertial navigation system (GPS / INS) that combines a micro-electromechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filter algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 766 can enable the vehicle 700 to estimate heading by directly observing and correlating velocity changes from the GPS to the IMU sensor 766 without the need for input from a magnetic sensor. In some examples, the IMU sensor 766 and the GNSS sensor 758 can be combined into a single integrated unit.

[0157] The vehicle may include a microphone 796 positioned in and / or around the vehicle 700. The microphone 796 may be used for, among other things, emergency vehicle detection and identification.

[0158] The vehicle may further include any number of camera types, including a stereo camera 768, a wide angle camera 770, an infrared camera 772, a surround camera 774, a long-range and / or mid-range camera 798, and / or other camera types. These cameras may be used to capture image data around the entire periphery of the vehicle 700. The type of camera used depends on the embodiment and the requirements of the vehicle 700, and any combination of camera types may be used to provide the necessary coverage around the vehicle 700. In addition, the number of cameras may vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and not limitation, the cameras may support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras described herein may include a 10GbE ... Fig. 7A and Figure 7B Described in more detail.

[0159] The vehicle 700 may further include a vibration sensor 742. The vibration sensor 742 may measure vibrations of components of the vehicle, such as axles. For example, changes in vibration may indicate changes in the road surface. In another example, when two or more vibration sensors 742 are used, the difference between the vibrations may be used to determine friction or slippage of the road surface (e.g., when there is a vibration difference between a powered drive shaft and a free-spinning shaft).

[0160] The vehicle 700 may include an ADAS system 738. In some examples, the ADAS system 738 may include a SoC. The ADAS system 738 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.

[0161] The ACC system may use a RADAR sensor 760, a LIDAR sensor 764, and / or a camera. The ACC system may include a longitudinal ACC and / or a lateral ACC. The longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the vehicle 700, and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle in front. The lateral ACC performs distance keeping and suggests that the vehicle 700 change lanes when necessary. The lateral ACC is related to other ADAS applications such as LCA and CWS.

[0162] CACC uses information from other vehicles, which can be received from other vehicles indirectly via a wireless link or through a network connection (e.g., through the Internet) via a network interface 724 and / or a wireless antenna 726. A direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while an indirect link can be an infrastructure-to-vehicle (I2V) communication link. Typically, the V2V communication concept provides information about the vehicle immediately ahead (e.g., the vehicle immediately ahead of the vehicle 700 and in the same lane as it), while the I2V communication concept provides information about traffic farther ahead. The CACC system may include either or both of the I2V and V2V information sources. Given information about the vehicle ahead of the vehicle 700, CACC can be more reliable, and it is possible to improve the smoothness of traffic flow and reduce road congestion.

[0163] The FCW system is designed to alert the driver to hazards so that the driver can take corrective action. The FCW system uses a front camera and / or RADAR sensor 760 coupled to a dedicated processor, DSP, FPGA and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker and / or vibrating components. The FCW system can provide warnings in the form of, for example, sound, visual warnings, vibrations and / or rapid brake pulses.

[0164] The AEB system detects an impending front collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within specified time or distance parameters. The AEB system can use a front camera and / or RADAR sensor 760 coupled to a dedicated processor, DSP, FPGA and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, the AEB system can automatically apply the brakes in an effort to prevent or at least mitigate the effects of the predicted collision. The AEB system may include technologies such as dynamic brake support and / or collision approach braking.

[0165] The LDW system provides visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle 700 crosses a lane marking. When the driver indicates an intention to leave the lane, by activating a turn signal, the LDW system is not activated. The LDW system may use a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.

[0166] The LKA system is a variation of the LDW system. If the vehicle 700 begins to leave the lane, the LKA system provides steering input or braking to correct the vehicle 700.

[0167] The BSW system detects and warns the driver of vehicles in the car's blind spot. The BSW system can provide visual, auditory and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses a turn signal. The BSW system can use a rear-facing camera and / or RADAR sensor 760 coupled to a dedicated processor, DSP, FPGA and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker and / or vibration components.

[0168] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the range of the rear camera while the vehicle 700 is in reverse. Some RCTW systems include AEB to ensure that the vehicle brakes are applied to avoid a crash. The RCTW system can use one or more rear RADAR sensors 760 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.

[0169] Conventional ADAS systems may be prone to false positive results, which may annoy and distract the driver, but are typically not catastrophic because the ADAS system alerts the driver and allows the driver to decide whether the safety condition really exists and take action accordingly. However, in the autonomous vehicle 700, in the case of conflicting results, the vehicle 700 itself must decide whether to pay attention to the results from the main computer or the auxiliary computer (e.g., the first controller 736 or the second controller 736). For example, in some embodiments, the ADAS system 738 can be a backup and / or auxiliary computer for providing perception information to the backup computer rationality module. The backup computer rationality monitor can run redundant and diverse software on hardware components to detect faults in perception and dynamic driving tasks. The output from the ADAS system 738 can be provided to the supervisory MCU. If the outputs from the main computer and the auxiliary computer conflict, the supervisory MCU must determine how to coordinate the conflict to ensure safe operation.

[0170] In some examples, the master computer can be configured to provide a confidence score to the supervisory MCU, indicating the master computer's confidence in the selected result. If the confidence score exceeds a threshold, the supervisory MCU can follow the direction of the master computer, regardless of whether the auxiliary computer provides conflicting or inconsistent results. In the case where the confidence score does not meet the threshold and the master computer and the auxiliary computer indicate different results (e.g., conflicts), the supervisory MCU can arbitrate between these computers to determine the appropriate result.

[0171] The supervisory MCU may be configured to run a neural network that is trained and configured to determine conditions under which the auxiliary computer provides a false alarm based on outputs from the primary computer and the auxiliary computer. Thus, the neural network in the supervisory MCU may learn when the output of the auxiliary computer may be trusted and when it may not. For example, when the auxiliary computer is a RADAR-based FCW system, the neural network in the supervisory MCU may learn when the FCW system is identifying a metal object that is not actually dangerous, such as a drainage grate or manhole cover that triggers an alarm. Similarly, when the auxiliary computer is a camera-based LDW system, the neural network in the supervisory MCU may learn to ignore the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In embodiments that include a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or GPU suitable for running the neural network with associated memory. In a preferred embodiment, the supervisory MCU may include and / or be included as a component of the SoC 704.

[0172] In other examples, the ADAS system 738 may include an auxiliary computer that uses traditional computer vision rules to perform ADAS functions. In this way, the auxiliary computer can use classic computer vision rules (if-then), and the presence of a neural network in the supervisory MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity make the entire system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functions. For example, if there is a software vulnerability or error in the software running on the main computer and the non-identical software code running on the auxiliary computer provides the same overall result, the supervisory MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the main computer does not cause a substantial error.

[0173] In some examples, the output of the ADAS system 738 can be fed to the perception block of the main computer and / or the dynamic driving task block of the main computer. For example, if the ADAS system 738 indicates a forward collision warning due to an object immediately ahead, the perception block can use this information in identifying the object. In other examples, the auxiliary computer can have its own neural network that is trained and thus reduces the risk of false positives as described herein.

[0174] The vehicle 700 may further include an infotainment SoC 730 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more discrete components. The infotainment SoC 730 may include a combination of hardware and software that may be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), phone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total distance covered, brake fuel level, oil level, door open / closed, air filter information, etc.) to the vehicle 700. For example, the infotainment SoC 730 may include a radio, a disc player, a navigation system, a video player, USB and Bluetooth connectivity, an onboard computer, onboard entertainment, Wi-Fi, steering wheel audio controls, hands-free voice controls, a head-up display (HUD), an HMI display 734, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. The infotainment SoC 730 may further be used to provide information (e.g., visual and / or auditory) to a user of the vehicle, such as information from an ADAS system 738, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0175] The infotainment SoC 730 may include GPU functionality. The infotainment SoC 730 may communicate with other devices, systems, and / or components of the vehicle 700 via a bus 702 (e.g., a CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 730 may be coupled to a supervisory MCU so that in the event of a failure of a main controller 736 (e.g., a main and / or backup computer of the vehicle 700), the GPU of the infotainment system may perform some self-driving functions. In such an example, the infotainment SoC 730 may place the vehicle 700 in a driver safety parking mode as described herein.

[0176] The vehicle 700 may further include an instrument cluster 732 (e.g., a digital instrument panel, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 732 may include a controller and / or a supercomputer (e.g., a separate controller or a supercomputer). The instrument cluster 732 may include a set of instruments, such as a speedometer, a fuel level, an oil pressure, a tachometer, an odometer, a turn indicator, a shift position indicator, a seat belt warning light, a parking brake warning light, an engine fault light, an airbag (SRS) system information, a lighting control, a safety system control, navigation information, and the like. In some examples, information may be displayed and / or shared between the infotainment SoC 730 and the instrument cluster 732. In other words, the instrument cluster 732 may be included as part of the infotainment SoC 730, or vice versa.

[0177] Fig.7D For a cloud-based server and Fig. 7A Schematic diagram of a system for communicating between an example autonomous vehicle 700. System 776 may include a server 778, a network 790, and a vehicle including vehicle 700. Server 778 may include multiple GPUs 784 (A) -784 (H) (collectively referred to as GPUs 784 here), PCIe switches 782 (A) -782 (D) (collectively referred to as PCIe switches 782 here), and / or CPUs 780 (A) -780 (B) (collectively referred to as CPUs 780 here). GPUs 784, CPUs 780, and PCIe switches may be interconnected with high-speed interconnects and / or PCIe connections 786 such as, for example and without limitation, NVLink interfaces 788 developed by NVIDIA. In some examples, GPUs 784 are connected via NVLink and / or NVSwitch SoCs, and GPUs 784 and PCIe switches 782 are connected via PCIe interconnects. Although eight GPUs 784, two CPUs 780, and two PCIe switches are illustrated, this is not intended to be limiting. Depending on the embodiment, each of the servers 778 may include any number of GPUs 784, CPUs 780, and / or PCIe switches. For example, each of the servers 778 may include eight, sixteen, thirty-two, and / or more GPUs 784.

[0178] Server 778 may receive image data over network 790 and from a vehicle, the image data representing images showing unexpected or changed road conditions, such as recently begun road work. Server 778 may transmit neural network 792, updated neural network 792, and / or map information 794, including information about traffic and road conditions, over network 790 and to the vehicle. Updates to map information 794 may include updates to HD map 722, such as information about construction sites, potholes, curves, flooding, or other obstacles. In some examples, neural network 792, updated neural network 792, and / or map information 794 may have been generated from new training and / or data received from any number of vehicles in the environment and / or based on experience from training performed at a data center (e.g., using server 778 and / or other servers).

[0179] Server 778 can be used to train a machine learning model (e.g., a neural network) based on training data. The training data can be generated by the vehicle, and / or can be generated in simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., when the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., when the neural network does not require supervised learning). Training can be performed according to any one or more classes of machine learning techniques, including but not limited to the following classes: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, joint learning, transfer learning, feature learning (including principal components and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including alternate dictionary learning), rule-based machine learning, anomaly detection, and any variants or combinations thereof. Once the machine learning model is trained, the machine learning model can be used by the vehicle (e.g., transmitted to the vehicle via network 790), and / or the machine learning model can be used by server 778 to remotely monitor the vehicle.

[0180] In some examples, server 778 can receive data from the vehicle and apply the data to the latest real-time neural network for real-time intelligent reasoning. Server 778 may include a deep learning supercomputer and / or a dedicated AI computer powered by GPU 784, such as DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 778 may include a deep learning infrastructure of a data center powered only by CPUs.

[0181] The deep learning infrastructure of server 778 may be capable of fast real-time inference, and may use this capability to assess and verify the health of the processors, software, and / or associated hardware in vehicle 700. For example, the deep learning infrastructure may receive periodic updates from vehicle 700, such as a sequence of images and / or objects located in the sequence of images that vehicle 700 has located (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure may run its own neural network to identify objects and compare them to the objects identified by vehicle 700, and if the results do not match and the infrastructure concludes that the AI ​​in vehicle 700 has failed, then server 778 may transmit a signal to vehicle 700 instructing the fail-safe computer of vehicle 700 to take control, notify passengers, and complete a safe parking maneuver.

[0182] For reasoning, the server 778 may include a GPU 784 and one or more programmable reasoning accelerators (e.g., NVIDIA's TensorRT 3). The combination of GPU-powered servers and reasoning acceleration can make real-time responses possible. In other examples, such as when performance is not so important, CPU, FPGA, and other processor-powered servers can be used for reasoning.

[0183] Example computing device

[0184] Figure 8 8 is a block diagram of an example computing device 800 suitable for implementing some embodiments of the present disclosure. The computing device 800 may include an interconnect system 802 that directly or indirectly couples the following devices: a memory 804, one or more central processing units (CPUs) 806, one or more graphics processing units (GPUs) 808, a communication interface 810, an input / output (I / O) port 812, an input / output component 814, a power supply 816, one or more presentation components 818 (e.g., a display), and one or more logic units 820. In at least one embodiment, the computing device 800 may include one or more virtual machines (VMs), and / or any of its components may include virtual components (e.g., virtual hardware components). For non-limiting examples, the one or more GPUs 808 may include one or more vGPUs, the one or more CPUs 806 may include one or more vCPUs, and / or the one or more logic units 820 may include one or more virtual logic units. Thus, computing device 800 may include discrete components (eg, a complete GPU dedicated to computing device 800 ), virtual components (eg, a portion of a GPU dedicated to computing device 800 ), or a combination thereof.

[0185] although Figure 8The various blocks of are shown as being connected via an interconnect system 802 having wires, but this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component 818 such as a display device may be considered an I / O component 814 (e.g., if the display is a touch screen). As another example, the CPU 806 and / or the GPU 808 may include memory (e.g., the memory 804 may represent a storage device other than the memory of the GPU 808, the CPU 806, and / or other components). In other words, Figure 8 The computing devices described herein are illustrative only. No distinction is made between categories such as "workstations," "servers," "laptops," "desktops," "tablets," "client devices," "mobile devices," "handheld devices," "game consoles," "electronic control units (ECUs)," "virtual reality systems," and / or other device or system types, as all are considered Figure 8 In some embodiments, one or more aspects of parameter coordination function 120 and / or stitching module 150 may be at least partially executed by one or more of CPU 806 and / or GPU 808.

[0186] The interconnection system 802 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnection system 802 may include one or more links or bus types, such as an industry standard architecture (ISA) bus, an extended industry standard architecture (EISA) bus, a video electronics standard association (VESA) bus, a peripheral component interconnect (PCI) bus, a peripheral component interconnect express (PCIe) bus, and / or another type of bus or link. In some embodiments, there is a direct connection between components. As an example, the CPU 806 may be directly connected to the memory 804. In addition, the CPU 806 may be directly connected to the GPU 808. In the case where there is a direct or point-to-point connection between components, the interconnection system 802 may include a PCIe link to perform the connection. In these examples, the computing device 800 does not need to include a PCI bus.

[0187] Memory 804 may include any of a variety of computer-readable media. Computer-readable media may be any available media that can be accessed by computing device 800. Computer-readable media may include volatile and nonvolatile media and removable and non-removable media. By way of example and not limitation, computer-readable media may include computer storage media and communication media.

[0188] Computer storage media may include volatile and nonvolatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 804 may store computer-readable instructions (e.g., representing programs and / or program elements, such as an operating system). Computer storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other storage technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by computing device 800. As used herein, computer storage media does not include the signals themselves.

[0189] Computer storage media may contain computer readable instructions, data structures, program modules, and / or other data types in a modulated data signal such as a carrier wave or other transmission mechanism, and include any information delivery media. The term "modulated data signal" may refer to a signal that has one or more of its characteristics set or changed in such a way that information is encoded into the signal. By way of example and not limitation, computer storage media may include wired media such as a wired network or a direct wired connection, and wireless media such as sound, RF, infrared, and other wireless media. Any combination of the above should also be included within the scope of computer readable media.

[0190] The CPU 806 may be configured to execute at least some of the computer-readable instructions so as to control one or more components of the computing device 800 to perform one or more of the methods and / or processes described herein. Each of the CPUs 806 may include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of simultaneously processing a large number of software threads. The CPU 806 may include any type of processor, and may include different types of processors, depending on the type of computing device 800 implemented (e.g., a processor with fewer cores for mobile devices and a processor with more cores for servers). For example, depending on the type of computing device 800, the processor may be an Advanced RISC Mechanism (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors such as mathematical coprocessors, the computing device 800 may also include one or more CPUs 806.

[0191] In addition to or in place of the CPU 806, the GPU 808 may also be configured to execute at least some computer-readable instructions to control one or more components of the computing device 800 to perform one or more methods and / or processes described herein. One or more GPUs 808 may be integrated GPUs (e.g., with one or more CPUs 806) and / or one or more GPUs 808 may be discrete GPUs. In an embodiment, one or more GPUs 808 may be coprocessors of one or more CPUs 806. The computing device 800 may use the GPU 808 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, the GPU 808 may be used for general-purpose computing on the GPU (GPGPU). The GPU 808 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU 808 may generate pixel data for outputting an image in response to a rendering command (e.g., a rendering command from the CPU 806 received via a host interface). The GPU 808 may include a graphics memory such as a display memory for storing pixel data or any other suitable data (e.g., GPGPU data). Display memory may be included as part of memory 804. GPU 808 may include two or more GPUs operating in parallel (e.g., via a link). The link may connect the GPUs directly (e.g., using NVLINK) or may connect the GPUs through a switch (e.g., using NVSwitch). When combined together, each GPU 808 may generate pixel data or GPGPU data for different portions of output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU may include its own memory or may share memory with other GPUs.

[0192] In addition to or in lieu of the CPU 806 and / or the GPU 808, the logic unit 820 may be configured to execute at least some computer-readable instructions to control one or more components of the computing device 800 to perform one or more methods and / or processes described herein. In an embodiment, the CPU 806, the GPU 808, and / or the logic unit 820 may discretely or jointly perform any combination of methods, processes, and / or portions thereof. The one or more logic units 820 may be a part of and / or integrated into the one or more CPUs 806 and / or the one or more GPUs 808, and / or the one or more logic units 820 may be discrete components of the CPU 806 and / or the GPU 808 or otherwise external thereto. In an embodiment, the one or more logic units 820 may be a processor of the one or more CPUs 806 and / or the one or more GPUs 808.

[0193] Examples of logic unit 820 include one or more processing cores and / or components thereof, such as a data processing unit (DPU), a tensor core (TC), a tensor processing unit (TPU), a pixel vision core (PVC), a vision processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multiprocessor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application specific integrated circuit (ASIC), a floating point unit (FPU), an input / output (I / O) element, a peripheral component interconnect (PCI) or a peripheral component interconnect express (PCIe) element, etc.

[0194] The communication interface 810 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 800 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communications. The communication interface 810 may include components and functions that enable communication over any of a number of different networks, such as wireless networks (e.g., Wi-Fi, Z-wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communicating over Ethernet or InfiniBand), low power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, the logic unit 820 and / or the communication interface 810 may include one or more data processing units (DPUs) to transmit data received over a network and / or over the interconnect system 802 directly to one or more GPUs 808 (e.g., memory in a GPU 808).

[0195] The I / O port 812 can enable the computing device 800 to be logically coupled to other devices including I / O components 814, presentation components 818, and / or other components, some of which can be built into (e.g., integrated into) the computing device 800. Illustrative I / O components 814 include microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dishes, browsers, printers, wireless devices, and the like. The I / O components 814 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by the user. In some instances, the input can be transmitted to an appropriate network element for further processing. The NUI can implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, gesture recognition on and adjacent to the screen, air gestures, head and eye tracking, and touch recognition associated with the display of the computing device 800 (as described in more detail below). The computing device 800 may include a depth camera such as a stereo camera system, an infrared camera system, an RGB camera system, touch screen technology, and a combination of these for gesture detection and recognition. Additionally, computing device 800 may include an accelerometer or gyroscope to enable motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope may be used by computing device 800 to render immersive augmented reality or virtual reality.

[0196] The power supply 816 may include a hardwired power supply, a battery power supply, or a combination thereof. The power supply 816 may provide power to the computing device 800 to enable the components of the computing device 800 to operate.

[0197] The presentation component 818 may include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. The presentation component 818 may receive data from other components (e.g., the GPU 808, the CPU 806, the DPU, etc.) and output the data (e.g., as an image, video, sound, etc.). In some embodiments, the visualization 165 may be presented on a display of the presentation component 818.

[0198] Sample Data Center

[0199] Fig. 9 An example data center 900 is shown, which may be used in at least one embodiment of the present disclosure. The data center 900 may include a data center infrastructure layer 910, a framework layer 920, a software layer 930, and an application layer 940.

[0200] like Fig. 9As shown, the data center infrastructure layer 910 may include a resource coordinator 912, grouped computing resources 914, and node computing resources ("node CRs") 916 (1)-916 (N), where "N" represents any complete positive integer. In at least one embodiment, the node CRs 916 (1)-916 (N) may include, but are not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state drives or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power modules and cooling modules, etc. In some embodiments, one or more of the node CRs 916 (1)-916 (N) may correspond to a server having one or more of the above computing resources. Furthermore, in some embodiments, the nodes CR916(1)-916(N) may include one or more virtual components, such as vGPUs, vCPUs, etc., and / or one or more of the nodes CR916(1)-916(N) may correspond to virtual machines (VMs). In some embodiments, one or more aspects of the parameter coordination function 120 and / or the stitching module 150 may be at least partially performed by one or more of the CR916(1)-916(N).

[0201] In at least one embodiment, the computing resources 914 of the grouping can include the separate grouping (not shown) of the node CR916 contained in one or more racks, or many racks (not shown) contained in the data center of each geographical location. The separate grouping of the node CR916 in the computing resources 914 of the grouping can include the computing, network, memory or storage resources that can be configured or assigned to support the grouping of one or more workloads. In at least one embodiment, several node CR916 including CPU, GPU, DPU and / or other processors can be grouped in one or more racks to provide computing resources to support one or more workloads. One or more racks can also include any number of power modules, cooling modules and / or network switches in any combination.

[0202] Resource coordinator 912 may configure or otherwise control one or more nodes CR 916(1)-916(N) and / or grouped computing resources 914. In at least one embodiment, resource coordinator 912 may include a software design infrastructure (SDI) management entity for data center 900. Resource coordinator 912 may include hardware, software, or some combination thereof.

[0203] In at least one embodiment, Fig. 9 As shown, the framework layer 920 may include a job scheduler 933, a configuration manager 934, a resource manager 936, and a distributed file system 938. The framework layer 920 may include a framework that supports the software 932 of the software layer 930 and / or one or more applications 942 of the application layer 940. The software 932 or the application 942 may include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 920 may be, but is not limited to, a free and open source software network application framework, such as Apache Spark, which may utilize the distributed file system 938 for large-scale data processing (e.g., "big data"). TM (hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 933 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 900. In at least one embodiment, the configuration manager 934 may be able to configure different layers, such as the software layer 930 and the framework layer 920 including Spark and a distributed file system 938 for supporting large-scale data processing. The resource manager 936 can manage cluster or group computing resources mapped to or allocated to support the distributed file system 938 and the job scheduler 933. In at least one embodiment, the cluster or group computing resources may include grouped computing resources 914 at the data center infrastructure layer 910. The resource manager 936 can coordinate with the resource coordinator 912 to manage these mapped or allocated computing resources.

[0204] In at least one embodiment, software 932 included in software layer 930 may include software used by at least a portion of node CRs 916(1)-916(N), grouped computing resources 914, and / or distributed file system 938 of framework layer 920. One or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.

[0205] In at least one embodiment, one or more applications 942 included in the application layer 940 may include one or more types of applications used by at least a portion of the node CRs 916(1)-916(N), the grouped computing resources 914, and / or the distributed file system 938 of the framework layer 920. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.

[0206] In at least one embodiment, any of configuration manager 934, resource manager 936, and resource coordinator 912 can implement any number and type of self-modification actions based on any number and type of data acquired in any technically feasible manner. The self-modification actions can relieve the data center operator of data center 900 from making potentially bad configuration decisions and can avoid underutilized and / or underperforming portions of the data center.

[0207] The data center 900 may include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, a machine learning model may be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to the data center 900. In at least one embodiment, by using weight parameters calculated by one or more training techniques, the resources described above with respect to the data center 900 may be used to infer or predict information using a trained machine learning model corresponding to one or more neural networks, such as but not limited to those described herein.

[0208] In at least one embodiment, the data center 900 can use a CPU, an application-specific integrated circuit (ASIC), a GPU, an FPGA, and / or other hardware (or virtual computing resources corresponding thereto) to use the above resources to perform training and / or reasoning. In addition, one or more of the above software and / or hardware resources can be configured as a service to allow users to train or perform information reasoning, such as image recognition, speech recognition, or other artificial intelligence services.

[0209] Example network environment

[0210] A network environment suitable for implementing embodiments of the present disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be configured to communicate with the client device, server, or other device types. Figure 8 The backend device 900 may be implemented on one or more instances of the computing device 800 of the embodiment of the present invention—for example, each device may include similar components, features and / or functions of the computing device 800. In addition, in the case of implementing a backend device (e.g., a server, NAS, etc.), the backend device may be included as part of the data center 900, examples of which are described herein with respect to Fig. 9 Describe in more detail.

[0211] The components of the network environment can communicate with each other through the network, which can be wired, wireless, or both. The network can include multiple networks, or a network of multiple networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. In the case where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connections.

[0212] Compatible network environments may include one or more peer-to-peer network environments (in which case the server may not be included in the network environment), and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the functionality described herein with respect to the server may be implemented on any number of client devices.

[0213] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, and the like. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework for supporting software at the software layer and / or one or more applications at the application layer. The software or application may include network-based service software or application programs, respectively. In an embodiment, one or more client devices may use network-based service software or application programs (e.g., by accessing the service software and / or application programs via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a type of free and open source software network application framework, such as a distributed file system that may be used for large-scale data processing (e.g., "big data").

[0214] A cloud-based network environment can provide cloud computing and / or cloud storage that performs any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these various functions can be distributed across multiple locations from a central or core server (e.g., one or more data centers that can be distributed across states, regions, countries, the world, etc.). If the connection to the user (e.g., client device) is relatively close to an edge server, the core server can assign at least a portion of the functionality to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0215] The client device may include Figure 8 At least some of the components, features, and functionality of the described example computing device 800. By way of example and not limitation, the client device may be embodied as a personal computer (PC), a laptop computer, a mobile device, a smartphone, a tablet computer, a smart watch, a wearable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality head mounted display, a global positioning system (GPS) or device, a video player, a camera, a surveillance device or system, a vehicle, a vessel, an aircraft, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or system, an entertainment system, an in-vehicle computer system, an embedded system controller, a remote control, an appliance, a consumer electronic device, a workstation, an edge device, any combination of these described devices, or any other suitable device.

[0216] The present disclosure may be described in the general context of machine-usable instructions or computer codes executed by a computer or other machine such as a personal digital assistant or other handheld device, including computer-executable instructions such as program modules. Typically, program modules including routines, programs, objects, components, data structures, etc. refer to codes that perform specific tasks or implement specific abstract data types. The present disclosure may be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The present disclosure may also be practiced in a distributed computing environment in which tasks are performed by remote processing devices linked through a communication network.

[0217] As used herein, the description of "and / or" about two or more elements should be interpreted as referring to only one element or a combination of elements. For example, "element A, element B and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B and C. In addition, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0218] The subject matter of the present disclosure is described in detail herein to meet statutory requirements. However, the description itself is not intended to limit the scope of the present disclosure. On the contrary, the inventors have contemplated that the claimed subject matter may also be embodied in other ways to include steps different from the steps described herein in conjunction with other current or future technologies or combinations of similar steps. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of the method employed, these terms should not be interpreted as implying any particular order among or between the various steps disclosed herein, unless the order of the steps is explicitly described.

Claims

1. One or more processors, including one or more processing units, wherein the one or more processing units are configured to: determining a first parameter value associated with rendering a first image and a second parameter value associated with rendering a second image, wherein the first image at least partially overlaps the second image at an image boundary; calculating a boundary parameter value based at least on the first parameter value and the second parameter value; adjusting at least a portion of the first image using a first parametric gain function based at least on the first parameter value and the boundary parameter value; adjusting at least a portion of the second image using a second parametric gain function based at least on the second parameter value and the boundary parameter value; and The first image is spliced ​​with the second image at the image boundary.

2. One or more processors according to claim 1, wherein the first parameter value and the second parameter value are associated with rendering parameters, and the rendering parameters include at least one of the following: white balance, tone mapping, exposure time, brightness, contrast, gamma, chroma, noise reduction, saturation, sharpness, color filter array (CFA) pattern, lens shading correction, lens distortion correction, focal length correction, barrel distortion correction or pincushion distortion.

3. The one or more processors of claim 1 , wherein the one or more processing units are further configured to: The boundary parameter value is calculated based on an average value calculated using the first parameter value and the second parameter value.

4. The one or more processors of claim 1, wherein the one or more processing units are further configured to: determining the first parameter value based on metadata corresponding to the first image; and The second parameter value is determined based on metadata corresponding to the second image.

5. The one or more processors of claim 1, wherein the one or more processing units are further configured to: determining the first parameter value based on metadata associated with characteristics of a first camera used to capture the first image; and The second parameter value is determined based on metadata associated with characteristics of a second camera used to capture the second image.

6. The one or more processors of claim 1 , wherein at least one of the first parameter gain function and the second parameter gain function comprises: Linear function, nonlinear curve function, S-curve function or logistic curve function.

7. The one or more processors of claim 1, wherein the one or more processing units are further configured to: At least one of the first parametric gain function and the second parametric gain function is adjusted based on ambient lighting condition data from a sensor.

8. The one or more processors of claim 1, wherein the one or more processing units are further configured to generate at least one of: A surround-view stitched image based at least on the first image and the second image; stitching an image based on at least a fisheye view of the first image and the second image; and A panoramic view stitching image is generated based on at least the first image and the second image.

9. The one or more processors of claim 1, wherein the first image is captured by a first camera and the second image is captured by a second camera, the first camera and the second camera having at least partially overlapping fields of view.

10. The one or more processors of claim 1, wherein the processor is included in at least one of: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; A system for performing simulation operations; Systems for performing digital twin operations; A system for performing light transport simulations; A system for performing collaborative content creation of three-dimensional assets; Systems for performing deep learning operations; Systems for performing remote operations; A system for performing real-time streaming; Systems for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; Systems for performing conversational AI operations; A system implementing one or more language models; A system implementing one or more Large Language Models (LLMs); Systems for generating synthetic data; Systems for generating synthetic data using AI; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.

11. A system comprising: One or more processing units for: determining a first parameter value associated with a metadata parameter of a first image in the plurality of images; calculating a first boundary value based at least on the first parameter value and a second parameter value associated with a metadata parameter of a second image in the plurality of images; calculating a second boundary value based at least on the first parameter value and a third parameter value associated with a metadata parameter of a third image in the plurality of images; adjusting at least a first portion of the first image using a first gain function based at least on the first parameter value and the first boundary value; as well as At least a second portion of the first image is adjusted using a second gain function based at least on the first parameter value and the second boundary value.

12. The system of claim 11, wherein the one or more processing units are further configured to: splicing the first image with the second image based on an overlapping image boundary between the first image and the second image; and The first image is spliced ​​with the third image based on an overlapping image boundary between the first image and the third image.

13. The system of claim 11, wherein the first parameter value and the second parameter value are associated with rendering parameters, the rendering parameters comprising at least one of: white balance, tone mapping, exposure time, brightness, contrast, gamma, chroma, noise reduction, saturation, sharpness, color filter array (CFA) pattern, lens shading correction, lens distortion correction, focal length correction, barrel distortion correction, or pincushion distortion.

14. The system of claim 11, wherein the one or more processing units are further configured to: calculating the first boundary value based on an average value calculated using the first parameter value and the second parameter value; and The second boundary value is calculated based on an average value calculated using the first parameter value and the third parameter value.

15. The system of claim 11, wherein the one or more processing units are further configured to: determining the first parameter value based on metadata corresponding to the first image; determining the second parameter value based on metadata corresponding to the second image; and The third parameter value is determined based on metadata corresponding to the third image.

16. The system of claim 11, wherein the one or more processing units are further configured to: determining the first parameter value based on metadata associated with characteristics of a first camera used to capture the first image; determining the second parameter value based on metadata associated with characteristics of a second camera used to capture the second image; and The third parameter value is determined based on metadata associated with characteristics of a third camera used to capture the second image.

17. The system of claim 11, wherein at least one of the first parameter gain function or the second parameter gain function comprises: Linear function, nonlinear curve function, S-curve function or logistic curve function.

18. The system of claim 11, wherein the one or more processing units are further configured to: At least one of the first parametric gain function and the second parametric gain function is adjusted based on ambient lighting condition data from a sensor.

19. The system of claim 11, wherein the one or more processing units are included in at least one of the following: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; A system for performing simulation operations; Systems for performing digital twin operations; A system for performing light transport simulations; A system for performing collaborative content creation of three-dimensional assets; Systems for performing deep learning operations; Systems for performing remote operations; A system for performing real-time streaming; Systems for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; Systems for performing conversational AI operations; A system implementing one or more language models; A system implementing one or more Large Language Models (LLMs); Systems for generating synthetic data; Systems for generating synthetic data using AI; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.

20. A method comprising: Stitching a first image with a second image that at least partially overlaps the first image at an image boundary, the stitching being based at least on adjusting a metadata parameter of the first image by blending the metadata parameter from a first value of the metadata parameter at a first location within the first image to a second value of the metadata parameter at a second location at the image boundary, wherein the second value is calculated based at least on the first value of the metadata parameter and a third value of the metadata parameter associated with the second image.

Citation Information

Patent Citations

  • Method for programmable timeouts of tree traversal mechanisms in hardware

    US10885698B2

  • Image stitching with color harmonization for surround view systems and applications

    US20240112472A1