Neural network routing planning

A neural network-based environment field representation addresses the challenge of dynamic path calculation by efficiently encoding reach distances, facilitating geometrically and semantically feasible navigation.

JP7856489B2Active Publication Date: 2026-05-11NVIDIA CORP
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NVIDIA CORP
Filing Date
2022-06-01
Publication Date
2026-05-11

Smart Images

  • Figure 0007856489000005
    Figure 0007856489000005
  • Figure 0007856489000006
    Figure 0007856489000006
  • Figure 0007856489000007
    Figure 0007856489000007
Patent Text Reader

Abstract

To provide a processor, system and machine readable medium that improve techniques for calculating environment paths.SOLUTION: A processor includes one or more circuits for using one or more neural networks to calculate a plurality of paths through which an autonomous device is going to traverse. In a process for calculating the plurality of paths using an implicit environment function, the one or more circuits are configured to: acquire at least first location, set of locations, and final location; calculate a set of distances at least partially based on the set of locations and the final location; and calculate the plurality of paths at least partially based on the set of distances. The plurality of paths are configured to form a path to the final location from the first location.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] At least one embodiment relates to a processing resource used to compute a path through an environment using a neural network. For example, at least one embodiment relates to a processor or computing resource used to compute the distance to a location in an environment using a neural network to compute a path according to various novel techniques described herein. [Background technology]

[0002] Calculating paths through an environment is a crucial task in many contexts. In various cases, calculating paths to a target through an environment can be difficult, such as when the target and environment undergo various changes. Calculating paths to a target through an environment can also require significant computing resources. Therefore, techniques for calculating paths through an environment can be improved. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (for example, Standard No. J3016-201806, issued June 15, 2018; Standard No. J3016-201609, issued September 30, 2016; and previous and newer versions of this standard) [Brief explanation of the drawing]

[0004] [Figure 1] This figure shows an example of an implicit environment function for the environment, based on at least one embodiment. [Figure 2]This figure shows one example of an implicit environment function for different target locations in the environment, based on at least one embodiment. [Figure 3] This figure shows one example of an implicit environment function for different target locations and environments, based on at least one embodiment. [Figure 4] This figure shows an example of an agent that uses an implicit environment function to navigate to a target, with at least one implementation. [Figure 5] This figure shows an example of an environment field for a multi-user navigation environment, based on at least one embodiment. [Figure 6] This figure shows one example of the use of a negative environment field for a 3D environment, based on at least one embodiment. [Figure 7] This figure shows an example of the results obtained using an implicit environment function, based on at least one embodiment. [Figure 8] This figure shows an example of the results of agent navigation, based on at least one embodiment. [Figure 9] This figure shows an example of the results obtained by fitting a given human sequence to an explored trajectory in a 3D indoor environment, according to at least one embodiment. [Figure 10] This figure shows an example of the results obtained by fitting a given human sequence to an explored trajectory in a 3D indoor environment, according to at least one embodiment. [Figure 11] This figure shows an example of a process for calculating multiple paths using an implicit environment function, based on at least one embodiment. [Figure 12A] This figure shows the inference and / or training logic according to at least one embodiment. [Figure 12B] This figure shows the inference and / or training logic according to at least one embodiment. [Figure 13] This figure shows the training and deployment of a neural network in at least one embodiment. [Figure 14] A diagram showing an exemplary data center system according to at least one embodiment. [Figure 15A] A diagram showing an example of an autonomous vehicle according to at least one embodiment. [Figure 15B] A diagram showing an example of camera locations and fields of view for the autonomous vehicle of FIG. 15A according to at least one embodiment. [Figure 15C] A block diagram showing an exemplary system architecture for the autonomous vehicle of FIG. 15A according to at least one embodiment. [Figure 15D] A diagram showing a system for communication between a (one or more) cloud-based server and the autonomous vehicle of FIG. 15A according to at least one embodiment. [Figure 16] A block diagram showing a computer system according to at least one embodiment. [Figure 17] A block diagram showing a computer system according to at least one embodiment. [Figure 18] A diagram showing a computer system according to at least one embodiment. [Figure 19] A diagram showing a computer system according to at least one embodiment. [Figure 20A] A diagram showing a computer system according to at least one embodiment. [Figure 20B] A diagram showing a computer system according to at least one embodiment. [Figure 20C] A diagram showing a computer system according to at least one embodiment. [Figure 20D] A diagram showing a computer system according to at least one embodiment. [Figure 20E] A diagram showing a shared programming model according to at least one embodiment. [Figure 20F] A diagram showing a shared programming model according to at least one embodiment. [Figure 21] A diagram showing an exemplary integrated circuit and related graphics processor according to at least one embodiment. [Figure 22A] A diagram showing an exemplary integrated circuit and related graphics processor according to at least one embodiment. [Figure 22B] A diagram showing an exemplary integrated circuit and related graphics processor according to at least one embodiment. [Figure 23A] A diagram showing additional exemplary graphics processor logic according to at least one embodiment. [Figure 23B] A diagram showing additional exemplary graphics processor logic according to at least one embodiment. [Figure 24] A diagram showing a computer system according to at least one embodiment. [Figure 25A] A diagram showing a parallel processor according to at least one embodiment. [Figure 25B] A diagram showing a partition unit according to at least one embodiment. [Figure 25C] A diagram showing a processing cluster according to at least one embodiment. [Figure 25D] A diagram showing a graphics multiprocessor according to at least one embodiment. [Figure 26] A diagram showing a multi-graphics processing unit (GPU) system according to at least one embodiment. [Figure 27] A diagram showing a graphics processor according to at least one embodiment. [Figure 28] A block diagram showing a processor microarchitecture for a processor according to at least one embodiment. [Figure 29] A diagram showing a deep learning application processor according to at least one embodiment. [Figure 30]A block diagram showing an exemplary neuromorphic processor with at least one embodiment. [Figure 31] This figure shows at least a portion of a graphics processor according to one or more embodiments. [Figure 32] This figure shows at least a portion of a graphics processor according to one or more embodiments. [Figure 33] This figure shows at least a portion of a graphics processor according to one or more embodiments. [Figure 34] This is a block diagram of the graphics processing engine of a graphics processor, according to at least one embodiment. [Figure 35] This is a block diagram of at least a portion of a graphics processor core, according to at least one embodiment. [Figure 36A] This figure shows thread execution logic including an array of processing elements for a graphics processor core, according to at least one embodiment. [Figure 36B] This figure shows thread execution logic including an array of processing elements for a graphics processor core, according to at least one embodiment. [Figure 37] This figure shows a parallel processing unit ("PPU") according to at least one embodiment. [Figure 38] This figure shows a general-purpose processing cluster ("GPC") in at least one embodiment. [Figure 39] This figure shows a memory partition unit of a parallel processing unit ("PPU") according to at least one embodiment. [Figure 40] This figure shows a streaming multiprocessor according to at least one embodiment. [Figure 41] This is an exemplary data flow diagram for an advanced computing pipeline, with at least one implementation example. [Figure 42] This is a system diagram for an exemplary system for training, adapting, instantiating, and deploying machine learning models in an advanced computing pipeline, with at least one embodiment. [Figure 43] This figure includes an illustrative diagram of an advanced computing pipeline 4210A for processing imaging data, according to at least one embodiment. [Figure 44A] This figure includes an illustrative data flow diagram of a virtual device supporting an ultrasonic device, according to at least one embodiment. [Figure 44B] This figure includes an illustrative data flow diagram of a virtual device supporting a CT scanner, according to at least one embodiment. [Figure 45A] This is a data flow diagram for the process of training a machine learning model, with at least one example. [Figure 45B] This diagram illustrates an exemplary client-server architecture for extending annotation tools with a pre-trained annotation model, based on at least one embodiment. [Modes for carrying out the invention]

[0005] In at least one embodiment, a neural network calculates the distance between a position and a target position along a feasible trajectory through the environment, based on the environment, the position, and the target position. In at least one embodiment, the position refers to any suitable location, area, and / or region of the environment. In at least one embodiment, the target position refers to a position in the environment where one or more agents are to navigate. In at least one embodiment, the environment, also called the scene, refers to any suitable two-dimensional (2D) or three-dimensional (3D) environment that may have various obstacles. In at least one embodiment, the environment can correspond to any suitable real-world environment or a simulated environment. In at least one embodiment, given the environment, a first position, and a second position, the neural network calculates the reaching distance from the first position to the second position, where the reaching distance refers to the distance from the first position to the second position along a feasible trajectory through the environment. In at least one embodiment, feasible trajectories refer to trajectories that are geometrically feasible and / or semantically reasonable, as will be described in more detail below. In at least one embodiment, feasible trajectories refer to any suitable trajectory through the environment that does not collide with or otherwise intersect with any obstacles and / or inaccessible areas of the environment.

[0006] In at least one embodiment, the neural network represents the environment through an environment field that encodes the reach distances of various locations in the environment to a target location. In at least one embodiment, the neural network represents the environment internally through the environment field. In at least one embodiment, the environment field is a representation that includes a set of values ​​indicating the reach distances for each location in the environment to a target location. In at least one embodiment, the environment field represents the environment, and each value in the environment field corresponds to a location in the environment. In at least one embodiment, a particular value in the environment field corresponds to a particular location in the environment, and the particular value is the reach distance value from the particular location in the environment to the target location in the environment.

[0007] In at least one embodiment, the environment field may be represented through a visual representation, such as an image, where different reach values ​​are assigned different color values. In at least one embodiment, one or more systems visualize the environment field by at least using a neural network to calculate the reach value for each position in the environment to a target position, and by utilizing the calculated reach values ​​to visualize the environment field. In at least one embodiment, the environment field visualization is continuous. In at least one embodiment, the various environment field visualizations described herein may include individual sections, but the environment field visualization may include any preferred continuous visualization showing a continuous color gradient, continuous shading, and / or reach values.

[0008] In at least one embodiment, the environment field is used to guide the dynamic behavior of an agent within the environment. In at least one embodiment, the environment is represented by an image, where each position corresponds to a pixel in the image. In at least one embodiment, the positions in the environment correspond to any preferred location, region, or area within the environment. In at least one embodiment, the environment is structured as a grid, where positions correspond to a set of pixels corresponding to grid cells. In at least one embodiment, the positions correspond to any preferred set of pixels. In at least one embodiment, the image is any preferred image, such as a red-green-blue (RGB) image, a black / white (B / W) image, a grayscale image, an RGB-depth (RGBD) image, and / or variations thereof. In at least one embodiment, the agent refers to any preferred entity, such as an autonomous device, robot, human, or other entity navigating or otherwise moving through the environment. In at least one embodiment, a neural implicit function represents the environment internally through the environment field. In at least one embodiment, the neural implicit function is a neural network that approximates or models a function such as a function that outputs the distance to a target location relative to a position in the environment. In at least one embodiment, the neural implicit function that approximates or models a function that outputs the distance to a target location relative to a position in the environment is called the implicit environment function.

[0009] In at least one embodiment, one or more systems utilize a neural implicit function representing the environment through an environment field to guide an agent navigating a scene, such as navigating an agent to reach a given target in a physically reasonable and / or feasible manner. In at least one embodiment, for each position in the scene, the environment field captures the distance from the position to a given target position along a geometrically feasible and / or semantically reasonable trajectory. In at least one embodiment, a geometrically feasible trajectory refers to a trajectory through the environment that does not collide with any obstacles in the environment. In at least one embodiment, a semantically reasonable trajectory refers to a trajectory through the environment that is physically appropriate and should be followed by one or more agents. In at least one embodiment, for example, a semantically unreasonable trajectory might be a trajectory through an environment that requires a human agent to crawl under an obstacle, in this example, crawling may not be a physically appropriate action for the human agent. In at least one embodiment, one or more systems use neural implicit functions to compute a plurality of paths that an agent is to traverse or otherwise navigate. In at least one embodiment, one or more systems use neural implicit functions to navigate the agent by repeatedly moving the agent to the next location with the minimum reach distance until the agent reaches a target location.

[0010] In at least one embodiment, the neural implicit function can be generalized to any target location and any scene / environment, and given different target locations and / or scenes / environments, generalization to different environmental fields may be possible using a single neural implicit function. In at least one embodiment, the neural implicit function can be trained to determine or otherwise compute a continuous environmental field using discretely sampled training data. In at least one embodiment, querying the distance to reach between an arbitrary location and a target location requires only a fast network forward path, enabling efficient trajectory prediction.

[0011] In at least one embodiment, to enforce the semantic validity of the environmental field in an indoor environment, one or more systems define accessible regions that point to areas where an agent is deemed reasonable to appear. In at least one embodiment, one or more systems use one or more neural network models, such as a variational autoencoder, to determine accessible regions for various environments and designate locations outside the accessible regions, also called inaccessible regions, as obstacles. In at least one embodiment, one or more systems utilize neural implicit functions that represent the environment through the environmental field in various human trajectory modeling in 3D environments, such as various indoor, outdoor, or other suitable environments, and in agent navigation in 2D environments, such as various mazes or other suitable environments.

[0012] In at least one embodiment, a neural implicit function is trained using training data generated based on one or more environments. In at least one embodiment, accessible and inaccessible regions are defined for an environment by one or more systems using various neural network models and / or other suitable systems. In at least one embodiment, training data is generated by one or more systems using any suitable path planning algorithm, method, system and / or variations thereof. In at least one embodiment, one or more systems generate training data by processing the environment using one or more path planning algorithms to calculate one or more reach distances from one or more accessible locations in an environment having defined accessible and inaccessible regions to one or more target locations in the environment, wherein the training data comprises one or more of the one or more reach distances, the one or more accessible locations, and / or the one or more target locations; and train a neural network using the training data.

[0013] The preceding and following descriptions include numerous specific details in order to provide a more complete understanding of at least one embodiment. However, it will be apparent to those skilled in the art that the inventive concept may be implemented without one or more of these specific details.

[0014] In at least one embodiment, the techniques described herein achieve a variety of technical advantages, including, but are not limited to, the ability to use a neural network to represent the environment in order to facilitate guidance for agent behavior / navigation in the environment; the ability to use a neural network to represent the environment through an environment field that ensures geometric and semantic feasibility; the ability to use a neural network to calculate one or more reach distances to one or more locations in the environment to a target in a single forward path; and a variety of other technical advantages.

[0015] In at least one embodiment, one or more systems train a neural implicit function to learn to represent an environment through a continuous environment field, where each value at each position in the continuous environment field represents the distance to be reached from a particular position in the environment to a target position. In at least one embodiment, one or more systems model the environment field as waveform propagation to calculate the distance to be reached. In at least one embodiment, waveform propagation refers to the phenomenon in which the distance between an initial position and a target is equivalent to the minimum amount of time required by a waveform starting from the initial position, propagating along its front boundary, and reaching the target.

[0016] In at least one embodiment, one or more systems represent all locations reachable by an agent as accessible regions and all other locations as obstacles, also called inaccessible regions. In at least one embodiment, for a diffuse waveform (e.g., a closed curve in the environment) denoted by τ, the waveform diffuses based on a speed function defined at each location, denoted by f(x). In at least one embodiment, the current position of the diffuse waveform denoted by τ is modeled by an arrival time function denoted by u(x) with respect to a location denoted by x, starting from a target position. In at least one embodiment, when f(x) > 0, one or more systems construct an arrival time function as a continuous shortest path problem through equations such as the Eikonal equation or any preferred equation, which is expressed by the following equation, although any variation thereof may be used. ||∇u(x)||f(x)=1,x∈Ω Here, Ω represents the feasible region (e.g., the accessible region), ∇ represents the gradient, ||·|| represents the Euclidean norm, and x e In the target location indicated by u(x e ) = 0. In at least one embodiment, one or more systems specify an environmental layout, where, in the presence of an obstacle, the waveform expansion speed indicated by f(x) is set to an infinitesimally small positive value or any preferred value, since the waveform cannot pass through the obstacle, and within the accessible region, the waveform expansion speed is set to a constant value of 1 or any preferred value.

[0017] In at least one embodiment, one or more systems solve the Iconal equations, such as those described herein, through a neural implicit function that models a continuous energy surface. In at least one embodiment, the neural implicit function, also called an implicit environment function, neural network model, neural network function, machine learning function, implicit neural representation, and / or a variation thereof, is a neural network that approximates one or more functions. In at least one embodiment, one or more systems solve the Iconal equations, such as x∈R 2 A neural implicit function is trained to represent the environment through the environment field as a mapping from location coordinates indicated by to arrival times indicated by u(x)∈R. In at least one embodiment, one or more systems train a neural implicit function using training data containing discretely sampled pairs of {x, u(x)} to represent the environment through a continuous field.

[0018] In at least one embodiment, one or more systems utilize one or more path planning algorithms, methods, and / or systems, such as Breath First Searching (BFS), Dijkstra's algorithm, Fast Marching Method (FMM), Rapidly Exploring Random Tree (RRT), Probabilistic Roadmap (PRM), and / or variations thereof, to generate training data for neural implicit functions. In at least one embodiment, one or more path planning algorithms, methods, and / or systems determine the reach from each accessible location in the input environment to the target location, based on the input environment and the target location. In at least one embodiment, the training data includes the environment, the target location, and the reach for each accessible location in the environment. In at least one embodiment, an accessible location refers to a location in the environment that exists within an accessible region.

[0019] In at least one embodiment, one or more systems utilize a path planning method to solve u(x) in one or more environments, such as an environment structured as a grid. In at least one embodiment, one or more path planning methods start with an individual map beginning at a target location and iteratively expand and update neighboring feasible grid cells according to speed using an iconal update criterion or any preferred criterion until the starting location is reached. In at least one embodiment, one or more path planning methods include methods such as an FMM implementation that take an individual grid (e.g., a grid corresponding to an environment) with a given target location as input and output the reach distance from each accessible location (e.g., a location in the environment) to the target location. In at least one embodiment, one or more systems train a neural implicit function to regress the reach distance at each grid cell, given cell coordinates (e.g., normalized to [-1,1]) as input. In at least one embodiment, Figures 1 to 3 show examples of implicit environment functions.

[0020] Figure 1 shows an example 100 of an implicit environment function for an environment according to at least one embodiment. In at least one embodiment, one or more processes, functions, and / or operations shown in Figures 1 to 11 are performed or otherwise executed by any preferred processing system or unit, such as a graphics processing unit (GPU), a parallel processing unit (PPU), a central processing unit (CPU), and / or a variation thereof, and in any preferred order, such as sequential, parallel, and / or a variation thereof. In at least one embodiment, a training framework 102 trains an implicit environment function 106 using training data 104 to determine an implicit environment function 108. In at least one embodiment, one or more systems input a query position 110 to the implicit environment function 108, which represents the environment through an environment field visualized by an environment field visualization 112 to determine a reach distance 114.

[0021] In at least one embodiment, the training framework 102 is a set of one or more hardware and / or software computing resources having instructions that, when executed, perform one or more neural network training processes, functions, and / or actions. In at least one embodiment, the training framework 102 is as described with respect to Figure 13. In at least one embodiment, the training framework 102 is a framework such as PyTorch, TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework 102 is a software program running on computer hardware, an application running on computer hardware, and / or a variation thereof. In at least one embodiment, the training framework 102 trains an implicit environment function 106 using training data 104 to determine an implicit environment function 108.

[0022] In at least one embodiment, the training data 104 is generated by one or more path planning algorithms, methods, and / or systems based on one or more environments. In at least one embodiment, the training data 104 is a collection of data containing one or more reach distances from one or more accessible locations in an environment to specific target locations in the environment. In at least one embodiment, the training data 104 is implemented through one or more data structures, such as an array or list of data, or any preferred set, that encodes the reach distances and the locations corresponding to the reach distances. In at least one embodiment, the locations in an environment correspond to any preferred locations, regions, or areas in the environment. In at least one embodiment, the environment is represented by an image, and each location corresponds to a pixel and / or set of pixels in the image. In at least one embodiment, the environment is represented by a set of points (e.g., point cloud data), and each location corresponds to a point and / or set of points in the set of points. In at least one embodiment, the training data 104 is generated by one or more path planning algorithms, methods, and / or systems that output the reach for each accessible location in the environment to the target location, based on the input environment and the target location (for example, through images, sets of points, and / or variations thereof).

[0023] In at least one embodiment, the training framework 102 trains an implicit environment function 106 using the training data 104 to represent an environment with a specific target location (e.g., the environment and target location used to generate the training data 104) through the environment field. In at least one embodiment, the implicit environment function 106 is any suitable neural network model, algorithm, representation, function, and / or a variation thereof. In at least one embodiment, the implicit environment function includes a variety of neural network models, such as perceptron models, radial basis networks (RBNs), autoencoders (AEs), Boltzmann machines (BMs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), deep convolutional networks (DCNs), extreme learning machines (ELMs), deep residual networks (DRNs), support vector machines (SVMs), and / or variations thereof.

[0024] In at least one embodiment, the training framework 102 trains the implicit environment function 106 by inputting the locations of the training data 104 into the implicit environment function 106, causing the implicit environment function 106 to calculate the reach for the locations, comparing the calculated reach for the locations with the reach of the training data 104, also known as the ground truth reach, corresponding to the locations, calculating a loss using one or more loss functions based on the comparison, and updating the implicit environment function 106 based on the calculated loss. In at least one embodiment, the training framework 102 updates one or more weights, biases, and / or structural connections (e.g., the architecture and / or configuration of one or more components) of the implicit environment function 106 so that the calculated loss is minimized. In at least one embodiment, the training framework 102 calculates the loss using any preferred loss function, such as cross-entropy loss, binary cross-entropy loss, softmax loss, logistic loss, focus loss, and / or variations thereof.

[0025] In at least one embodiment, the implicit environment function 106 is trained when the calculated loss for the implicit environment function 106 falls below a defined threshold which may be any preferred value. In at least one embodiment, the training framework 102 trains the implicit environment function 106 to represent the reach for a position to a specific target position in the environment used to generate the training data 104. In at least one embodiment, the training framework 102 trains the implicit environment function 106 to determine an implicit environment function 108, which is the implicit environment function 106 after one or more training processes. In at least one embodiment, the implicit environment function 108 is the trained implicit environment function 106. In at least one embodiment, the implicit environment function 108 is specific to the environment and target position used to generate the training data 104.

[0026] In at least one embodiment, the implicit environment function 108 represents the environment (for example, the environment used to generate the training data 104) through the environment field. In at least one embodiment, the implicit environment function 108 represents the environment internally through the environment field representation. In at least one embodiment, the environment field is a representation of the environment, where each position in the environment field corresponds to a position in the environment, and each position in the environment field includes, or indicates, a reach value to a particular target position. In at least one embodiment, the reach value for a particular position in the environment field indicates the reach from the corresponding position in the environment to the target position in the environment.

[0027] In at least one embodiment, the implicit environment function 108 represents the environment through an environment field through one or more data structures and / or sets of data that encode the location of the environment and the reach of the location to a particular target location in the environment. In at least one embodiment, the environment field is visualized through an environment field visualization 112. In at least one embodiment, the environment field visualization 112 is an image, where each position in the environment field visualization 112 corresponds to a position in the environment field. In at least one embodiment, different colors, shades, patterns, and / or variations thereof of the environment field visualization 112 correspond to different reach values. In at least one embodiment, referring to Figure 1, the environment field visualization 112 includes brighter shading for lower reach values, darker shading for higher reach values, and completely dark shading for inaccessible areas. In at least one embodiment, the environmental field visualization can represent the reach values ​​of the environmental field in any preferred manner, such as a color gradient (e.g., lighter colors for lower reach values, darker colors for higher reach values, or any preferred manner), specific colors (e.g., specific colors for lower reach values, different specific colors for higher reach values), and / or variations thereof.

[0028] In at least one embodiment, one or more systems input a query location 110 to an implicit environment function 108 to determine the reach distance 114. In at least one embodiment, the query location 110 is a location in the environment and is implemented through one or more data objects, types, and / or variations thereof that encode the location. In at least one embodiment, the query location 110 includes coordinate data indicating the location. In at least one embodiment, the implicit environment function 108 processes the query location 110 to calculate the reach distance 114. In at least one embodiment, the reach distance 114 is a reach distance value and is implemented through one or more data objects, types, and / or variations thereof that encode the reach distance value. In at least one embodiment, the reach distance 114 includes a reach distance value from the location in the environment indicated by the query location 110 to a specific target location in the environment (for example, the environment and the specific target location used to generate the training data 104). In at least one embodiment, one or more systems utilize the determined reach distance 114 for various navigation tasks, as will be described in more detail with respect to Figures 4 to 11. In at least one embodiment, one or more systems utilize the determined reach distance value to calculate multiple paths that an entity, such as an autonomous device, is to traverse.

[0029] In at least one embodiment, one or more systems train an implicit function to predict different environmental fields for different target locations in the same environment, which is sometimes called a conditional implicit function. Figure 2 shows an example of an implicit environmental function for different target locations in an environment, according to at least one embodiment. In at least one embodiment, a training framework 202 trains an implicit environmental function 206 using training data 204 to determine an implicit environmental function 208. In at least one embodiment, one or more systems input a query location 210 and a target location 212 into the implicit environmental function 208, which represents the environment through environmental fields visualized by environmental field visualizations 214A-214B to determine the reach distance 216. In at least one embodiment, the training framework 202, training data 204, implicit environment function 206, implicit environment function 208, query position 210, target position 212, environment field visualizations 214A-214B, and reach distance 216 are as described with respect to Figure 1.

[0030] In at least one embodiment, the training framework 202 is a set of one or more hardware and / or software computing resources having instructions that, when executed, perform one or more neural network training processes, functions, and / or operations. In at least one embodiment, the training framework 202 trains an implicit environment function 206 using training data 204 to determine an implicit environment function 208. In at least one embodiment, the training data 204 is generated by one or more path planning algorithms, methods, and / or systems based on one or more environments. In at least one embodiment, the training data 204 is a set of data containing one or more reach distances to one or more accessible locations in an environment to one or more target locations in the environment. In at least one embodiment, the training data 204 is implemented through one or more data structures, such as an array or list of data, or any preferred set, that encode the reach distances and the corresponding locations and target locations.

[0031] In at least one embodiment, the training data 204 is generated by one or more path planning algorithms, methods, and / or systems that output one or more reach distances to each accessible location in the environment to the one or more target locations, based on the input environment (e.g., through images, sets of points, and / or variations thereof) and one or more target locations. In at least one embodiment, the training data 204 is generated by one or more systems by inputting the environment and the one or more randomly sampled target locations to one or more path planning algorithms, methods, and / or systems in order to determine one or more reach distances from each accessible location in the environment to one or more randomly sampled target locations.

[0032] In at least one embodiment, the training framework 202 trains an implicit environment function 206 using the training data 204 to represent an environment with an arbitrary preferred target position (for example, the environment used to generate the training data 204) through the environment field. In at least one embodiment, the implicit environment function 206 is any preferred neural network model, algorithm, representation, function, and / or a variation thereof. In at least one embodiment, the implicit environment function 206 includes a variety of neural network models, such as perceptron models, radial basis networks (RBNs), autoencoders (AEs), Boltzmann machines (BMs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), deep convolutional networks (DCNs), extreme learning machines (ELMs), deep residual networks (DRNs), support vector machines (SVMs), and / or variations thereof.

[0033] In at least one embodiment, the training framework 202 trains the implicit environment function 206 by inputting the positions of training data 204 and target positions into the implicit environment function 206, causing the implicit environment function 206 to calculate the reach for the position to the target position, comparing the calculated reach for the position to the target position with the reach of the training data 204, also known as the ground truth reach, corresponding to the position and the target position, calculating a loss using one or more loss functions based on the comparison, and updating the implicit environment function 206 based on the calculated loss. In at least one embodiment, the training framework 202 updates one or more weights, biases, and / or structural connections (e.g., the architecture and / or configuration of one or more components) of the implicit environment function 206 so that the calculated loss is minimized. In at least one embodiment, the training framework 202 calculates the loss using any suitable loss function, such as cross-entropy loss, binary cross-entropy loss, softmax loss, logistic loss, focus loss, and / or variations thereof.

[0034] In at least one embodiment, the implicit environment function 206 is trained when the calculated loss for the implicit environment function 206 falls below a defined threshold which may be any preferred value. In at least one embodiment, the training framework 202 trains the implicit environment function 206 to represent the distance to reach a position to any preferred target position in the environment used to generate the training data 204. In at least one embodiment, the training framework 202 trains the implicit environment function 206 to determine an implicit environment function 208, which is the implicit environment function 206 after one or more training processes. In at least one embodiment, the implicit environment function 208 is the trained implicit environment function 206. In at least one embodiment, the implicit environment function 208 is specific to the environment used to generate the training data 204.

[0035] In at least one embodiment, the implicit environment function 208 represents an environment (e.g., the environment used to generate training data 204) through an environment field. In at least one embodiment, the implicit environment function 208 represents an environment having any preferred target position internally through an environment field representation. In at least one embodiment, the implicit environment function 208 determines the environment field representation based on an input target position (e.g., target position 212). In at least one embodiment, the implicit environment function 208 represents the environment through an environment field through one or more data structures and / or sets of data that encode the position of the environment and the reach for the position to a specific target position in the environment (e.g., input via target position 212).

[0036] In at least one embodiment, the environment field is visualized through environment field visualizations 214A–214B. In at least one embodiment, the environment field visualization is an image, and each position in the environment field visualization corresponds to a position in the environment field. In at least one embodiment, different colors, shades, patterns, and / or variations thereof of the environment field visualization correspond to different reach values. In at least one embodiment, referring to Figure 2, the environment field visualizations 214A–214B include brighter shading for lower reach values, darker shading for higher reach values, and completely dark shading for inaccessible areas. In at least one embodiment, referring to Figure 2, the implicit environment function 208 determines an environment field to be visualized by environment field visualization 214A based on a first target position value (indicated by, for example, target position 212) corresponding to a target position at the upper left corner of the environment, and / or determines a different environment field to be visualized by environment field visualization 214B based on a second target position value (indicated by, for example, target position 212) corresponding to a target position at the lower right corner of the environment.

[0037] In at least one embodiment, one or more systems input the query location 210 and the target location 212 into an implicit environment function 208 to calculate the reach distance 216. In at least one embodiment, one or more systems concatenate the query location 210 and the target location 212. In at least one embodiment, the query location 210 is a location in the environment and is implemented through one or more data objects, types, and / or variations thereof that encode the location. In at least one embodiment, the query location 210 includes coordinate data indicating the location. In at least one embodiment, the target location 212 is a target location in the environment and is implemented through one or more data objects, types, and / or variations thereof that encode the location. In at least one embodiment, the target location 212 includes coordinate data indicating the target location.

[0038] In at least one embodiment, the implicit environment function 208 processes the query location 210 and the target location 212 to determine the reach distance 216 (for example, the environment field value at the query location 210). In at least one embodiment, the reach distance 216 is a reach distance value and is implemented through one or more data objects, types, and / or variations thereof that encode the reach distance value. In at least one embodiment, the reach distance 216 includes a reach distance value from a location in the environment indicated by the query location 210 to a specific target location indicated by the target location 212 in the environment (for example, the environment used to generate the training data 204). In at least one embodiment, one or more systems utilize the determined reach distance 216 for various navigation tasks, as will be described in more detail with respect to Figures 4 to 11.

[0039] In at least one embodiment, one or more systems train an implicit environment function, also called a context-aligned implicit function, which should generalize to any environment and / or any target location, in order to extend the implicit function to any environment. In at least one embodiment, given a scene context (e.g., a representation of the environment, such as an image or a set of points), one or more systems extract scene context features using a fully convolutional environment encoder. In at least one embodiment, one or more systems forward the concatenation of target coordinates, query location coordinates, and scene context features aligned to the query location to an implicit function, and regress the reach value for the query location. In at least one embodiment, the fully convolutional environment encoder expands the receptive field at the query location so that the implicit function has enough context information to predict the reach to the target. In at least one embodiment, scene environment feature matching causes the implicit function to focus on the query location while ignoring other scene context information. In at least one embodiment, one or more systems utilize various neural network models, such as hypernetworks, to encode arbitrary environments and objectives.

[0040] In at least one embodiment, one or more systems train a context-matched implicit function using different environments (e.g., mazes) along with the distances to be reached calculated by one or more path-planning algorithms, methods, and / or systems as training data. In at least one embodiment, during inference, the implicit environment function is used to predict or otherwise compute an environment field for an environment with a given target position and to explore feasible trajectories from an arbitrary starting position to the target position.

[0041] Figure 3 shows an example of an implicit environment function 300 for different target locations and environments according to at least one embodiment. In at least one embodiment, a training framework 302 trains an implicit environment function 306 using training data 304 to determine an implicit environment function 308. In at least one embodiment, one or more systems input a query location 316, a target location 318, and a feature 314 determined by an encoder 312 based on an environment image 310 into the implicit environment function 308, which then represents the environment through an environment field visualized by environment field visualizations 320A-320B to determine the reach distance 322. In at least one embodiment, the training framework 302, training data 304, implicit environment function 306, implicit environment function 308, query location 316, target location 318, environment field visualizations 320A-320B, and reach distance 322 are as described with respect to Figures 1 and 2.

[0042] In at least one embodiment, the training framework 302 is a set of one or more hardware and / or software computing resources having instructions that, when executed, perform one or more neural network training processes, functions, and / or actions. In at least one embodiment, the training framework 302 trains an implicit environment function 306 using training data 304 to determine an implicit environment function 308. In at least one embodiment, the training data 304 is a set of data containing one or more reach distances to one or more accessible locations in one or more environments to one or more target locations in said one or more environments. In at least one embodiment, the training data 304 is implemented through one or more data structures, such as an array or list of data, or any preferred set, that encodes the reach distances, locations, target locations, and / or environments.

[0043] In at least one embodiment, the training data 304 is generated by one or more path planning algorithms, methods, and / or systems that output one or more reach distances to each accessible location in the environment to one or more target locations, based on an input environment (e.g., through images, sets of points, environmental features, and / or variations thereof) and one or more target locations. In at least one embodiment, the training data 304 is generated by one or more systems by inputting the environment and the one or more randomly sampled target locations into one or more path planning algorithms, methods, and / or systems to determine one or more reach distances to each accessible location in the environment to one or more randomly sampled target locations. In at least one embodiment, the training data 304 is generated by one or more path planning algorithms, methods, and / or systems based on one or more environments. In at least one embodiment, for example, the training data 304 includes one or more reach distances for each accessible location in the first environment to one or more target locations in the first environment, one or more reach distances for each accessible location in the second environment to one or more target locations in the second environment, and so on for any number of environments. In at least one embodiment, one or more systems use various neural network models and / or other suitable systems to define accessible and / or inaccessible regions of the environment.

[0044] In at least one embodiment, the training framework 302 trains an implicit environment function 306 using training data 304 to represent any suitable environment with any suitable target position through the environment field. In at least one embodiment, the implicit environment function 306 is any suitable neural network model, algorithm, representation, function, and / or a variation thereof. In at least one embodiment, the implicit environment function 306 includes a variety of neural network models, such as perceptron models, radial basis networks (RBNs), autoencoders (AEs), Boltzmann machines (BMs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), deep convolutional networks (DCNs), extreme learning machines (ELMs), deep residual networks (DRNs), support vector machines (SVMs), and / or variations thereof.

[0045] In at least one embodiment, the training framework 302 trains an implicit environment function 306 by determining the features of each environment used to generate training data 304. In at least one embodiment, the training framework 302 determines or generates features from the environment, also called environment features, scene features, scene context features, and / or variations thereof, through an encoder (e.g., encoder 312). In at least one embodiment, the encoder is a neural network that processes an input and outputs a representation of the input, also called the input features, which exhibits various features of the input. In at least one embodiment, the input features may include indications of various details or other characteristics of the input, such as edges, several patterns, objects, obstacles, accessible / inaccessible regions, other features, and / or variations thereof. In at least one embodiment, the features are represented through a feature map, a set of feature vectors, or any preferred representation. In at least one embodiment, the encoder is a convolutional neural network (CNN), a recurrent neural network (RNN), and / or a variation thereof. In at least one embodiment, the encoder is a convolution-based encoder. In at least one embodiment, the training framework 302 inputs the environment to the encoder (e.g., via an image, a set of points, or other representation) to determine the features of the environment, the features of the environment being represented through any preferred representation such as a feature map output by the encoder, or one or more feature vectors.

[0046] In at least one embodiment, the training framework 302 performs various feature matching processes on the determined features based on the query location. In at least one embodiment, the training framework 302 performs one or more feature matching processes on features determined from the environment based on the query location, which attenuate or otherwise enhance the portion of the features corresponding to the query location. In at least one embodiment, the feature map represents the features of each location in the environment, and the training framework 302 matches the feature map by attenuating or otherwise enhancing a particular aspect of the feature map corresponding to the query location (e.g., the features of the query location) based on the query location in the environment.

[0047] In at least one embodiment, the training framework 302 trains the implicit environment function 306 by inputting the locations of training data 304, target locations, and environmental features (e.g., features that can be matched based on location) into the implicit environment function 306; causing the implicit environment function 306 to calculate the reach for the location to the target location; comparing the calculated reach for the location to the target location with the reach of the training data 304, also known as the ground truth reach, corresponding to the location, the target location, and the environment; calculating a loss using one or more loss functions based on the comparison; and updating the implicit environment function 306 based on the calculated loss. In at least one embodiment, the training framework 302 updates one or more weights, biases, and / or structural connections (e.g., the architecture and / or configuration of one or more components) of the implicit environment function 306 so that the calculated loss is minimized. In at least one embodiment, the training framework 302 calculates the loss using any preferred loss function, such as cross-entropy loss, binary cross-entropy loss, softmax loss, logistic loss, focus loss, and / or variations thereof.

[0048] In at least one embodiment, the implicit environment function 306 is trained when the calculated loss for the implicit environment function 306 falls below a defined threshold which may be any preferred value. In at least one embodiment, the training framework 302 trains the implicit environment function 306 to represent the reach for a position to any preferred target position for any preferred environment. In at least one embodiment, the training framework 302 trains the implicit environment function 306 to determine the implicit environment function 308, which is the implicit environment function 306 after one or more training processes. In at least one embodiment, the implicit environment function 308 is the trained implicit environment function 306. In at least one embodiment, the implicit environment function 308 handles environments which may or may not be used to generate training data 304.

[0049] In at least one embodiment, the implicit environment function 308 represents an environment through an environment field. In at least one embodiment, the implicit environment function 308 represents any suitable environment having any suitable target position internally through an environment field representation. In at least one embodiment, the implicit environment function 308 determines the environment field representation based on an input target position (e.g., target position 318) and a feature of the environment (e.g., feature 314). In at least one embodiment, the implicit environment function 308 represents the environment through an environment field through one or more data structures and / or sets of data that encode the position of the environment (e.g., the environment corresponding to feature 314) and the reach for the position to a specific target position in the environment (e.g., input via target position 318).

[0050] In at least one embodiment, the environment field is visualized through environment field visualizations 320A–320B. In at least one embodiment, the environment field visualization is an image, and each position in the environment field visualization corresponds to a position in the environment field. In at least one embodiment, different colors, shades, patterns, and / or variations thereof of the environment field visualization correspond to different reach values. In at least one embodiment, referring to Figure 3, the environment field visualizations 320A–320B include brighter shading for lower reach values, darker shading for higher reach values, and completely dark shading for inaccessible areas. In at least one embodiment, referring to Figure 3, the implicit environment function 308 determines an environment field to be visualized by environment field visualization 320A based on features of the environment (e.g., indicated by feature 314) corresponding to the environment image 310A determined by encoder 312 and a first target position value (e.g., indicated by target position 318) corresponding to a target position in the lower right corner of the environment, and / or determines a different environment field to be visualized by environment field visualization 320B based on features of the environment (e.g., indicated by feature 314) corresponding to the environment image 310B determined by encoder 312 and a second target position value (e.g., indicated by target position 318) corresponding to a target position in the upper right corner of the environment.

[0051] In at least one embodiment, the environment image 310 includes an image or other representation of the environment, such as a set of points. In at least one embodiment, the representation of the environment is referred to as the environment context and / or scene context. In at least one embodiment, one or more systems input the environment to an encoder 312 (e.g., via an image or other representation of the environment image 310 representing the environment) in order to determine the features of the environment, and the features of the environment are represented through any preferred representation, such as a feature map output by the encoder 312 or one or more feature vectors. In at least one embodiment, one or more systems perform various feature matching processes on the determined features based on a query location (e.g., indicated by a query location 316). In at least one embodiment, one or more systems perform one or more feature matching processes on features determined from the environment (e.g., represented via an image or other representation) based on the query location, which attenuates or otherwise enhances the portion of the features corresponding to the query location. In at least one embodiment, the feature map represents the features of each location in the environment, and one or more systems match the feature map by attenuating or enhancing a particular aspect of the feature map corresponding to the query location (e.g., the features of the query location) based on the query location in the environment.

[0052] In at least one embodiment, feature 314 includes a feature (e.g., a feature map) output by encoder 312 based on an environment (e.g., an environment represented by a representation of environment image 310). In at least one embodiment, feature 314 includes a feature (e.g., a matched feature map) output by encoder 312 based on an environment (e.g., an environment represented by a representation of environment image 310) processed through one or more matching processes based on query location 316.

[0053] In at least one embodiment, one or more systems input feature 314, query location 316, and target location 318 to an implicit environment function 308 to calculate the reach distance 322. In at least one embodiment, one or more systems concatenate the input feature 314, query location 316, and target location 318. In at least one embodiment, query location 316 is a location in the environment and is implemented through one or more data objects, types, and / or variations thereof that encode the location. In at least one embodiment, query location 316 includes coordinate data indicating the location. In at least one embodiment, target location 318 is a target location in the environment and is implemented through one or more data objects, types, and / or variations thereof that encode the location. In at least one embodiment, target location 318 includes coordinate data indicating the target location.

[0054] In at least one embodiment, the implicit environment function 308 processes the feature 314, the query location 316, and the target location 318 to determine the reach 322 (for example, the environment field value at the query location 316). In at least one embodiment, the reach 322 is a reach value, which is implemented through one or more data objects, types, and / or variations thereof that encode the reach value. In at least one embodiment, the reach 322 includes a reach value for the environment corresponding to the feature 314 (for example, the environment used to generate the feature 314), from a location in the environment indicated by the query location 316 to a specific target location indicated by the target location 318 in the environment.

[0055] In at least one embodiment, the query location 316 indicates one or more locations in the environment corresponding to the feature 314 (for example, the environment used to generate the feature 314), and the reach 322 includes the reach value for each of the one or more locations to a specific target location indicated by the target location 318 in the environment. In at least one embodiment, the feature 314 includes one or more features, each aligned to a specific location of the query location 316. In at least one embodiment, the implicit environment function 308 calculates any number of reach values ​​for any number of query locations to a specific target location in a single forward path. In at least one embodiment, one or more systems utilize the determined reach 322 to determine multiple paths for various navigation tasks, as will be described in more detail with respect to Figures 4 to 11.

[0056] Figure 4 shows an example 400 of an agent navigating to a target using an implicit environment function, according to at least one embodiment. In at least one embodiment, agent 404 navigates to target 408 using an implicit environment function, such as those described elsewhere in this disclosure. In at least one embodiment, agent 404 is any preferred entity capable of navigating an environment. In at least one embodiment, agent 404 is an autonomous device associated with one or more systems that implement or otherwise utilize one or more implicit environment functions. In at least one embodiment, the autonomous device is a device such as an autonomous vehicle, an autonomous aircraft (e.g., a drone), a robot, and / or a variant thereof. In at least one embodiment, agent 404 comprises one or more systems having one or more hardware and / or software resources that implement or otherwise utilize one or more implicit environment functions. In at least one embodiment, agent 404 communicates with one or more systems having one or more hardware and / or software resources that implement or otherwise utilize one or more implicit environment functions. In at least one embodiment, the implicit environment functions are trained through one or more processes, such as those described with reference to Figures 1 to 3.

[0057] In at least one embodiment, environment 402 is any preferred environment, such as various indoor environments, outdoor environments, maze environments, and / or variations thereof. In at least one embodiment, environment 402 is a physical environment. In at least one embodiment, agent 404 is a virtual entity, and environment 402 is a virtual environment or otherwise a simulated environment. In at least one embodiment, environment 402 includes various inaccessible areas, shown in Figure 4 by fully darkened shaded areas. In at least one embodiment, the inaccessible areas correspond to areas inaccessible to agent 404. In at least one embodiment, for example, agent 404 is an autonomous vehicle, and the inaccessible areas correspond to obstacles, objects, and / or variations thereof that the autonomous vehicle cannot pass over or through, such as buildings or other vehicles. In at least one embodiment, environment 402 includes a target 408.

[0058] In at least one embodiment, target 408 indicates a location in environment 402, also called a location. In at least one embodiment, target 408 points to a location in environment 402 to which agent 404 is to navigate. In at least one embodiment, one or more systems define a navigation task indicating that agent 404 is to navigate to target 408 through environment 402. In at least one embodiment, one or more systems input features of environment 402 into an implicit environment function. In at least one embodiment, one or more systems determine the features of environment 402 by processing a representation of environment 402. In at least one embodiment, environment 402 is represented in any preferred form, such as through an image, a set of points, a point cloud, a set of voxels, and / or variations thereof. In at least one embodiment, the environment is represented through any preferred 2D representation, such as an image, a set of points, and / or a modified form thereof, or through any preferred 3D representation, such as an image with depth information, a point cloud, a set of voxels, and / or a modified form thereof.

[0059] In at least one embodiment, one or more systems determine the features by inputting a representation of the environment 402 to one or more encoders that output features of the environment 402. In at least one embodiment, one or more systems each perform various feature matching processes on the features to determine one or more matched features matched to a particular location in the environment 402. In at least one embodiment, one or more systems input the features of the environment and a location indication of the target 408 into an implicit environment function. In at least one embodiment, one or more systems input one or more location indications of the environment 402, one or more features (e.g., the features matched to the one or more locations), and a location indication of the target 408 into an implicit environment function. In at least one embodiment, the location indication refers to any preferred indication of the location, such as coordinate data or other indications.

[0060] In at least one embodiment, the shadow environment function outputs a reach value for each location in the environment 402 up to the target 408 (e.g., each accessible location). In at least one embodiment, the shadow environment function outputs a reach value for each location input to the shadow environment function. In at least one embodiment, one or more systems visualize the environment field for the environment 402 using the reach values ​​by visualizing each reach value for each location in the environment 402 through a particular color scheme, shading scheme, and / or variations thereof. In at least one embodiment, for example, one or more systems visualize the environment field by assigning different shades to particular reach values, such as brighter shading for lower reach values, darker shading for higher reach values, and completely dark shading for inaccessible areas.

[0061] In at least one embodiment, one or more systems utilize reach values ​​to guide agent 404 to target 408. In at least one embodiment, one or more systems determine reach values ​​for each location accessible to agent 404 from agent 404's current location. In at least one embodiment, one or more systems define a step size for agent 404 that corresponds to the distance of each step of movement that agent 404 can take, and the one or more systems determine reach values ​​for each location accessible to agent 404 through the step from agent 404's current location, with the step size corresponding to the defined step size. In at least one embodiment, a step refers to a step of movement or navigation of an entity. In at least one embodiment, the step size refers to how far the entity moves with each step. In at least one embodiment, the step size can be any preferred distance value. In at least one embodiment, one or more systems represent the environment 402 through a grid, each grid cell represents a location, and agent 404 can move to any grid cell that is immediately accessible from agent 404's current grid cell. In at least one embodiment, the size of the grid cell can be any preferred value. In at least one embodiment, agent 404 can move from the center of a grid cell to the center of an adjacent accessible grid cell. In at least one embodiment, one or more systems determine the reach value for each grid cell accessible to agent 404 from agent 404's current grid cell.

[0062] In at least one embodiment, one or more systems determine a reachable value for each location accessible through steps from the agent 404's current position. In at least one embodiment, the size of agent 404's steps corresponds to the distance traveled through one or more movements of agent 404 and can be any preferred distance value. In at least one embodiment, for example, agent 404 is an autonomous robot with robotic legs, wheels, and / or other hardware for movement, and the size of the steps corresponds to the distance traveled through one or more steps of the robotic legs, the distance traveled through one or more rotations of the wheels, and / or variations thereof. In at least one embodiment, one or more systems determine a location with the minimum reachable value of a set of locations accessible through steps from agent 404's current position, determine a path to the location, and cause agent 404 to navigate to the location using the path. In at least one embodiment, one or more systems determine the location having the minimum reach value in a set of locations by comparing each reach value at each location in the set of locations and determining which location in the set of locations has the minimum reach value. In at least one embodiment, locations are determined using the minimum reach value, the maximum reach value, a reach value below a threshold, or any preferred reach value. In at least one embodiment, one or more systems utilize reach values ​​to calculate multiple paths (e.g., paths 406A–406E) that an entity (e.g., an agent) is going to traverse or otherwise navigate.

[0063] In at least one embodiment, one or more systems cause the agent to navigate to one or more locations until the agent is within a step size from the target and can navigate to the target, and the one or more systems cause the agent to navigate to the target. In at least one embodiment, for example, for a navigation task, one or more systems determine a first location having a minimum reach value for a first set of locations accessible through steps from the agent's current location, determine a first path to the first location, and cause the agent to navigate to the first location via the first path, the one or more systems then determine a second location having a minimum reach value for a second set of locations accessible through steps from the agent's new current location, determine a second path to the second location, and cause the agent to navigate to the second location via the second path, and so on until the agent is within a step size from the target location, and the one or more systems cause the agent to navigate to the target location to complete the navigation task.

[0064] In at least one embodiment, referring to Figure 4, path 406A represents a path to a location having the minimum reach value of a set of locations accessible through steps from agent 404's current location. In at least one embodiment, referring to Figure 4, one or more systems cause agent 404 to navigate to a first location via path 406A. In at least one embodiment, referring to Figure 4, one or more systems then determine a second location having the minimum reach value of a set of locations accessible through steps from agent 404's first location, and determine path 406B from the first location to the second location. In at least one embodiment, referring to Figure 4, one or more systems cause agent 404 to navigate to the second location via path 406B. In at least one embodiment, referring to Figure 4, one or more systems then determine a third location having the minimum reach value of a set of locations accessible through steps from agent 404's second location, and determine path 406C from the second location to the third location. In at least one embodiment, referring to Figure 4, one or more systems cause agent 404 to navigate to a third location via path 406C. In at least one embodiment, referring to Figure 4, one or more systems then determine a fourth location having the minimum reachable distance value of a set of locations accessible through steps from agent 404's third location, and determine path 406D from the third location to the fourth location. In at least one embodiment, referring to Figure 4, one or more systems cause agent 404 to navigate to the fourth location via path 406D. In at least one embodiment, referring to Figure 4, one or more systems then determine that target 408 is accessible through steps from agent 404's fourth location, and determine path 406E from the fourth location to target 408. In at least one embodiment, referring to Figure 4, one or more systems cause agent 404 to navigate to target 408 via path 406E.

[0065] In at least one embodiment, one or more systems navigate the agent through the environment by determining one or more reach values ​​for one or more locations accessible from the agent's current location, one or more reach values ​​for one or more locations accessible through steps of a predefined step size from the agent's current location, representing the environment through a grid, and determining one or more reach values ​​for one or more grid cells accessible from the agent's current grid cell, and / or variations thereof. In at least one embodiment, one or more systems use implicit environment functions to compute multiple paths, also called trajectories and / or paths, for the agent to traverse or otherwise navigate to a target.

[0066] Figure 5 shows an example of an environment field for a multi-person navigation environment according to at least one embodiment. In at least one embodiment, Figure 5 shows the environment field from a top-down view, also known as a bird's-eye view, of the environment. In at least one embodiment, scenes 502A, 502B, environment field visualization 504A, and environment field visualization 504B are as described elsewhere in this disclosure.

[0067] In at least one embodiment, Scene 502A and / or Scene 502B are top-down views of an environment. In at least one embodiment, Scene 502A and / or Scene 502B are top-down views of a 3D environment. In at least one embodiment, one or more systems capture Scene 502A and / or Scene 502B using a depth camera or other suitable hardware capable of capturing depth information. In at least one embodiment, the images of Scene 502A and / or Scene 502B include depth information associated with each pixel of the image.

[0068] In at least one embodiment, to model human gait, one or more systems define a floor area as an accessible region for a human in a bird's-eye view of a 3D scene. In at least one embodiment, one or more systems determine the depth of each pixel in the bird's-eye rendering and label the pixel with the greatest depth as the floor. In at least one embodiment, one or more systems indicate all pixels belonging to the floor area as accessible and other pixels as obstacles, also called inaccessible. In at least one embodiment, one or more systems utilize a bird's-eye view image with the indicated accessible and inaccessible regions to train an implicit environment function. In at least one embodiment, one or more systems utilize an arbitrary target location and an arbitrary environment to train an implicit environment function to determine the environment field of any bird's-eye view image.

[0069] In at least one embodiment, one or more systems, during inference, given a starting position and a target position, search for a feasible trajectory that guides a human to navigate from the starting position to the target while avoiding collisions with obstacles in the room. In at least one embodiment, one or more systems define a human step size that may approximate the average human step size and move the human one step in the direction that results in the maximum reach reduction. In at least one embodiment, one or more systems repeat this process until the human reaches the target. In at least one embodiment, given the searched trajectory in a bird's-eye view, one or more systems align an existing human walking sequence on top of it by rotating the human posture tangentially to the trajectory at each time step.

[0070] In at least one embodiment, one or more systems utilize an implicit environment function to navigate multiple people in a multi-person environment. In at least one embodiment, Figure 5 illustrates a navigation task in a multi-person environment. In at least one embodiment, Scene 502A shows a scene at a first time, where person 506A is to navigate to a target. In at least one embodiment, one or more systems represent person 508A as an obstacle. In at least one embodiment, one or more systems input the features of Scene 502A into an implicit environment function to determine a reach value and generate an environment field visualization 504A based on the reach value. In at least one embodiment, one or more systems navigate person 506A based on the environment field visualization 504A. In at least one embodiment, Scene 502B shows a scene at a second time, where person 506B (corresponding, for example, person 506A) is moving as part of the navigation task, and person 508B (corresponding, for example, person 508A) is also moving. In at least one embodiment, one or more systems represent a human 508B as an obstacle. In at least one embodiment, one or more systems input features of scene 502B into an implicit environment function to determine reach values ​​and generate an environment field visualization 504B based on the reach values. In at least one embodiment, one or more systems navigate human 506B based on the environment field visualization 504B. In at least one embodiment, the environment field changes dynamically based on changes to the environment. In at least one embodiment, the environment field can be generated any number of times for any number of changes in the environment by inputting features corresponding to the current state of the environment. In at least one embodiment, during training of the implicit environment function, one or more systems represent a random space as an obstacle to mimic any scene with any obstacles that may be encountered during inference.

[0071] Figure 6 shows an example 600 of the use of an implicit environment field for a 3D environment according to at least one embodiment. In at least one embodiment, one or more systems utilize an implicit environment function to process the 3D environment 602 to determine an accessible region visualization 604, an environment field visualization 606, and a trajectory visualization 608. In at least one embodiment, the environment field visualization 606 is as described elsewhere in this disclosure.

[0072] In at least one embodiment, the 3D environment 602 is any preferred environment, such as various indoor environments, outdoor environments, maze environments, and / or variations thereof. In at least one embodiment, the 3D environment 602 is a physical environment. In at least one embodiment, the 3D environment 602 is an environment in which an entity is to navigate to a target as part of a navigation task.

[0073] In at least one embodiment, one or more systems define an accessible region about a 3D environment, also called a 3D scene. In at least one embodiment, one or more systems train a generative model to generate an accessible region about a human in the 3D environment. In at least one embodiment, one or more systems learn the distribution of human torso locations in a 3D scene, scene context-dependent, through a variational autoencoder (VAE). In at least one embodiment, one or more systems generate features from a scene point cloud as scene context features using one or more layers, such as a pooling layer of a network like PointNet. In at least one embodiment, one or more systems use the encoder to map the context features, along with human torso location observations, to a normal distribution. In at least one embodiment, one or more systems utilize a decoder that reconstructs human torso locations given a concatenation of sampled noise from a normal distribution with the scene context features. In at least one embodiment, one or more systems train a VAE using a reconstruction objective regarding the location of a human torso, along with objectives such as a Kullback-Leibler (KL) divergence objective. In at least one embodiment, one or more systems sample random noise from a standard normal distribution during inference to generate feasible locations for a human torso in order to create an accessible area in a 3D scene. In at least one embodiment, one or more systems project all generated locations onto a bird's-eye view and remove locations that hit furniture or walls in order to further suppress noisy generation by the VAE model and filter out locations that could potentially lead to collisions with furniture in the room.In at least one embodiment, the accessible area visualization 604 is a visualization of the generated accessible area, and the VAE predicts reasonable locations suitable for a person to sit, walk, etc., while avoiding collisions with furniture in the room. In at least one embodiment, the generated accessible area is uniformly spread throughout the 3D room and reasonable for all kinds of human actions. In at least one embodiment, for example, a point on an obstacle such as a chair or sofa is a feasible location for the torso of a seated person, while a point in the air is a possible location for the torso of a walking person.

[0074] In at least one embodiment, one or more systems train an implicit environment function to model a complete 3D space. In at least one embodiment, one or more systems train an implicit environment function to model a specific 3D environment. In at least one embodiment, one or more systems utilize an arbitrary target position in a fixed scene environment for training the implicit environment function. In at least one embodiment, the input to the implicit function is a concatenation of target coordinates and query position coordinates. In at least one embodiment, the output is the distance reached from the query position to the target. In at least one embodiment, one or more systems discretize the 3D scene into a voxel grid of any suitable size (e.g., a 64x64x64 voxel grid), mark all voxel cells within the generated accessible area as accessible, and mark other voxel cells as obstacles, also called inaccessible. In at least one embodiment, one or more systems determine a trajectory for navigating an entity from a starting location to a target by querying an implicit function for reach at all possible locations the entity can reach, and by navigating the entity to a position with the minimum reach until the entity reaches the target. In at least one embodiment, one or more systems set the step size to an average human step size and determine all possible directions a human can take in 3D space. In at least one embodiment, the environment field visualization 606 is a visualization of the environment field determined by an implicit environment function trained to model a particular 3D environment. In at least one embodiment, the environment field visualization 606 shows brighter shading for lower reach values ​​and darker shading for higher reach values. In at least one embodiment, in the environment field, the closer a point is to the target, the smaller its reach value.

[0075] In at least one embodiment, trajectory visualization 608 is a visualization of a trajectory calculated by one or more systems for an entity to navigate to a target. In at least one embodiment, one or more systems calculate a trajectory that includes multiple paths. In at least one embodiment, one or more systems optimize the trajectory based on a given pose sequence to ensure that the human is navigated toward a physically attainable location, which results in a trajectory that appears reasonable and consistent with a given human pose sequence. In at least one embodiment, one or more systems use a step size calculated from adjacent locations in the pose sequence instead of a predefined average step size at each step. In at least one embodiment, for each possible location that the human can reach within one step size, one or more systems check whether the human is well supported and whether the human is not colliding with other objects at these locations, and using these constraints and / or other constraints, move the human to the location with the minimum reach value. In at least one embodiment, one or more systems check whether the torso, left knee, and / or right knee joints in a seated posture, and the ankle joints in a standing posture, have a non-positive signed distance to the scene surface in order to determine whether the posture is well supported. In at least one embodiment, one or more systems check whether all joints of a posture, excluding various support joints, have a non-negative signed distance in order to determine whether the posture will collide with other objects. In at least one embodiment, one or more systems determine a trajectory visualized in trajectory visualization 608 for an entity (e.g., a human) to navigate to a target, based on various entity postures. In at least one embodiment, the trajectory successfully avoids collisions with obstacles in the room (e.g., furniture) while guiding the entity (e.g., a human) toward the target. In at least one embodiment, the entity posture (e.g., a human posture) is deemed reasonable at each location along the trajectory.

[0076] Figure 7 shows an example 700 of results using implicit environment functions in at least one embodiment. In at least one embodiment, Tables 702, 704, 706, and 708 show various results using implicit environment functions, including those described herein.

[0077] In at least one embodiment, one or more systems use any preferred learning rate (e.g., 5 × 10⁻¹⁰). -5 The conditional VAE and all implicit functions are trained using a solver such as the Adam solver with a learning rate of 1 / 2. In at least one embodiment, one or more systems utilize a schedule such as a Cyclical Annealing Schedule to balance the KL divergence objective and the reconstruction objective in the VAE.

[0078] In at least one embodiment, one or more systems evaluate an implicit environment function (IEF) for a dataset such as the Minigrid and Gridworld maze datasets, which include 2D mazes with randomly placed obstacles. In at least one embodiment, the implicit environment function determines an environment field that accurately encodes the reach from any accessible point to a target while successfully detecting all obstacles in the maze. In at least one embodiment, one or more systems use the success rate (e.g., the percentage of all explored paths that successfully lead the agent to the target) as an evaluation metric.

[0079] In at least one embodiment, Table 1702 shows the results of an ablation study on the network architecture. In at least one embodiment, the 11x11 maze is from a dataset such as the MiniGrid dataset, while all other mazes are from datasets such as the GridWorld dataset. In at least one embodiment, "IF" indicates an implicit function. In at least one embodiment, the context-aligned implicit functions perform comparably to the hypernetwork on smaller maze sizes (e.g., 8x8 and 11x11 mazes). In at least one embodiment, the hypernetwork fails on larger maze sizes (e.g., 16x16 and 28x28 mazes), where the context-aligned implicit functions perform better than the hypernetwork.

[0080] In at least one embodiment, one or more systems evaluate an implicit environment function with a network such as a value iteration network (VIN) on a dataset such as a minigrid dataset. In at least one embodiment, one or more systems learn different implicit functions for mazes of different sizes, utilizing the same train and test split as VIN. In at least one embodiment, Table 2 704 shows the success rate and average time required to explore paths in different ways. In at least one embodiment, a batch size of 1 is used, or any preferred value. In at least one embodiment, the implicit environment function achieves a comparable success rate in much less time for larger mazes compared to VIN. In at least one embodiment, the implicit environment function requires only one network forwarding path to obtain the reach value for all locations in the maze and can guide agents at any location to reach the goal.

[0081] In at least one embodiment, one or more systems utilize an implicit environment function to model long-term dynamic human motion with respect to a dataset such as a Proximal Relationships with Object eXclusion (PROX) dataset. In at least one embodiment, the implicit environment function can predict long-term dynamic human motion over any number of video frames. In at least one embodiment, the implicit environment field represents the environment through an environment field that encodes the reach between any points to a target position. In at least one embodiment, areas closer to the target have lower reach values, and vice versa. In at least one embodiment, the environment field changes dynamically based on the positions of people in the environment. In at least one embodiment, inaccessible areas correspond to areas with larger reach values ​​(e.g., above a certain threshold). In at least one embodiment, the implicit environment function can model the environment field in a scene and generalize to a dynamically changing environment.

[0082] In at least one embodiment, Table 3 706 shows a quantitative comparison with sampling-based path planning methods, such as fast search random trees (RRT) and probabilistic roadmaps (PRM). In at least one embodiment, AFF represents attitude-dependent search. In at least one embodiment, one or more systems evaluate metrics such as the distance between the end of the searched path and the target, and the average time cost for a single trajectory search. In at least one embodiment, one or more systems set the number of sampling points and nearest neighbors in the PRM to any preferred value, such as 500 and 5, respectively. In at least one embodiment, one or more systems set the maximum iterations for the trajectory search in the RRT to any preferred value, such as 500 iterations. In at least one embodiment, the negative environment field is more efficient because each step can be predicted by a single fast network forward path.

[0083] In at least one embodiment, one or more systems quantitatively compare trajectory predictions with a network such as PathNet using a distance metric and a ratio of valid trajectories to all trajectories on the trajectory. In at least one embodiment, one or more systems define a trajectory as valid if it is both well supported and not collision-free in the scene. In at least one embodiment, one or more systems utilize training and test trajectories, such as those of a scene-context-based long-term human motion prediction (HMP) method in which each trajectory consists of 30 frames. In at least one embodiment, PathNet in HMP requires torso location in the first 10 frames of each trajectory as input and predicts torso location for the subsequent 20 frames. In at least one embodiment, one or more systems predict any number of locations (e.g., 30 locations) that guide a human from a starting position in the first frame to a target in the last frame. In at least one embodiment, Table 4 708 shows the quantitative evaluation. In at least one embodiment, the implicit environment function is more effective in navigating a human towards a target position. In at least one embodiment, after using a posture-dependent trajectory search process, the implicit environment function can better fit the human posture to the scene, as shown in the last column of Table 4 708.

[0084] In at least one embodiment, the implicit environment field represents the environment through an environment field that encodes the reach between pairs of points in either 2D or 3D space. In at least one embodiment, the learned environment field is a continuous energy surface that allows an agent to navigate a 2D maze in a dynamically changing scene environment. In at least one embodiment, the environment field is extended to a 3D scene to model dynamic human motion in an indoor environment.

[0085] In at least one embodiment, one or more systems utilize a network such as a hypernetwork to encode different maze environments and target locations. In at least one embodiment, the hypernetwork refers to a neural network that generates a network for the main network, also called a hyponetwork. In at least one embodiment, for each maze map, one or more systems use the hypernetwork to predict the parameters of the hyponetwork (e.g., an implicit environment function). In at least one embodiment, the hyponetwork takes a concatenation of target coordinates and query location coordinates as input and outputs the distance to reach between them. In at least one embodiment, the hypernetwork includes any preferred layers, such as 11 convolutional layers followed by 7 fully connected layers. In at least one embodiment, the hyponetwork includes any preferred layers, such as 7 fully connected layers, each of which is followed by periodic activation.

[0086] In at least one embodiment, instead of using the reach calculated by the FMM as supervision to train the implicit environment function, one or more systems utilize the reciprocal of the reach from a feasible point to a target, normalize it to the range [0,1], and train the implicit environment function to predict negative values ​​instead of very small values ​​for locations occupied by obstacles. In at least one embodiment, the implicit environment function can discriminate between feasible locations and those with obstacles.

[0087] In at least one embodiment, during training of the VAE model, one or more systems also augment more points in the scene, in addition to the points on the human trajectory being trained. In at least one embodiment, one or more systems utilize a random standing human posture and include all locations that can give this posture (e.g., the posture may be supported by the floor and not collide with other objects). In at least one embodiment, one or more systems train a conditional VAE model for all the trajectories being trained in a variety of different scenes, but a separate implicit function is trained for each 3D indoor scene. In at least one embodiment, one or more models are 5 × 10 -5 The system is trained using a solver such as the Adam solver, which uses a learning rate of . In at least one embodiment, one or more systems implement one or more models using a framework such as the PyTorch framework, or any preferred framework.

[0088] Figure 8 shows an example 800 of the results of agent navigation according to at least one embodiment. In at least one embodiment, path 802 shows various mazes, a ground truth path (shown as, for example, a straight line), and a explored path (shown as, for example, a dotted line). In at least one embodiment, learned level set 804 shows various mazes and explored paths. In at least one embodiment, learned IEF 806 shows various mazes and an environment field determined by an implicit environment function. In at least one embodiment, the implicit environment function determines an environment field that captures the reach from any point to a target and guides the agent from the starting position to the target.

[0089] Figures 9 and 10 illustrate an example of fitting a given human sequence to a explored trajectory in a 3D indoor environment, according to at least one embodiment. In at least one embodiment, given a human sequence, one or more systems calculate the human step size at each time step and dynamically determine the best next move in terms of step size.

[0090] In at least one embodiment, one or more systems utilize a given human posture to avoid moving the human to a location that violates physical constraints in the room (e.g., a location that cannot adequately support the human or would lead to collision with another object). In at least one embodiment, one or more systems consider vertices belonging to body parts instead of joints when it is determined that the location can give other preferred representations, such as a human body mesh, or a skeleton, skeletal mesh, and / or their deformed forms. In at least one embodiment, one or more systems utilize a human body mesh, a skeleton, skeletal mesh, and / or any preferred representation of a human. In at least one embodiment, for example, to check whether a seated human is adequately supported, one or more systems check whether vertices belonging to the gluteal region have non-positively signed distance values ​​to the object surface, and to check whether the human is colliding with another object, one or more systems check whether vertices belonging to the legs and thighs have non-negatively signed distance values.

[0091] In at least one embodiment, Figures 9 and 10 visualize the environment field in the last two rows. In at least one embodiment, the reach decreases as the point gets closer to the target position (e.g., less shading) and vice versa. In at least one embodiment, the arrow points to the target position (e.g., the position of the human torso in the last time step).

[0092] Figure 11 shows an example of process 1100 for calculating multiple paths using an implicit environment function, according to at least one embodiment. In at least one embodiment, all or part of process 1100 (or any other processes described herein, or variations and / or combinations thereof) is carried out under the control of one or more computer systems consisting of computer-executable instructions and is implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) that is collectively executed by hardware, software, or a combination thereof on one or more processors. In at least one embodiment, the code is stored in a computer-readable storage medium in the form of a computer program comprising multiple computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-temporary computer-readable medium. In at least one embodiment, at least some computer-readable instructions available for carrying out process 1100 are not stored using only temporary signals (e.g., transient electrical or electromagnetic transmissions that propagate). In at least one embodiment, a non-temporary computer-readable medium does not necessarily include non-temporary data storage circuitry (e.g., buffers, caches, and queues) within a temporary signal transceiver.

[0093] In at least one embodiment, process 1100 is carried out by one or more systems, such as those described herein. In at least one embodiment, one or more systems include any preferred system having a set of one or more hardware and / or software resources having instructions that, when executed, perform various implicit environment function training operations, implicit environment function processing operations, neural network functions, navigation operations (e.g., causing an entity to navigate to a location), environment processing functions, and / or other operations, such as those described herein. In at least one embodiment, process 1100 is carried out by one or more systems associated with an autonomous device.

[0094] In at least one embodiment, a system performing at least a portion of process 1100 includes, in 1102, executable code for obtaining at least a first location, a set of locations, and a final location. In at least one embodiment, a location is a location in an environment, and may be referred to as a position. In at least one embodiment, the environment is any preferred environment, such as a 2D environment, a 3D environment, an indoor environment, an outdoor environment, a simulated environment, a maze, and / or variations thereof. In at least one embodiment, the environment comprises an autonomous device. In at least one embodiment, the autonomous device is a system such as an autonomous vehicle, an autonomous robot, and / or variations thereof. In at least one embodiment, the autonomous device is to navigate in the environment from a first location to a final location, also referred to as a target position or location, as part of a navigation task. In at least one embodiment, the first location is the location of the autonomous device. In at least one embodiment, a set of locations is one or more locations in the environment. In at least one embodiment, a set of locations includes one or more locations in an accessible area of ​​the environment.

[0095] In at least one embodiment, a system performing at least a portion of process 1100 includes executable code in 1104 that causes at least one or more neural networks to calculate a set of distances based at least partially on a set of locations and a final location. In at least one embodiment, one or more neural networks include an implicit environment function. In at least one embodiment, the system inputs the set of locations and the final location into the implicit environment function. In at least one embodiment, the system inputs features of the environment into the implicit environment function. In at least one embodiment, the system generates features of the environment by inputting a representation of the environment into one or more neural networks, such as an encoder. In at least one embodiment, the representation of the environment is any preferred representation, such as an image, point cloud data, a set of points, and / or a variation thereof. In at least one embodiment, the representation may be captured using various image capture hardware, such as a camera, a depth camera, a sensor device, and / or a variation thereof. In at least one embodiment, the system performs various alignment processes on the generated features.

[0096] In at least one embodiment, one or more neural networks output a set of distances corresponding to a set of locations. In at least one embodiment, the set of distances is the reachable distance corresponding to the set of locations. In at least one embodiment, the reachable distance is also called a distance value, distance, reachable distance value, and / or variations thereof. In at least one embodiment, a first distance in the set of distances corresponds to a first location in the set of locations, indicating the distance of a feasible path from the first location to the final location. In at least one embodiment, the set of locations is relative to the final location. In at least one embodiment, the set of distances for a set of locations is relative to the final location. In at least one embodiment, a feasible path refers to a path that is semantically and / or geometrically feasible. In at least one embodiment, a feasible path refers to any suitable path that an autonomous device can navigate. In at least one embodiment, an implicit environment function outputs a set of distances in a single forward path.

[0097] In at least one embodiment, a system performing at least a portion of process 1100 includes, in 1106, executable code for calculating a plurality of paths based at least partially on a set of distances, the plurality of paths forming a path from a first location to a final location. In at least one embodiment, the system determines a subset of locations accessible from the first location by an autonomous device. In at least one embodiment, a location is accessible to the autonomous device if the autonomous device can navigate to the location from the autonomous device's current location. In at least one embodiment, an accessible location is a location to which the autonomous device can navigate in a single step. In at least one embodiment, the system calculates a size for the step of the autonomous device, corresponding to a distance value of how far the autonomous device travels in each step. In at least one embodiment, the system defines the step size as any preferred value.

[0098] In at least one embodiment, the system determines a subset of distances from a set of distances corresponding to a subset of locations. In at least one embodiment, the system determines a second location from a subset of locations corresponding to the minimum distance from the subset of distances. In at least one embodiment, the second location corresponds to the minimum distance, maximum distance, or any preferred distance from the subset of distances. In at least one embodiment, the second location corresponds to a location in a specific direction from the first location of the autonomous device. In at least one embodiment, the system calculates a first path including the path from the first location to the second location. In at least one embodiment, the system causes the autonomous device to navigate to the second location using the first path.

[0099] In at least one embodiment, the system sequentially determines a path for the autonomous device until the autonomous device can navigate to a final location. In at least one embodiment, for example, the system determines a second subset of locations accessible from a second location of the autonomous device, obtains a second subset of distances corresponding to the second subset of locations, determines a third location from the second subset of locations (e.g., the minimum distance or any preferred distance), calculates a second path from the second location to the third location, causes the autonomous device to navigate to the third location using the second path, and so on until the final location is accessible from the current location of the autonomous device, the system causing the autonomous device to navigate to the final location.

[0100] In at least one embodiment, the system trains one or more neural networks to compute a plurality of paths that an autonomous device is to traverse. In at least one embodiment, the system acquires an environment and a location, causes one or more algorithms to determine one or more reach values ​​for one or more locations in the environment to the location, and trains one or more neural networks using at least the one or more reach values. In at least one embodiment, the system trains one or more neural networks by having one or more neural networks process one or more locations to compute one or more predicted reach values, and by updating the one or more neural networks based on the difference between the one or more predicted reach values ​​and the one or more reach values ​​output by one or more algorithms. In at least one embodiment, the system updates or trains one or more neural networks to minimize the difference between predicted reach values ​​and reach values ​​output by one or more algorithms (for example, by updating them so that the difference is less than a predefined threshold). In at least one embodiment, one or more algorithms include various path planning algorithms, various FMM algorithms, or any preferred algorithm. In at least one embodiment, the system trains one or more neural networks so that the trained one or more neural networks can predict reach values ​​that are similar to or the same as the reach values ​​output by one or more algorithms. In at least one embodiment, the system utilizes predicted reach values ​​to determine a set of paths that an autonomous device is to traverse.

[0101] In at least one embodiment, one or more processes of process 1100, and those described with respect to process 1100, may be carried out in any preferred order, including sequential, parallel, and / or variations thereof. In at least one embodiment, process 1100 may include various processes described elsewhere in this disclosure.

[0102] Reasoning and training logic Figure 12A shows the inference and / or training logic 1215 used to perform inference and / or training operations related to one or more embodiments. Further details regarding the inference and / or training logic 1215 are provided below in conjunction with Figures 12A and / or 12B.

[0103] In at least one embodiment, the inference and / or training logic 1215 may include, but not limited to, code and / or data storage 1201 for storing forward and / or output weights and / or input / output data, and / or other parameters, for constituting neurons or layers of a neural network used for training and / or inference in one or more embodiments. In at least one embodiment, the training logic 1215 may include, or be coupled to, code and / or data storage 1201 for storing graph code or other software for controlling timing and / or sequence, and weight and / or other parameter information should be loaded into the code and / or data storage 1201 to constitute logic including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, the code, such as graph code, loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which such code corresponds. In at least one embodiment, the code and / or data storage 1201 stores the weight parameters and / or input / output data of each layer of the neural network being trained or used in conjunction with one or more embodiments during the forward propagation of input / output data and / or weight parameters during training and / or inference using the embodiments of one or more embodiments. In at least one embodiment, any portion of the code and / or data storage 1201 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.

[0104] In at least one embodiment, any portion of the code and / or data storage 1201 may be inside or outside one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or code and / or data storage 1201 may be cache memory, dynamic randomly addressable memory ("DRAM"), static randomly addressable memory ("SRAM"), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the selection of whether the code and / or code and / or data storage 1201 is inside or outside the processor, or whether it includes DRAM, SRAM, flash, or any other storage type may depend on available storage, on-chip vs. off-chip, latency requirements of the training and / or inference functions being performed, batch size of data used in neural network inference and / or training, or any combination of these factors.

[0105] In at least one embodiment, the inference and / or training logic 1215 may include code and / or data storage 1205 for storing backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network used to train and / or infer in one or more embodiments, but not limited to. In at least one embodiment, the code and / or data storage 1205 stores weight parameters and / or input / output data for each layer of the neural network used to train or in conjunction with one or more embodiments during backward propagation of input / output data and / or weight parameters during training and / or inference using one or more embodiments. In at least one embodiment, the training logic 1215 may include, or be coupled to, code and / or data storage 1205 for storing graph code or other software for controlling timing and / or sequence, in which weight and / or other parameter information should be loaded to constitute logic including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)).

[0106] In at least one embodiment, code such as graph code triggers the loading of weights or other parameter information into the processor ALU based on the architecture of the corresponding neural network. In at least one embodiment, any portion of the code and / or data storage 1205 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of the code and / or data storage 1205 may be inside or outside one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 1205 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the selection of whether the code and / or data storage 1205 is, for example, internal or external to the processor, or whether it includes DRAM, SRAM, flash memory, or any other type of storage, may depend on the available storage, on-chip vs. off-chip, latency requirements of the training and / or inference functions being performed, the batch size of the data used in the neural network inference and / or training, or any combination of these factors.

[0107] In at least one embodiment, the code and / or data storage 1201 and the code and / or data storage 1205 may be separate storage structures. In at least one embodiment, the code and / or data storage 1201 and the code and / or data storage 1205 may be a combined storage structure. In at least one embodiment, the code and / or data storage 1201 and the code and / or data storage 1205 may be partially combined and partially separate. In at least one embodiment, any portion of the code and / or data storage 1201 and the code and / or data storage 1205 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.

[0108] In at least one embodiment, the inference and / or training logic 1215 may include, but not limited to, one or more arithmetic logic units ("ALUs") 1210, including integer and / or floating-point units, for performing logical and / or mathematical operations that are at least partially based on or shown by training and / or inference code (e.g., graph code), the result of which activations (e.g., output values ​​from layers or neurons in a neural network) stored in activation storage 1220, and these activations are functions of input / output and / or weight parameter data stored in code and / or data storage 1201 and / or code and / or data storage 1205. In at least one embodiment, the activation stored in the activation storage 1220 is generated according to linear algebra and / or matrix-based mathematics performed by (one or more) ALUs 1210 in response to the execution of an instruction or other code, and the weight values ​​stored in the code and / or data storage 1205 and / or data storage 1201 are used as operands along with other values ​​such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in the code and / or data storage 1205 or the code and / or data storage 1201, or in other on-chip or off-chip storage.

[0109] In at least one embodiment, the (one or more) ALU 1210 are contained within one or more processors or other hardware logic devices or circuits, but in another embodiment, the (one or more) ALU 1210 may be outside of the processor or other hardware logic devices or circuits (e.g., coprocessors) that use them. In at least one embodiment, the ALU 1210 may be contained within an execution unit of a processor, or otherwise contained within a bank of ALUs accessible by execution units of a processor, either within the same processor or distributed across different types of processors (e.g., a central processing unit, a graphics processing unit, a fixed-function unit, etc.). In at least one embodiment, the code and / or data storage 1201, the code and / or data storage 1205, and the activation storage 1220 may share a processor or other hardware logic devices or circuits, but in another embodiment, they may be in different processors or other hardware logic devices or circuits, or in some combination of the same processor or other hardware logic devices or circuits and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of the activated storage 1220 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Furthermore, inference and / or training code may be stored together with other code accessible to the processor or other hardware logic or circuitry, and may be fetched and / or processed using the processor's fetch, decode, schedule, execute, retirement, and / or other logic circuits.

[0110] In at least one embodiment, the activated storage 1220 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the activated storage 1220 may be located entirely or partially within or outside one or more processors or other logic circuits. In at least one embodiment, the selection of whether the activated storage 1220 is located, for example, inside or outside the processor, or whether it includes DRAM, SRAM, flash memory, or any other storage type may depend on the available storage, on-chip vs. off-chip, latency requirements of the training and / or inference functions being performed, batch size of data used in neural network inference and / or training, or any combination of these factors.

[0111] In at least one embodiment, the inference and / or training logic 1215 shown in Figure 12A may be used in conjunction with an application-specific integrated circuit ("ASIC"), such as a TensorFlow® processing unit from Google, an inference processing unit (IPU) from Graphcore®, or a Nervana® (e.g., "Lake Crest") processor from Intel Corp. In at least one embodiment, the inference and / or training logic 1215 shown in Figure 12A may be used in conjunction with other hardware, such as a central processing unit ("CPU") hardware, a graphics processing unit ("GPU") hardware, or a field-programmable gate array ("FPGA").

[0112] Figure 12B shows inference and / or training logic 1215 according to at least one embodiment. In at least one embodiment, the inference and / or training logic 1215 may include, but is not limited to, hardware logic in which computational resources are dedicated or, otherwise, used only in conjunction with weight values ​​or other information corresponding to one or more layers of neurons in a neural network. In at least one embodiment, the inference and / or training logic 1215 shown in Figure 12B may be used in conjunction with application-specific integrated circuits (ASICs), such as a TensorFlow® processing unit from Google, an inference processing unit (IPU) from Graphcore®, or a Nervana® (e.g., "Lake Crest") processor from Intel Corp. In at least one embodiment, the inference and / or training logic 1215 shown in Figure 12B may be used in conjunction with other hardware such as a central processing unit (CPU) hardware, a graphics processing unit (GPU) hardware, or a field-programmable gate array (FPGA). In at least one embodiment, the inference and / or training logic 1215 may include, but are not limited to, code and / or data storage 1201 and code and / or data storage 1205, which may be used to store code (e.g., graph code), weight values, and / or other information including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In at least one embodiment shown in Figure 12B, each of the code and / or data storage 1201 and code and / or data storage 1205 is associated with a dedicated computing resource, such as computing hardware 1202 and computing hardware 1206, respectively.In at least one embodiment, each of the computation hardware 1202 and computation hardware 1206 comprises one or more ALUs that perform mathematical functions, such as linear algebraic functions, only on information stored in code and / or data storage 1201 and code and / or data storage 1205, respectively, and the results are stored in activation storage 1220.

[0113] In at least one embodiment, each of the code and / or data storages 1201 and 1205 and the corresponding compute hardware 1202 and 1206 correspond to different layers of a neural network, so that the activation resulting from one storage / compute pair 1201 / 1202 of the code and / or data storage 1201 and compute hardware 1202 is provided as input to the next storage / compute pair 1205 / 1206 of the code and / or data storage 1205 and compute hardware 1206 to mirror the conceptual organization of the neural network. In at least one embodiment, the storage / compute pairs 1201 / 1202 and 1205 / 1206 may correspond to two or more neural network layers. In at least one embodiment, additional storage / computation pairs (not shown) may be included in the inference and / or training logic 1215 after or in parallel with the storage / computation pairs 1201 / 1202 and 1205 / 1206.

[0114] In at least one embodiment, one or more systems shown in Figures 12A-12B are used to implement one or more implicit environment functions. In at least one embodiment, one or more systems shown in Figures 12A-12B are used to use one or more neural networks, such as one or more implicit environment functions, to compute multiple paths that an entity, such as an autonomous device, is to traverse. In at least one embodiment, one or more systems shown in Figures 12A-12B are used to implement one or more systems and / or processes, such as those described with respect to Figures 1-11.

[0115] Neural network training and implementation Figure 13 shows the training and deployment of a deep neural network in at least one embodiment. In at least one embodiment, an untrained neural network 1306 is trained using a training dataset 1302. In at least one embodiment, the training framework 1304 is the PyTorch framework, while in other embodiments, the training framework 1304 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework 1304 trains the untrained neural network 1306 and enables it to be trained using the processing resources described herein to produce a trained neural network 1308. In at least one embodiment, the weights may be randomly selected or selected by pre-training using a deep belief network. In at least one embodiment, training may be performed in a supervised, partially supervised, or unsupervised manner.

[0116] In at least one embodiment, an untrained neural network 1306 is trained using supervised learning, and the training dataset 1302 contains inputs paired with desired outputs for inputs, or the training dataset 1302 contains inputs with known outputs, and the outputs of the neural network 1306 are manually scored. In at least one embodiment, the untrained neural network 1306 is trained in a supervised manner, processing inputs from the training dataset 1302 and comparing the resulting outputs with a set of expected or desired outputs. In at least one embodiment, errors are then backpropagated through the untrained neural network 1306. In at least one embodiment, a training framework 1304 adjusts the weights controlling the untrained neural network 1306. In at least one embodiment, the training framework 1304 includes tools for monitoring how well the untrained neural network 1306 is converging toward a model such as a trained neural network 1308 that is suitable for generating correct answers in results such as Result 1314 based on input data such as a new dataset 1312. In at least one embodiment, the training framework 1304 iteratively trains the untrained neural network 1306, adjusting the weights to improve the output of the untrained neural network 1306 using a loss function and a modulating algorithm such as stochastic gradient descent. In at least one embodiment, the training framework 1304 trains the untrained neural network 1306 until it achieves a desired accuracy. In at least one embodiment, the trained neural network 1308 can then be introduced to implement any number of machine learning operations.

[0117] In at least one embodiment, an untrained neural network 1306 is trained using unsupervised learning, and the untrained neural network 1306 attempts to train itself using unlabeled data. In at least one embodiment, the unsupervised learning training dataset 1302 includes input data that has no associated output data or “ground truth” data. In at least one embodiment, the untrained neural network 1306 can learn grouping within the training dataset 1302 and can determine how individual inputs relate to the untrained dataset 1302. In at least one embodiment, unsupervised training may be used to generate a self-organizing map in a trained neural network 1308 that is capable of performing useful operations in reducing the dimensionality of a new dataset 1312. In at least one embodiment, unsupervised training may also be used to perform anomaly detection, which enables the identification of data points in the new dataset 1312 that deviate from the normal pattern of the new dataset 1312.

[0118] In at least one embodiment, semi-supervised learning may be used, which is a technique that includes a mixture of labeled and unlabeled data in the training dataset 1302. In at least one embodiment, the training framework 1304 may be used to implement incremental learning, such as through a transfer learning technique. In at least one embodiment, incremental learning allows the trained neural network 1308 to adapt to a new dataset 1312 without forgetting the knowledge that was taught within the trained neural network 1308 during the initial training.

[0119] In at least one embodiment, the training framework 1304 is a framework processed together with a software development toolkit such as the OpenVINO (Open Visual Inference and Neural Network Optimization) toolkit. In at least one embodiment, the OpenVINO toolkit is a toolkit such as the one developed by Intel Corporation in Santa Clara, California.

[0120] In at least one embodiment, OpenVINO is a toolkit for facilitating the development of applications, specifically neural network applications, for a variety of tasks and operations, such as human vision emulation, speech recognition, natural language processing, recommendation systems, and / or variations thereof. In at least one embodiment, OpenVINO supports neural networks such as convolutional neural networks (CNNs), recurrent and / or attention-based neural networks, and / or various other neural network models. In at least one embodiment, OpenVINO supports various software libraries such as OpenCV, OpenCL, and / or variations thereof.

[0121] In at least one embodiment, OpenVINO supports neural network models for a variety of tasks and operations, including classification, segmentation, object detection, face recognition, speech recognition, pose estimation (e.g., human and / or object), monocular depth estimation, image inpainting, style transfer, action recognition, colorization, and / or modified forms thereof.

[0122] In at least one embodiment, OpenVINO includes one or more software tools and / or modules for model optimization, also known as model optimizers. In at least one embodiment, the model optimizer is a command-line tool that facilitates the transition between training and deployment of a neural network model. In at least one embodiment, the model optimizer optimizes a neural network model for execution on various devices and / or processing units, such as GPUs, CPUs, PPUs, GPGPUs, and / or variations thereof. In at least one embodiment, the model optimizer generates an internal representation of the model and optimizes the model to generate an intermediate representation. In at least one embodiment, the model optimizer reduces the number of layers in the model. In at least one embodiment, the model optimizer removes layers of the model that are used for training. In at least one embodiment, the model optimizer performs various neural network operations, such as modifying the input to the model (e.g., resizing the input to the model), modifying the size of the input to the model (e.g., modifying the batch size of the model), modifying the model structure (e.g., modifying the layers of the model), normalization, standardization, quantization (e.g., converting the model weights from a first representation such as floating-point numbers to a second representation such as integers), and / or variations thereof.

[0123] In at least one embodiment, OpenVINO includes one or more software libraries for inference, also called inference engines. In at least one embodiment, the inference engine is a C++ library or a library in any preferred programming language. In at least one embodiment, the inference engine is used to infer input data. In at least one embodiment, the inference engine implements various classes to infer input data and produce one or more results. In at least one embodiment, the inference engine implements one or more API functions to process intermediate representations, set input and / or output formats, and / or run the model on one or more devices.

[0124] In at least one embodiment, OpenVINO provides various capabilities for heterogeneous execution of one or more neural network models. In at least one embodiment, heterogeneous execution, or heterogeneous computing, refers to one or more computing processes and / or systems that utilize one or more types of processors and / or cores. In at least one embodiment, OpenVINO provides various software capabilities for running a program on one or more devices. In at least one embodiment, OpenVINO provides various software capabilities for running a program and / or parts of a program on different devices. In at least one embodiment, OpenVINO provides various software capabilities for running, for example, a first part of the code on a CPU and a second part of the code on a GPU and / or FPGA. In at least one embodiment, OpenVINO provides various software capabilities for running one or more layers of a neural network on one or more devices (for example, running a first set of layers on a first device such as a GPU and a second set of layers on a second device such as a CPU).

[0125] In at least one embodiment, OpenVINO includes a variety of functionalities similar to those associated with CUDA programming models, such as various neural network model behaviors related to frameworks like TensorFlow, PyTorch, and / or their variations. In at least one embodiment, one or more CUDA programming model behaviors are performed using OpenVINO. In at least one embodiment, various systems, methods, and / or techniques described herein are implemented using OpenVINO.

[0126] In at least one embodiment, one or more systems shown in Figure 13 are used to implement one or more implicit environment functions. In at least one embodiment, one or more systems shown in Figure 13 are used to use one or more neural networks, such as one or more implicit environment functions, to compute multiple paths that an entity, such as an autonomous device, is to traverse. In at least one embodiment, one or more systems shown in Figure 13 are used to implement one or more systems and / or processes, such as those described with respect to Figures 1 to 11.

[0127] Data center Figure 14 shows an exemplary data center 1400 in which at least one embodiment may be used. In at least one embodiment, the data center 1400 includes a data center infrastructure layer 1410, a framework layer 1420, a software layer 1430, and an application layer 1440.

[0128] In at least one embodiment, as shown in Figure 14, the data center infrastructure layer 1410 may include a resource orchestrator 1412, grouped computing resources 1414, and node computing resources ("node CRs") 1416(1) to 1416(N), where "N" represents a positive integer (which may be a different integer "N" than that used in other figures). In at least one embodiment, nodes CR1416(1) to 1416(N) may include, but not limited to, any number of central processing units ("CPU") or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processors, etc.), memory and storage devices 1418(1) to 1418(N) (e.g., dynamic read-only memory, solid-state storage, or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules, and cooling modules, etc. In at least one embodiment, one or more nodes CR from among nodes CR1416(1) to 1416(N) may be servers having one or more of the computing resources described above.

[0129] In at least one embodiment, the grouped computing resources 1414 may include separate groupings of node CRs housed in one or more racks (not shown), or many racks housed in a data center at various geographical locations (also not shown). In at least one embodiment, the separate groupings of node CRs within the grouped computing resources 1414 may include grouped compute resources, network resources, memory resources, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped in one or more racks to provide compute resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.

[0130] In at least one embodiment, the resource orchestrator 1412 may configure or otherwise control one or more nodes CR1416(1) to 1416(N) and / or a grouped computing resource 1414. In at least one embodiment, the resource orchestrator 1412 may include a software design infrastructure ("SDI") management entity for the data center 1400. In at least one embodiment, the resource orchestrator 1212 may include hardware, software, or any combination thereof.

[0131] In at least one embodiment, as shown in Figure 14, the framework layer 1420 includes a job scheduler 1422, a configuration manager 1424, a resource manager 1426, and a distributed file system 1428. In at least one embodiment, the framework layer 1420 may include a framework for supporting the software 1432 of the software layer 1430 and / or one or more applications 1442 of the application layer 1440. In at least one embodiment, the software 1432 or (one or more) applications 1442 may each include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 1420 may be a type of free and open-source software web application framework, such as Apache Spark® ("Spark"), which can leverage the distributed file system 1428 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 1422 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 1400. In at least one embodiment, the configuration manager 1424 may be able to configure different layers, such as the software layer 1430 and the framework layer 1420, which includes Spark and a distributed file system 1428 to support large-scale data processing. In at least one embodiment, the resource manager 1426 may be able to manage clustered or grouped computing resources that are mapped or allocated to support the distributed file system 1428 and the job scheduler 1422. In at least one embodiment, the clustered or grouped computing resources may include a grouped computing resource 1414 in the data center infrastructure layer 1410.In at least one embodiment, the resource manager 1426 may work in conjunction with the resource orchestrator 1412 to manage these mapped or allocated computing resources.

[0132] In at least one embodiment, the software 1432 contained within the software layer 1430 may include software used by nodes CR1416(1) to 1416(N), grouped computing resources 1414, and / or at least a portion of the distributed file system 1428 of the framework layer 1420. In at least one embodiment, one or more types of software may include, but are not limited to, internet web page search software, email virus scanning software, database software, and streaming video content software.

[0133] In at least one embodiment, one or more applications 1442 contained within the application layer 1440 may include one or more types of applications used by nodes CR1416(1) to 1416(N), grouped computing resources 1414, and / or at least a portion of the distributed file system 1428 of the framework layer 1420. In at least one embodiment, one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive compute applications, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0134] In at least one embodiment, any of the configuration manager 1424, resource manager 1426, and resource orchestrator 1412 may implement any number and type of self-correcting actions based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-correcting actions may relieve the data center operator of data center 1400 of the task of determining potentially faulty configurations and potentially avoiding underutilized and / or underperforming portions of the data center.

[0135] In at least one embodiment, the data center 1400 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by computing weight parameters according to a neural network architecture using the software and computing resources described above with respect to the data center 1400. In at least one embodiment, a trained machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to the data center 1400 by using weight parameters computed through one or more training techniques described herein.

[0136] In at least one embodiment, the data center may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, or other hardware to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above may be configured as services that enable users to train information or perform inference on information, such as image recognition, speech recognition, or other artificial intelligence services.

[0137] The inference and / or training logic 1215 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 1215 are provided herein in conjunction with Figures 12A and / or 12B. In at least one embodiment, the inference and / or training logic 1215 may be used in the system of Figure 14 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0138] In at least one embodiment, one or more systems shown in Figure 14 are used to implement one or more implicit environment functions. In at least one embodiment, one or more systems shown in Figure 14 are used to use one or more neural networks, such as one or more implicit environment functions, to compute multiple paths that an entity, such as an autonomous device, is to traverse. In at least one embodiment, one or more systems shown in Figure 14 are used to implement one or more systems and / or processes, such as those described with respect to Figures 1 to 11.

[0139] Autonomous vehicles Figure 15A shows an example of an autonomous vehicle 1500 according to at least one embodiment. In at least one embodiment, the autonomous vehicle 1500 (alternatively referred to herein as "vehicle 1500") may be a passenger vehicle, including, but not limited to, a car, truck, bus, and / or another type of vehicle that accommodates one or more passengers. In at least one embodiment, vehicle 1500 may be a semi-tractor-trailer truck used for transporting cargo. In at least one embodiment, vehicle 1500 may be an aircraft, a robotic vehicle, or another type of vehicle.

[0140] Autonomous vehicles can be described in terms of automation levels as defined by the National Highway Traffic Safety Administration ("NHTSA"), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers ("SAE")'s "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (e.g., Standard No. J3016-201806, issued June 15, 2018; Standard No. J3016-201609, issued September 30, 2016; and previous and newer versions of this standard). In at least one embodiment, vehicle 1500 may be capable of functionality at one or more of the autonomous driving levels from Level 1 to Level 5. For example, in at least one embodiment, vehicle 1500 may be capable of conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on the embodiment.

[0141] In at least one embodiment, the vehicle 1500 may include components such as a chassis, a vehicle body, wheels (e.g., two, four, six, eight, eighteen, etc.), tires, axles, and other components of the vehicle, but are not limited. In at least one embodiment, the vehicle 1500 may include a propulsion system 1550, such as an internal combustion engine, a hybrid power plant, a fully electric engine, and / or another propulsion system type, but are not limited. In at least one embodiment, the propulsion system 1550 may be connected to the drivetrain of the vehicle 1500, and the drivetrain may include a transmission to enable the propulsion of the vehicle 1500, but are not limited. In at least one embodiment, the propulsion system 1550 may be controlled in response to receiving a signal from (one or more) throttles / accelerators 1552.

[0142] In at least one embodiment, a steering system 1554, which may include a steering wheel, is used to steer the vehicle 1500 (for example, along a desired path or route) when the propulsion system 1550 is operating (for example, when the vehicle 1500 is moving). In at least one embodiment, the steering system 1554 may receive signals from one or more steering actuators 1556. In at least one embodiment, the steering wheel may be optional for fully automated (level 5) functionality. In at least one embodiment, a brake sensor system 1546 may be used to actuate the vehicle brakes in response to receiving signals from one or more brake actuators 1548 and / or brake sensors.

[0143] In at least one embodiment, but not limited to, one or more controllers 1536 which may include one or more system-on-chip ("SoC") (not shown in Figure 15A) and / or one or more graphics processing units ("GPU") provide signals (for example, representing commands) to one or more components and / or systems of the vehicle 1500. For example, in at least one embodiment, one or more controllers 1536 may send signals to operate the vehicle brakes via one or more brake actuators 1548, signals to operate the steering system 1554 via one or more steering actuators 1556, and signals to operate the propulsion system 1550 via one or more throttle / accelerators 1552. In at least one embodiment, the controller 1536 (one or more) may include one or more onboard (e.g., integrated) computing devices that process sensor signals and output operational commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving the vehicle 1500. In at least one embodiment, the controller 1536 (one or more) may include a first controller for autonomous driving functions, a second controller for functional safety functions, a third controller for artificial intelligence functions (e.g., computer vision), a fourth controller for infotainment functions, a fifth controller for emergency redundancy, and / or other controllers. In at least one embodiment, a single controller may address two or more of the above functions, two or more controllers may address a single function, and / or any combination thereof.

[0144] In at least one embodiment, one or more controllers 1536 provide signals for controlling one or more components and / or systems of the vehicle 1500 in response to sensor data (e.g., sensor inputs) received from one or more sensors. In at least one embodiment, the sensor data may include, but are not limited to, one or more global navigation satellite system ("GNSS") sensors 1558 (e.g., one or more global positioning system sensors), one or more radar sensors 1560, one or more ultrasonic sensors 1562, one or more lithium-ion sensors 1564, or one or more inertial measurement units ("IMUs"). The unit includes: 1566 sensors (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetic compasses, one or more magnetometers, etc.), 1596 microphones, 1568 stereo cameras, 1570 wide-angle cameras (e.g., fisheye cameras), 1572 infrared cameras, and 1574 ambient cameras (e.g., 360-degree cameras). ), long-range cameras (not shown in Figure 15A), (one or more) medium-range cameras (not shown in Figure 15A), (one or more) speed sensors 1544 (for measuring the speed of vehicle 1500), (one or more) vibration sensors 1542, (one or more) steering sensors 1540, (one or more) brake sensors (for example, as part of a brake sensor system 1546), and / or other sensor types may be received.

[0145] In at least one embodiment, one or more of the controllers 1536 may receive inputs (represented, for example, by input data) from the instrument cluster 1532 of the vehicle 1500 and provide outputs (represented, for example, by output data, display data, etc.) via the human-machine interface ("HMI") display 1534, an audible annunciator, a loudspeaker, and / or other components of the vehicle 1500. In at least one embodiment, the outputs may include information such as vehicle speed, vehicle speed, time, map data (e.g., a high-definition map (not shown in Figure 15A)), location data (e.g., the location of the vehicle 1500, such as on a map), direction, location of other vehicles (e.g., an occupied grid), and information about objects and the status of objects sensed by the controllers 1536. For example, in at least one embodiment, the HMI display 1534 may display information regarding the presence of one or more objects (e.g., road signs, warning signs, changes in traffic signals, etc.) and / or information regarding driving operations that the vehicle has performed, is performing, or will perform (e.g., changing lanes, exiting at Exit 34B 3.22 km (2 miles) ahead, etc.).

[0146] In at least one embodiment, the vehicle 1500 further includes a network interface 1524, which may use (one or more) wireless antennas 1526 and / or (one or more) modems to communicate over one or more networks. For example, in at least one embodiment, the network interface 1524 may enable communication over Long-Term Evolution ("LTE"), Wideband Code Division Multiple Access ("WCDMA®"), Universal Mobile Telecommunications System ("UMTS"), Global System for Mobile Communication ("GSM"), IMT-CDMA Multi-Carrier ("CDMA2000") networks, and the like. Furthermore, in at least one embodiment, one or more wireless antennas 1526 may enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using one or more local area networks such as Bluetooth, Bluetooth Low Energy ("LE"), Z-wave, ZigBee, and / or one or more low-power wide-area networks ("LPWAN") such as LoRaWAN, SigFox, etc.

[0147] The inference and / or training logic 1215 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 1215 are provided herein in conjunction with Figures 12A and / or 12B. In at least one embodiment, the inference and / or training logic 1215 may be used in the system of Figure 15A for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0148] Figure 15B shows an example of camera locations and fields of view for the autonomous vehicle 1500 of Figure 15A, according to at least one embodiment. In at least one embodiment, the cameras and their respective fields of view are illustrative and not limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included, and / or the cameras may be located in different locations on the vehicle 1500.

[0149] In at least one embodiment, the camera type for a camera may include, but is not limited to, a digital camera that can be adapted for use with components and / or systems of the vehicle 1500. In at least one embodiment, one or more cameras may operate at Automotive Safety Integrity Level ("ASIL") B and / or another ASIL. In at least one embodiment, the camera type may be capable of any image capture rate, depending on the embodiment, such as 60 frames per second (fps), 1220 fps, 240 fps, etc. In at least one embodiment, the camera may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In at least one embodiment, the color filter array may include a red, clear, clear, clear ("RCCC") color filter array, a red, clear, clear, blue ("RCCB") color filter array, a red, blue, green, clear ("RBGC") color filter array, a Foveon X3 color filter array, a Bayer sensor ("RGGB") color filter array, a monochrome sensor color filter array, and / or other types of color filter arrays. In at least one embodiment, a clear pixel camera may be used to increase light sensitivity, such as a camera with RCCC, RCCB, and / or RBGC color filter arrays.

[0150] In at least one embodiment, one or more of the cameras may be used to implement advanced driver assistance system ("ADAS") functions (for example, as part of a redundant or fail-safe design). For example, in at least one embodiment, a multi-function mono-camera may be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. In at least one embodiment, one or more of the cameras (for example, all of the cameras) may simultaneously record and provide image data (for example, video).

[0151] In at least one embodiment, one or more cameras may be mounted in a mounting assembly, such as a custom-designed (3D printed) assembly, to eliminate stray light and reflections from within the vehicle (e.g., reflections reflected from the dashboard to the windshield) that could interfere with the camera image data capture ability. Referring to a door mirror mounting assembly, in at least one embodiment, the door mirror assembly may be custom 3D printed so that the camera mounting plate matches the shape of the door mirror. In at least one embodiment, one or more cameras may be integrated into the door mirror. In at least one embodiment, for a side-view camera, one or more cameras may be integrated into the four pillars at each corner of the cabin.

[0152] In at least one embodiment, a camera having a field of view including a portion of the environment in front of the vehicle 1500 (e.g., a front camera) may be used for a perimeter view to help identify the path and obstacles ahead and to assist in providing information essential for generating an occupied grid and / or determining a preferred vehicle path, with the help of one or more of the controllers 1536 and / or control SoCs. In at least one embodiment, the front camera may be used to implement many ADAS functions similar to LIDAR, including, but not limited to, emergency braking, pedestrian detection, and collision avoidance. In at least one embodiment, the front camera may also be used for ADAS functions and systems including, but not limited to, lane departure warning ("LDW"), autonomous cruise control ("ACC"), and / or traffic sign recognition.

[0153] In at least one embodiment, various cameras, including a monocular camera platform including, for example, a CMOS ("complementary metal oxide semiconductor") color imager, may be used in a front configuration. In at least one embodiment, a wide-angle camera 1570 may be used to perceive objects entering the view from the periphery (e.g., pedestrians, crossing traffic, or bicycles). Although only one wide-angle camera 1570 is shown in Figure 15B, in other embodiments, there may be any number of wide-angle cameras (including zero) on the vehicle 1500. In at least one embodiment, any number of long-range cameras 1598 (e.g., long-view stereo camera pairs) may be used for depth-based object detection, particularly for objects for which the neural network has not yet been trained. In at least one embodiment, the long-range cameras 1598 may also be used for object detection and classification, as well as basic object tracking.

[0154] In at least one embodiment, any number of stereo cameras 1568 may also be included in the front configuration. In at least one embodiment, one or more of the (one or more) stereo cameras 1568 may include an integrated control unit with a scalable processing unit, which may provide a programmable logic ("FPGA") and a multi-core microprocessor having an integrated Controller Area Network ("CAN") or Ethernet interface on a single chip. In at least one embodiment, such a unit may be used to generate a 3D map of the environment of the vehicle 1500, including distance estimation for all points in the image. In at least one embodiment, one or more of the (one or more) stereo cameras 1568 may include, but are not limited to, one or more compact stereo vision sensors, which may include, but are not limited to, two camera lenses (one on the left and one on the right) and an image processing chip capable of measuring the distance from a vehicle 1500 to a target object and using the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. In at least one embodiment, other types of (one or more) stereo cameras 1568 may be used in addition to or as an alternative to those described herein.

[0155] In at least one embodiment, a camera having a field of view that includes a portion of the environment to the sides of the vehicle 1500 (e.g., a side-view camera) may be used for the surrounding view and provide information used to create and update the occupy grid and generate side collision warnings. For example, in at least one embodiment, one or more surrounding cameras 1574 (e.g., four surrounding cameras shown in Figure 15B) may be positioned on the vehicle 1500. In at least one embodiment, one or more surrounding cameras 1574 may include, but are not limited to, any number and combination of wide-angle cameras, one or more fisheye cameras, one or more 360-degree cameras, and / or similar cameras. For example, in at least one embodiment, four fisheye cameras may be positioned in front of, behind, and to the sides of the vehicle 1500. In at least one embodiment, the vehicle 1500 may use three surrounding cameras 1574 (e.g., left, right, and rear) and may utilize one or more other cameras (e.g., a front camera) as a fourth surrounding view camera.

[0156] In at least one embodiment, a camera having a field of view that includes a portion of the environment behind the vehicle 1500 (e.g., a rear-view camera) may be used for parking assistance, surrounding view, rear collision warning, and creation and updating of the occupancy grid. In at least one embodiment, a wide variety of cameras may be used, including, but not limited to, cameras that are also suitable as (one or more) front cameras as described herein (e.g., a long-range camera 1598 and / or (one or more) medium-range cameras 1576, (one or more) stereo cameras 1568, (one or more) infrared cameras 1572, etc.).

[0157] The inference and / or training logic 1215 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 1215 are provided herein in conjunction with Figures 12A and / or 12B. In at least one embodiment, the inference and / or training logic 1215 may be used in the system of Figure 15B for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0158] Figure 15C is a block diagram illustrating an exemplary system architecture for the autonomous vehicle 1500 of Figure 15A, according to at least one embodiment. In at least one embodiment, each of the components, features, and systems of the vehicle 1500 in Figure 15C is shown as being connected via a bus 1502. In at least one embodiment, the bus 1502 may include, but is not limited to, a CAN data interface (alternatively referred to herein as the “CAN bus”). In at least one embodiment, the CAN may be an internal network of the vehicle 1500 used to assist in the control of various features and functionalities of the vehicle 1500, such as brake activation, acceleration, brake control, steering, and windshield wipers. In at least one embodiment, the bus 1502 may be configured to have tens or even hundreds of nodes, each having its own unique identifier (e.g., a CAN ID). In at least one embodiment, bus 1502 may be read to find the steering angle, ground speed, engine revolutions per minute ("RPM"), button position, and / or other vehicle status indicators. In at least one embodiment, bus 1502 may be an ASIL B compliant CAN bus.

[0159] In at least one embodiment, the FlexRay and / or Ethernet protocols may be used in addition to or as an alternative to CAN. In at least one embodiment, there may be any number of buses forming bus 1502, which may include, but are not limited to, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet buses, and / or zero or more other types of buses using different protocols. In at least one embodiment, two or more buses may be used to perform different functions and / or for redundancy. For example, a first bus may be used for collision avoidance functionality and a second bus may be used for operation control. In at least one embodiment, each bus of bus 1502 may communicate with any of the components of vehicle 1500, and two or more buses of bus 1502 may communicate with corresponding components. In at least one embodiment, each of any number of system-on-chip ("SoC") 1504 (such as SoC 1504(A) and SoC 1504(B)), each of the (one or more) controllers 1536, and / or each computer in the vehicle may have access to the same input data (e.g., inputs from sensors in the vehicle 1500) and may be connected to a common bus such as a CAN bus.

[0160] In at least one embodiment, the vehicle 1500 may include one or more controllers 1536, such as those described herein with respect to Figure 15A. In at least one embodiment, one or more controllers 1536 may be used for a variety of functions. In at least one embodiment, one or more controllers 1536 may be coupled to any of the various other components and systems of the vehicle 1500 and may be used for controlling the vehicle 1500, the artificial intelligence of the vehicle 1500, infotainment for the vehicle 1500, and / or other functions.

[0161] In at least one embodiment, the vehicle 1500 may include any number of SoCs 1504. In at least one embodiment, each of the SoCs 1504 may include, but not limited to, a central processing unit ("CPU") 1506, a graphics processing unit ("GPU") 1508, one or more processors 1510, one or more caches 1512, one or more accelerators 1514, one or more data stores 1516, and / or other components and features not shown. In at least one embodiment, one or more SoCs 1504 may be used to control the vehicle 1500 in various platforms and systems. For example, in at least one embodiment, one or more SoCs 1504 may be combined in a system (for example, a system in a vehicle 1500) that has a high-definition ("HD") map 1522 that can receive map refreshes and / or updates from one or more servers (not shown in Figure 15C) via a network interface 1524.

[0162] In at least one embodiment, one or more CPUs 1506 may include a CPU cluster or CPU complex (alternatively referred to herein as “CCPLEX”). In at least one embodiment, one or more CPUs 1506 may include multiple cores and / or Level 2 ("L2") caches. For example, in at least one embodiment, one or more CPUs 1506 may include eight cores in a coherent multiprocessor configuration. In at least one embodiment, one or more CPUs 1506 may include four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2 megabyte (MB) L2 cache). In at least one embodiment, one or more CPUs 1506 (e.g., CCPLEX) may be configured to support concurrent cluster operation, which allows any combination of clusters of one or more CPUs 1506 to be active at any given time.

[0163] In at least one embodiment, one or more of the (one or more) CPUs 1506 may implement a power management capability, which includes, but is not limited to, one or more of the following features: individual hardware blocks may be automatically clock-gated when idle to conserve dynamic power; each core clock may be gated when such core is not actively executing instructions by executing an interrupt-wait ("WFI") / event-wait ("WFE") instruction; each core may be independently power-gated; each core cluster may be independently clock-gated when all cores are clock-gated or power-gated; and / or each core cluster may be independently power-gated when all cores are power-gated. In at least one embodiment, one or more CPUs 1506 may further implement an extended algorithm for managing power states, where acceptable power states and expected wake-up times are specified, and the hardware / microcode determines which power state is best to enter for the core, cluster, and CCPLEX. In at least one embodiment, the processing core may support a simple power state entry sequence in software where the work is offloaded to the microcode.

[0164] In at least one embodiment, one or more GPUs 1508 may include an integrated GPU (alternatively referred to herein as an “iGPU”). In at least one embodiment, one or more GPUs 1508 may be programmable and efficient for parallel workloads. In at least one embodiment, one or more GPUs 1508 may use an extended tensor instruction set. In at least one embodiment, one or more GPUs 1508 may include one or more streaming microprocessors, each streaming microprocessor may include a Level 1 ("L1") cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In at least one embodiment, one or more GPUs 1508 may include at least eight streaming microprocessors. In at least one embodiment, one or more GPU1508s may use one or more compute application programming interfaces (APIs). In at least one embodiment, one or more GPU1508s may use one or more parallel computing platforms and / or programming models (for example, NVIDIA's CUDA model).

[0165] In at least one embodiment, one or more of the (one or more) GPU1508s may be power-optimized for best performance in automotive and embedded use cases. For example, in at least one embodiment, the (one or more) GPU1508s may be fabricated on Fin field-effect transistor ("FinFET") circuit elements. In at least one embodiment, each streaming microprocessor may incorporate several mixed-precision processing cores divided into multiple blocks. For example, but not limited to, 64 PF32 cores and 32 PF64 cores may be divided into four processing blocks. In at least one embodiment, each processing block may be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, 2 mixed-precision NVIDIA Tensor cores for deep learning matrix arithmetic, a Level 0 ("L0") instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. In at least one embodiment, the streaming microprocessor may include independent parallel integer and floating-point data paths to efficiently execute workloads involving a mixture of computation and addressing calculations. In at least one embodiment, the streaming microprocessor may include independent thread scheduling capabilities to enable finer-grained synchronization and coordination between parallel threads. In at least one embodiment, the streaming microprocessor may include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.

[0166] In at least one embodiment, one or more of the (one or more) GPU1508s may include high-bandwidth memory ("HBM") and / or a 16GB HBM2 memory subsystem to provide a peak memory bandwidth of approximately 900 GB / s in some examples. In at least one embodiment, in addition to or as an alternative to HBM memory, synchronous graphics random-access memory ("SGRAM"), such as graphics double data rate type five synchronous random-access memory ("GDDR5").

[0167] In at least one embodiment, one or more GPUs 1508 may include unified memory technology. In at least one embodiment, address translation service (ATS) support may be used to enable one or more GPUs 1508 to directly access the page tables of one or more CPUs 1506. In at least one embodiment, when a GPU in the GPU 1508 memory management unit (MMU) encounters a miss, an address translation request may be sent to one or more CPUs 1506. In at least one embodiment, in response, two of the CPUs 1506 may look up virtual-physical mappings for addresses in their page tables and send the translation back to one or more GPUs 1508. In at least one embodiment, the unified memory technology enables a single, unified virtual address space for the memory of both the (one or more) CPU 1506 and the (one or more) GPU 1508, thereby simplifying the programming of the (one or more) GPU 1508 and the porting of applications to the (one or more) GPU 1508.

[0168] In at least one embodiment, one or more GPUs 1508 may include any number of access counters that can track how often one or more GPUs 1508 access the memory of other processors. In at least one embodiment, one or more access counters can help ensure that memory pages are moved to the physical memory of the processor that accesses the page most frequently, thereby improving the efficiency of memory ranges shared between processors.

[0169] In at least one embodiment, one or more of the (one or more) SoC 1504 may include any number of caches 1512, including those described herein. For example, in at least one embodiment, the (one or more) caches 1512 may include a Level 3 ("L3") cache that is available to both the (one or more) CPUs 1506 and the (one or more) GPUs 1508 (for example, connected to the (one or more) CPUs 1506 and the (one or more) GPUs 1508). In at least one embodiment, the (one or more) caches 1512 may include a write-back cache that can track the state of the line, for example by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, the L3 cache may include 4 MB or more of memory, depending on the embodiment, but smaller cache sizes may be used.

[0170] In at least one embodiment, one or more of the (one or more) SoC1504 may include one or more accelerators 1514 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, the (one or more) SoC1504 may include a hardware acceleration cluster which may include an optimized hardware accelerator and / or a large on-chip memory. In at least one embodiment, the large on-chip memory (e.g., 4 MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, the hardware acceleration cluster may complement the (one or more) GPU1508 and be used to offload some of the tasks of the (one or more) GPU1508 (e.g., to free up more cycles of the (one or more) GPU1508 to perform other tasks). In at least one embodiment, accelerator 1514 may be used for a target workload that is stable enough to accept acceleration (e.g., perception, convolutional neural networks ("CNN"), recurrent neural networks ("RNN"), etc.). In at least one embodiment, the CNN may include region-based, i.e., regional convolutional neural networks ("RCNN"), and fast RCNNs (such as those used for object detection), or other types of CNNs.

[0171] In at least one embodiment, one or more accelerators 1514 (e.g., a hardware acceleration cluster) may include one or more deep learning accelerators ("DLAs"). In at least one embodiment, one or more DLAs may include, but not limited to, one or more Tensor processing units ("TPUs"), which may be configured to provide additional, 10 trillion operations per second for deep learning applications and inference. In at least one embodiment, a TPU may be configured to perform image processing functions (e.g., for CNNs, RCNNs, etc.) and may be an accelerator optimized for that purpose. In at least one embodiment, one or more DLAs may be further optimized for specific neural network types and sets of floating-point operations, as well as for inference. In at least one embodiment, a design of one or more DLAs may provide more performance per millisecond than a typical general-purpose GPU, and generally far exceed the performance of a CPU. In at least one embodiment, one or more TPUs may implement several functions, including, for example, a single-instance convolution function supporting INT8, INT16, and FP16 data types for both features and weights, as well as a post-processing function. In at least one embodiment, one or more DLAs may rapidly and efficiently run neural networks, particularly CNNs, on processed or unprocessed data for any of a variety of functions, including, but not limited to, CNNs for object recognition and detection using data from camera sensors, CNNs for distance estimation using data from camera sensors, CNNs for emergency vehicle detection and identification and detection using data from microphones, CNNs for face recognition and vehicle owner identification using data from camera sensors, and / or CNNs for security and / or safety-related events.

[0172] In at least one embodiment, one or more DLAs may perform any function of one or more GPUs 1508, and for example, by using inference accelerators, the designer may target either one or more DLAs or one or more GPUs 1508 for any function. For example, in at least one embodiment, the designer may concentrate CNN and floating-point arithmetic processing on one or more DLAs and offload other functions to one or more GPUs 1508 and / or one or more accelerators 1514.

[0173] In at least one embodiment, one or more accelerators 1514 may include a programmable vision accelerator ("PVA"), which may be referred to herein as a computer vision accelerator instead. In at least one embodiment, the PVA may be designed and configured to accelerate computer vision algorithms for advanced driver-assistance systems ("ADAS") 1538, autonomous driving, augmented reality ("AR") applications, and / or virtual reality ("VR") applications. In at least one embodiment, the PVA may provide a balance between performance and flexibility. For example, in at least one embodiment, each PVA may include, for example, any number of reduced instruction set computer ("RISC") cores, direct memory access ("DMA"), and / or any number of vector processors.

[0174] In at least one embodiment, the RISC core may interact with an image sensor (e.g., an image sensor of any camera described herein), one or more image signal processors, and so on. In at least one embodiment, each RISC core may include any amount of memory. In at least one embodiment, the RISC core may use one of several protocols, depending on the embodiment. In at least one embodiment, the RISC core may run a real-time operating system ("RTOS"). In at least one embodiment, the RISC core may be implemented using one or more integrated circuit devices, application-specific integrated circuits ("ASICs"), and / or memory devices. For example, in at least one embodiment, the RISC core may include an instruction cache and / or tightly coupled RAM.

[0175] In at least one embodiment, the DMA may enable components of the PVA to access system memory independently of one or more CPUs 1506. In at least one embodiment, the DMA may support any number of features used to provide optimization to the PVA, including, but not limited to, support for multidimensional addressing and / or circular addressing. In at least one embodiment, the DMA may support up to six or more addressing dimensions, these addressing dimensions may include, but not limited to, block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.

[0176] In at least one embodiment, a vector processor is a programmable processor that can be designed to efficiently and flexibly perform programming for computer vision algorithms and can provide signal processing capabilities. In at least one embodiment, the PVA may include a PVA core and two vector processing subsystem partitions. In at least one embodiment, the PVA core may include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. In at least one embodiment, the vector processing subsystem may act as the primary processing engine of the PVA and may include a vector processing unit ("VPU"), an instruction cache, and / or vector memory (e.g., "VMEM"). In at least one embodiment, the VPU core may include a digital signal processor, such as a single instruction, multiple data ("SIMD") or very long instruction word ("VLIW") digital signal processor. In at least one embodiment, a combination of SIMD and VLIW can improve throughput and speed.

[0177] In at least one embodiment, each vector processor may include an instruction cache and be coupled to dedicated memory. As a result, in at least one embodiment, each vector processor may be configured to operate independently of other vector processors. In at least one embodiment, the vector processors included in a particular PVA may be configured to employ data parallelism. For example, in at least one embodiment, multiple vector processors included in a single PVA may execute a common computer vision algorithm, but on different regions of an image. In at least one embodiment, the vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on a single image, or even execute different algorithms on a contiguous image or portion of an image. In at least one embodiment, among other things, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each PVA. In at least one embodiment, the PVA may include additional error correction code ("ECC") memory to improve the overall safety of the system.

[0178] In at least one embodiment, one or more accelerators 1514 may include a computer vision network on-chip and static random-access memory ("SRAM") to provide high-bandwidth, low-latency SRAM for one or more accelerators 1514. In at least one embodiment, the on-chip memory may include, for example, at least 4 MB of SRAM including, but not limited to, eight field-configurable memory blocks, which may be accessible by both the PVA and DLA. In at least one embodiment, each pair of memory blocks may include an advanced peripheral bus ("APB") interface, configurable circuit elements, a controller, and a multiplexer. In at least one embodiment, any type of memory may be used. In at least one embodiment, the PVA and DLA may access the memory via a backbone that provides high-speed access to the memory. In at least one embodiment, the backbone may include a computer vision network on a chip that interconnects the PVA and DLA to memory (for example, using an APB).

[0179] In at least one embodiment, the computer vision network on chip may include an interface in which both the PVA and DLA determine to provide ready and enable signals before any control signals / addresses / data are transmitted. In at least one embodiment, the interface may provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-type communication for continuous data transfer. In at least one embodiment, the interface may conform to the International Organization for Standardization (ISO) 26262 or the International Electrotechnical Commission (IEC) 61508 standard, but other standards and protocols may be used.

[0180] In at least one embodiment, one or more of the (one or more) SoC1504s may include a real-time ray tracing hardware accelerator. In at least one embodiment, the real-time ray tracing hardware accelerator may be used to quickly and efficiently determine the location and extent of objects (e.g., in a world model) to generate real-time visualization simulations for RADAR signal interpretation, sound propagation synthesis and / or analysis, SONAR system simulation, general waveform propagation simulation, comparison with LIDAR data for localization and / or other functions, and / or other uses.

[0181] In at least one embodiment, one or more accelerators 1514 can have diverse uses for autonomous driving. In at least one embodiment, PVA can be used for key processing stages in ADAS and autonomous vehicles. In at least one embodiment, the capabilities of PVA are well matched for algorithmic domains requiring predictable processing at low power and low latency. In other words, PVA performs well even on small datasets for semi-dense or dense regular computations that may require predictable runtime along with low latency and low power. In at least one embodiment, such as in a vehicle 1500, PVA can be designed to run conventional computer vision algorithms, as they may be efficient in object detection and integer calculations.

[0182] For example, according to at least one embodiment of the technology, PVA is used to implement computer stereo vision. In at least one embodiment, a semi-global matching-based algorithm may be used in some examples, but is not limited to this. In at least one embodiment, an application for Level 3–5 autonomous driving uses motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.) on the fly. In at least one embodiment, PVA may implement computer stereo vision functionality for input from two monocular cameras.

[0183] In at least one embodiment, PVA may be used to perform high-density optical flow. For example, in at least one embodiment, PVA may process raw RADAR data (e.g., using a 4D Fast Fourier Transform) to provide processed RADAR data. In at least one embodiment, PVA may be used for time-of-flight depth processing by processing raw time-of-flight data to provide processed time-of-flight data.

[0184] In at least one embodiment, DLA may be used to power any type of network for improving control and driving safety, including, for example, a neural network that outputs a measure of confidence for each object detection. In at least one embodiment, confidence may be expressed or interpreted as the probability of each detection compared to other detections, or as providing its relative “weight.” In at least one embodiment, the confidence measure allows the system to make further determinations about which detections should be considered true positive detections rather than false positive detections. In at least one embodiment, the system may set a threshold for confidence and consider only detections exceeding the threshold as true positive detections. In embodiments where an automatic emergency braking (“AEB”) system is used, a false positive detection would cause the vehicle to automatically apply the emergency brakes, which is obviously undesirable. In at least one embodiment, a highly reliable detection may be considered a trigger for the AEB. In at least one embodiment, DLA may power a neural network to regress confidence values. In at least one embodiment, the neural network may take as its input at least a subset of parameters, including, among other things, the dimensions of the bounding box, ground plane estimates obtained (e.g., from another subsystem), outputs from one or more IMU sensors 1566 correlated with the orientation of the vehicle 1500, distance, and 3D location estimates of objects obtained from the neural network and / or other sensors (e.g., one or more LiDAR sensors 1564 or one or more RADAR sensors 1560).

[0185] In at least one embodiment, one or more of the (one or more) SoC1504 may include (one or more) data stores 1516 (e.g., memory). In at least one embodiment, the (one or more) data stores 1516 may be on-chip memory of the (one or more) SoC1504, which may store neural networks to be run on the (one or more) GPUs 1508 and / or DLA. In at least one embodiment, the capacity of the (one or more) data stores 1516 may be large enough to store multiple instances of the neural network for redundancy and safety. In at least one embodiment, the (one or more) data stores 1516 may include (one or more) L2 or L3 caches.

[0186] In at least one embodiment, one or more of the (one or more) SoC1504 may include any number of (one or more) processors 1510 (e.g., embedded processors). In at least one embodiment, the (one or more) processors 1510 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions and related security enforcement. In at least one embodiment, the boot and power management processor may be part of the boot sequence of the (one or more) SoC1504 and may provide runtime power management services. In at least one embodiment, the boot power and management processor may provide clock and voltage programming, assistance with system low-power state transitions, management of thermal and temperature sensors of the (one or more) SoC1504, and / or management of the power state of the (one or more) SoC1504. In at least one embodiment, each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and (one or more) SoC 1504 may use the ring oscillator to detect the temperatures of (one or more) CPU 1506, (one or more) GPU 1508, and / or (one or more) accelerator 1514. In at least one embodiment, if it is determined that the temperature exceeds a threshold, the boot and power management processor may enter a temperature fault routine to put (one or more) SoC 1504 into a low-power state and / or put the vehicle 1500 into chauffeur-to-safe stop mode (e.g., to safely stop the vehicle 1500).

[0187] In at least one embodiment, the (one or more) processor 1510 may further include a set of embedded processors that can act as an audio processing engine, which may be an audio subsystem enabling full hardware support for multi-channel audio via multiple interfaces and a wide range of flexible audio I / O interfaces. In at least one embodiment, the audio processing engine is a dedicated processor core having a digital signal processor with dedicated RAM.

[0188] In at least one embodiment, the (one or more) processors 1510 may further include an always-on processor engine that can provide the hardware features necessary to support low-power sensor management and startup use cases. In at least one embodiment, the always-on processor engine may include, but is not limited to, a processor core, tightly coupled RAM, supporting peripherals (e.g., timer and interrupt controllers), various I / O controller peripherals, and routing logic.

[0189] In at least one embodiment, the (one or more) processor 1510 may further include a safety cluster engine, which may include, but is not limited to, a dedicated processor subsystem for addressing safety management for automotive applications. In at least one embodiment, the safety cluster engine may include, but is not limited to, two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In safety mode, the two or more cores may, in at least one embodiment, operate in lockstep mode and function as a single core with comparison logic for detecting any differences between their operations. In at least one embodiment, the (one or more) processor 1510 may further include a real-time camera engine, which may include, but is not limited to, a dedicated processor subsystem for addressing real-time camera management. In at least one embodiment, the processor 1510 (one or more) may further include a high dynamic range signal processor, which may include, but is not limited to, an image signal processor that is a hardware engine that is part of a camera processing pipeline.

[0190] In at least one embodiment, the (one or more) processor 1510 may include a video image synthesizer, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to produce a final image for a player window. In at least one embodiment, the video image synthesizer may perform lens distortion correction on the (one or more) wide-angle camera 1570, the (one or more) ambient camera 1574, and / or the (one or more) cabin surveillance camera sensors. In at least one embodiment, the (one or more) cabin surveillance camera sensors are preferably monitored by a neural network running on another instance of the SoC 1504, which is configured to identify and respond to events within the cabin. In at least one embodiment, the in-cabin system may perform lip-reading to activate cellular services, make phone calls, write emails, change the vehicle's destination, activate or change the vehicle's infotainment system and settings, and provide voice-activated web surfing. In at least one embodiment, some functions are available to the driver when the vehicle is operating in autonomous mode and are unavailable in other cases.

[0191] In at least one embodiment, the video image synthesizer may include extended temporal noise reduction for both spatial and temporal noise reduction. For example, in at least one embodiment, if motion occurs in the video, the noise reduction appropriately weights the spatial information and reduces the weight of the information provided by adjacent frames. In at least one embodiment, if the image or part of the image does not contain motion, the temporal noise reduction performed by the video image synthesizer may use information from previous images to reduce noise in the current image.

[0192] In at least one embodiment, the video image synthesizer may also be configured to perform stereo rectification on the input stereo lens frame. In at least one embodiment, the video image synthesizer may be further used for user interface compositing when the operating system desktop is in use, and (one or more) GPUs 1508 are not required to continuously render new surfaces. In at least one embodiment, when (one or more) GPUs 1508 are powered on, active, and performing 3D rendering, the video image synthesizer may be used to offload (one or more) GPUs 1508 to improve performance and responsiveness.

[0193] In at least one embodiment, one or more of the SoC1504 may further include a Mobile Industry Processor Interface ("MIPI") camera serial interface, a high-speed interface, and / or a video input block that can be used for the camera and associated pixel input functions to receive video and input from the camera. In at least one embodiment, one or more of the SoC1504 may further include one or more input / output controllers, which may be controlled by software and may be used to receive I / O signals not committed to a specific role.

[0194] In at least one embodiment, one or more of the (one or more) SoCs 1504 may further include a wide range of peripheral interfaces for enabling communication with peripherals, audio encoders / decoders ("codecs"), power management, and / or other devices. In at least one embodiment, the (one or more) SoCs 1504 may be used to process data from cameras (e.g., connected via Gigabit Multimedia Serial Link and Ethernet Channel), data from sensors (e.g., one or more LiDAR sensors 1564, one or more RADAR sensors 1560, etc., which may be connected via Ethernet Channel), data from bus 1502 (e.g., vehicle speed, steering wheel position, etc.), data from one or more GNSS sensors 1558 (e.g., connected via Ethernet Bus or CAN Bus), and the like. In at least one embodiment, one or more of the (one or more) SoCs 1504 may further include a dedicated high-performance, high-capacity storage controller, which may include its own DMA engine and may be used to free the (one or more) CPUs 1506 from routine data management tasks.

[0195] In at least one embodiment, one or more SoC1504s may be an end-to-end platform with a flexible architecture spanning automation levels 3–5, thereby providing a comprehensive functional safety architecture that leverages and efficiently utilizes computer vision and ADAS techniques for diversity and redundancy, and can provide a platform for a flexible and reliable driving software stack, along with deep learning tools. In at least one embodiment, one or more SoC1504s may be faster, more reliable, and more energy-efficient and space-efficient than conventional systems. For example, in at least one embodiment, one or more accelerators 1514, when combined with one or more CPUs 1506, one or more GPUs 1508, and one or more data stores 1516, may provide a fast and efficient platform for Level 3–5 autonomous vehicles.

[0196] In at least one embodiment, a computer vision algorithm may run on a CPU, and this algorithm may be configured using a high-level programming language such as C to execute a wide variety of processing algorithms across a wide variety of visual data. However, in at least one embodiment, the CPU often fails to meet the performance requirements of many computer vision applications, such as requirements related to execution time and power consumption. In at least one embodiment, many CPUs are unable to execute complex object detection algorithms in real time, as used in in-vehicle ADAS applications and in actual Level 3-5 autonomous vehicles.

[0197] The embodiments described herein enable multiple neural networks to be implemented simultaneously and / or sequentially, allowing the results to be combined to enable Level 3–5 autonomous driving capabilities. For example, in at least one embodiment, a CNN running on a DLA or a separate GPU (e.g., one or more GPU1520s) may include text and word recognition, enabling the neural network to read and understand traffic signs, including signs that the neural network has not been specifically trained on. In at least one embodiment, the DLA may further include a neural network that can identify and interpret signs, provide a semantic understanding of the signs, and pass that semantic understanding to a route planning module running on a CPU complex.

[0198] In at least one embodiment, multiple neural networks may be operating simultaneously with respect to Level 3, 4, or 5 driving. For example, in at least one embodiment, a warning sign with an electric light and the text "Caution: Flashing light indicates icy condition" may be interpreted independently or collectively by several neural networks. In at least one embodiment, such a warning sign itself may be identified as a traffic sign by a first introduced neural network (e.g., a trained neural network), and the text "Flashing light indicates icy condition" may be interpreted by a second introduced neural network, which informs the vehicle's route planning software (preferably running on the CPU complex) that an icy condition is present when a flashing light is detected. In at least one embodiment, the flashing light may be identified by running a third introduced neural network over multiple frames, which informs the vehicle's route planning software of the presence (or absence) of the flashing light. In at least one embodiment, all three neural networks may run simultaneously within the DLA and / or on one or more GPU1508s, etc.

[0199] In at least one embodiment, a CNN for facial recognition and vehicle owner identification may use data from camera sensors to identify the presence of an authorized driver and / or owner of vehicle 1500. In at least one embodiment, an always-on sensor processing engine may be used to unlock the vehicle when the owner approaches the driver's door and turns on the lights, and in security mode to disable such vehicle when the owner leaves such vehicle. In this way, (one or more) SoCs 1504 provide security against theft and / or vehicle hijacking.

[0200] In at least one embodiment, a CNN for emergency vehicle detection and identification may use data from microphone 1596 to detect and identify emergency vehicle sirens. In at least one embodiment, one or more SoCs 1504 use a CNN to classify environmental and urban sounds, as well as visual data. In at least one embodiment, a CNN operating on DLA is trained to identify the relative speed at which an emergency vehicle is approaching (for example, by using the Doppler effect). In at least one embodiment, a CNN may also be trained to identify emergency vehicles specific to the area in which the vehicle is operating, as identified by one or more GNSS sensors 1558. In at least one embodiment, when operating in Europe, the CNN attempts to detect European sirens, and when in North America, the CNN attempts to identify only North American sirens. In at least one embodiment, when an emergency vehicle is detected, a control program may be used to execute an emergency vehicle safety routine, slow down the vehicle, pull over to the side of the road, stop the vehicle, and / or idle the vehicle in conjunction with one or more ultrasonic sensors 1562 until the emergency vehicle has passed.

[0201] In at least one embodiment, the vehicle 1500 may include (one or more) CPUs 1518 (e.g., (one or more) individual CPUs, or (one or more) dCPUs), which may be coupled to (one or more) SoCs 1504 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, the (one or more) CPUs 1518 may include, for example, an x86 processor. The (one or more) CPUs 1518 may be used to perform any of a variety of functions, including, for example, mediating potentially inconsistent results between ADAS sensors and (one or more) SoCs 1504, and / or monitoring the status and health of (one or more) controllers 1536 and / or the infotainment system ("Infotainment SoC") 1530 on the chip.

[0202] In at least one embodiment, the vehicle 1500 may include (one or more) GPUs 1520 (e.g., (one or more) individual GPUs, or (one or more) dGPUs), and the (one or more) GPUs 1520 may be coupled to the (one or more) SoCs 1504 via a high-speed interconnect (e.g., NVIDIA's NVLINK channel). In at least one embodiment, the (one or more) GPUs 1520 may provide additional artificial intelligence functionality, such as by running redundant and / or different neural networks, and may be used to train and / or update neural networks based at least in part on input from the vehicle 1500's sensors (e.g., sensor data).

[0203] In at least one embodiment, vehicle 1500 may further include a network interface 1524, which may include, but is not limited to, one or more wireless antennas 1526 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). In at least one embodiment, the network interface 1524 may be used to enable wireless connectivity to Internet cloud services (e.g., with one or more servers and / or other network devices) with other vehicles and / or computing devices (e.g., passenger client devices). In at least one embodiment, a direct link may be established between vehicle 150 and other vehicles for communication with other vehicles, and / or an indirect link may be established (e.g., over a network and via the Internet). In at least one embodiment, the direct link may be provided using an inter-vehicle communication link. In at least one embodiment, the vehicle-to-vehicle communication link may provide vehicle 1500 with information about nearby vehicles (e.g., vehicles in front of, to the side of, and / or behind vehicle 1500). In at least one embodiment, such aforementioned functionality may be part of the cooperative adaptive driving control functionality of vehicle 1500.

[0204] In at least one embodiment, the network interface 1524 may include an SoC that provides modulation and demodulation functionality, enabling one or more controllers 1536 to communicate over a wireless network. In at least one embodiment, the network interface 1524 may include a radio frequency front end for baseband-to-radio frequency up-conversion and radio frequency-to-baseband down-conversion. In at least one embodiment, frequency conversion may be carried out in any technically feasible manner. For example, frequency conversion may be carried out through a well-known process and / or using a superheterodyne process. In at least one embodiment, the radio frequency front end functionality may be provided by a separate chip. In at least one embodiment, the network interface may include wireless functionality for communication over LTE, WCDMA®, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0205] In at least one embodiment, the vehicle 1500 may further include one or more data stores 1528, which may include, but are not limited to, off-chip (e.g., not on one or more SoCs 1504) storage. In at least one embodiment, the data stores 1528 may include one or more storage elements, which may include, but are not limited to, RAM, SRAM, dynamic random-access memory ("DRAM"), video random-access memory ("VRAM"), flash memory, hard disks, and / or other components and / or devices capable of storing at least one bit of data.

[0206] In at least one embodiment, the vehicle 1500 may further include (one or more) GNSS sensors 1558 (e.g., GPS and / or auxiliary GPS sensors) to assist mapping, perception, occupy grid generation, and / or route planning functions. In at least one embodiment, any number of GNSS sensors 1558 may be used, including, for example, a GPS using a USB connector with an Ethernet-to-serial (e.g., RS-232) bridge.

[0207] In at least one embodiment, the vehicle 1500 may further include one or more RADAR sensors 1560. In at least one embodiment, one or more RADAR sensors 1560 may be used by the vehicle 1500 for long-range vehicle detection, even in darkness and / or severe weather conditions. In at least one embodiment, the functional safety level of the RADAR may be ASIL B. In at least one embodiment, one or more RADAR sensors 1560 may, in some examples, use a CAN bus and / or bus 1502 for control (e.g., to transmit data generated by one or more RADAR sensors 1560) and for accessing object tracking data, along with access to an Ethernet channel for accessing raw data. In at least one embodiment, a wide variety of RADAR sensor types may be used. For example, but not limited to, one or more RADAR sensors 1560 may be preferred for forward, rear, and side RADAR use. In at least one embodiment, one or more of the (one or more) RADAR sensors 1560 are pulsed Doppler RADAR sensors.

[0208] In at least one embodiment, the (one or more) RADAR sensors 1560 may include different configurations, such as narrow-field long-range, wide-field short-range, and short-range lateral coverage. In at least one embodiment, the long-range RADAR may be used for adaptive driving control functionality. In at least one embodiment, the long-range RADAR system may provide a wide field of view achieved by two or more independent scans, such as within a 250 m (meter) range. In at least one embodiment, the (one or more) RADAR sensors 1560 may help distinguish between static and moving objects and may be used by the ADAS system 1538 for emergency braking assistance and forward collision warning. In at least one embodiment, the (one or more) sensors 1560 included in the long-range RADAR system may include, but are not limited to, multiple (e.g., six or more) fixed RADAR antennas, as well as monostatic and multimodal RADARs with high-speed CAN and FlexRay interfaces. In at least one embodiment, if there are six antennas, the four central antennas may create a focused beam pattern designed to record the area around vehicle 1500 at a faster speed with minimal interference from traffic in adjacent lanes. In at least one embodiment, two additional antennas may expand the field of view, which may allow for the rapid detection of vehicles entering or leaving the lane of vehicle 1500.

[0209] In at least one embodiment, the medium-range RADAR system may, as an example, include a range of up to 160 m (forward) or 80 m (rearward) and a field of view of up to 42 degrees (forward) or 150 degrees (rearward). In at least one embodiment, the short-range RADAR system may include, but not limited to, any number of RADAR sensors 1560 designed to be mounted on both ends of the rear bumper. When mounted on both ends of the rear bumper, in at least one embodiment, the RADAR sensor system may create two beams that constantly monitor blind spots in the rearward and adjacent to the vehicle. In at least one embodiment, the short-range RADAR system may be used in an ADAS system 1538 for blind spot detection and / or lane change assistance.

[0210] In at least one embodiment, the vehicle 1500 may further include one or more ultrasonic sensors 1562. In at least one embodiment, one or more ultrasonic sensors 1562, which can be positioned in front of, behind, and / or to the side of the vehicle 1500, may be used for parking assistance and / or to create and update the occupancy grid. In at least one embodiment, a wide variety of one or more ultrasonic sensors 1562 may be used, and different ultrasonic sensors 1562 may be used for different detection ranges (e.g., 2.5m, 4m). In at least one embodiment, one or more ultrasonic sensors 1562 may operate at functional safety level ASIL B.

[0211] In at least one embodiment, the vehicle 1500 may include one or more LiDAR sensors 1564. In at least one embodiment, one or more LiDAR sensors 1564 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, one or more LiDAR sensors 1564 may operate at functional safety level ASIL B. In at least one embodiment, the vehicle 1500 may include multiple LiDAR sensors 1564 (e.g., two, four, six, etc.), and these LiDAR sensors 1564 may use an Ethernet channel (e.g., to provide data to a Gigabit Ethernet switch).

[0212] In at least one embodiment, one or more LiDAR sensors 1564 may be capable of providing a list of objects and their distances over a 360-degree field of view. In at least one embodiment, one or more commercially available LiDAR sensors 1564 may have an advertised range of approximately 100m, for example, with an accuracy of 2cm to 3cm and support for a 100Mbps Ethernet connection. In at least one embodiment, one or more non-protruding LiDAR sensors may be used. In such embodiments, one or more LiDAR sensors 1564 may include small devices that can be incorporated into the front, rear, side, and / or corner locations of a vehicle 1500. In at least one embodiment, one or more LiDAR sensors 1564 may, in such embodiments, provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees over a range of 200m, even for low-reflectivity objects. In at least one embodiment, one or more front-mounted LiDAR sensors 1564 may be configured for a horizontal field of view between 45 and 135 degrees.

[0213] In at least one embodiment, LiDAR technology such as 3D flash LiDAR may also be used. In at least one embodiment, the 3D flash LiDAR uses a laser flash as a transmission source to illuminate the area around the vehicle 1500 up to approximately 200 m. In at least one embodiment, the flash LiDAR unit includes, but is not limited to, a receptor which records the transit time of the laser pulse and the reflected light on each pixel, which corresponds to the range from the vehicle 1500 to the object. In at least one embodiment, the flash LiDAR enables the generation of highly accurate and distortion-free ambient images with each laser flash. In at least one embodiment, four flash LiDAR sensors may be introduced, one on each side of the vehicle 1500. In at least one embodiment, the 3D flash LiDAR system includes, but is not limited to, a solid-state 3D staring array LiDAR camera (e.g., a non-scanning LiDAR device) with no moving parts other than a fan. In at least one embodiment, a flash LiDAR device may use a Class I (eye-safe) laser pulse of 5 nanoseconds per frame and capture reflected laser light as a 3D-range point cloud and position-synchronized (co-registered) intensity data.

[0214] In at least one embodiment, the vehicle 1500 may further include one or more IMU sensors 1566. In at least one embodiment, one or more IMU sensors 1566 may be located in the center of the rear axle of the vehicle 1500. In at least one embodiment, one or more IMU sensors 1566 may include, for example, one or more accelerometers, one or more magnetometers, one or more gyroscopes, magnetic compasses, multiple magnetic compasses, and / or other sensor types. In at least one embodiment, such as in a 6-axis application, one or more IMU sensors 1566 may include, for example, an accelerometer and a gyroscope. In at least one embodiment, such as in a 9-axis application, one or more IMU sensors 1566 may include, for example, an accelerometer, a gyroscope, and a magnetometer.

[0215] In at least one embodiment, the (one or more) IMU sensors 1566 may be implemented as a small, high-performance GPS-Aided Inertial Navigation System ("GPS / INS") that combines micro-electro-mechanical systems ("MEMS") inertial sensors, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. In at least one embodiment, the (one or more) IMU sensors 1566 enable the vehicle 1500 to estimate its bearing without requiring input from magnetic sensors by directly observing changes in velocity and correlating them from GPS to the (one or more) IMU sensors 1566. In at least one embodiment, the (one or more) IMU sensors 1566 and the (one or more) GNSS sensors 1558 may be combined in a single integrated unit.

[0216] In at least one embodiment, the vehicle 1500 may include one or more microphones 1596 placed inside and / or around the vehicle 1500. In at least one embodiment, one or more microphones 1596 may, among other things, be used for emergency vehicle detection and identification.

[0217] In at least one embodiment, the vehicle 1500 may further include any number of camera types, including (one or more) stereo cameras 1568, (one or more) wide-angle cameras 1570, (one or more) infrared cameras 1572, (one or more) ambient cameras 1574, (one or more) long-range cameras 1598, (one or more) medium-range cameras 1576, and / or other camera types. In at least one embodiment, cameras may be used to capture image data around the entire periphery of the vehicle 1500. In at least one embodiment, the type of camera used depends on the vehicle 1500. In at least one embodiment, any combination of camera types may be used to provide the required coverage around the vehicle 1500. In at least one embodiment, the number of cameras introduced may vary depending on the embodiment. For example, in at least one embodiment, the vehicle 1500 may include six cameras, seven cameras, ten cameras, twelve cameras, or any other number of cameras. In at least one embodiment, the camera may, as an example, support Gigabit Multimedia Serial Link ("GMSL") and / or Gigabit Ethernet communication, though not limited to this example. In at least one embodiment, each camera may be as described more in detail previously herein with respect to Figures 15A and 15B.

[0218] In at least one embodiment, the vehicle 1500 may further include one or more vibration sensors 1542. In at least one embodiment, one or more vibration sensors 1542 may measure vibrations of components of the vehicle 1500, such as one or more axles. For example, in at least one embodiment, a change in vibration may indicate a change in the road surface. In at least one embodiment, when two or more vibration sensors 1542 are used, the difference in vibration may be used to determine the amount of friction or slip of the road surface (for example, when the difference in vibration is between a power-driven axle and a free-rotating axle).

[0219] In at least one embodiment, the vehicle 1500 may include an ADAS system 1538. In at least one embodiment, the ADAS system 1538 may include a SoC in some examples, but is not limited to. In at least one embodiment, the ADAS system 1538 may include, but not limited to, any number and combination of autonomous / adaptive / automatic cruise control ("ACC") systems, cooperative adaptive cruise control ("CACC") systems, forward crash warning ("FCW") systems, automatic emergency braking ("AEB") systems, lane departure warning ("LDW") systems, lane keep assist ("LKA") systems, blind spot warning ("BSW") systems, rear cross-traffic warning ("RCTW") systems, collision warning ("CW") systems, lane centering ("LC") systems, and / or other systems, features, and / or functionalities.

[0220] In at least one embodiment, the ACC system may use (one or more) RADAR sensors 1560, (one or more) LIDAR sensors 1564, and / or any number of cameras. In at least one embodiment, the ACC system may include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, the longitudinal ACC system monitors and controls the distance of vehicle 1500 to another vehicle directly in front and automatically adjusts the speed of vehicle 1500 to maintain a safe distance from the vehicle ahead. In at least one embodiment, the lateral ACC system enforces distance maintenance and advises vehicle 1500 to change lanes when necessary. In at least one embodiment, the lateral ACC is relevant to other ADAS applications such as LC and CW.

[0221] In at least one embodiment, the CACC system uses information from other vehicles, which may be received from other vehicles via a network interface 1524 and / or (one or more) wireless antennas 1526, either wirelessly or indirectly via a network connection (e.g., via the Internet). In at least one embodiment, direct links may be provided by vehicle-to-vehicle ("V2V") communication links, and indirect links may be provided by infrastructure-to-vehicle ("I2V") communication links. Generally, V2V communication provides information about the immediately preceding vehicle (e.g., a vehicle in the same lane immediately before vehicle 1500), and I2V communication provides information about traffic further ahead. In at least one embodiment, the CACC system may include either or both I2V and V2V information sources. In at least one embodiment, if information about the vehicle in front of vehicle 1500 is available, the CACC system may become more reliable, which could improve the smoothness of traffic flow and reduce congestion on the road.

[0222] In at least one embodiment, the FCW system is designed to alert the driver about hazardous materials so that such a driver can take corrective action. In at least one embodiment, the FCW system uses a front camera and / or (one or more) radar sensors 1560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to provide driver feedback, such as a display, speaker, and / or vibration components. In at least one embodiment, the FCW system may provide warnings in the form of sound, visual warnings, vibration, and / or quick brake pulses.

[0223] In at least one embodiment, the AEB system may detect an impending forward collision with another vehicle or other object and automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, the AEB system may use one or more front cameras and / or one or more RADAR sensors 1560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when the AEB system detects a hazard, the AEB system may first alert the driver to take corrective action to avoid a collision, and if the driver does not take corrective action, the AEB system may automatically apply the brakes to prevent the anticipated collision or at least mitigate its impact. In at least one embodiment, the AEB system may include techniques such as dynamic brake support and / or pre-crash braking.

[0224] In at least one embodiment, the LDW system provides visual, auditory, and / or tactile warnings, such as vibration of the steering wheel or seat, to alert the driver when the vehicle 1500 crosses a lane marker. In at least one embodiment, the LDW system does not activate when the driver indicates an intentional lane departure, such as by activating the turn signal. In at least one embodiment, the LDW system may use a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to provide driver feedback, such as a display, speaker, and / or vibration component. In at least one embodiment, the LKA system is a variation of the LDW system. In at least one embodiment, the LKA system provides steering input or brake control to correct the vehicle 1500 if the vehicle 1500 begins to move out of its lane.

[0225] In at least one embodiment, the BSW system detects vehicles in the vehicle's blind spot and warns the driver about those vehicles. In at least one embodiment, the BSW system may provide visual, auditory, and / or tactile alerts to indicate that merging or changing lanes is not safe. In at least one embodiment, the BSW system may provide additional warnings when the driver uses the turn signal. In at least one embodiment, the BSW system may use one or more rear-facing cameras and / or one or more RADAR sensors 1560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to driver feedback, such as a display, speaker, and / or vibration component.

[0226] In at least one embodiment, the RCTW system may provide visual, auditory, and / or tactile notifications when an object is detected outside the range of the rear camera while the vehicle 1500 is reversing. In at least one embodiment, the RCTW system includes an AEB system to ensure that the vehicle brakes are applied to avoid a crash. In at least one embodiment, the RCTW system may use one or more rear-facing radar sensors 1560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to provide driver feedback, such as a display, speaker, and / or vibration components.

[0227] In at least one embodiment, conventional ADAS systems may be prone to producing false positive results, which can be annoying and distracting to the driver, but this is usually not a major issue as conventional ADAS systems alert the driver, allowing the driver to determine whether safety conditions truly exist and act accordingly. In at least one embodiment, the vehicle 1500 itself determines, in the event of conflicting results, whether to follow the result from a primary computer (e.g., a first controller among the controllers 1536) or a secondary computer (e.g., a second controller among the controllers 1536). For example, in at least one embodiment, the ADAS system 1538 may be a backup and / or secondary computer for providing perceptual information to a backup computer rationality module. In at least one embodiment, the backup computer rationality monitor may run a variety of redundant software on hardware components to detect failures in perceptual and dynamic driving tasks. In at least one embodiment, the output from the ADAS system 1538 may be provided to a supervisory MCU. In at least one embodiment, if the output from the primary computer and the output from the secondary computer are inconsistent, the supervising MCU determines how to reconcile the inconsistency to ensure safe operation.

[0228] In at least one embodiment, the primary computer may be configured to provide the supervising MCU with a reliability score indicating the reliability of the primary computer in the selected outcome. In at least one embodiment, if the reliability score exceeds a threshold, the supervising MCU may follow the instructions of the primary computer, regardless of whether the secondary computer provides contradictory or inconsistent results. In at least one embodiment, if the reliability score does not meet the threshold, and the primary and secondary computers produce different results (e.g., contradictory), the supervising MCU may mediate between the computers to determine an appropriate outcome.

[0229] In at least one embodiment, the supervising MCU may be configured to operate one or more neural networks trained and configured to determine, at least in part, the conditions under which a secondary computer provides a false alarm, based on the output from the primary computer and the output from the secondary computer. In at least one embodiment, one or more neural networks in the supervising MCU may learn when the output from the secondary computer can be trusted and when it cannot. For example, in at least one embodiment, when the secondary computer is a radar-based FCW system, one or more neural networks in the supervising MCU may learn when the FCW system identifies a metallic object that is not actually a hazard, such as a drain grate or manhole cover, which triggers an alarm. In at least one embodiment, when the secondary computer is a camera-based LDW system, the neural network in the supervising MCU may learn to disable the LDW when a cyclist or pedestrian is present and lane departure is actually the safest operation. In at least one embodiment, the supervisor MCU may include at least one DLA or GPU suitable for running (one or more) neural networks together with associated memory. In at least one embodiment, the supervisor MCU may comprise and / or be included as a component of (one or more) SoC1504.

[0230] In at least one embodiment, the ADAS system 1538 may include a secondary computer that implements ADAS functionality using conventional computer vision rules. In at least one embodiment, the secondary computer may use conventional computer vision rules (if-then), and the presence of (one or more) neural networks in the supervising MCU may improve reliability, safety, and performance. For example, in at least one embodiment, diverse implementations and intentional non-identities make the entire system more fault-tolerant, particularly against failures caused by software (or software-hardware interface) functionality. For example, in at least one embodiment, if there is a software bug or error in the software running on the primary computer, and non-identical software code running on the secondary computer provides a consistent overall result, the supervising MCU may have greater confidence that the overall result is correct and that the bug in the software or hardware on the primary computer has not caused a critical error.

[0231] In at least one embodiment, the output of the ADAS system 1538 may be fed to the perception block and / or the dynamic driving task block of the primary computer. For example, in at least one embodiment, if the ADAS system 1538 indicates a forward crash warning due to an object immediately preceding, the perception block may use this information when identifying the object. In at least one embodiment, the secondary computer may have its own trained, and therefore false-positive-reducing, neural network, as described herein.

[0232] In at least one embodiment, the vehicle 1500 may further include an infotainment SoC 1530 (for example, an in-vehicle infotainment system (IVI)). Although shown and described as an SoC, the infotainment system SoC 1530 may not be an SoC in at least one embodiment and may include, but not limited to, two or more separate components. In at least one embodiment, the infotainment SoC 1530 may include, but is not limited to, a combination of hardware and software which may be used to provide the vehicle 1500 with audio (e.g., music, personal digital assistant, navigation commands, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., a navigation system, rear parking assist, wireless data system, vehicle-related information such as fuel level, total mileage, brake fuel level, oil level, door open / closed, air filter information, etc.). For example, the infotainment SoC 1530 may include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, car computer, in-car entertainment, Wi-Fi, steering wheel audio control, hands-free voice control, heads-up display ("HUD"), HMI display 1534, telematics devices, control panels (for example, for controlling and / or interacting with various components, features, and / or systems), and / or other components. In at least one embodiment, the infotainment SoC 1530 may be further used to provide (one or more) users of the vehicle 1500 with (for example, visual and / or auditory) information, such as information from the ADAS system 1538, autonomous driving information such as planned vehicle operation and trajectory, ambient information (for example, intersection information, vehicle information, road information, etc.), and / or other information.

[0233] In at least one embodiment, the infotainment SoC 1530 may include any amount and type of GPU functionality. In at least one embodiment, the infotainment SoC 1530 may communicate with other devices, systems, and / or components of the vehicle 1500 via bus 1502. In at least one embodiment, the infotainment SoC 1530 may be coupled to a supervisory MCU so that if one or more primary controllers 1536 (e.g., the primary and / or backup computers of the vehicle 1500) fail, the GPU of the infotainment system may perform some self-driving functions. In at least one embodiment, the infotainment SoC 1530 may put the vehicle 1500 into driver-safe stop mode as described herein.

[0234] In at least one embodiment, the vehicle 1500 may further include an instrument cluster 1532 (e.g., a digital dashboard, electronic instrument cluster, digital instrument panel, etc.). In at least one embodiment, the instrument cluster 1532 may include, but is not limited to, a controller and / or supercomputer (e.g., a separate controller or supercomputer). In at least one embodiment, the instrument cluster 1532 may include, but is not limited to, a set of instrumentation such as a speedometer, fuel level, oil pressure, tachometer, odometer, direction indicator, shift lever position indicator, (one or more) seat belt warning lights, (one or more) parking brake warning lights, (one or more) engine fault lights, auxiliary restraint system (e.g., airbag) information, light control, safety system control, navigation information, etc. In some examples, the information may be displayed and / or shared between the infotainment SoC 1530 and the instrument cluster 1532. In at least one embodiment, the instrument cluster 1532 may be included as part of the infotainment SoC 1530, and vice versa.

[0235] The inference and / or training logic 1215 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 1215 are provided herein in conjunction with Figures 12A and / or 12B. In at least one embodiment, the inference and / or training logic 1215 may be used in the system of Figure 15C for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0236] Figure 15D is a diagram of a system for communication between one or more cloud-based servers and the autonomous vehicle 1500 of Figure 15A, according to at least one embodiment. In at least one embodiment, the system may include, but is not limited to, one or more servers 1578, one or more networks 1590, and any number and types of vehicles, including the vehicle 1500. In at least one embodiment, the one or more servers 1578 may include, but is not limited to, a plurality of GPUs 1584(A) to 1584(H) (collectively referred to herein as GPU 1584), PCIe switches 1582(A) to 1582(D) (collectively referred to herein as PCIe switches 1582), and / or CPUs 1580(A) to 1580(B) (collectively referred to herein as CPU 1580). In at least one embodiment, the GPU 1584, the CPU 1580, and the PCIe switch 1582 may be interconnected by a high-speed interconnect, such as, for example, an NVLink interface 1588 and / or PCIe connection 1586 developed by NVIDIA. In at least one embodiment, the GPU 1584 is connected via an NVLink and / or NVSwitch SoC, and the GPU 1584 and PCIe switch 1582 are connected via a PCIe interconnect. Eight GPUs 1584, two CPUs 1580, and four PCIe switches 1582 are shown, but this is not limited to them. In at least one embodiment, each of (one or more) servers 1578 may include, but not limited to, any number of GPUs 1584, CPUs 1580, and / or PCIe switches 1582 in any combination. For example, in at least one embodiment, one or more servers 1578 may each include eight, sixteen, thirty-two, and / or more GPUs 1584.

[0237] In at least one embodiment, one or more servers 1578 may receive from a vehicle via one or more networks 1590 image data representing images showing unexpected or altered road conditions, such as recently started road construction. In at least one embodiment, one or more servers 1578 may transmit to a vehicle via one or more networks 1590 map information 1594, which includes updated or unupdated neural networks 1592 and / or, but not limited to, information about traffic and road conditions. In at least one embodiment, updates to the map information 1594 may include updates to the HD map 1522, which includes, but not limited to, information about construction sites, potholes, detours, floods, and / or other obstacles. In at least one embodiment, the neural network 1592 and / or map information 1594 may arise from new training and / or experience represented in data received from any number of vehicles in the environment, and / or from training performed in a data center (for example, using one or more servers 1578 and / or other servers).

[0238] In at least one embodiment, one or more servers 1578 may be used to train a machine learning model (e.g., a neural network) based at least in part on training data. In at least one embodiment, the training data may be generated by the vehicle and / or generated in a simulation (e.g., using a game engine). In at least one embodiment, any amount of training data may be tagged and / or otherwise preprocessed (e.g., if the relevant neural network benefits from supervised learning). In at least one embodiment, any amount of training data may not be tagged and / or preprocessed (e.g., if the relevant neural network does not require supervised learning). In at least one embodiment, once the machine learning model is trained, the machine learning model may be used by the vehicle (e.g., transmitted to the vehicle via one or more networks 1590) and / or the machine learning model may be used by one or more servers 1578 to remotely monitor the vehicle.

[0239] In at least one embodiment, one or more servers 1578 may receive data from a vehicle and apply the data to a state-of-the-art real-time neural network for real-time intelligent inference. In at least one embodiment, one or more servers 1578 may include one or more deep learning supercomputers and / or dedicated AI computers powered by GPUs 1584, such as DGX and DGX Station Machines developed by NVIDIA. However, in at least one embodiment, one or more servers 1578 may include a deep learning infrastructure using a CPU-powered data center.

[0240] In at least one embodiment, the deep learning infrastructure of one or more servers 1578 may be capable of high-speed real-time inference and may use that capability to assess and verify the health of the processor, software, and / or associated hardware in the vehicle 1500. For example, in at least one embodiment, the deep learning infrastructure may receive periodic updates from the vehicle 1500, such as a series of images and / or objects located in that series of images (e.g., via computer vision and / or other machine learning object classification techniques). In at least one embodiment, the deep learning infrastructure may operate its own neural network to identify objects and compare them to objects identified by the vehicle 1500. If the results do not match and the deep learning infrastructure concludes that the AI ​​in the vehicle 1500 is malfunctioning, the one or more servers 1578 may send a signal to the vehicle 1500 instructing the vehicle 1500's fail-safe computer to take control, notify passengers, and complete a safe parking operation.

[0241] In at least one embodiment, a server 1578 may include a GPU 1584 and a programmable inference accelerator (e.g., an NVIDIA TensorRT3 device). In at least one embodiment, a combination of a GPU-driven server and inference acceleration may enable real-time response. In at least one embodiment, a server driven by a CPU, FPGA, and other processors may be used for inference, for example, when performance is not critical. In at least one embodiment, a hardware structure 1215 is used to carry out one or more embodiments. Details relating to the hardware structure (x) 1215 are provided herein in conjunction with Figures 12A and / or 12B.

[0242] In at least one embodiment, one or more systems shown in Figures 15A to 15D are used to implement one or more implicit environment functions. In at least one embodiment, one or more systems shown in Figures 15A to 15D are used to use one or more neural networks, such as one or more implicit environment functions, to compute multiple paths that an entity, such as an autonomous device, is to traverse. In at least one embodiment, one or more systems shown in Figures 15A to 15D are used to implement one or more systems and / or processes, such as those described with respect to Figures 1 to 11.

[0243] Computer system Figure 16 is a block diagram showing an exemplary computer system, which may be a system having interconnected devices and components, a system-on-a-chip (SOC), or any combination thereof, formed together with a processor that may include an execution unit for executing instructions, according to at least one embodiment. In at least one embodiment, computer system 1600 may include components such as processor 1602 for employing an execution unit including logic for implementing algorithms for process data, as described herein, but not limited to embodiments described herein. In at least one embodiment, computer system 1600 may include a processor such as the PENTIUM® processor family, Xeon®, Itanium®, XScale®, and / or StrongARM®, Intel® Core®, or Intel® Nervana® microprocessors, available from Intel Corporation in Santa Clara, California, but other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.) may also be used. In at least one embodiment, the computer system 1600 may run a version of the WINDOWS® operating system available from Microsoft Corporation in Redmond, Washington, but other operating systems (e.g., UNIX® and Linux®), embedded software, and / or graphical user interfaces may also be used.

[0244] The embodiments may be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor ("DSP"), a system-on-a-chip, a network computer ("NetPC"), a set-top box, a network hub, a wide area network ("WAN") switch, or any other system capable of implementing one or more instructions according to at least one embodiment.

[0245] In at least one embodiment, the computer system 1600 may include, but is not limited to, a processor 1602, which may include, but is not limited to, one or more execution units 1608 for performing machine learning model training and / or inference by the techniques described herein. In at least one embodiment, the computer system 1600 is a single-processor desktop or server system, but in another embodiment, the computer system 1600 may be a multiprocessor system. In at least one embodiment, the processor 1602 may include, but is not limited to, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor implementing a combination of instruction sets, or any other processor device such as a digital signal processor. In at least one embodiment, the processor 1602 may be coupled to a processor bus 1610, and the processor bus 1610 may transmit data signals between the processor 1602 and other components in the computer system 1600.

[0246] In at least one embodiment, the processor 1602 may include, but is not limited to, a level 1 ("L1") internal cache memory ("cache") 1604. In at least one embodiment, the processor 1602 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory may reside outside the processor 1602. Other embodiments may also include a combination of both internal and external caches, depending on the specific implementation and requirements. In at least one embodiment, the register file 1606 may store different types of data in various registers, including, but is not limited to, integer registers, floating-point registers, status registers, and instruction pointer registers.

[0247] In at least one embodiment, without limitation, an execution unit 1608 including logic for performing integer operations and floating-point operations also exists in the processor 1602. In at least one embodiment, the processor 1602 may also include a microcode (u-code) read-only memory (ROM) that stores microcode for several macro instructions. In at least one embodiment, the execution unit 1608 may include logic for handling a packed instruction set 1609. In at least one embodiment, by including the packed instruction set 1609 in the instruction set of a general-purpose processor together with associated circuit elements for executing instructions, operations used by many multimedia applications can be performed using packed data in the processor 1602. In at least one embodiment, many multimedia applications can be accelerated and executed more efficiently by using the full width of the processor's data bus to perform operations on packed data, which can eliminate the need to transfer smaller units of data across the processor's data bus to perform one or more operations one data element at a time.

[0248] In at least one embodiment, the execution unit 1608 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, the computer system 1600 may include, without limitation, a memory 1620. In at least one embodiment, the memory 1620 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, or other memory device. In at least one embodiment, the memory 1620 may store one or more instructions 1619 and / or data 1621 represented by data signals executable by the processor 1602.

[0249] In at least one embodiment, a system logic chip may be coupled to a processor bus 1610 and memory 1620. In at least one embodiment, the system logic chip may include, but is not limited to, a memory controller hub ("MCH") 1616, and the processor 1602 may communicate with the MCH 1616 via the processor bus 1610. In at least one embodiment, the MCH 1616 may provide a high-bandwidth memory path 1618 to memory 1620 for instruction and data storage, as well as for the storage of graphics commands, data, and textures. In at least one embodiment, the MCH 1616 may direct data signals between the processor 1602, memory 1620, and other components in the computer system 1600, and bridge data signals between the processor bus 1610, memory 1620, and system I / O interface 1622. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH1616 may be coupled to memory 1620 through a high-bandwidth memory path 1618, and a graphics / video card 1612 may be coupled to the MCH1616 via an Accelerated Graphics Port ("AGP") interconnect 1614.

[0250] In at least one embodiment, computer system 1600 may use system I / O interface 1622 as a proprietary hub interface bus for coupling MCH 1616 to an I / O controller hub (ICH) 1630. In at least one embodiment, ICH 1630 may provide direct connections to several I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripheral devices to memory 1620, the chipset, and processor 1602. Examples may include, without limitation, an audio controller 1629, a firmware hub ("Flash BIOS") 1628, a wireless transceiver 1626, data storage 1624, a legacy I / O controller 1623 that includes user input and a keyboard interface 1625, a serial expansion port such as a Universal Serial Bus ("USB") port 1627, and a network controller 1634. In at least one embodiment, data storage 1624 may comprise a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0251] In at least one embodiment, FIG. 16 shows a system that includes interconnected hardware devices or "chips", but in other embodiments, FIG. 16 may show an exemplary SoC. In at least one embodiment, the devices shown in FIG. 16 may be interconnected by a proprietary interconnect, a standard interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 1600 are interconnected using a Compute Express Link (CXL) interconnect.

[0252] The inference and / or training logic 1215 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 1215 are provided herein in conjunction with Figures 12A and / or 12B. In at least one embodiment, the inference and / or training logic 1215 may be used in the system of Figure 16 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0253] In at least one embodiment, one or more systems shown in Figure 16 are used to implement one or more implicit environment functions. In at least one embodiment, one or more systems shown in Figure 16 are used to use one or more neural networks, such as one or more implicit environment functions, to compute multiple paths that an entity, such as an autonomous device, is to traverse. In at least one embodiment, one or more systems shown in Figure 16 are used to implement one or more systems and / or processes, such as those described with respect to Figures 1 to 11.

[0254] Figure 17 is a block diagram showing an electronic device 1700 for utilizing the processor 1710 according to at least one embodiment. In at least one embodiment, the electronic device 1700 may be, for example, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop, a tablet, a mobile device, a telephone, an embedded computer, or any other suitable electronic device.

[0255] In at least one embodiment, the electronic device 1700 may include a processor 1710 communicably coupled to any preferred number or type of components, peripherals, modules, or devices, but is not limited to. In at least one embodiment, the processor 1710 is I 2The devices are coupled using buses or interfaces such as the C-bus, System Management Bus ("SMBus"), Low Pin Count (LPC) bus, Serial Peripheral Interface ("SPI"), High Definition Audio ("HDA") bus, Serial Advance Technology Attachment ("SATA") bus, Universal Serial Bus ("USB") (versions 1, 2, 3, etc.), or Universal Asynchronous Receiver / Transmitter ("UART") bus. In at least one embodiment, Figure 17 shows a system including interconnected hardware devices or "chips," while in other embodiments, Figure 17 may show an exemplary SoC. In at least one embodiment, the devices shown in Figure 17 may be interconnected by proprietary interconnects, standard interconnects (e.g., PCIe), or any combination thereof. In at least one embodiment, one or more components of Figure 17 are interconnected using a Compute Express Link (CXL) interconnect.

[0256] In at least one embodiment, Figure 17 shows a display 1724, a touchscreen 1725, a touchpad 1730, a Near Field Communication ("NFC") unit 1745, a sensor hub 1740, a thermal sensor 1746, an Express Chipset ("EC") 1735, a Trusted Platform Module ("TPM") 1738, a BIOS / firmware / flash memory ("BIOS,FW flash") 1722, a DSP 1760, a drive 1720 such as a Solid State Disk ("SSD") or Hard Disk Drive ("HDD"), a Wireless Local Area Network ("WLAN") unit 1750, a Bluetooth unit 1752, and a Wireless Wide Area Network ("WWAN") unit. The components may include a network (1756), a Global Positioning System (GPS) unit (1755), a camera such as a USB 3.0 camera ("USB 3.0 camera") (1754), and / or a Low Power Double Data Rate ("LPDDR") memory unit ("LPDDR3") (1715) implemented, for example, in the LPDDR3 standard. Each of these components may be implemented in any preferred manner.

[0257] In at least one embodiment, other components may be communicatively coupled to the processor 1710 through components described herein. In at least one embodiment, an accelerometer 1741, an ambient light sensor ("ALS") 1742, a compass 1743, and a gyroscope 1744 may be communicatively coupled to a sensor hub 1740. In at least one embodiment, a thermal sensor 1739, a fan 1737, a keyboard 1736, and a touchpad 1730 may be communicatively coupled to the EC 1735. In at least one embodiment, a speaker 1763, headphones 1764, and a microphone ("mic") 1765 may be communicatively coupled to an audio unit ("audio codec and Class D amplifier") 1762, and the audio unit 1762 may be communicatively coupled to a DSP 1760. In at least one embodiment, the audio unit 1762 may include, for example, an audio coder / decoder ("codec") and a Class D amplifier. In at least one embodiment, a SIM card ("SIM") 1757 may be communicatively coupled to the WWAN unit 1756. In at least one embodiment, components such as the WLAN unit 1750 and the Bluetooth unit 1752, as well as the WWAN unit 1756, may be implemented in a Next Generation Form Factor ("NGFF").

[0258] The inference and / or training logic 1215 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 1215 are provided herein in conjunction with Figures 12A and / or 12B. In at least one embodiment, the inference and / or training logic 1215 may be used in the system of Figure 17 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0259] In at least one embodiment, one or more systems shown in Figure 17 are used to implement one or more implicit environment functions. In at least one embodiment, one or more systems shown in Figure 17 are used to use one or more neural networks, such as one or more implicit environment functions, to compute multiple paths that an entity, such as an autonomous device, is to traverse. In at least one embodiment, one or more systems shown in Figure 17 are used to implement one or more systems and / or processes, such as those described with respect to Figures 1 to 11.

[0260] Figure 18 shows a computer system 1800 according to at least one embodiment. In at least one embodiment, the computer system 1800 is configured to implement various processes and methods described throughout this disclosure.

[0261] In at least one embodiment, the computer system 1800 includes, but is not limited to, at least one central processing unit ("CPU") 1802, the at least one central processing unit 1802 being connected to a communications bus 1810 implemented using any preferred protocol, such as PCI ("Peripheral Component Interconnect"), Peripheral Component Interconnect Express ("PCI-Express"), AGP ("Accelerated Graphics Port"), Hypertransport, or any other bus or (one or more) point-to-point communication protocols. In at least one embodiment, the computer system 1800 includes, but is not limited to, main memory 1804 and control logic (implemented, for example, as hardware, software, or a combination thereof), the data being stored in the main memory 1804, which may take the form of random-access memory ("RAM"). In at least one embodiment, the network interface subsystem ("Network Interface") 1822 provides an interface to other computing devices and networks for receiving data from other systems and transmitting data to other systems, together with the computer system 1800.

[0262] In at least one embodiment, the computer system 1800 includes, in at least one embodiment, an input device 1808, a parallel processing system 1812, and a display device 1806, the display device 1806 of which may be implemented using a conventional cathode ray tube ("CRT"), a liquid crystal display ("LCD"), a light-emitting diode ("LED") display, a plasma display, or other suitable display technology. In at least one embodiment, user input is received from the input device 1808, such as a keyboard, mouse, touchpad, or microphone. In at least one embodiment, each module described herein may be on a single semiconductor platform to form a processing system.

[0263] The inference and / or training logic 1215 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 1215 are provided herein in conjunction with Figures 12A and / or 12B. In at least one embodiment, the inference and / or training logic 1215 may be used in the system of Figure 18 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0264] In at least one embodiment, one or more systems shown in Figure 18 are used to implement one or more implicit environment functions. In at least one embodiment, one or more systems shown in Figure 18 are used to use one or more neural networks, such as one or more implicit environment functions, to compute multiple paths that an entity, such as an autonomous device, is to traverse. In at least one embodiment, one or more systems shown in Figure 18 are used to implement one or more systems and / or processes, such as those described with respect to Figures 1 to 11.

[0265] Figure 19 shows a computer system 1900 according to at least one embodiment. In at least one embodiment, the computer system 1900 may include, but is not limited to, a computer 1910 and a USB stick 1920. In at least one embodiment, the computer 1910 may include, but is not limited to, any number and type of processors (not shown) and memory (not shown). In at least one embodiment, the computer 1910 may include, but is not limited to, a server, a cloud instance, a laptop, and a desktop computer.

[0266] In at least one embodiment, the USB stick 1920 includes, but is not limited to, a processing unit 1930, a USB interface 1940, and a USB interface logic 1950. In at least one embodiment, the processing unit 1930 may be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, the processing unit 1930 may include, but is not limited to, any number and type of processing cores (not shown). In at least one embodiment, the processing unit 1930 comprises an application-specific integrated circuit ("ASIC") optimized to perform any amount and type of operations related to machine learning. For example, in at least one embodiment, the processing unit 1930 is a tensor processing unit ("TPC") optimized to perform machine learning inference operations. In at least one embodiment, the processing unit 1930 is a vision processing unit ("VPU") optimized to perform machine vision and machine learning inference operations.

[0267] In at least one embodiment, the USB interface 1940 may be any type of USB connector or USB socket. For example, in at least one embodiment, the USB interface 1940 is a USB 3.0 Type-C socket for data and power. In at least one embodiment, the USB interface 1940 is a USB 3.0 Type-A connector. In at least one embodiment, the USB interface logic 1950 may include any amount and type of logic that enables the processing unit 1930 to interface with a device (e.g., a computer 1910) via the USB connector 1940.

[0268] The inference and / or training logic 1215 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 1215 are provided herein in conjunction with Figures 12A and / or 12B. In at least one embodiment, the inference and / or training logic 1215 may be used in the system of Figure 19 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0269] In at least one embodiment, one or more systems shown in Figure 19 are used to implement one or more implicit environment functions. In at least one embodiment, one or more systems shown in Figure 19 are used to use one or more neural networks, such as one or more implicit environment functions, to compute multiple paths that an entity, such as an autonomous device, is to traverse. In at least one embodiment, one or more systems shown in Figure 19 are used to implement one or more systems and / or processes, such as those described with respect to Figures 1 to 11.

[0270] Figure 20A shows an exemplary architecture in which multiple GPUs 2010(1)–2010(N) are communicably coupled to multiple multicore processors 2005(1)–2005(M) via high-speed links 2040(1)–2040(N) (e.g., bus, point-to-point interconnect). In at least one embodiment, the high-speed links 2040(1)–2040(N) support communication throughput of 4 GB / s, 30 GB / s, 80 GB / s, or higher. In at least one embodiment, but not limited to, various interconnect protocols may be used, including PCIe 4.0 or 5.0 and NVLink 2.0. In various figures, "N" and "M" represent positive integers whose values ​​may differ from figure to figure.

[0271] Furthermore, in at least one embodiment, two or more of the GPUs 2010 are interconnected via high-speed links 2029(1) to 2029(2), and the high-speed links 2029(1) to 2029(2) can be implemented using the same or different protocols / links as those used for the high-speed links 2040(1) to 2040(N). Similarly, two or more of the multi-core processors 2005 can be connected via the high-speed link 2028, and the high-speed link 2028 can be a symmetric multi-processor (SMP) bus operating at 20 GB / second, 30 GB / second, 120 GB / second, or more. Alternatively, all communications between the various system components shown in FIG. 20A can be realized using the same protocol / link (e.g., via a common interconnect fabric).

[0272] In at least one embodiment, each multi-core processor 2005 is communicatively coupled to the processor memories 2001(1) to 2001(M) via the memory interconnects 2026(1) to 2026(M), respectively, and each GPU 2010(1) to 2010(N) is communicatively coupled to the GPU memories 2020(1) to 2020(N) via the GPU memory interconnects 2050(1) to 2050(N), respectively. In at least one embodiment, the memory interconnects 2026 and 2050 can utilize the same or different memory access technologies. By way of example, and without limitation, the processor memories 2001(1) to 2001(M) and the GPU memory 2020 can be volatile memories such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or non-volatile memories such as 3D XPoint or Nano-Ram. In at least one embodiment, a portion of the processor memory 2001 can be a volatile memory, and another portion can be a non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).

[0273] As described herein, various multicore processors 2005 and GPUs 2010 may be physically coupled to specific memories 2001, 2020, respectively, and / or a unified memory architecture may be implemented in which a virtual system address space (also called the “effective address” space) is distributed across various physical memories. For example, processor memories 2001(1) to 2001(M) may each have a system memory address space of 64GB, and GPU memories 2020(1) to 2020(N) may each have a system memory address space of 32GB, resulting in a total of 256GB of addressable memory when M=2 and N=4. Other values ​​for N and M are possible.

[0274] Figure 20B provides additional details about the interconnection between a multicore processor 2007 and a graphics acceleration module 2046 in one exemplary embodiment. In at least one embodiment, the graphics acceleration module 2046 may include one or more GPU chips integrated on a line card coupled to the processor 2007 via a high-speed link 2040 (e.g., PCIe bus, NVLink, etc.). In at least one embodiment, the graphics acceleration module 2046 may alternatively be integrated on a package or chip with the processor 2007.

[0275] In at least one embodiment, the processor 2007 comprises a plurality of cores 2060A–2060D, each having translation lookaside buffers ("TLB") 2061A–2061D and one or more caches 2062A–2062D. In at least one embodiment, the cores 2060A–2060D may include various other components not shown for executing instructions and processing data. In at least one embodiment, the caches 2062A–2062D may comprise a Level 1 (L1) cache and a Level 2 (L2) cache. Furthermore, one or more shared caches 2056 may be included in the caches 2062A–2062D and shared by the set of cores 2060A–2060D. For example, one embodiment of processor 2007 includes 24 cores, each having its own L1 cache, 12 shared L2 caches, and 12 shared L3 caches. In this embodiment, one or more L2 and L3 caches are shared by two adjacent cores. In at least one embodiment, processor 2007 and graphics acceleration module 2046 are connected to system memory 2014, which may include processor memories 2001(1) to 2001(M) in Figure 20A.

[0276] In at least one embodiment, coherence is maintained over intercore communication on the coherence bus 2064 for data and instructions stored in various caches 2062A-2062D, 2056 and system memory 2014. In at least one embodiment, for example, each cache may have associated cache coherence logic / circuit elements to communicate over the coherence bus 2064 in response to detected reads or writes to a particular cache line. In at least one embodiment, a cache snooping protocol is implemented over the coherence bus 2064 to snoop cache accesses.

[0277] In at least one embodiment, proxy circuit 2025 connects graphics acceleration module 2046 to coherence bus 2064 in a communicable manner, enabling graphics acceleration module 2046 to participate in the cache coherence protocol as a peer of cores 2060A-2060D. Specifically, in at least one embodiment, interface 2035 provides connectivity to proxy circuit 2025 via high-speed link 2040, and interface 2037 connects graphics acceleration module 2046 to high-speed link 2040.

[0278] In at least one embodiment, the accelerator integration circuit 2036 provides cache management, memory access, context management, and interrupt management services on behalf of the multiple graphics processing engines 2031(1) to 2031(N) of the graphics acceleration module 2046. In at least one embodiment, each of the graphics processing engines 2031(1) to 2031(N) may comprise a separate graphics processing unit (GPU). In at least one embodiment, the graphics processing engines 2031(1) to 2031(N) may alternatively comprise different types of graphics processing engines within the GPU, such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a blit engine. In at least one embodiment, the graphics acceleration module 2046 may be a GPU having multiple graphics processing engines 2031(1) to 2031(N), or the graphics processing engines 2031(1) to 2031(N) may be individual GPUs integrated on a common package, line card, or chip.

[0279] In at least one embodiment, the accelerator integration circuit 2036 includes a memory management unit (MMU) 2039 for performing various memory management functions, such as virtual-physical memory translation (also known as effective-to-real memory translation), and a memory access protocol for accessing system memory 2014. In at least one embodiment, the MMU 2039 may also include a translation lookaside buffer (TLB) (not shown) for caching virtual / effective-physical / real address translations. In at least one embodiment, the cache 2038 can store commands and data for efficient access by graphics processing engines 2031(1) to 2031(N). In at least one embodiment, data stored in cache 2038 and graphics memory 2033(1)-2033(M) is kept coherent with core caches 2062A-2062D, 2056 and system memory 2014, possibly using fetch unit 2044. As stated, this can be achieved via proxy circuit 2025 instead of cache 2038 and memory 2033(1)-2033(M) (for example, by sending updates related to modifying / accessing cache lines on processor caches 2062A-2062D, 2056 to cache 2038 and receiving updates from cache 2038).

[0280] In at least one embodiment, a set of registers 2045 stores context data for threads executed by graphics processing engines 2031(1) to 2031(N), and a context management circuit 2048 manages thread contexts. For example, the context management circuit 2048 may perform save and restore operations to save and restore the contexts of various threads during context switching (for example, the first thread is saved and the second thread is stored so that a second thread can be executed by the graphics processing engine). For example, during context switching, the context management circuit 2048 may store the current register values ​​in a specified area of ​​memory (for example, identified by a context pointer). The context management circuit 2048 may then restore the register values ​​when returning to the context. In at least one embodiment, an interrupt management circuit 2047 receives and processes interrupts received from system devices.

[0281] In at least one embodiment, virtual / effective addresses from the graphics processing engine 2031 are translated by the MMU 2039 to real / physical addresses in system memory 2014. In at least one embodiment, the accelerator integration circuit 2036 supports multiple (e.g., 4, 8, or 16) graphics accelerator modules 2046 and / or other accelerator devices. In at least one embodiment, the graphics accelerator module 2046 may be dedicated to a single application running on the processor 2007 or may be shared among multiple applications. In at least one embodiment, there exists a virtualized graphics execution environment in which the resources of the graphics processing engines 2031(1) to 2031(N) are shared among multiple applications or virtual machines (VMs). In at least one embodiment, the resources may be subdivided into "slices," which are allocated to different VMs and / or applications based on processing requirements and priorities related to the VMs and / or applications.

[0282] In at least one embodiment, the accelerator integration circuit 2036 acts as a bridge to the system for the graphics acceleration module 2046 and provides address translation and system memory cache services. Furthermore, in at least one embodiment, the accelerator integration circuit 2036 may provide virtualization facilities for the host processor to manage the virtualization, interrupts, and memory management of the graphics processing engines 2031(1) to 2031(N).

[0283] In at least one embodiment, the hardware resources of the graphics processing engines 2031(1)–2031(N) are explicitly mapped to the real address space visible to the host processor 2007, so that any host processor can directly address these resources using effective address values. In at least one embodiment, one function of the accelerator integration circuit 2036 is to physically isolate the graphics processing engines 2031(1)–2031(N) so that they appear as independent units to the system.

[0284] In at least one embodiment, one or more graphics memories 2033(1) to 2033(M) are each coupled to each of the graphics processing engines 2031(1) to 2031(N), where N=M. In at least one embodiment, the graphics memories 2033(1) to 2033(M) store instructions and data being processed by each of the graphics processing engines 2031(1) to 2031(N). In at least one embodiment, the graphics memories 2033(1) to 2033(M) may be volatile memory such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or non-volatile memory such as 3D XPoint or Nano-Ram.

[0285] In at least one embodiment, biasing techniques may be used to reduce data traffic over the high-speed link 2040, ensuring that the data stored in graphics memory 2033(1)-2033(M) is most frequently used by the graphics processing engines 2031(1)-2031(N) and preferably not used (or at least not frequently used) by cores 2060A-2060D. Similarly, in at least one embodiment, the biasing mechanism attempts to keep data required by the cores (preferably not required by the graphics processing engines 2031(1)-2031(N)) in caches 2062A-2062D, 2056 and system memory 2014.

[0286] Figure 20C shows another exemplary embodiment in which the accelerator integration circuit 2036 is integrated into the processor 2007. In this embodiment, the graphics processing engines 2031(1) to 2031(N) communicate directly over the high-speed link 2040 to the accelerator integration circuit 2036 via interfaces 2037 and 2035 (which, in this case, may be any form of bus or interface protocol). In at least one embodiment, the accelerator integration circuit 2036 may perform operations similar to those described with respect to Figure 20B, but potentially at higher throughput given its proximity to the coherence bus 2064 and caches 2062A to 2062D, 2056. In at least one embodiment, the accelerator integration circuit supports different programming models, including a dedicated process programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization), and these programming models may include a programming model controlled by the accelerator integration circuit 2036 and a programming model controlled by the graphics acceleration module 2046.

[0287] In at least one embodiment, the graphics processing engines 2031(1) to 2031(N) may be dedicated to a single application or process under a single operating system. In at least one embodiment, a single application may direct other application requests to the graphics processing engines 2031(1) to 2031(N) to provide virtualization within a VM / partition.

[0288] In at least one embodiment, the graphics processing engines 2031(1) to 2031(N) may be shared by multiple VM / application partitions. In at least one embodiment, the sharing model may use a system hypervisor to virtualize the graphics processing engines 2031(1) to 2031(N) to allow access by each operating system. In at least one embodiment, for a single-partition system without a hypervisor, the graphics processing engines 2031(1) to 2031(N) are owned by the operating system. In at least one embodiment, the operating system may virtualize the graphics processing engines 2031(1) to 2031(N) to provide access to each process or application.

[0289] In at least one embodiment, the graphics acceleration module 2046 or the individual graphics processing engines 2031(1) to 2031(N) select a process element using a process handle. In at least one embodiment, the process element is stored in system memory 2014 and is addressable using the effective address-actual address translation technique described herein. In at least one embodiment, the process handle may be an implementation-specific value provided to the host process when registering the host process context with the graphics processing engines 2031(1) to 2031(N) (i.e., calling system software to add a process element to the process element link list). In at least one embodiment, the lower 16 bits of the process handle may be the offset of the process element in the process element link list.

[0290] Figure 20D shows an exemplary accelerator integration slice 2090. In at least one embodiment, the “slice” comprises a designated portion of the processing resources of the accelerator integration circuit 2036. In at least one embodiment, the application’s effective address space 2082 in system memory 2014 stores a process element 2083. In at least one embodiment, the process element 2083 is stored in response to a GPU call 2081 from an application 2080 running on processor 2007. In at least one embodiment, the process element 2083 contains the process state of the corresponding application 2080. In at least one embodiment, the work descriptor (WD) 2084 contained in the process element 2083 may be a single job requested by the application or may contain a pointer to a queue of jobs. In at least one embodiment, the WD 2084 is a pointer to a job request queue in the application’s effective address space 2082.

[0291] In at least one embodiment, the graphics acceleration module 2046 and / or individual graphics processing engines 2031(1) to 2031(N) may be shared by all or a subset of processes in the system. In at least one embodiment, infrastructure may be included for setting process states and sending WD2084 to the graphics acceleration module 2046 to start jobs in a virtualized environment.

[0292] In at least one embodiment, the dedicated process programming model is implementation-specific. In at least one embodiment, in this model, a single process owns either the graphics acceleration module 2046 or an individual graphics processing engine 2031. In at least one embodiment, when the graphics acceleration module 2046 is owned by a single process, the hypervisor initializes the accelerator integration circuit 2036 for the owning partition, and when the graphics acceleration module 2046 is assigned, the operating system initializes the accelerator integration circuit 2036 for the owning process.

[0293] In at least one embodiment, during operation, the WD fetch unit 2091 in the accelerator integrated slice 2090 fetches the next WD2084, which contains instructions for work to be performed by one or more graphics processing engines of the graphics acceleration module 2046. In at least one embodiment, as shown, the data from WD2084 is stored in register 2045 and may be used by the MMU 2039, interrupt management circuit 2047, and / or context management circuit 2048. For example, one embodiment of the MMU 2039 includes a segment / page walk circuit element for accessing the segment / page table 2086 in the OS virtual address space 2085. In at least one embodiment, the interrupt management circuit 2047 may process an interrupt event 2092 received from the graphics acceleration module 2046. In at least one embodiment, when performing graphics operations, the effective address 2093 generated by the graphics processing engines 2031(1) to 2031(N) is translated to a real address by the MMU 2039.

[0294] In at least one embodiment, register 2045 may be duplicated for each graphics processing engine 2031(1)–2031(N) and / or graphics acceleration module 2046 and initialized by the hypervisor or operating system. In at least one embodiment, each of these duplicated registers may be included in the accelerator integration slice 2090. Exemplary registers that may be initialized by the hypervisor are shown in Table 1. [Table 1]

[0295] Table 2 shows exemplary registers that can be initialized by the operating system. [Table 2]

[0296] In at least one embodiment, each WD2084 is specific to a particular graphics acceleration module 2046 and / or graphics processing engine 2031(1)-2031(N). In at least one embodiment, the WD2084 may contain all the information required by the graphics processing engine 2031(1)-2031(N) to perform the work, or it may be a pointer to a memory location set up by the application for a command queue of work to be completed.

[0297] Figure 20E provides additional details of one exemplary embodiment of the shared model. This embodiment includes a hypervisor real address space 2098 where the process element list 2099 is stored. In at least one embodiment, the hypervisor real address space 2098 is accessible via a hypervisor 2096 that virtualizes the graphics acceleration module engine for the operating system 2095.

[0298] In at least one embodiment, a shared programming model allows all or a subset of processes from all or a subset of partitions in the system to use the graphics acceleration module 2046. In at least one embodiment, there are two programming models in which the graphics acceleration module 2046 is shared by multiple processes and partitions: time-slice sharing and graphics-directed sharing.

[0299] In at least one embodiment, in this model, the system hypervisor 2096 owns the graphics acceleration module 2046 and makes its functionality available to all operating systems 2095. In at least one embodiment, for the graphics acceleration module 2046 to support virtualization by the system hypervisor 2096, the graphics acceleration module 2046 may be subject to several requirements, including: (1) application job requests must be autonomous (i.e., no state needs to be maintained between jobs), or the graphics acceleration module 2046 must provide a context saving and restoration mechanism; (2) the graphics acceleration module 2046 must guarantee that application job requests will be completed within a specified amount of time, including any translation failures, or the graphics acceleration module 2046 must provide the ability to preempt job processing; and (3) when the graphics acceleration module 2046 is operating in a specified shared programming model, fairness between processes must be guaranteed.

[0300] In at least one embodiment, application 2080 is required to make a system call to operating system 2095, accompanied by a graphics acceleration module type, a work descriptor (WD), an authority mask register (AMR) value, and a context save / restore area pointer (CSRP). In at least one embodiment, the graphics acceleration module type describes the acceleration function of the system call. In at least one embodiment, the graphics acceleration module type may be a system-specific value. In at least one embodiment, the WD may be specifically formatted for graphics acceleration module 2046 and may be in the form of a command for graphics acceleration module 2046, an effective address pointer to a user-defined structure, an effective address pointer to a queue of commands, or any other data structure for describing the work to be performed by graphics acceleration module 2046.

[0301] In at least one embodiment, the AMR value is the AMR state to be used for the current process. In at least one embodiment, the value passed to the operating system is the same as that of the application setting the AMR. In at least one embodiment, if the accelerator integration circuit 2036 (not shown) implementation and the graphics acceleration module 2046 implementation do not support the User Authority Mask Override Register (UAMOR), the operating system may apply the current UAMOR value to the AMR value before passing the AMR in a hypervisor call. In at least one embodiment, the hypervisor 2096 may optionally apply the current Authority Mask Override Register (AMOR) value before putting the AMR into process element 2083. In at least one embodiment, CSRP is one of the registers 2045 that contains the effective addresses of areas in the application's effective address space 2082 for the graphics acceleration module 2046 to save and restore context state. In at least one embodiment, this pointer is optional when no state needs to be saved between jobs or when a job is preempted. In at least one embodiment, the context save / restore area may be pinned system memory.

[0302] Upon receiving a system call, operating system 2095 may verify that application 2080 is registered and authorized to use graphics acceleration module 2046. In at least one embodiment, operating system 2095 then calls hypervisor 2096 with the information shown in Table 3. [Table 3]

[0303] In at least one embodiment, upon receiving a hypervisor call, hypervisor 2096 verifies that operating system 2095 is registered and authorized to use graphics acceleration module 2046. In at least one embodiment, hypervisor 2096 then places process element 2083 in the process element link list for the corresponding graphics acceleration module 2046 type. In at least one embodiment, the process element may include the information shown in Table 4. [Table 4]

[0304] In at least one embodiment, the hypervisor initializes multiple accelerator integration slice 2090 registers 2045.

[0305] As shown in Figure 20F, in at least one embodiment, a unified memory is used that is addressable via a common virtual memory address space used to access physical processor memories 2001(1) to 2001(N) and GPU memories 2020(1) to 2020(N). In this implementation, operations performed on GPUs 2010(1) to 2010(N) utilize the same virtual / effective memory address space to access processor memories 2001(1) to 2001(M), and vice versa, thereby simplifying programmability. In at least one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 2001(1), a second portion to a second processor memory 2001(N), a third portion to GPU memory 2020(1), and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes called the effective address space) is distributed across processor memory 2001 and GPU memory 2020, respectively, allowing either processor or GPU to access either physical memory, where virtual addresses are mapped to physical memory.

[0306] In at least one embodiment, bias / coherence management circuit elements 2094A-2094E within one or more of the MMUs 2039A-2039E ensure cache coherence between the cache of one or more host processors (e.g., 2005) and the cache of the GPU 2010, implement bias techniques, and indicate physical memory where certain types of data should be stored. In at least one embodiment, multiple instances of bias / coherence management circuit elements 2094A-2094E are shown in Figure 20F, although the bias / coherence circuit elements may be implemented within the MMU of one or more host processors 2005 and / or within the accelerator integration circuit 2036.

[0307] One embodiment allows GPU memory 2020 to be mapped as part of system memory and accessed using shared virtual memory (SVM) technology, but without incurring the performance drawbacks associated with full system cache coherence. In at least one embodiment, the ability of GPU memory 2020 to be accessed as system memory without cumbersome cache coherence overhead provides a beneficial operating environment for GPU offloading. In at least one embodiment, this mechanism allows software on the host processor 2005 to set operands and access computation results without the overhead of conventional I / O DMA data copying. In at least one embodiment, such conventional copying involves driver calls, interrupts, and memory-mapped I / O (MMIO) access, all of which are inefficient compared to simple memory access. In at least one embodiment, the ability to access GPU memory 2020 without cache coherence overhead may be essential for the execution time of offloaded computations. In at least one embodiment, for example, when there is significant streaming write memory traffic, cache coherence overhead can significantly reduce the effective write bandwidth seen by the GPU2010. In at least one embodiment, the efficiency of operand configuration, the efficiency of result access, and the efficiency of GPU computation can be helpful in determining the effectiveness of GPU offloading.

[0308] In at least one embodiment, the selection between GPU bias and host processor bias is driven by a bias tracker data structure. In at least one embodiment, for example, a bias table may be used, which may be a page-granular structure containing 1 or 2 bits per GPU-enabled memory page (for example, it may be controlled at the memory page granularity). In at least one embodiment, the bias table may be implemented in a stolen memory range of one or more GPU memory 2020, with or without a bias cache (for example, for caching frequently used / recently used entries in the bias table) in GPU 2010. Alternatively, in at least one embodiment, the entire bias table may be maintained within the GPU.

[0309] In at least one embodiment, a bias table entry associated with each access to GPU-biased memory 2020 is accessed before the actual access to GPU memory, causing the following behavior: In at least one embodiment, a local request from GPU 2010 finding its page in GPU bias is forwarded directly to the corresponding GPU memory 2020. In at least one embodiment, a local request from a GPU finding its page in host bias is forwarded to processor 2005 (for example, via the fast link described above). In at least one embodiment, a request from processor 2005 finding the requested page in host processor bias completes the request like a normal memory read. Alternatively, a request targeting a GPU-biased page may be forwarded to GPU 2010. In at least one embodiment, the GPU may then move the page to host processor bias if the page is not currently in use. In at least one embodiment, the bias state of a page can be modified by either a software-based mechanism, a hardware-assisted software-based mechanism, or, for a limited set of cases, simply by a hardware-based mechanism.

[0310] In at least one embodiment, one mechanism for changing the bias state employs an API call (e.g., OpenCL) which calls the GPU's device driver, which changes the bias state and, for some transitions, sends a message to the GPU (or queues a command descriptor) instructing the GPU to perform a cache-flushing operation on the host. In at least one embodiment, the cache-flushing operation is used for transitions from host processor 2005 bias to GPU bias, but not for transitions in the opposite direction.

[0311] In at least one embodiment, cache coherency is maintained by temporarily rendering GPU-biased pages that cannot be cached by the host processor 2005. In at least one embodiment, to access these pages, processor 2005 may request access from GPU 2010, which may or may not immediately grant access. In at least one embodiment, it is therefore beneficial to ensure that GPU-biased pages are those that are required by the GPU but not by the host processor 2005, and vice versa, in order to reduce communication between processor 2005 and GPU 2010.

[0312] One or more hardware structures 1215 are used to carry out one or more embodiments. Details relating to one or more hardware structures 1215 may be provided herein in conjunction with Figures 12A and / or 12B.

[0313] In at least one embodiment, one or more systems shown in Figures 20A to 20F are used to implement one or more implicit environment functions. In at least one embodiment, one or more systems shown in Figures 20A to 20F are used to use one or more neural networks, such as one or more implicit environment functions, to compute multiple paths that an entity, such as an autonomous device, is to traverse. In at least one embodiment, one or more systems shown in Figures 20A to 20F are used to implement one or more systems and / or processes, such as those described with respect to Figures 1 to 11.

[0314] Figure 21 shows exemplary integrated circuits and associated graphics processors that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to those shown, at least one embodiment may include other logic and circuitry, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.

[0315] Figure 21 is a block diagram showing an exemplary system-on-chip integrated circuit 2100 that may be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, the integrated circuit 2100 includes one or more application processors 2105 (e.g., CPUs), at least one graphics processor 2110, and additionally, an image processor 2115 and / or a video processor 2120, any of which may be modular IP cores. In at least one embodiment, the integrated circuit 2100 includes a USB controller 2125, a UART controller 2130, an SPI / SDIO controller 2135, and I 2 2S / I 2 The integrated circuit includes peripherals or bus logic, including a 2C controller 2140. In at least one embodiment, the integrated circuit 2100 may include a display device 2145 coupled to one or more of the following: a High-Definition Multimedia Interface (HDMI®) controller 2150 and a Mobile Industry Processor Interface (MIPI) display interface 2155. In at least one embodiment, storage may be provided by a flash memory subsystem 2160, which includes flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 2165 for access to SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits may additionally include an embedded security engine 2170.

[0316] The inference and / or training logic 1215 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 1215 are provided herein in conjunction with Figures 12A and / or 12B. In at least one embodiment, the inference and / or training logic 1215 may be used in the integrated circuit 2100 for inference or prediction operations based at least in part on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0317] In at least one embodiment, one or more systems shown in Figure 21 are used to implement one or more implicit environment functions. In at least one embodiment, one or more systems shown in Figure 21 are used to use one or more neural networks, such as one or more implicit environment functions, to compute multiple paths that an entity, such as an autonomous device, is to traverse. In at least one embodiment, one or more systems shown in Figure 21 are used to implement one or more systems and / or processes, such as those described with respect to Figures 1 to 11.

[0318] Figures 22A and 22B show exemplary integrated circuits and associated graphics processors that may be fabricated using one or more IP cores according to various embodiments described herein. In addition to those shown, at least one embodiment may include other logic and circuitry, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.

[0319] Figures 22A and 22B are block diagrams illustrating exemplary graphics processors for use in a SoC according to embodiments described herein. Figure 22A shows an exemplary graphics processor 2210 of a system-on-chip integrated circuit that may be fabricated using one or more IP cores, according to at least one embodiment. Figure 22B shows an additional exemplary graphics processor 2240 of a system-on-chip integrated circuit that may be fabricated using one or more IP cores, according to at least one embodiment. In at least one embodiment, the graphics processor 2210 of Figure 22A is a low-power graphics processor core. In at least one embodiment, the graphics processor 2240 of Figure 22B is a higher-performance graphics processor core. In at least one embodiment, each of the graphics processors 2210 and 2240 may be a variation of the graphics processor 2110 of Figure 21.

[0320] In at least one embodiment, the graphics processor 2210 includes a vertex processor 2205 and one or more fragment processors 2215A-2215N (e.g., 2215A, 2215B, 2215C, 2215D-2215N-1, and 2215N). In at least one embodiment, the graphics processor 2210 can execute different shader programs via separate logic, thereby optimizing the vertex processor 2205 to perform operations for a vertex shader program, and the one or more fragment processors 2215A-2215N to perform fragment (e.g., pixel) shading operations for a fragment or pixel shader program. In at least one embodiment, the vertex processor 2205 performs the vertex processing stage of the 3D graphics pipeline and generates primitive and vertex data. In at least one embodiment, one or more fragment processors 2215A-2215N use primitive and vertex data generated by the vertex processor 2205 to create a frame buffer that is displayed on a display device. In at least one embodiment, one or more fragment processors 2215A-2215N are optimized to execute fragment shader programs such as those provided in the OpenGL API, and the OpenGL API may be used to perform operations similar to those of pixel shader programs such as those provided in the Direct 3D API.

[0321] In at least one embodiment, the graphics processor 2210 additionally includes one or more memory management units (MMUs) 2220A-2220B, one or more caches 2225A-2225B, and one or more circuit interconnects 2230A-2230B. In at least one embodiment, one or more MMUs 2220A-2220B provide virtual-physical address mappings for the graphics processor 2210, including vertex processors 2205 and / or one or more fragment processors 2215A-2215N, which may reference vertex or image / texture data stored in memory, in addition to vertex or image / texture data stored in one or more caches 2225A-2225B. In at least one embodiment, one or more MMUs 2220A-2220B may be synchronized with other MMUs in the system, including one or more MMUs associated with one or more application processors 2105, image processors 2115, and / or video processors 2120 in Figure 21, thereby allowing each processor 2105-2120 to participate in a shared or unified virtual memory system. In at least one embodiment, one or more circuit interconnects 2230A-2230B enable the graphics processor 2210 to interface with other IP cores in the SoC, either via the SoC's internal bus or via a direct connection.

[0322] In at least one embodiment, as shown in Figure 22B, the graphics processor 2240 includes one or more shader cores 2255A-2255N (for example, 2255A, 2255B, 2255C, 2255D, 2255E, 2255F-2255N-1, and 2255N), and one or more shader cores 2255A-2255N provide a unified shader core architecture in which a single core, or type, or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 2240 includes an intercore task manager 2245 that acts as a thread dispatcher for dispatching execution threads to one or more shader cores 2255A-2255N, and a tiling unit 2258 for accelerating tiling operations for tile-based rendering, where rendering operations for a scene are subdivided in image space, for example, to take advantage of local space coherence within the scene or to optimize the use of an internal cache.

[0323] The inference and / or training logic 1215 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 1215 are provided herein in conjunction with Figures 12A and / or 12B. In at least one embodiment, the inference and / or training logic 1215 may be used in integrated circuits 22A and / or 22B for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0324] In at least one embodiment, one or more systems shown in Figures 22A and 22B are used to implement one or more implicit environment functions. In at least one embodiment, one or more systems shown in Figures 22A and 22B are used to use one or more neural networks, such as one or more implicit environment functions, to compute multiple paths that an entity, such as an autonomous device, is to traverse. In at least one embodiment, one or more systems shown in Figures 22A and 22B are used to implement one or more systems and / or processes, such as those described with respect to Figures 1 to 11.

[0325] Figures 23A and 23B show additional exemplary graphics processor logic according to embodiments described herein. Figure 23A shows a graphics core 2300, which in at least one embodiment may be included within the graphics processor 2110 of Figure 21, and in at least one embodiment may be unified shader cores 2255A-2255N, as in Figure 22B. Figure 23B shows a highly parallel general-purpose graphics processing unit ("GPGPU") 2330, which in at least one embodiment is suitable for deployment on a multi-chip module.

[0326] In at least one embodiment, the graphics core 2300 includes a shared instruction cache 2302, a texture unit 2318, and a cache / shared memory 2320, which are common to the execution resources within the graphics core 2300. In at least one embodiment, the graphics core 2300 may include multiple slices 2301A-2301N, or partitions for each core, and the graphics processor may include multiple instances of the graphics core 2300. In at least one embodiment, slices 2301A-2301N may include support logic including local instruction caches 2304A-2304N, thread schedulers 2306A-2306N, thread dispatchers 2308A-2308N, and sets of registers 2310A-2310N. In at least one embodiment, slices 2301A to 2301N may include a set of additional function units (AFUs) 2312A to 2312N, floating-point units (FPUs) 2314A to 2314N, integer arithmetic logic units (ALUs) 2316A to 2316N, address computational units (ACUs) 2313A to 2313N, double-precision floating-point units (DPFPUs) 2315A to 2315N, and matrix processing units (MPUs) 2317A to 2317N.

[0327] In at least one embodiment, the FPU2314A-2314N can perform single-precision (32-bit) and half-precision (16-bit) floating-point operations, and the DPFPU2315A-2315N can perform double-precision (64-bit) floating-point operations. In at least one embodiment, the ALU2316A-2316N can perform variable-precision integer operations with 8-bit, 16-bit, and 32-bit precision and can be configured for mixed-precision operations. In at least one embodiment, the MPU2317A-2317N can also be configured for mixed-precision matrix operations, including half-precision floating-point operations and 8-bit integer operations. In at least one embodiment, the MPU2317-2317N can perform various matrix operations to accelerate machine learning application frameworks, including enabling support for accelerated general-purpose matrix-to-matrix multiplication (GEMM). In at least one embodiment, AFU2312A~2312N can perform additional logical operations not supported by the floating-point unit or integer unit, including trigonometric function operations (e.g., sine, cosine, etc.).

[0328] The inference and / or training logic 1215 is used to perform inference and / or training operations related to one or more embodiments. Details relating to the inference and / or training logic 1215 are provided herein in conjunction with Figures 12A and / or 12B. In at least one embodiment, the inference and / or training logic 1215 may be used in the graphics core 2300 for inference or prediction operations, at least in part, based on weight parameters calculated using the neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0329] Figure 23B shows a general-purpose processing unit (GPGPU) 2330, which in at least one embodiment may be configured to enable highly parallel compute operations to be performed by an array of graphics processing units. In at least one embodiment, the GPGPU 2330 may be directly linked to other instances of the GPGPU 2330 to create a multi-GPU cluster to improve training speed for deep neural networks. In at least one embodiment, the GPGPU 2330 includes a host interface 2332 for enabling connectivity with a host processor. In at least one embodiment, the host interface 2332 is a PCI Express interface. In at least one embodiment, the host interface 2332 may be a vendor-specific communication interface or communication fabric. In at least one embodiment, the GPGPU 2330 receives commands from the host processor and uses the global scheduler 2334 to distribute the execution threads associated with those commands across a set of compute clusters 2336A–2336H. In at least one embodiment, the compute clusters 2336A–2336H share a cache memory 2338. In at least one embodiment, the cache memory 2338 can act as a higher-level cache for the cache memory within the compute clusters 2336A–2336H.

[0330] In at least one embodiment, the GPGPU 2330 includes memory 2344A-2344B coupled to compute clusters 2336A-2336H via a set of memory controllers 2342A-2342B. In at least one embodiment, memory 2344A-2344B may include various types of memory devices, including graphics random access memory such as dynamic random access memory (DRAM) or synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory.

[0331] In at least one embodiment, compute clusters 2336A to 2336H each include a set of graphics cores, such as the graphics core 2300 in Figure 23A, and the set of graphics cores may include multiple types of integer and floating-point logic units capable of performing computational operations at varying accuracies, including those suitable for machine learning computations. For example, in at least one embodiment, at least a subset of floating-point units in each of compute clusters 2336A to 2336H may be configured to perform 16-bit or 32-bit floating-point operations, and different subsets of floating-point units may be configured to perform 64-bit floating-point operations.

[0332] In at least one embodiment, multiple instances of GPGPU2330 may be configured to operate as a compute cluster. In at least one embodiment, the communication used for synchronization and data exchange by compute clusters 2336A-2336H varies across embodiments. In at least one embodiment, multiple instances of GPGPU2330 communicate via a host interface 2332. In at least one embodiment, GPGPU2330 includes an I / O hub 2339, which couples GPGPU2330 with a GPU link 2340 that enables direct connections to other instances of GPGPU2330. In at least one embodiment, the GPU link 2340 is coupled with a dedicated GPU-GPU bridge that enables communication and synchronization between multiple instances of GPGPU2330. In at least one embodiment, the GPU link 2340 is coupled with a high-speed interconnect to send and receive data to and from other GPGPUs or parallel processors. In at least one embodiment, multiple instances of the GPGPU 2330 are located on separate data processing systems and communicate via a ...

Claims

1. It is a processor, Equipped with one or more circuits, The one or more circuits use one or more neural networks to calculate multiple paths from different starting positions to one or more destination positions that the autonomous device is to traverse, at least in part on the outputs of one or more of the one or more neural networks, which include the distances from the different starting positions to the one or more destination positions. The aforementioned one or more neural networks include an implicit environment function that represents the environment via an environment field. The distance is the distance to be reached, which is obtained as the output of the implicit environment function, which has been trained to represent the environment field as a mapping from location coordinates to an arrival time function u(x). The aforementioned arrival time function u(x) is a processor obtained by solving a continuous shortest path problem that represents the aforementioned arrival time function u(x).

2. The aforementioned one or more circuits shall be at least, To obtain the first location, the set of locations, and the final location, The one or more neural networks are made to calculate a set of distances based at least partially on the set of locations and the final location. Calculating the plurality of paths based at least partially on the set of distances such that the plurality of paths form a path from the first location to the final location. The processor according to claim 1, which uses the one or more neural networks to compute the plurality of paths.

3. The first location described above is the location of the autonomous device, A subset of locations from the set of locations is accessible to the autonomous device from the first location. The processor according to claim 2.

4. The aforementioned one or more circuits include at least, Obtaining a subset of distances from the set of distances corresponding to the subset of locations, Selecting a second location from the subset of locations based at least partially on the subset of distances, Calculating the first route among the plurality of routes, including the route from the first location to the second location. The processor according to claim 3, which is for calculating the first path by

5. The processor according to claim 4, wherein the second location corresponds to the minimum distance among the subset of distances.

6. The processor according to claim 2, wherein one or more neural networks calculate the set of distances in a single forward pass.

7. The processor according to claim 2, wherein the distances in the set of distances correspond to the distance along the path from one of the locations in the set of locations to the final location.

8. A machine-readable medium, The machine-readable medium stores a set of instructions. If the set of instructions is executed by one or more processors, the one or more processors will use one or more neural networks to calculate multiple paths from different starting positions to one or more destination positions that the autonomous device is to traverse, based at least in part on one or more outputs of the one or more neural networks, which include the distances from the different starting positions to the one or more destination positions. The aforementioned one or more neural networks include an implicit environment function that represents the environment via an environment field. The distance is the distance to be reached, which is obtained as the output of the implicit environment function, which has been trained to represent the environment field as a mapping from location coordinates to an arrival time function u(x). The aforementioned arrival time function u(x) is a machine-readable medium obtained by solving a continuous shortest path problem that represents the aforementioned arrival time function u(x).

9. If the set of instructions is executed by one or more processors, the one or more processors shall: To acquire the characteristics of the environment, Selecting a location in the aforementioned environment, In order to obtain multiple distances corresponding to multiple locations in the aforementioned environment, at least the features and the locations are input to one or more neural networks. The machine-readable medium according to claim 8, further comprising an instruction to perform the following.

10. If the set of instructions is executed by the one or more processors, the one or more processors will: Selecting a set of locations accessible to the autonomous device, Obtaining a set of distances from among the multiple distances corresponding to the set of locations, Selecting a first location from the set of locations based at least partially on the set of distances, wherein the first route among the plurality of routes represents the route from the autonomous device to the first location. The machine-readable medium according to claim 9, further comprising an instruction to cause the medium to perform the following.

11. The machine-readable medium according to claim 10, wherein the set of instructions, when executed by the one or more processors, further comprises instructions causing the one or more processors to cause the autonomous device to navigate to the first location using the first path.

12. The machine-readable medium according to claim 11, wherein the autonomous device is an autonomous vehicle.

13. The machine-readable medium according to claim 9, wherein the features are generated by one or more encoders based on a representation of the environment.

14. The machine-readable medium according to claim 13, wherein the representation of the environment is an image or a point cloud.

15. It is a system, A computer comprising one or more computers having one or more processors, The one or more processors use one or more neural networks to calculate multiple paths from different starting positions to one or more destination positions that the autonomous device is to traverse, at least in part on the outputs of the one or more neural networks, which include the distances from the different starting positions to the one or more destination positions. The aforementioned one or more neural networks include an implicit environment function that represents the environment via an environment field. The distance is the distance to be reached, which is obtained as the output of the implicit environment function, which has been trained to represent the environment field as a mapping from location coordinates to an arrival time function u(x). The aforementioned arrival time function u(x) is a system obtained by solving a continuous shortest path problem that represents the aforementioned arrival time function u(x).

16. The aforementioned one or more processors further, Capturing the representation of the environment, Using one or more neural networks to calculate the multiple paths from the first location to the second location of the autonomous device in the environment described above. The system according to claim 15, which is for performing the following.

17. The system according to claim 16, wherein the one or more processors are further used to use the one or more neural networks to compute one or more distance values ​​for one or more locations in the environment, based at least in part on the representation of the environment.

18. The aforementioned one or more processors further, Calculating the size of the step of the autonomous device, Selecting a set of locations accessible from the first location of the autonomous device through the steps, Selecting a third location from the set of locations based at least partially on one or more of the aforementioned distance values. The system according to claim 17, which is for performing the following.

19. The system according to claim 16, wherein the representation of the environment is captured through one or more depth cameras.

20. The system according to claim 16, wherein the representation of the environment is a 2D representation or a 3D representation.

21. The system according to claim 15, wherein the autonomous device is an autonomous robot.

22. A machine-readable medium, The machine-readable medium stores a set of instructions. If the set of instructions is executed by one or more processors, then the one or more processors will have at least the following instructions: One or more neural networks are trained to calculate multiple paths from different starting positions to one or more destination positions that an autonomous device is scheduled to traverse, at least partially based on one or more outputs of the one or more neural networks, which include the distance from the different starting positions to the one or more destination positions. The aforementioned one or more neural networks include an implicit environment function that represents the environment via an environment field. The distance is the distance to be reached, which is obtained as the output of the implicit environment function, which has been trained to represent the environment field as a mapping from location coordinates to an arrival time function u(x). The aforementioned arrival time function u(x) is a machine-readable medium obtained by solving a continuous shortest path problem that represents the aforementioned arrival time function u(x).

23. If the set of instructions is executed by the one or more processors, the one or more processors will: To obtain the environment and location, One or more algorithms are used to determine one or more reachable distance values ​​for one or more locations in the environment to the said location, Training one or more neural networks using at least one or more of the aforementioned distance values. The machine-readable medium according to claim 22, further comprising an instruction to cause the machine-readable medium to perform the following actions.

24. If the set of instructions is executed by the one or more processors, the one or more processors will: The one or more neural networks are to process at least one or more locations in order to calculate one or more predicted range values, Updating the one or more neural networks based at least partially on the difference between the one or more predicted range values ​​and the one or more range values. The machine-readable medium according to claim 23, further comprising an instruction to cause the machine-readable medium to perform the following actions.

25. The machine-readable medium according to claim 23, wherein the one or more algorithms include one or more fast marching method (FMM) algorithms.

26. The machine-readable medium according to claim 23, wherein the first of the one or more reach values ​​corresponds to the first location among the one or more locations and indicates the distance along the path from the first location to the location.

27. The machine-readable medium according to claim 26, wherein the aforementioned path is a geometrically achievable path.

28. It is a processor, Equipped with one or more circuits, The one or more circuits train one or more neural networks to compute multiple paths from different starting positions to one or more destination positions that the autonomous device is to traverse, at least in part on the outputs of one or more of the one or more neural networks, which include the distance from the different starting positions to the one or more destination positions. The aforementioned one or more neural networks include an implicit environment function that represents the environment via an environment field. The distance is the distance to be reached, which is obtained as the output of the implicit environment function, which has been trained to represent the environment field as a mapping from location coordinates to an arrival time function u(x). The aforementioned arrival time function u(x) is a processor obtained by solving a continuous shortest path problem that represents the aforementioned arrival time function u(x).

29. The aforementioned one or more circuits may further include: One or more algorithms are used to process at least the environment and location set and the target location, Obtaining a set of distance values ​​based at least partially on the results of one or more of the aforementioned algorithms, Training one or more neural networks using the aforementioned set of distance values The processor according to claim 28, which is for performing the following.

30. The processor according to claim 29, wherein the one or more circuits are further for training the one or more neural networks to process at least the set of environment and location and the target location in order to compute the set of distance values.

31. The processor according to claim 29, wherein a first distance value from the set of distance values ​​represents the distance along a semantically feasible path from a first position from the set of positions to the target position.

32. The processor according to claim 29, wherein the one or more algorithms include one or more path planning algorithms.