Long-term perception for robotic systems and applications

By generating and storing robot sensor data as a vector database, and using a large language model to query and answer questions, the problem of robot perception systems being unable to capture dynamic environmental changes is solved, enabling long-term perception and extensive task execution capabilities.

CN121492005APending Publication Date: 2026-02-10NVIDIA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511111397.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-16
Filing Date
2025-08-08
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing robot perception systems cannot effectively capture dynamic changes in objects and phenomena in the environment, have long deployment times and short memory durations, which limits the robot's ability to perform tasks in large, dynamic environments.

Method used

By generating streams of video frames, location information, and sensor data, the data is transformed into embedded representations using visual language models and machine learning models, stored as a vector database memory for the machine, and queries and answers are performed using large language models and multimodal language models to generate navigation targets and action commands.

Benefits of technology

It enables robots to achieve long-term perception in dynamic environments, quickly retrieve and reason about memories that match queries, detect open set objects, and expand the time range and task types for task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121492005A_ABST
    Figure CN121492005A_ABST
Patent Text Reader

Abstract

The invention discloses long-term perception for robotic systems and applications. In various examples, techniques for performing a task include converting one or more sensing inputs obtained using one or more sensors of a machine into a plurality of segments. The technique further includes, for each segment included in the plurality of segments, generating a description for the segment via execution of a machine learning model; and storing the representation of the description in a data store in association with the segment. The technique further includes performing, by the machine, one or more actions based at least on the one or more queries on the data store.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 681,363, filed on August 9, 2024, the entire contents of which are incorporated herein by reference. Background Technology

[0003] Many autonomous or semi-autonomous mobile robots or other machine types perform tasks involving movement in large, dynamic, and / or semi-structured environments. During the deployment of a given robot (or another type of machine), the robot perceives (e.g., learns) objects, dynamic events, and phenomena in the corresponding environment via a series of sensors and sensing modules. For example, autonomous mobile robots (AMRs), humanoid robots, autonomous or semi-autonomous vehicles, etc., can use cameras, LiDAR sensors, RADAR sensors, inertial sensors, and / or other types of sensors to detect and / or identify items on shelves, other types of AMRs, robots, humans, manually operated equipment, and / or other targets or obstacles. In another example, outdoor robots can use Global Positioning System (GPS) positioning information, computer vision, and / or mapping techniques to monitor and / or explore outdoor environments.

[0004] However, conventional methods used for robot perception, navigation, and other tasks are associated with several drawbacks. First, robots typically navigate their environment by generating metric maps, scene graphs, and / or other representations that focus on static elements of the environment. These types of representations fail to capture dynamic changes in objects and phenomena within the environment.

[0005] Secondly, the deployment time and / or task execution time of robots are becoming increasingly longer. For example, navigation robots can perform inspections, anomaly detection, language-guided guidance, and / or other tasks in a single session lasting from hours to days. However, the duration of typical representations in spatiotemporal robot memory is much shorter (e.g., 1 to 2 minutes of video), which interferes with the robot's ability to recall key events and / or task-related context relevant to the task.

[0006] Third, conventional methods rely on a manually defined list of objects (e.g., chairs, people, tables, etc.) for mapping and / or perception. This "closed set" approach excludes the perception and / or recall of other objects not included in the list and limits the ability to perform tasks related to additional objects.

[0007] Therefore, more effective technologies are needed to improve robot perception and recall. Attached Figure Description

[0008] The system and method for long-term sensing for robotic systems and applications are described in detail below with reference to the accompanying drawings, wherein:

[0009] Figure 1 The illustration depicts a computing device configured to implement one or more aspects of various embodiments;

[0010] Figure 2 The illustration depicts a system for providing long-term sensing according to various embodiments, the system comprising: Figure 1 The processing engine and retrieval engine;

[0011] Figure 3A The illustrations show the components according to various embodiments. Figure 1 The processing engine generates example datasets associated with memory;

[0012] Figure 3B The illustrations show the components according to various embodiments. Figure 1 The processing engine generates example datasets associated with memory;

[0013] Figure 4 The illustrations show various embodiments. Figure 1 The operation of the search engine when generating results associated with the example question;

[0014] Figure 5 The illustration shows a flowchart illustrating a method for generating a memory representation for a machine, according to various embodiments;

[0015] Figure 6 The illustration shows a flowchart of a method for performing a task using a memory representation for a machine, according to various embodiments;

[0016] Figure 7A This is a block diagram of an example generative language model system suitable for implementing at least some embodiments of the present disclosure;

[0017] Figure 7B It is a block diagram of an example generative language model that includes a converter encoder-decoder suitable for implementing at least some embodiments of the present disclosure;

[0018] Figure 7C It is a block diagram of an example generative language model that includes a decoder-only converter architecture suitable for implementing at least some embodiments of the present disclosure;

[0019] Figure 8 This is a block diagram of an example computing device suitable for implementing at least some embodiments of the present disclosure;

[0020] Figure 9 This is a block diagram of an example data center applicable to implementing at least some embodiments of this disclosure;

[0021] Figure 10A These are illustrations of example autonomous vehicles according to some embodiments of the present disclosure;

[0022] Figure 10B According to some embodiments of this disclosure Figure 10A Examples of camera positions and fields of view for autonomous vehicles;

[0023] Figure 10C According to some embodiments of this disclosure Figure 10A A block diagram of an example system architecture for an example autonomous vehicle; and

[0024] Figure 10D This is based on some embodiments of the present disclosure for use in cloud-based servers and Figure 10A Here is a system diagram illustrating communication between autonomous vehicles. Detailed Implementation

[0025] Systems and applications relating to long-term perception for robotic systems and applications are disclosed. While examples of autonomous or semi-autonomous vehicles, machines, or robots 1000 (which may alternatively be referred to herein as "robot 1000," "self-robot 1000," "machine 1000," or "self-machine 1000") are also discussed, Figures 10A to 10D This disclosure is described with reference to examples thereof, but is not intended to be limiting. For example, the systems and methods described herein can be used by, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, airships, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, underwater vehicles, drones, robots, and / or other vehicle types. Furthermore, while this disclosure may be described with respect to robot perception, this is not intended to be limiting, and the systems and methods described herein can be used in augmented reality, virtual reality, mixed reality, security and supervision, autonomous or semi-autonomous machine applications, and / or any other technological space where perception and recall can be used.

[0026] As discussed in this paper, robots (e.g., AMRs, humanoid robots, forklifts, vehicles, drones, underwater vehicles, etc.) typically navigate their environments by generating metric maps, scene graphs, and / or other representations that focus on static elements of the environment. However, these types of representations fail to capture dynamic changes in objects and phenomena within the environment. Furthermore, robot deployment times are increasingly longer (e.g., hours to days at a time), while the duration of conventional representations in spatiotemporal robot memory is much shorter (e.g., 1 to 2 minutes of video). Consequently, conventional robot perception systems may fail to perceive and / or recall critical events relevant to a given task or context.

[0027] To address the aforementioned limitations, the disclosed technology provides long-term perception for robots (and other machine types) operating in dynamic environments. During the memory-building phase, the robot generates a stream of video frames, location information, temporal information, and / or other types of sensor data. This data is aggregated into “segments” corresponding to discrete time intervals (e.g., every 3 seconds). A visual language model (VLM) and / or another type of machine learning model is used to generate captions for each segment, and these captions are converted into embedded representations (e.g., using an embedding model). The embedded representations, location information, temporal information, and / or other information generated and / or collected during the memory-building phase are stored as corresponding “memories” of the robot in a vector database residing on the robot and / or in remote locations accessible to the robot.

[0028] During the subsequent query phase, Large Language Models (LLM), Visual Language Models (VLM), Multimodal Language Models (MMLM), and / or other types of machine learning models are used to transform the user's question to the robot into a set of queries against a vector database. These queries are used to retrieve relevant memories from the vector database, and the retrieved memories are added to the context used by the machine learning model. This process is iteratively repeated using a new set of queries generated from the updated context until the LLM / VLM / MMLM determines that the updated context is sufficient to answer the question. The LLM / VLM / MMLM then uses the information from the updated context to generate and output the answer to the question. The answer may include text related to the question, location information, time information, and / or time duration information. Some or all of this information may also be used to generate navigation goals for the robot, trajectories for the robot, commands to move the robot toward the navigation goals, and / or other outputs that cause the robot to perform one or more actions related to the question.

[0029] One technical advantage of the disclosed technique compared to prior methods is the continuous conversion of sensor data collected by the robot (or other machine type) into efficient memory representations, which can be used to retrieve objects, scenes, and / or dynamic events perceived by the robot. Therefore, the disclosed technique improves performance on tasks involving long time horizons compared to conventional methods with short memory durations. Another technical advantage of the disclosed technique is its ability to rapidly retrieve and reason about memories matching queries. Thus, the disclosed technique allows the robot to generate answers to any set of spatiotemporal questions and perform related tasks in a timely and feasible manner. Yet another technical advantage of the disclosed technique is its ability to detect, identify, and recall "open set" objects during query processing. Therefore, compared to conventional methods, the disclosed technique allows the robot to answer a wider range of questions and / or perform a wider range of tasks related to perception and / or retrieval, which specify "close set" objects for mapping and / or perception.

[0030] The examples above are by no means intended to be limiting. As those skilled in the art will appreciate, in general, the techniques used to perform conditional data sourcing and curation can be implemented and / or used with any suitable application.

[0031] The systems and methods described herein can be used for a variety of purposes, as examples but not limited to, systems associated with: machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and supervision, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, data center processing, conversational AI, generative AI, optical transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation of three-dimensional (3D) assets, cloud computing and / or any other suitable application.

[0032] The disclosed embodiments can be included in a variety of different systems, such as automotive systems (e.g., infotainment or plug-in gaming / streaming systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems implementing one or more language models (such as LLM / VLM / multimodal language models / other model types capable of handling text, audio, 3D data, and / or image data), systems for performing optical transmission simulations, systems for performing collaborative content creation of 3D assets, systems implemented at least partially using cloud computing resources, systems for performing generative AI operations, and / or other types of systems.

[0033] In some embodiments, the systems and methods described herein can be deployed in in-vehicle infotainment (IVI) systems or in-cabin experience (IX) applications. For example, an infotainment system within a vehicle (e.g., a car, truck, drone, engineering equipment, robot, semi-autonomous vehicle, or autonomous vehicle) may include one or more onboard processors (e.g., CPU, GPU, hardware-based deep learning accelerator (DLA), hardware-based programmable vision accelerator (PVA) (which may include one or more vector processing units (VPUs)), direct memory access (DMA) systems and / or pixel processing engines (PPE), hardware-based optical flow accelerators (OFA), SoCs, etc.), memory and / or storage devices (e.g., for storing control algorithms, sensor data, and one or more machine learning models), and memory and / or storage devices (e.g., for storing infotainment content, navigation data, and user preferences). The system can use these processors to execute one or more machine learning models (e.g., language models) to achieve features such as voice control, personalized media recommendations, dynamic navigation, and real-time communication with other services via network connectivity. In-vehicle infotainment systems can also use natural language processing (NLP) models to enable voice-based interaction. One or more machine learning models can be stored locally or accessed via one or more APIs connected to cloud services, enabling the system to process requests in real-time or near real-time.

[0034] In some embodiments, the systems and methods described herein can be deployed in robotic applications. For example, a robot or robotic system may include one or more onboard processors (e.g., CPU, GPU, hardware-based deep learning accelerator (DLA), hardware-based programmable vision accelerator (PVA) (which may include one or more vector processing units (VPU)), direct memory access (DMA) systems and / or pixel processing engines (PPE), hardware-based optical flow accelerators (OFA), SoCs, etc.) and memory and / or storage devices (e.g., for storing control algorithms, sensor data, and one or more machine learning models). The robotic system can use these processors to execute one or more machine learning models (e.g., language models) that enable the robotic system to perform complex tasks autonomously or semi-autonomously, such as interacting with and / or manipulating static and / or dynamic objects, or navigating the environment using sensors (such as cameras, LiDAR, RADAR, ultrasonic sensors, etc.). This system can use sensor fusion technology to combine data from multiple sensors (e.g., cameras, infrared, LiDAR, RADAR, accelerometers) to create a comprehensive model of the robot's surrounding environment. This data can be processed locally on the robot or sent to a remote server for computationally intensive tasks such as 3D mapping or SLAM (Simultaneous Localization and Mapping). In one or more embodiments, data from individual robots (e.g., sensor data, task status, or environmental conditions) can be uploaded to the cloud, where a centralized AI model can analyze optimized commands and distribute them throughout the fleet. In some embodiments, one or more machine learning models described herein (e.g., language models, VLM, LLM, MMLM, diffusion models, NeRF models, DNN, etc.) can be used to enable the robot to perceive and reason about its environment and / or communicate with one or more other robots and / or humans in the environment. In some embodiments, the robot may (e.g., using one or more network interface cards (NICs) and / or data processing units (DPUs)) communicate with one or more locally hosted servers / computing devices and / or with one or more remotely located servers / computing devices (e.g., in one or more data centers).

[0035] In some embodiments, simulated data (e.g., simulated sensor data from simulated sensors of virtual machines or simulated machines) can be used to execute the systems and methods described herein in a simulated environment (e.g., NVIDIA DriveSIM, NVIDIA ISAAC GYM, NVIDIA ISAAC SIM, etc.). For example, simulated sensor data can be used (e.g., processed using one or more machine learning models, neural networks, etc.) to identify, detect, and / or classify lane lines, road boundary lines, other lines, vertical structures / features, etc., in the simulated environment using individual points of curves and / or one or more curve fitting algorithms, and this information can be used to perform operations associated with virtual machines in the environment (e.g., control, navigation, planning, etc.). These simulated operations can be used to test their execution before deploying the underlying algorithms, systems, and / or processes to the real world. In some instances, simulation can be used to generate synthetic training data, e.g., training data including regions of interest and / or subregions of interest from the simulation. In some embodiments, other methods may be used besides simulation or as an alternative to simulation to generate synthetic training data. For example, synthetic training data can be generated using neural rendering fields (NERF), Gaussian sputtering techniques, diffusion models, electrostatic models (e.g., Poisson flow generation models (PFGM), etc.). The synthetic training data (in addition to or as a substitute for real-world data) can then be processed to determine geometry, curvature, semantic information, classification information, and / or other information relating to features of interest (such as lines, longitudinal features (e.g., poles), and / or other features in a driving environment, warehouse, etc.). In any example (such as where a simulated environment is used for testing, validation, training, etc.), the simulated environment and / or associated training data can be rendered or otherwise generated using one or more optical propagation algorithms (such as ray tracing and / or path tracing algorithms). In some embodiments, the simulated environment and / or one or more of its objects, features, or components can be generated or managed in a 3D content collaboration platform (e.g., NVIDIA's OMNIVERSE) for industrial digitization, generative physics AI, and / or other use cases, applications, or services. For example, a content collaboration platform or system may include a system that uses generic scene descriptors (USD) (e.g., OpenUSD) data to manage objects, features, scenes, etc., in simulated environments, digital environments, etc. The platform may include realistic physics simulations, such as using NVIDIA's PhysX SDK, to simulate real physics and physical interactions with simulations hosted by the platform.This platform can integrate OpenUSD with ray tracing / path tracing / light transport simulations (e.g., NVIDIA's RTX rendering technology) into software tools and simulation workflows for building, training, deploying, or testing AI systems, such as systems for testing, validating, training (e.g., machine learning models, neural networks, etc.) and / or for tasks related to automobiles, robots, machines, or other applications.

[0036] In some embodiments, a remote control or teleoperation system can be used to remotely operate or control a vehicle or other machine. For example, the systems and methods described herein can be used to identify lane lines, road boundary lines, longitudinal features, etc., which can be included in a visualization or mapping of the environment to assist a remote operator in controlling an autonomous or semi-autonomous machine through the environment (or to provide indications of waypoints or other control or navigation mappings).

[0037] System Overview

[0038] Figure 1 This is a block diagram illustrating a computing system 100 configured to implement one or more aspects of at least one embodiment. In at least one embodiment, the computing system 100 may include any type of computing device, including but not limited to server machines, server platforms, desktop machines, laptop machines, handheld devices / mobile devices, digital kiosks, in-vehicle infotainment systems, smart speakers or displays, televisions, and / or wearable devices. In at least one embodiment, the computing system 100 is a server machine operating in a data center or cloud computing environment that provides scalable computing resources as a service over a network. In one or more embodiments, the computing system 100 is included in and / or accessible by robots, autonomous vehicles, semi-autonomous vehicles, and / or other types of machines capable of performing perception, planning, control, prediction, and / or other tasks related to movement and / or navigation in large, dynamic, and / or semi-structured environments.

[0039] In various embodiments, the computing system 100 includes, but is not limited to, one or more processors 102 and one or more memories 104, which are coupled to a parallel processing subsystem 112 via a memory bridge 105 and a communication path 113. The memory bridge 105 is also coupled to an I / O (input / output) bridge 107 via a communication path 106, and the I / O bridge 107 is coupled to a switch 116.

[0040] In one embodiment, I / O bridge 107 is configured to receive user input from optional input devices 108 (such as, but not limited to, keyboards, mice, touchscreens, sensor data analysis (e.g., evaluating gestures, speech, or other information for one or more uses regarding the field of view or sensing field of one or more sensors), VR / MR / AR headsets, gesture recognition systems, steering wheels, mechanical, digital, or touch-sensitive buttons or input components, and / or microphones), and forward the input to one or more processors 102 for processing. In at least one embodiment, computing system 100 may be a server machine in a cloud computing environment. In such embodiments, computing system 100 may omit input device 108 and receive input as equivalent to commands (e.g., in response to one or more inputs from a remote computing device) and / or messages transmitted over a network and received via network adapter 118. In at least one embodiment, switch 116 is configured to provide connectivity between I / O bridge 107 and other components of computing system 100, such as network adapter 118 and various add-on cards 120 and 121.

[0041] In at least one embodiment, I / O bridge 107 is coupled to system disk 114, which can be configured to store content, applications, and data for use by one or more processors 102 and parallel processing subsystems 112. In one embodiment, system disk 114 provides non-volatile storage for applications and data and may include fixed or removable hard disk drives, flash memory devices, and CD-ROMs (optical disc read-only memory), DVD-ROMs (digital versatile discs), Blu-ray, HD-DVDs (high-definition DVDs), or other magnetic, optical, or solid-state storage devices. In various embodiments, other components (such as universal serial bus or other port connections, optical disc drives, digital versatile disc drives, movie recording devices, etc.) may also be connected to I / O bridge 107.

[0042] In various embodiments, memory bridge 105 may be a northbridge chip, and I / O bridge 107 may be a southbridge chip. Furthermore, communication paths 106 and 113, as well as other communication paths within computing system 100, may be implemented using any technically suitable protocol, including but not limited to AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol known in the art.

[0043] In at least one embodiment, the parallel processing subsystem 112 includes a graphics subsystem that delivers pixels to an optional display device 110, which can be any conventional cathode ray tube, liquid crystal display, light-emitting diode display, etc. In such an embodiment, the parallel processing subsystem 112 may include circuitry optimized for graphics and video processing, including, for example, video output circuitry. This circuitry may be incorporated into one or more parallel processing units (PPUs), also referred to herein as parallel processors, included within the parallel processing subsystem 112.

[0044] In at least one embodiment, the parallel processing subsystem 112 includes circuitry optimized for general and / or computational processing (e.g., optimized circuitry). Similarly, such circuitry may be included on one or more PPUs incorporated within the parallel processing subsystem 112, which are configured to perform such general and / or computational operations. In some other embodiments, one or more PPUs incorporated within the parallel processing subsystem 112 may be configured to perform graphics processing, general processing, and / or computational processing operations. One or more memories 104 include at least one device driver configured to manage the processing operations of one or more PPUs within the parallel processing subsystem 112. Furthermore, one or more memories 104 include a processing engine 122 and a retrieval engine 124, which may be executed by one or more processors and / or the parallel processing subsystem 112.

[0045] In various embodiments, the parallel processing subsystem 112 can be coupled with... Figure 1 One or more other components are integrated together to form a single system. For example, the parallel processing subsystem 112 may be integrated with one or more processors 102 and other interconnect circuitry on a single chip to form a system-on-a-chip (SoC).

[0046] One or more processors 102 may include any suitable processor implemented as a central processing unit (CPU), graphics processing unit (GPU), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), artificial intelligence (AI) accelerator, deep learning accelerator (DLA), parallel processing unit (PPU), data processing unit (DPU), vector or vision processing unit (VPU), programmable vision accelerator (PVA) (which may include one or more VPUs and / or direct memory access (DMA) systems), any other type of processing unit, or a combination of different processing units (such as one or more CPUs configured to operate in conjunction with one or more GPUs). In general, one or more processors 102 may include any technically feasible hardware unit capable of processing data and / or executing software applications. Further, in the context of this disclosure, the computing elements shown in computing system 100 may correspond to physical computing systems (e.g., systems in a data center or machine) and / or may correspond to virtual computing instances performed in a computing cloud.

[0047] In at least one embodiment, one or more processors 102 issue commands to control the operation of the PPU. In at least one embodiment, the communication path 113 is a Fast PCI link, in which a dedicated lane is allocated to each PPU. Other communication paths may also be used. The PPU advantageously implements a highly parallel processing architecture, and any amount of local parallel processing memory (PP memory) can be provided to the PPU.

[0048] It should be understood that the system shown herein is illustrative, and variations and modifications are possible. The connection topology, including the number and arrangement of bridges, the number of processors 102, and the number of parallel processing subsystems 112, can be modified as desired. For example, in at least one embodiment, one or more memories 104 may be connected directly to one or more processors 102 instead of via memory bridge 105, and other devices may communicate with one or more memories 104 via memory bridge 105 and processors 102. In other embodiments, parallel processing subsystems 112 may be connected to I / O bridge 107 or directly to one or more processors 102, rather than to memory bridge 105. In still other embodiments, I / O bridge 107 and memory bridge 105 may be integrated into a single chip, rather than existing as one or more discrete devices. In some embodiments, Figure 1 One or more of the components shown may be absent. For example, switch 116 may be eliminated, and network adapter 118 and add-on cards 120, 121 may be directly connected to I / O bridge 107. In some embodiments, Figure 1The one or more components shown can be implemented as virtualized resources in a virtual computing environment (such as a cloud computing environment). Specifically, in at least one embodiment, the parallel processing subsystem 112 can be implemented as a virtualized parallel processing subsystem. For example, the parallel processing subsystem 112 can be implemented as one or more virtual graphics processing units (vGPUs) that render graphics on one or more virtual machines (VMs) that execute on one or more server machines, and one or more GPUs and other physical resources of the one or more server machines are shared across one or more VMs.

[0049] Long-term sensing for robotic systems and applications

[0050] Figure 2 The figure illustrates a system for providing long-term sensing according to at least one embodiment, the system comprising: Figure 1 The processing engine 122 and retrieval engine 124. In some embodiments, the processing engine 122 and retrieval engine 124 include functionality for providing long-term perception for robots (and other machine types) operating in dynamic environments. Each of these components is described in further detail below.

[0051] Processing engine 122 converts individual streams 220(1)-220(Y) (each stream is individually referred to herein as stream 220) of input data 210 associated with sensors 252(1)-252(X) (each of which is individually referred to herein as sensor 252) on the machine into a set of memories 226(1)-226(Z) (each of which is individually referred to herein as memory 226) for the machine. Input data 210 may include streams 220 of the following: images, LiDAR data, RADAR data, audio data, ultrasonic data, inertial measurement unit (IMU) data, timestamp data, odometer data and / or other data collected and / or generated by cameras, LiDAR sensors, RADAR sensors, microphones, IMUs and / or other sensors 252 on the machine. Input data 210 may also include, or alternatively include, occupancy maps, visualizations, semantic labels (e.g., segmented maps, detected objects, boundary shapes, etc.), status data (e.g., location, heading, speed, battery level, etc.), guidance data (e.g., routes, trajectories, paths, headings, navigation targets, etc.) and / or other data derived from sensor data collected by sensor 252.

[0052] like Figure 2As shown, processing engine 122 performs aggregation 212 of the stream 220 of input data 210 into time-ordered segments 222(1)-222(Z) (each segment is individually referred to herein as segment 222). For example, processing engine 122 may execute on a machine (or at a remote location accessible to the machine) and process each stream 220 in real-time or near real-time as the corresponding input data 210 is collected and / or generated by sensor 252 and / or other components on the machine. Processing engine 122 may divide each stream 220 of input data 210 into discrete subsets, wherein each subset includes input data 210 collected and / or generated within a fixed and / or varying time interval (e.g., a number of seconds) between a certain start time and a certain end time. These discrete subsets of input data 210 can be consecutive (e.g., such that the time spanned by a given subset of input data 210 begins immediately after the end of the time spanned by a previous subset of input data 210) and / or overlapping (e.g., such that the start time of a given subset of input data 210 falls within the time spanned by one or more previous subsets of input data 210). Processing engine 122 can also group multiple streams 220 of input data 210 associated with the same start and end times into corresponding segments 222 (e.g., by storing the original input data 210, summary statistics associated with the original input data 210, a compressed representation of the original input data 210, and / or another representation of the input data 210 in the corresponding segment 222).

[0053] Processing engine 122 uses one or more machine learning models 214 to generate a set of one or more descriptive texts 224(1)-224(Z) (each of which is individually referred to herein as descriptive text 24) and a set of one or more embeddings 216(1)-216(Z) (each of which is individually referred to herein as embedding 216). For example, processing engine 122 may prompt and / or execute visual language (VILA) models, multimodal large language models (MMLM), visual language models (VLM), and / or other types of machine learning models to generate descriptive texts that describe semantic content and / or context associated with image data, 3D data, location data, temporal data, and / or other input data 210 included in each segment 222. The prompts input into the machine learning model may include (but are not limited to) context associated with the input data 210 (e.g., the specific environment in which the machine is deployed, roles and / or use cases associated with perception performed through the machine, etc.), instructions for describing and / or interpreting the input data 210 in the corresponding segment 222 (e.g., style and / or format associated with the explanatory text), and / or attention to certain aspects of the input data 210 (e.g., specific objects, events, movements, patterns, analysis types, etc.). The processing engine 122 may also use text embedding models, image embedding models, video embedding models, multimodal embedding models, and / or other types of embedding models to transform each explanatory text and / or some or all of the data from the corresponding segment 222 into corresponding embeddings in a low-dimensional latent vector space.

[0054] Processing engine 122 converts the data associated with fragment 222, descriptive text 224, and / or embedding 216 into corresponding memories 226 for the machine. Processing engine 122 additionally stores these memories 226 in a data store 208 accessible to the machine, processing engine 122, and / or retrieval engine 124. For example, processing engine 122 may generate different memories 226 for each fragment 222 and store the generated memories 226 in a vector database, relational database, key-value store, and / or other types of data stores 208 residing on the machine and / or accessible to the machine. Thus, each memory 226 may reflect objects, events, and / or phenomena encountered and / or perceived by the machine during a corresponding time interval.

[0055] like Figure 2As shown, each memory 226 includes one or more embeddings 216, one or more locations 230, and one or more times 232. As mentioned above, the embeddings 216 may include a latent vector representation of the descriptive text 224 generated by the machine learning model 214 and / or data in the corresponding segment 222. The locations 230 may include coordinates, paths, and / or other representations of two-dimensional (2D) and / or 3D location information of the machine during the time interval spanned by the corresponding segment 222. The times 232 may include a start time and / or an end time associated with the time interval spanned by the corresponding segment. The times 232 may also, or alternatively, include timestamps for some or all of the input data 210 stored in the memory 226.

[0056] The embedding 216, position 230, and / or time 232 in a given memory 226 can be mapped to the descriptive text 224 and / or input data 210 in the corresponding segment. For example, the embedding 216, position 230, and / or time 232 can be stored in association with the corresponding descriptive text 224 and / or input data 210, the identifier for the corresponding descriptive text 224 and / or input data 210, the position of the corresponding descriptive text 224 and / or input data 210 (e.g., in data repository 208 and / or other data sources), and / or other information that can be used to retrieve the corresponding descriptive text 224 and / or input data 210.

[0057] Therefore, processing engine 122 can generate a queryable memory representation V, which includes multiple memories 226. During the generation of a given memory 226, processing engine 122 can process input data 210 spanning t seconds from time i. Aggregated into the corresponding fragment 222, one or more machine learning models 214 are used to generate some or all of the one or more descriptive text L from the aggregated input data 210. i:i+t And one or more additional machine learning models 214 with corresponding embedding functions E are used to generate one or more embeddings 216 of the descriptive text 224. and / or one or more embeddings 216 of the input data 210 Then, the processing engine 122 can process the generated embedding 216 and the position 230 associated with the machine during the same t seconds. and time 232 The memory 226 is stored in this memory. The processing engine 122 can repeat this process to generate memory 226 within an additional time i while the machine is being deployed.

[0058] Figure 3A The illustrations show the relationship between various embodiments and Figure 1The processing engine 122 generates a set of sample data associated with memory 226(1). This data includes frames 302 of video captured by one or more sensors 252 on the machine (e.g., one or more cameras). The data also includes corresponding descriptive text 224(1) that describes the semantic content of a given segment 222 of the video including frame 302. Figure 3A As shown, the descriptive text 224 includes the location (e.g., university campus), objects (e.g., paths, sidewalks, trees, buildings, etc.), color, weather, and / or other visual attributes of the video segment 222. The descriptive text 224, frame 302, and / or other data associated with the segment 222 can be converted into one or more embeddings 216. These embeddings 216 can be stored and / or mapped to the location 230, time 232, descriptive text 224, frame 302, and / or other information associated with the segment 222, along with the other information associated with the segment 222, to facilitate subsequent retrieval and use of this information by the retrieval engine 124.

[0059] Figure 3B The illustrations show the relationship between various embodiments and Figure 1 A set of example data associated with memory 226(2) generated by the processing engine 122. With Figure 3A Similar to memory 226(1), the data includes frames 304 of video captured from one or more sensors 252 on the machine (e.g., one or more cameras). The data also includes corresponding descriptive text 224(2) that describes the semantic content of a given segment 222 of the video including frame 304.

[0060] More specifically, frame 304 depicts the same environment as frame 302, but is captured later than frame 302. Further, explanatory text 224(2) describes people, objects, camera movement, and / or other events present in frame 304 and its corresponding segments but missing in frame 302 and its corresponding segments. Similarly, location 230 and time 232 in memory 226(2) may differ from the location and time in memory 226(1). Therefore, Figure 3A and Figure 3B The memories 226(1) and 226(2) can reflect the changes in the environment and phenomena perceived by the machine over time.

[0061] Back Figure 2In the discussion, retrieval engine 124 uses memory 226 in data repository 208 to perform one or more actions 250. In some embodiments, action 250 includes generating an answer to a question 234 from a user, wherein question 234 relates to an object, event, and / or other entity of interest perceived and / or encountered by the machine. For example, retrieval engine 124 may execute on the machine and receive question 234 from the user in the form of text, voice, one or more gestures, and / or other input. Question 234 may include (but is not limited to) spatial questions relating to the location of a given entity of interest (e.g., the nearest chair or bathroom), temporal questions relating to the time when a given event occurs (e.g., when a pile of boxes falls) and / or the duration of an event or activity (e.g., how long the machine stayed inside the building), descriptive questions relating to details observed by the machine (e.g., how busy a street or sidewalk is, the side of a street or driveway where vehicles are traveling, etc.), and / or binary questions that can be answered with "yes" or "no" (e.g., whether the machine encountered a person or object).

[0062] To answer a given question 234, retrieval engine 124 uses one or more machine learning models 202 to transform question 234 into a set of queries 238(1)-238(N) in data repository 208 (where each query is individually referred to herein as query 238). Retrieval engine 124 may also use one or more machine learning models 202 to transform each query 238 into one or more embeddings 240(1)-240(N) (where each embedding is individually referred to herein as embedding 240). Retrieval engine 124 performs a lookup in data repository 208 using each generated query 238 and / or corresponding embedding 240 to generate results 206, which include a set of one or more matching memories 242(1)-242(N) (where each memory is individually referred to herein as matching memory 242). Each set of matching memories 242 includes one or more memories 226 in data repository 208 that are semantically similar to and / or related to the corresponding query 238. The retrieval engine 124 updates the context 236 associated with the question 234 using the retrieved matching memory 242, and repeats this process one or more times using the same question 234 and the updated context 236 until the machine learning model 202 determines that the question 234 can be answered using the context 236. The retrieval engine 124 then uses the same machine learning model 202 and / or one or more additional machine learning models 202 to generate an answer to the question 234 and / or perform other actions 250 related to the question 234 using the collected context 236.

[0063] The operations of search engine 124 can be performed by p(A|Q,H)1:K () represents, where Q represents question 234, A represents the answer and / or action to be performed in response to question 234, and H represents... 1:K This represents the history of memory 226 generated by processing engine 122 during machine deployment over a period of K minutes. To compute A, retrieval engine 124 retrieves a subset of the history. This subset includes memory 242, which is related to question 234.

[0064] Therefore, the operation of search engine 124 can be broken down into the following:

[0065] p(A∣H 1:K ,Q)=p(A∣R * ,Q)≈p(A∣R,Q) (1)

[0066] In the above equation, It is the "best" subset of memory 242, which can be used to generate the answer to problem 234. Because R cannot be computed... * Therefore, search engine 124 uses a sampling strategy F:V→R to sample a subset from the memory representation V, where F(V)={h∣h∈H} 1:K}

[0067] To estimate R * To ensure that the answers derived from R and H are consistent, retrieval engine 124 can minimize the size of R while ensuring that the answer can be predicted from both the history H and the subset R:

[0068]

[0069] Given that the history H is relatively long (e.g., hours or days), the memory representation V and the sampling strategy F make the computation more tractable.

[0070] More specifically, the retrieval engine 124 uses a machine learning model 202 to generate function calls f and queries 238q, based on existing matching memories 242R of question 234Q and one or more groups. 0: Retrieve a set of up to m matching memories 242. Each retrieved memory 226 may include location, time, descriptive text, and / or other information that can be added to R and used as additional context 236:

[0071] R i:i+m =f(q), where q = LLM(R) 0:i ,Q) (3)

[0072] In the above equation, LLM represents the operation of machine learning model 202 when generating a given query 238 based on the existing context 236 and question 234.

[0073] In some embodiments, the functions called by the machine learning model 202 include (but are not limited to) the following:

[0074] Text retrieval: f l (object)

[0075] • Location retrieval: f p (x,y,z)

[0076] • Time-based retrieval: f t ("HH:MM:SS")

[0077] During each iteration, machine learning model 202 can formulate one or more queries 238 on memory 226 to help answer question 234. Once a certain number of matching memories 242 are retrieved using query 238, machine learning model 202 evaluates whether question 234 can be answered with updated context 236 (e.g., based on instructions included in hints from retrieval engine 124).

[0078] If machine learning model 202 determines that question 234 cannot be answered using the current context 236, then machine learning model 202 uses the current context 236 and question 234 as input for the next iteration to generate one or more additional queries 238, retrieve one or more matching memories 242 for corresponding groups, and update the context 236 with the retrieved matching memories 242. If question 234 can be answered using the updated context 236, then machine learning model 202 summarizes the relevant information from context 236 and uses the summarized information to generate an answer. The output can be formatted (e.g., formatted as a JSON object) using keys for text, location, time, duration, and / or other types of answers. This structured output can be additionally used to generate navigation goals and / or perform other actions 250 related to question 234.

[0079] In one or more embodiments, machine learning model 202 includes (but is not limited to) one or more LLMs, VLMs, multimodal language models, and other methods described herein. Figures 7A to 7CThe machine learning models described herein and / or other types of machine learning models capable of processing text and / or other representations of question 234. For example, but not limited to, any of the various machine learning models and / or neural networks described herein can include any type of machine learning model, such as those using linear regression, logistic regression, decision trees, support vector machines (SVM), Naive Bayes, k-nearest neighbors (Knn) (K stands for clustering), random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoder neural networks, artificial neural networks (ANN), convolutional neural networks (CNN), recurrent neural networks (RNN), perceptrons, long / short-term memory (LSTM) networks, multilayer perceptron (MLP) networks, deep stacked networks (DSN), generative pre-trained (GPT) models or networks, feedforward networks, radial basis function ANNs, self-organizing maps (SOM), Kohonen maps, Hopfield maps, etc. One or more machine learning models and / or other types of machine learning models, including (opfield) networks, Boltzmann machines, deep belief neural networks, deconvolutional neural networks, generative adversarial networks (GANs), liquid machines, modular neural networks, sequence-to-sequence models, networks using transformer architectures, diffusion models (e.g., diffusion probability models, score-based generative models, etc.), neural rendering field (NeRF) models, Kolmogorov-Arnold networks (KANs), models with encoder-only architectures, models with decoder-only architectures, models with encoder-decoder architectures, generative machine learning models, language models, large language models (LLMs), visual language models (VLMs), multimodal language models (MMLMs), etc.

[0080] The retrieval engine 124 can input prompts into these machine learning models 202, which include a question 234 and instructions for converting the question 234 into a query 238 to a vector database and / or another type of data repository 208 where memories 226 are stored. The prompt can specify that the question 234 can be converted into a query 238 relating to text, semantic content, location, time, and / or other information included in the memories 226. The prompt can also, or alternatively, specify a role associated with retrieving matching memories 242 to answer the question 234 (e.g., “You are a five-star agent who determines whether you have enough information to answer the question based on the robot’s memories. Everything the robot sees is stored and can be queried based on the tools you have. These tools can retrieve and query your history and provide responses to the user. It is up to you to decide whether to call a retrieval function to help answer the question or a response function to provide a response.”). The prompt may also, or alternatively, include instructions for invoking tools and / or attaching machine learning models 202 to generate an embedding 240 of query 238, to retrieve memory 226 from data repository 208 based on query 238, and / or to update context 236 based on the retrieved memory 226; the type of reasoning to be applied (e.g., “contextual reasoning: step-by-step thinking about the context, summarizing the question, and whether it is sufficient to answer the user’s question”; “tool reasoning: based on contextual reasoning, deciding whether to trigger a response to the user or to invoke tools for more information”); the format of the memory 226 returned by the tool, the context 236, and / or the response to be generated based on the returned memory 226 (e.g., JavaScript Object Notation (JSON) schema); and / or other instructions related to the task of retrieving and processing memory 226 from data repository 208 that can be used to answer question 234. The hint may additionally include one or more example questions, example queries 238 generated in response to each example question, example context 236 generated and / or updated based on example queries 238, and / or example reasoning related to example queries 238 and / or context 236.

[0081] Given the input prompts, the machine learning model 202 iteratively generates a query 208 that can be used to answer question 234, uses the generated query 238 and / or corresponding embedding 240 to generate a result 206 including matching memory 242 from data repository 208 and / or reasoning associated with matching memory 242, uses the retrieved matching memory 242 to update the context 236, and determines whether the updated context 236 can be used to answer question 234.

[0082] Once machine learning model 202 determines that context 236 can be used to answer question 234, machine learning model 202 generates one or more responses that include one or more answers to question 234. For example, a first LLM, VLM, and / or other type of machine learning model that generates query 238 and updates context 236 based on corresponding matching memory 242 can invoke a second LLM, VLM, and / or other type of machine learning model to generate one or more answers. The input to the second machine learning model may include context 236 and / or other information from the first machine learning model that is considered relevant to answering question 234. The input to the second machine learning model may also include a prompt that specifies the role associated with generating an answer to question 234 (e.g., “You are a robot capable of answering specific types of questions related to your memory. As a robot, you have seen many things. The user asks you a question, and an external system will retrieve fragments of your memory as context in the form of explanatory text. The question will start from the current time and location, but the user wants to know about something in the past. Using this information, please answer the following question “{Question}””). The prompt may also, or alternatively, specify the type of answer that can be generated (e.g., text, location (x, y, z), time (in minutes), duration (in minutes), binary (yes / no), etc.); the format of the answer (e.g., JSON with specific fields / elements for the corresponding type of answer); the type of reasoning to be applied (e.g., "Type reasoning: Enter your reasoning for this type of question here", "Answer reasoning: Enter your reasoning for the answer to the question here"); and / or other instructions related to the task of answering question 234 using memory 226 in data repository 208. The prompt may also include one or more example questions, example context 236 generated based on example query 238, example reasoning related to one or more types of one or more example questions and answers to one or more example questions, and / or one or more example answers to one or more example questions. Based on the input prompt and information deemed relevant to answering question 234, the second machine learning model can generate answers in the form of text, audio, images, video, and / or other formats.

[0083] Figure 4 The illustrations show various embodiments. Figure 1 The operation of search engine 124 in generating result 206 associated with example question 234. For example... Figure 4As shown, question 234 specifies the time as 08:00:44 on January 16, 2023, the location as [-55.74, 86.31, -2.11], and the request as the nearest place to sit. Question 234 is matched against a memory that includes a representation of frame 402 of a video depicting a room with a table and chairs. For example, the memory may include other representations of embedded and / or descriptive text describing the room depicted in frame 402. The memory may also, or alternatively, include frame 402, the embedding of frame 402, and / or other representations of the image data included in frame 402.

[0084] The retrieved memories were used to generate results 206, which included answers to question 234. Results 206 included type inferences such as "The user wants to know the nearest place to sit down at the current time and location." Results 206 also included answers such as "I remember seeing a room with tables and chairs at 7:57:51 AM on January 16, 2023. This room looked like a common area or a cafeteria. Based on my current location, I can infer that this room is likely nearby." Results 206 could be additionally used to generate answers including the location of the room and / or the location associated with the memory.

[0085] Back Figure 2 The discussion proceeds as follows: after the machine learning model generates one or more answers to question 234, the retrieval engine 124 performs one or more actions 250 based on one or more answers. For example, the retrieval engine 124 may generate text, synthesize speech, one or more gestures, and / or other outputs that deliver one or more answers to the user on the machine and / or other systems. The retrieval engine 124 may also, or alternatively, generate a target, a path to the target, and / or other actions 250 to be performed based on question 234. The retrieval engine 124 may also, or alternatively, generate commands that cause the machine to navigate to the target based on the generated actions 250.

[0086] It should be understood that the arrangements and other arrangements described herein are merely illustrative examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used to supplement or replace the arrangements and elements shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that can be implemented as discrete or distributed components, or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities can be performed by hardware, firmware, and / or software. For example, a processor executing instructions stored in memory can be used to perform these functions. In some embodiments, the systems, methods, and processes described herein can use... Figures 10A-10D Example of autonomous vehicles 1000 Figure 8 Example computing devices 800 and / or Figure 9 Example data center 900 uses components, features, and / or functions similar to those of other components, features, and / or functions to perform the operation.

[0087] Now for reference Figure 5 and Figure 6 Each block of methods 500 and 600 described herein includes a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be performed using a processor that executes instructions stored in memory. These methods can also be embodied as computer-usable instructions stored on a computer storage medium. These methods can be provided by standalone applications, services, or managed services (independently or in combination with other managed services) or plug-ins of other products, to name a few. Furthermore, by way of example, regarding... Figure 1 The systems described herein are for methods 500 and 600. However, these methods may be additionally or alternatively performed by any system or any combination of systems, including but not limited to the systems described herein.

[0088] Figure 5 This is a flowchart illustrating a method 500 for generating a memory representation for a machine, according to some embodiments of the present disclosure. Figure 5 As shown, method 500 begins with operation 502, where processing engine 122 aggregates one or more sensor inputs obtained using one or more sensors of the machine into segments spanning time intervals. For example, processing engine 122 may acquire a subset of one or more streams of audio data, video data, location information, LiDAR data, RADAR data, ultrasonic data, IMU data, timestamp data, odometer data, and / or other data collected and / or generated by sensors on the machine within the time interval. Processing engine 122 may also include, or alternatively, generating and / or acquiring occupancy maps, visualizations, semantic labels (e.g., segmented maps, detected objects, boundary shapes, etc.), status data (e.g., location, heading, speed, etc.), guidance data (e.g., routes, trajectories, paths, headings, navigation targets, etc.), and / or other supplementary data derived from sensor data within the same time interval. Processing engine 122 can then store the sensor data and / or supplementary data in segments corresponding to the time intervals.

[0089] In operation 504, processing engine 122 generates descriptive text and / or embeddings of the descriptive text for the segment by executing one or more machine learning models. For example, processing engine 122 may use VLM, LLM, and / or other types of machine learning models to generate descriptive text that describes the semantic content of some or all of the data in the segment. Processing engine 122 may also use embedding models to convert the descriptive text and / or some or all of the data in the segment into embeddings in a low-dimensional latent vector space.

[0090] In operation 506, processing engine 122 stores the memories, including embeddings, descriptive text, and / or fragments, in a data repository. For example, processing engine 122 may store embeddings, descriptive text, and / or fragments in a vector database.

[0091] In operation 508, processing engine 122 determines whether to continue generating memory. For example, processing engine 122 may determine that memory generation should continue while deploying the machine and / or collecting sensor input. When processing engine 122 determines that memory generation should continue, processing engine 122 repeats operations 502, 504, and 506 to generate additional memory from fragments of sensor input (e.g., as sensor input is collected by the machine). Processing engine 122 also repeats operation 508 to determine whether to continue generating memory. Processing engine 122 may continue to generate memory for the machine in real-time or near real-time (e.g., as sensor input is generated and / or received) until the machine is no longer deployed, sensor input is no longer collected, and / or other conditions are met.

[0092] Figure 6 This is a flowchart illustrating a method 600 for performing a task using a memory representation for a machine, according to some embodiments of the present disclosure. Figure 6 As shown, method 600 begins with operation 602, where retrieval engine 124 receives a question from the user. For example, retrieval engine 124 may receive a question from the user in the form of text, audio, video, gesture-based input, tactile input, and / or other types of input. The question may include (but is not limited to) spatial questions relating to the location of a given entity of interest (e.g., the nearest chair or bathroom), temporal questions relating to the time when a given event occurs (e.g., when a pile of boxes falls) and / or the duration of an event or activity (e.g., how long the machine stayed inside the building), descriptive questions relating to details observed by the machine (e.g., the level of traffic on a street or sidewalk, the side of a street or driveway where vehicles are traveling, etc.), and / or binary questions that can be answered with "yes" or "no" (e.g., whether the machine encountered a person).

[0093] In operation 604, retrieval engine 124 generates a set of queries based on the question and / or the context associated with that question. For example, retrieval engine 124 may input the question, an empty context, and a hint instructing the machine learning model to generate queries against a data repository that can be used to answer the question into an LLM, VLM, and / or other type of machine learning model. Based on the input question, context, and hints, the machine learning model may generate one or more queries that can be used to retrieve text, location information, time information, and / or other information related to the question.

[0094] In operation 606, retrieval engine 124 uses queries to retrieve a set of memories from a data repository. For example, retrieval engine 124 can use the embedding representation of each query to perform a lookup on a vector database. This lookup can be used to retrieve one or more memories that are semantically similar to the query.

[0095] In operation 608, retrieval engine 124 uses the retrieved memory to update the context. For example, retrieval engine 124 may use LLM, VLM and / or other types of machine learning models to add information summaries from the retrieved memory, key fragments of data from the retrieved memory (e.g., location, time, duration, etc.) and / or other information from the retrieved memory that is considered relevant to the question to the context.

[0096] In operation 610, retrieval engine 124 determines whether context can be used to answer the question. For example, retrieval engine 124 may provide hints about LLM, VLM, and / or another machine learning model to assess whether context can be used to answer the question.

[0097] If retrieval engine 124 determines in operation 610 that the context cannot be used to answer the question, retrieval engine 124 repeats operations 604, 606, and 608 to generate additional queries and also updates the context using the memories retrieved from the data repository using the additional queries. After updating the context with a new set of additional memories, retrieval engine 124 repeats operation 610 to reassess whether the context can be used to answer the question.

[0098] If retrieval engine 124 determines in operation 610 that the question can be answered using context, then retrieval engine performs operation 612, in which retrieval engine 124 outputs an answer and / or performs other actions based on the context. For example, retrieval engine 124 may translate the context into a conversational answer and deliver the conversational answer in the form of text, audio, images, gestures, and / or other outputs. Retrieval engine 124 may also, or alternatively, generate a navigation target associated with the conversational answer, a trajectory for the machine to move toward the navigation target, commands to navigate the machine to the navigation target, and / or other outputs that cause the machine to perform one or more actions related to the conversational answer.

[0099] Example language model

[0100] In at least some embodiments, language models such as Large Language Models (LLM), Visual Language Models (VLM), Multimodal Language Models (MMLM), and / or other types of generative artificial intelligence (AI) can be implemented. These models may be able to understand, summarize, translate, and / or otherwise generate text (e.g., natural language text, code, etc.), images, videos, computer-aided design (CAD) assets, OMNIVERSE and / or METAVERSE file information (e.g., USD formats such as OpenUSD), and / or the like based on context provided in input prompts or queries. In embodiments, these language models may be considered “large” because they are trained on massive datasets and have architectures with a large number of learnable network parameters (weights and biases)—e.g., millions or billions of parameters. LLM / VLM / MMLM / etc. can be implemented for summarizing textual data, analyzing data (e.g., text, images, videos, etc.), extracting insights from data (e.g., text, images, videos, etc.), and generating new text / images / videos / etc. in a user-specified style, tone, and / or format. In some embodiments, the LLM / VLM / MMLM / etc. disclosed herein may be specifically designed for text processing, while in others, a multimodal LLM may be implemented to accept, understand, and / or generate text and / or other types of content, such as images, audio, 2D and / or 3D data (e.g., USD format) and / or video. For example, a Visual Language Model (VLM) or more specifically a Multimodal Language Model (MMLM) may be implemented to accept images, video, audio, text, 3D designs (e.g., CAD) and / or other input data types and / or generate or output images, video, audio, text, 3D designs and / or other output data types.

[0101] Various types of LLM / VLM / MMLM / etc. architectures can be implemented in various embodiments. For example, different architectures can be implemented using different techniques to understand and generate outputs (e.g., text, audio, video, images, 2D and / or 3D design or asset data, etc.). In some embodiments, LLM / VLM / MMLM / etc. architectures (e.g., recurrent neural networks (RNNs) or long short-term memory networks (LSTMs)) can be used, while in other embodiments, converter architectures (e.g., architectures relying on self-attention and / or cross-attention (e.g., between contextual data and textual data) mechanisms) can be used to understand and recognize relationships between words or tokens and / or contextual data (e.g., other text, video, images, design data, USD, etc.). One or more generative processing pipelines including LLM / VLM / MMLM / etc. may also include one or more diffusion blocks (e.g., noise reduction blocks). The LLM / VLM / MMLM / etc. of this disclosure may include encoder and / or decoder blocks. For example, discriminative or encoder-only models (e.g., BERT (Bidirectional Encoder Representations from Transformers)) can be implemented for tasks involving language understanding (e.g., classification, sentiment analysis, question answering, and named entity recognition). As another example, generative or decoder-only models (e.g., GPT (Generative Pretrained Transformer)) can be implemented for tasks involving language and content generation (e.g., text completion, story generation, and dialogue generation). LLM / VLM / MMLM / etc., including encoder and decoder components (e.g., T5 (Text-to-Text Transformer)), can be implemented to understand and generate content, such as for translation and summarization. These examples are not intended to be limiting and any architecture type (including, but not limited to, those described herein) can be implemented depending on the specific implementation and the task performed using LLM / VLM / MMLM / etc.

[0102] In various embodiments, LLM / VLM / MMLM / etc. can be trained using unsupervised learning, whereby LLM / VLM / MMLM / etc. learns patterns from a large amount of unlabeled text / audio / video / image / design / USD / etc. data. Due to the extensive training, in these embodiments, the model may not require task-specific or domain-specific training. An LLM / VLM / MMLM / etc. extensively pre-trained on a large amount of unlabeled data can be referred to as a base model and can excel at various tasks, such as question answering, summarizing, filling in missing information, translation, and image / video / design / USD / data generation. Some LLM / VLM / MMLM / etc. can be customized for specific use cases using techniques such as cue tuning, fine-tuning, retrieval augmentation generation (RAG), adding adapters (e.g., custom neural networks and / or neural network layers to tune or adjust cues or labels to bias the language model towards a specific task or domain), and / or using optimization models for specific tasks and / or other fine-tuning or customization techniques within a specific domain.

[0103] In some embodiments, the LLM / VLM / MMLM / etc. disclosed herein can be implemented using various model alignment techniques. For example, in some embodiments, guardrails can be implemented to identify incorrect or unwanted inputs (e.g., prompts) and / or outputs of the model. In this process, the system can use guardrails and / or other model alignment techniques to prevent the processing of specific unwanted inputs using LLM / VLM / MMLM / etc., and / or to prevent the output or presentation of information generated by LLM / VLM / MMLM / etc. (e.g., displays, audio outputs, etc.). In some embodiments, one or more additional models (or layers thereof) can be implemented to identify problems with the model's inputs and / or outputs. For example, these "protective" models can be trained to identify "safe" or otherwise okay or desired inputs and / or outputs and / or "unsafe" or otherwise unwanted inputs and / or outputs for a particular application / implementation. Therefore, the LLM / VLM / MMLM / etc. disclosed herein are unlikely to output language / text / audio / video / design data / USD data / etc. that may be offensive, vulgar, inappropriate, insecure, out of scope, and / or unwanted for a particular application / implementation.

[0104] In some embodiments, an LLM / VLM / etc. can be configured or able to access or use one or more plugins, application programming interfaces (APIs), databases, data stores, repositories, etc. For example, for certain tasks or operations where the model is not ideally suited, the model may have instructions for accessing one or more plugins (e.g., third-party plugins) to help process the current input (e.g., as a result of training, and / or based on instructions in a given prompt). In such an example, when at least part of the prompt relates to restaurants or weather, the model can access one or more restaurant or weather plugins (e.g., via one or more APIs) to retrieve relevant information. Another example is that if at least part of the response requires mathematical computation, the model can access one or more mathematical plugins or APIs to help solve the problem, and then the response from the plugins and / or APIs can be used in the model's output. This process can be repeated (e.g., recursively) an arbitrary number of iterations, using any number of plugins and / or APIs, until a response to each query / question / request / process / action / etc. can be generated in response to the input prompt. Therefore, models can rely not only on their own knowledge gained from training on large datasets, but also on the expertise or optimized properties of one or more external resources (such as APIs, plugins, etc.).

[0105] In some embodiments, multiple language models (e.g., LLM / VLM / MMLM / etc., multiple instances of the same language model, and / or multiple hints provided to the same language model or instances of the same language model) can be implemented, executed, or accessed (e.g., using one or more plugins, user interfaces, APIs, databases, data stores, repositories, etc.) to provide output in response to the same query or in response to separate parts of a query. In at least one embodiment, the same input query and hints (e.g., a set of constraints, condition generators, etc.) can be provided to multiple language models (e.g., language models with different architectures, language models trained on different (e.g., updated) data corpora). In one or more embodiments, the language models can be different versions of the same base model. In one or more embodiments, at least one language model can be instantiated as multiple agents, for example, providing more than one hint to constrain, guide, or otherwise influence the style, content, or character of the provided output. In one or more exemplary non-limiting embodiments, the same language model can be required to provide output corresponding to different roles, perspectives, characters, or different knowledge bases, as defined by the provided hints.

[0106] In any such embodiment, the outputs of two or more (e.g., each) language models, two or more versions of at least one language model, two or more instantiated proxies of at least one language model, and / or provided to two or more prompts for at least one language model can be further processed, such as aggregated, compared, or filtered, or used to determine (and provide) a consensus response. In one or more embodiments, the output from one language model (or version, instance, or proxy) can be provided as input to another language model for further processing and / or validation. In one or more embodiments, the language model can be required to generate or otherwise obtain output about the input source material, wherein the output is associated with the input source material. This association may include, for example, generating captions or text portions embedded (e.g., as metadata) within the input source text or image. In one or more embodiments, the output of the language model can be used to determine the validity of the input source material for further processing or inclusion in a dataset. For example, the language model can be used to evaluate the presence (or absence) of a target word in a text portion or the presence (or absence) of an object in an image, wherein the text or image is annotated to indicate such presence (or absence). Alternatively, the determination from the language model can be used to determine whether the source material should be included in the curatorial dataset, for example, but not limited to this.

[0107] In some embodiments, one or more converter engines (TEs) can be implemented. Converter engines can use micro-tensor scaling to optimize performance and accuracy, such as enabling 16-bit floating-point (FP16), 8-bit floating-point (FP8), and / or 4-bit floating-point (FP4) AI processing. For example, a converter engine can use 16-bit or 8-bit floating-point precision and 8-bit or 4-bit floating-point data formats, combined with software algorithms, to improve AI performance and capabilities. By reducing mathematical operations to 8 bits or 4 bits, TEs can train larger networks faster without compromising accuracy. For example, TEs can include libraries for accelerating converter models on processing devices such as GPUs to provide better performance in training and inference with lower memory utilization. When TEs are combined with other technologies, such as high-speed interconnects between nodes (e.g., using NVLink switches) and tensor cores (which enable mixed-precision computation, such as microscale precision support), server clusters can be more capable of training large networks at high speeds. Therefore, it can support tensor core precision for FP64, TF32, BF16, FP16, FP8, INT8, FP6 and FP4, as well as CUDA core precision for FP64, FP32, FP16 and BF16.

[0108] Figure 7A This is a block diagram of an example generative language model system 700 suitable for implementing at least some embodiments of the present disclosure. Figure 7A In the example shown, the generative language model system 700 includes a retrieval-enhanced generation (RAG) component 792, an input processor 705, a tokenizer 710, an embedding component 720, a plug-in / API 795, and a generative language model (LM) 730 (which may include LLM, VLM, multimodal LM, etc.).

[0109] At a high level, the input processor 705 can receive input 701, which includes text and / or other types of input data (e.g., audio data, video data, image data, sensor data (e.g., LiDAR, RADAR, ultrasound, etc.), 3D design data, CAD data, generic scene descriptor (USD) data (e.g., OpenUSD, etc.), depending on the architecture of the generative LM 730 (e.g., LLM / VLM / MMLM, etc.). In some embodiments, input 701 includes plain text in the form of one or more sentences, paragraphs, and / or documents. Additionally or alternatively, input 701 may include numerical sequences, pre-computed embeddings (e.g., word or sentence embeddings), and / or structured data (e.g., tabular format, JSON, or XML). In the generative LM In some implementations of 730 capable of handling multimodal input, input 701 can combine text (or text that may be omitted) with image data, audio data, video data, design data, USD data, and / or other types of input data (e.g., but not limited to the data described herein). Taking raw input text as an example, input processor 705 can prepare the raw input text in various ways. For example, input processor 705 can perform various types of text filtering to remove noise from relevant text content (e.g., special characters, punctuation marks, HTML tags, stop words, portions of images, portions of audio, etc.). In examples involving stop words (common words that often have little semantic meaning), input processor 705 can remove stop words to reduce noise and allow the generative LM 730 to focus on more meaningful content. Input processor 705 can apply text normalization, for example, by converting all characters to lowercase, removing accent marks, and / or handling special cases (such as abbreviations or abbreviations) to ensure consistency. These are just a few examples; other types of input processing can be applied.

[0110] In some embodiments, RAG component 792 (which may include one or more RAG models, and / or may be performed using generative LM 730 itself) may be used to retrieve additional information to be used as part of input 701 or a prompt. RAGs can be used to enhance input to LLM / VLM / MMLM / etc. with external knowledge to make the answer to a specific question or query or request more relevant, for example, where specific knowledge is required. RAG component 792 may obtain this additional information (e.g., basic information such as basic text / images / videos / audio / USD / CAD / etc.) from one or more external sources, and then feed it along with the prompt to LLM / VLM / MMLM / etc. to improve the accuracy of the model's response or output.

[0111] For example, in some embodiments, in addition to the data retrieved using the RAG component 792, input 701 may also be generated using query or model inputs (e.g., questions, requests, etc.). In some embodiments, the input processor 705 may analyze the input 701 and communicate with the RAG component 792 (or in some embodiments, the RAG component 792 may be part of the input processor 705) to identify relevant text and / or other data to provide to the generative LM 730 as additional context or information sources, typically from which responses, answers, or outputs 790 are identified. For example, when the input indicates that a user is interested in the required tire pressure for a particular brand and model of vehicle, the RAG component 792 may use the RAG model, for example, to perform a vector search in the embedding space to retrieve tire pressure information or its corresponding text from a digital (embedded) version of the owner's manual for that particular vehicle brand and model. Similarly, when a user revisits the chatbot related to a specific product sale or service, the RAG component 792 can retrieve previously stored conversation history (or at least its summary) and provide the previous conversation history, along with the current inquiry / request, as part of the generative LM 730 as input 701.

[0112] RAG component 792 can use various RAG techniques. For example, it can use naive RAG ( The document is indexed, chunked, and applied to an embedding model to generate embeddings corresponding to chunks. User queries can also be applied to this embedding model and / or another embedding model of the RAG component 792, and the embeddings of the chunks can be compared with the embeddings of the query to identify the most similar / relevant embeddings to the query. These most similar / relevant embeddings can be provided to the generative LM 730 to generate output.

[0113] In some embodiments, more advanced RAG techniques can be used. For example, chunks can undergo pre-retrieval processes (e.g., routing, rewriting, metadata analysis, expansion, etc.) before being passed to the embedding model. Furthermore, post-retrieval processes (e.g., re-ranking, hint compression, etc.) can be performed on the output of the embedding model before generating the final embedding, which is then used for comparison with the input query.

[0114] As a further example, modular RAG techniques can be used, such as those similar to Naive RAG and / or Advanced RAG, but also including features such as hybrid search, recursive retrieval and query engines, StepBack methods, subqueries and hypothetical document embeddings.

[0115] As another example, Graph RAG can use a knowledge graph as a source of context or factual information. Graph RAG can be implemented using a graph database as a source of contextual information sent to LLM / VLM / MMLM / etc. Instead of providing the model with data chunks extracted from larger documents (which may result in a lack of context, factual accuracy, linguistic accuracy, etc.) (or anything other than providing the model with data chunks extracted from larger documents), Graph RAG can also provide structured entity information to LLM / VLM / MMLM / etc. by combining structured entity text descriptions with their many attributes and relationships, thus giving the model deeper insights. In implementing Graph RAG, the systems and methods described herein use graphs as content stores and extract relevant document chunks, requiring LLM / VLM / MMLM / etc. to use them to answer questions. In such embodiments, the knowledge graph may contain relevant textual content and metadata about the knowledge graph, or it may be integrated with a vector database. In some embodiments, Graph RAG can use the graph as a subject matter expert, where descriptions of concepts and entities relevant to the query / hint can be extracted and passed to the model as semantic context. These descriptions may include relationships between concepts. In other examples, the graph can be used as a database where a portion of a query / hint can be mapped to a graph query, the graph query can be executed, and LLM / VLM / MMLM / etc. can aggregate the results. In such examples, the graph can store relevant factual information and can be used for queries (natural language queries) and entity links to graph query tools (NL to graph query tools). In some embodiments, the graph RAG (e.g., using a graph database) can be combined with standard (e.g., vector database) RAGs and / or other RAG types to benefit from a variety of approaches.

[0116] In any embodiment, the RAG component 792 can implement plugins, APIs, user interfaces, and / or other functions to perform RAG. For example, LLM / VLM / MMLM / etc. can use graph RAG plugins to run queries on knowledge graphs to extract relevant information to feed into the model, and can use standard or vector RAG plugins to run queries on vector databases. For example, the graph database can interact with the plugin's REST interface, thus decoupling the graph database from the vector database and / or the embedded model.

[0117] The tokenizer 710 can segment (e.g., processed) text data into smaller units (tags) for subsequent analysis and processing. Depending on the implementation, the tags can represent individual words, sub-words, characters, audio / video / images, etc. Word-based tokenization divides the text into individual words, treating each word as a separate tag. Sub-word tokenization breaks words down into smaller meaningful units (e.g., prefixes, suffixes, stems), enabling the generative LM 730 to understand morphological changes and process words outside the vocabulary more effectively. Character-based tokenization represents each character as a separate tag, enabling the generative LM 730 to process text at a fine-grained level. The choice of tokenization strategy can depend on factors such as the language being processed, the task at hand, and / or the characteristics of the training dataset. Therefore, the tokenizer 710 can transform (e.g., processed) text into a structured format according to the tokenization scheme implemented in a particular embodiment.

[0118] Embedding component 720 can use any known embedding technique to transform discrete tokens into semantically meaningful (e.g., dense, continuous vector) representations. For example, embedding component 720 can use pre-trained word embeddings (e.g., Word2Vec, GloVe, or FastText), one-hot encoding, Term Frequency-Inverse Document Frequency (TF-IDF) encoding, one or more embedding layers of a neural network, and / or others.

[0119] In some implementations where input 701 includes image data / video data, etc., input processor 701 may resize the data to a standard size compatible with the format of the corresponding input channel and / or normalize pixel values ​​to a common range (e.g., 0 to 1) to ensure consistent representation, and embedding component 720 may encode the image data using any known technique (e.g., using one or more convolutional neural networks (CNNs) to extract visual features). In some implementations where input 701 includes audio data, input processor 701 may resample the audio file to a consistent sampling rate for uniform processing, and embedding component 720 may use any known technique to extract and encode audio features, such as in the form of a spectrogram (e.g., a Mel spectrogram). In some implementations where input 701 includes video data, input processor 701 may extract frames or apply resizing to extracted frames, and embedding component 720 may extract features such as optical flow embedding or video embedding and / or encode temporal information or frame sequences. In some implementations where input 701 includes multimodal data, the embedded component 720 can use techniques such as early fusion (concatenation), late fusion (sequential processing), and attention-based fusion (e.g., self-attention, cross-attention) to fuse representations of different types of data (e.g., text, images, audio, data, video, design, etc.).

[0120] Other components of the generative LM 730 and / or generative LM system 700 may use different types of neural network architectures depending on the implementation scheme. For example, a transducer-based architecture (e.g., the architecture used in models such as GPT) may be implemented, and it may include a self-attention mechanism that weights the importance of different words or tokens in the input sequence and / or a feedforward network that processes the output of the self-attention layer, applying a nonlinear transformation to the input representation and extracting higher-level features. Some non-limiting example architectures include transducers (e.g., encoder-decoder, decoder-only, multimodal), RNNs, LSTMs, fusion models, diffusion models, cross-modal embedding models that learn a joint embedding space, graph neural networks (GNNs), hybrid architectures that combine different types of adversarial networks (such as generative adversarial networks or GANs or adversarial autoencoders (AAEs) for joint distribution learning), etc. Therefore, depending on the implementation scheme and architecture, the embedded component 720 can apply the encoded representation of the input 701 to the generative LM 730, and the generative LM 730 can process the encoded representation of the input 701 to generate an output 790, which may include response text and / or other types of data.

[0121] As described herein, in some embodiments, the generative LM 730 may be configured to access or use (or be able to access or use) plugins / APIs 795 (which may include one or more plugins, application programming interfaces (APIs), databases, data stores, repositories, etc.). For example, for certain tasks or operations where the generative LM 730 is not ideally suited, the model may have instructions (e.g., as a result of training, and / or based on instructions in a given prompt, such as instructions retrieved using RAG component 792) to access one or more plugins / APIs 795 (e.g., third-party plugins) to help process the current input. In such an example, when at least part of the prompt is related to a restaurant or weather, the model may access one or more restaurant or weather plugins (e.g., via one or more APIs), sending at least part of the prompt related to a particular plugin / API 795 to the plugin / API 795, which can process the information and return an answer to the generative LM 730, which can then use the response to generate output 790. This process can be repeated (e.g., recursively) an arbitrary number of iterations and repeated with an arbitrary number of plugins / APIs 795 until an output 790 that resolves each query / question / request / process / action / etc. from input 701 is generated. Therefore, the model can rely not only on its own knowledge acquired from training on a large dataset and / or from data retrieved using the RAG component 792, but also on the expertise or optimized properties of one or more external resources (e.g., plugins / APIs 795).

[0122] Figure 7B This is a block diagram of an example implementation scheme, where the generative LM 730 includes a converter encoder-decoder. For example, suppose the input text (e.g., “Who discovered gravity”) is tokenized (e.g., by...) Figure 7A The tokenizer 710) is used for tokens such as words, and each token is encoded (e.g., by...). Figure 7A The embedding component 720 is a corresponding embedding (e.g., of size 512). Since these token embeddings do not typically represent the position of the tokens in the input sequence, positional encoding can be added to each token embedding using any known technique to encode the order relation and context of the tokens in the input sequence. Thus, (e.g., the resulting) embeddings can be applied to one or more encoders 735 of the generative LM 730.

[0123] In the example implementation, encoder 735 forms an encoder stack, where each encoder includes a self-attention layer and a feedforward network. In the example converter architecture, each token (e.g., a word) flows through a separate path. Therefore, each encoder can accept a sequence of vectors, pass each vector through the self-attention layer, then through the feedforward network, and then up to the next encoder in the stack. Any known self-attention technique can be used. For example, to compute a self-attention score for each token (word), a query vector, a key vector, and a value vector can be created for each token. The self-attention score for a token pair can be computed by taking the dot product of the query vector and the corresponding key vector, normalizing the resulting score, multiplying by the corresponding value vector, and summing the weighted value vectors. The encoder can apply multi-head attention, where the attention mechanism is applied multiple times in parallel with different learned weight matrices. Any number of encoders can be cascaded to generate a context vector encoding the input. Attention projection layer 740 can transform the context vector into attention vectors (keys and values) for decoder 745.

[0124] In the example implementation, decoder 745 forms a decoder stack, where each decoder includes a self-attention layer, an encoder-decoder self-attention layer that uses attention vectors (keys and values) from the encoder to focus on relevant parts of the input sequence, and a feedforward network. Similar to encoder 735, in the example converter architecture, each token (e.g., a word) flows through a separate path in decoder 745. During the first pass, decoder 745, classifier 750, and generation mechanism 755 can generate a first token, and generation mechanism 755 can apply the generated token as input during a second pass. This process can be repeated cyclically, generating tokens (e.g., words) and adding them to the output of the previous pass, and in subsequent passes applying token embeddings of positionally encoded composite sequences as input to decoder 745, generating one token at a time (called autoregression) until a symbol or token representing the end of the response is predicted. In each decoder, the self-attention layer is typically restricted to focusing only on earlier positions in the output sequence by applying a masking technique (e.g., setting future positions to negative infinity) before the softmax operation. In the example implementation, the encoder-decoder attention layer operates similarly to the (e.g., multi-head) self-attention operation in encoder 735, except that it creates its queries from the layers below it and obtains keys and values ​​(e.g., matrices) from the output of encoder 735.

[0125] Therefore, decoder 745 can output some decoded (e.g., vector) representation of the input applied during a particular pass. Classifier 750 can include a multi-class classifier comprising one or more neural network layers and a softmax operation that transforms logit probabilities into probabilities, the neural network layers projecting the decoded (e.g., vector) representation onto corresponding dimensions (e.g., one dimension for each supported word or token in the output vocabulary). Thus, generation mechanism 755 can select or sample words or tokens based on corresponding predicted probabilities (e.g., selecting the word with the highest predicted probability) and append it to the output of the previous pass, thereby generating each word or token sequentially. Generation mechanism 755 can repeat this process, triggering successive decoder inputs and corresponding predictions until a symbol or token representing the end of the response is selected or sampled, at which point generation mechanism 755 can output the generated response.

[0126] Figure 7C This is a block diagram of an example implementation where the generative LM 730 includes a decoder-only converter architecture. For example, Figure 7C The decoder 760 can be used with Figure 7B The decoder 745 operates similarly, except... Figure 7C Each decoder 760 omits the encoder-decoder self-attention layer (because there is no encoder in this implementation). Therefore, decoders 760 can form a decoder stack, where each decoder includes a self-attention layer and a feedforward network. Furthermore, instead of encoding the input sequence, a symbol or tag indicating the end of the input sequence (or the beginning of the output sequence) can be appended to the input sequence, and the resulting sequence (e.g., a corresponding embedding with positional encoding) can be applied to decoder 760. Figure 7B Similar to decoder 745, each tag (e.g., a word) can flow through a separate path in decoder 760, and decoder 760, classifier 765, and generation mechanism 770 can use autoregression to generate one tag at a time sequentially until a symbol or tag indicating the end of the response is predicted. Classifier 765 and generation mechanism 770 can be combined with... Figure 7B The classifier 750 and the generation mechanism 755 operate similarly, wherein the generation mechanism 770 selects or samples each consecutive output label based on the corresponding predicted probability and appends it to the output of the previous iteration, generating each label sequentially until a symbol or label representing the end of the response is selected or sampled. The architectures described herein, and others, are merely examples, and other suitable architectures may be implemented within the scope of this disclosure.

[0127] Example computing device

[0128] Figure 8This is a block diagram of an example computing device 800 suitable for implementing some embodiments of the present disclosure. The computing device 800 may include an interconnect system 802 directly or indirectly coupled to: a memory 804, one or more central processing units (CPUs) 806, one or more graphics processing units (GPUs) 808, a communication interface 810, input / output (I / O) ports 812, input / output components 814, a power supply 816, one or more presentation components 818 (e.g., one or more displays), and one or more logic units 820. In at least one embodiment, one or more computing devices 800 may include one or more virtual machines (VMs), and / or any component thereof may include virtual components (e.g., virtual hardware components). For a non-limiting example, one or more GPUs 808 may include one or more vGPUs, one or more CPUs 806 may include one or more vCPUs, and / or one or more logic units 820 may include one or more virtual logic units. Thus, one or more computing devices 800 may include discrete components (e.g., a full GPU dedicated to computing device 800), virtual components (e.g., a portion of the GPU dedicated to computing device 800), or a combination thereof.

[0129] although Figure 8 The various blocks are shown as connected via interconnect system 802 using lines, but this is not intended to be limiting and is merely for clarity. For example, in some embodiments, presentation component 818 (such as a display device) may be considered I / O component 814 (e.g., if the display is a touchscreen). As another example, CPU 806 and / or GPU 808 may include memory (e.g., memory 804 may represent a storage device other than the memory of GPU 808, CPU 806, and / or other components). Therefore, Figure 8 The computing devices described are for illustrative purposes only. No distinction is made between such categories as “workstation,” “server,” “laptop computer,” “desktop computer,” “tablet computer,” “client device,” “mobile device,” “handheld device,” “game console,” “electronic control unit (ECU),” “virtual reality system,” and / or other device or system types, as all are considered within the scope of… Figure 8 Within the scope of computing devices.

[0130] Interconnect system 802 may represent one or more links or buses, such as address buses, data buses, control buses, or combinations thereof. Interconnect system 802 may include one or more bus or link types, such as Industry Standard Architecture (ISA) bus, Extended Industry Standard Architecture (EISA) bus, Video Electronics Standards Association (VESA) bus, Peripheral Component Interconnect (PCI) bus, Fast Peripheral Component Interconnect (PCIe) bus, and / or another type of bus or link. In some embodiments, there is a direct connection between components. As an example, CPU 806 may be directly connected to memory 804. Further, CPU 806 may be directly connected to GPU 808. In cases where there is a direct or point-to-point connection between components, interconnect system 802 may include a PCIe link to perform the connection. In these examples, a PCI bus is not required to be included in computing device 800.

[0131] The memory 804 may include any computer-readable medium from a variety of computer-readable media. A computer-readable medium may be any available medium accessible by the computing device 800. Computer-readable media may include volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, computer-readable media may include computer storage media and communication media.

[0132] Computer storage media may include volatile and non-volatile media and / or removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 804 may store computer-readable instructions (e.g., representing one or more programs and / or one or more program elements, such as an operating system). Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, Digital Universal Disc (DVD) or other optical disc storage, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible by computing device 800. As used herein, computer storage media does not include the signal itself.

[0133] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in modulated data signals such as carrier waves or other transmission mechanisms, and includes any information transmission medium. The term "modulated data signal" can refer to a signal whose one or more characteristics are set or altered in a manner that encodes information in the signal. By way of example and not limitation, computer storage media can include wired media (such as wired networks or direct wired connections) and wireless media (such as acoustic, RF, infrared, and other wireless media). Any combination of the above should also be included within the scope of computer-readable media.

[0134] CPU 806 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 800 to perform one or more of the methods and / or processes described herein. For example, CPU 806 may be configured to execute Figure 1 The CPU 806 may contain a processing engine 122 and / or a retrieval engine 124. Each CPU 806 may contain one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of handling numerous software threads simultaneously. The CPU 806 may contain any type of processor and may contain different types of processors depending on the type of computing device 800 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of computing device 800, the processor may be an advanced RISC machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors (such as math coprocessors), the computing device 800 may also include one or more CPUs 806.

[0135] In addition to or in lieu of one or more CPUs 806, one or more GPUs 808 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 800 to perform one or more of the methods and / or processes described herein. For example, GPU 808 may be configured to execute Figure 1The processing engine 122 and / or retrieval engine 124. One or more GPUs 808 may be integrated GPUs (e.g., one or more CPUs 806) and / or one or more GPUs 808 may be discrete GPUs. In embodiments, one or more GPUs 808 may be coprocessors of one or more CPUs 806. GPUs 808 may be used by computing device 800 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, GPUs 808 may be used for general-purpose computing on a GPU (GPGPU). GPUs 808 may include hundreds or thousands of cores capable of handling hundreds or thousands of software threads simultaneously. GPUs 808 may generate pixel data of an output image in response to rendering commands (e.g., rendering commands received from CPUs 806 via a host interface). GPUs 808 may include graphics memory (e.g., display memory) for storing pixel data or any other suitable data (e.g., GPGPU data). Display memory may be included as part of memory 804. GPUs 808 may include two or more GPUs operating in parallel (e.g., via links). The link can be directly connected to the GPU (e.g., using NVLINK) or connected via a switch (e.g., using NVSwitch). When combined, each GPU 808 can generate different portions of pixel data or GPGPU data for different outputs (e.g., a first GPU for a first image and a second GPU for an analog image). Each GPU can contain its own memory or can share memory with other GPUs.

[0136] In addition to or in lieu of CPU 806 and / or GPU 808, logic unit 820 may be configured to execute at least some of computer-readable instructions to control one or more components of computing device 800 to perform one or more of the methods and / or processes described herein. In embodiments, one or more CPUs 806, one or more GPUs 808, and / or one or more logic units 820 may execute any combination of methods, processes, and / or portions thereof discretely or jointly. One or more logic units 820 may be a portion of one or more CPUs 806 and / or GPUs 808 and / or integrated into one or more CPUs 806 and / or GPUs 808, and / or one or more logic units 820 may be discrete components or otherwise external to CPUs 806 and / or GPUs 808. In embodiments, one or more logic units 820 may be coprocessors of one or more CPUs 806 and / or GPUs 808.

[0137] Examples of logic unit 820 include one or more processing cores and / or components thereof, such as a data processing unit (DPU), tensor core (TC), tensor processing unit (TPU), pixel vision core (PVC), vision processing unit (VPU), graphics processing cluster (GPC), texture processing cluster (TPC), streaming multiprocessor (SM), tree lateral unit (TTU), artificial intelligence accelerator (AIA), deep learning accelerator (DLA), programmable vision accelerator (PVA) (which may include one or more direct memory access (DMA) systems), one or more vision or vector processing units (VPU), and one or more pixel processing engines (PPE) (e.g.) Examples include 2D arrays of processing elements (each of which communicates north, south, east, and west with one or more other processing elements in the array), one or more decoupled accelerators or units (e.g., decoupled lookup table (DLUT) accelerators or units), vision processing units (VPUs), optical flow accelerators (OFAs), field-programmable gate arrays (FPGAs), neuromorphic chips, quantum processing units (QPUs), associative processing units (APUs), arithmetic logic units (ALUs), application-specific integrated circuits (ASICs), floating-point units (FPUs), input / output (I / O) elements, peripheral component interconnects (PCIs) or fast peripheral component interconnects (PCIe) elements, etc.

[0138] The communication interface 810 may include one or more receivers, transmitters, and / or transceivers enabling the computing device 800 to communicate with other computing devices via electronic communication networks (including wired and / or wireless communications). The communication interface 810 may include components and functions for enabling communication over any of a plurality of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or wirelessband), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, the logic unit 820 and / or the communication interface 810 may include one or more data processing units (DPUs) to directly transmit data received via a network and / or via interconnect system 802 to one or more GPUs 808 (e.g., memory of one or more GPUs 808).

[0139] I / O port 812 enables computing device 800 to be logically coupled to other devices including I / O component 814, one or more presentation components 818, and / or other components, some of which may be built into (e.g., integrated into) computing device 800. Illustrative I / O component 814 includes microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dish antennas, scanners, printers, wireless devices, etc. I / O component 814 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological input generated by the user. In some cases, input may be transmitted to appropriate network elements for further processing. NUI can implement any combination of voice recognition, pen recognition, facial recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition (as described in more detail below) associated with the display of computing device 800. Computing device 800 may include depth cameras for gesture detection and recognition, such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof. Additionally, the computing device 800 may include an accelerometer or gyroscope (e.g., as part of an inertial measurement unit (IMU)) that enables motion detection. In some examples, the computing device 800 may use the output of the accelerometer or gyroscope to render immersive augmented reality or virtual reality.

[0140] The power supply 816 may include a hardwired power supply, a battery power supply, or a combination thereof. The power supply 816 may provide power to the computing device 800 so that the components of the computing device 800 can operate.

[0141] The presentation component 818 may include a display (e.g., a monitor, touchscreen, television screen, head-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. The presentation component 818 may receive data from other components (e.g., GPU 808, CPU 806, etc.) and output the data (e.g., as images, videos, sounds, etc.).

[0142] Example Data Center

[0143] Figure 9 An example data center 900 that may be used in at least one embodiment of this disclosure is shown. The data center 900 may include a data center infrastructure layer 910, a framework layer 920, a software layer 930, and / or an application layer 940.

[0144] like Figure 9As shown, the data center infrastructure layer 910 may include a resource coordinator 912, grouped computing resources 914, and node computing resources (“nodes CRs”) 916(1)-916(N), where “N” represents any complete positive integer. In at least one embodiment, nodes CRs 916(1)-916(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules and / or cooling modules, etc. In some embodiments, one or more node CRs from nodes CRs 916(1)-916(N) may correspond to servers having one or more of the aforementioned computing resources. In addition, in some embodiments, nodes CRs 916(1)-916(N) may include one or more virtual components, such as vGPU, vCPU, etc., and / or one or more nodes CRs 916(1)-916(N) may correspond to virtual machines (VMs).

[0145] In at least one embodiment, the grouped computing resources 914 may include individual groups of node CRs 916 housed within one or more racks (not shown), or multiple racks housed within a data center in different geographical locations (also not shown). Individual groups of node CRs 916 within the grouped computing resources 914 may include grouped computing, networking, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several node CRs 916, including CPUs, GPUs, DPUs, and / or other processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.

[0146] Resource coordinator 922 may be configured or otherwise control one or more nodes CRs916(1)-916(N) and / or grouped computing resources 914. In at least one embodiment, resource coordinator 922 may include a Software Design Infrastructure (“SDI”) management entity for data center 900. Resource coordinator 922 may include hardware, software, or some combination thereof.

[0147] In at least one embodiment, such as Figure 9As shown, framework layer 920 may include job scheduler 928, configuration manager 934, resource manager 936, and / or distributed file system 938. Framework layer 920 may include a framework of software 932 supporting software layer 930 and / or one or more applications 942 of application layer 940. Software 932 or application 942 may respectively contain web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. Framework layer 920 may be, but is not limited to, free and open-source software web application frameworks (such as Apache Spark) that can utilize distributed file system 938 for large-scale data processing (e.g., "big data"). TM (Hereinafter referred to as "Spark") is a type of resource. In at least one embodiment, the job scheduler 928 may include Spark drivers to facilitate the scheduling of workloads supported by different layers of data center 900. The configuration manager 934 may be able to configure different layers, such as the software layer 930 and the framework layer 920 (which includes Spark and a distributed file system 938 for supporting large-scale data processing). The resource manager 936 may be able to manage compute resources mapped to or allocated to the distributed file system 938 and the job scheduler 928, or to clusters or groups of resources allocated to support the distributed file system 938 and the job scheduler 928. In at least one embodiment, the clustered or grouped compute resources may include grouped compute resources 914 in the data center infrastructure layer 910. The resource manager 936 may coordinate with the resource coordinator 912 to manage these mapped or allocated compute resources.

[0148] In at least one embodiment, the software 932 included in the software layer 930 may include software used in at least a portion of the nodes CRs 916(1)-916(N), the grouped computing resources 914, and / or the distributed file system 938 of the framework layer 920. One or more types of software may include, but are not limited to, internet webpage search software, email virus scanning software, database software, and streaming video content software.

[0149] In at least one embodiment, the application 942 included in the application layer 940 may include one or more types of applications used at least in part by nodes CRs 916(1)-916(N), grouped computing resources 914, and / or the distributed file system 938 of the framework layer 920. One or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in combination with one or more embodiments. In some embodiments, the application 942 includes... Figure 1 The processing engine 122 and / or the retrieval engine 124.

[0150] In at least one embodiment, any of the configuration manager 934, resource manager 936, and resource coordinator 912 can implement any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. Self-modification actions can free data center operators of data center 900 from making potentially poor configuration decisions and may prevent underutilization and / or poor performance of the data center.

[0151] According to one or more embodiments described herein, data center 900 may include tools, services, software, or other resources to train one or more machine learning models or to use one or more machine learning models to predict or infer information. For example, one or more machine learning models may be trained by using the software and / or computing resources described above with respect to data center 900 to compute weight parameters according to a neural network architecture. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 900 by using weight parameters computed through one or more training techniques (such as, but not limited to, those described herein).

[0152] In at least one embodiment, the data center 900 may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference using the aforementioned resources. Furthermore, one or more of the software and / or hardware resources described above may be configured to allow a user to train or perform services that infer information, such as image recognition, speech recognition, or other artificial intelligence services.

[0153] Example network environment

[0154] A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may be... Figure 8 This is implemented on one or more instances of computing devices 800—for example, each device may include similar components, features, and / or functions of one or more computing devices 800. Furthermore, in the case of implementing backend devices (e.g., servers, NAS, etc.), the backend devices may be included as part of a data center 900, examples of which are described in this document. Figure 9 To describe in more detail.

[0155] Components of a network environment can communicate with each other via a network, which can be wired, wireless, or both. A network can include multiple networks or one of multiple networks. For example, a network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. Where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connectivity.

[0156] A compatible network environment may include one or more peer-to-peer network environments (in which case the server may not be included in the network environment) and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the functionality described herein for the server can be implemented on any number of client devices.

[0157] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. The cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework supporting software at the software layer and / or application at the application layer. The software or application may respectively include network-based service software or applications. In embodiments, one or more client devices may use the network-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a free and open-source software network application framework that can use a distributed file system for large-scale data processing (e.g., "big data").

[0158] A cloud-based network environment can provide cloud computing and / or cloud storage for any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these different functions can be distributed across multiple locations from a central or core server (e.g., distributed across one or more data centers at the state, region, country, global, etc.). The core server may assign at least a portion of the functionality to the edge server if the connection to the user (e.g., a client device) is relatively close to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0159] One or more client devices may include the information described in this article. Figure 8 At least some of the components, features, and functions of one or more example computing devices 800 described. By way of example and not limitation, the client device may be implemented as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance equipment or system, vehicle, ship, spacecraft, virtual machine, drone, robot, handheld communication device, hospital equipment, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, electrical appliance, consumer electronics device, workstation, edge device, any combination of these depicted devices, or any other suitable device.

[0160] In some embodiments, the systems and methods described herein can be deployed in robotic applications. For example, a robot or robotic system may include one or more onboard processors (e.g., CPUs, GPUs, hardware-based deep learning accelerators (DLAs), hardware-based programmable vision accelerators (PVAs) that may include one or more vector processing units (VPUs), direct memory access (DMA) systems and / or pixel processing engines (PPEs), hardware-based optical flow accelerators (OFAs), SoCs, etc.) and memory and / or storage (e.g., for storing control algorithms, sensor data, and one or more machine learning models). The robotic system can use these processors to execute one or more machine learning models (e.g., language models), enabling it to perform complex tasks autonomously or semi-autonomously, such as interacting with and / or manipulating static and / or dynamic objects, or navigating its environment using sensors such as cameras, LiDAR, RADAR, and ultrasonic sensors. The system can use sensor fusion techniques to combine data from multiple sensors (e.g., cameras, infrared, LiDAR, RADAR, accelerometers) to create a comprehensive model of the robot's surrounding environment. This data can be processed locally on the robot or sent to a remote server to perform computationally intensive tasks, such as 3D mapping or SLAM (Simultaneous Localization and Mapping). In one or more embodiments, data from individual robots (e.g., sensor data, task status, or environmental conditions) can be uploaded to the cloud, where a centralized AI model can analyze and distribute optimized commands across the entire fleet. In some embodiments, the machine learning models described herein (e.g., language models, VLM, LLM, MMLM, diffusion models, NeRF models, DNN, etc.) can be used to allow robots to perceive and reason about their environment and / or communicate with one or more other robots and / or people in the environment. In some embodiments, the robot can communicate with one or more locally hosted servers / computing devices and / or one or more remotely located servers / computing devices (e.g., in one or more data centers) using one or more network interface cards (NICs) and / or data processing units (DPUs).

[0161] In some examples, the machine learning models described herein (e.g., deep neural networks, language models, LLM, VLM, multimodal language models, perceptual models, tracking models, fusion models, transformer models, diffusion models, encoder-only models, decoder-only models, encoder-decoder models, neural rendering field (NERF) models, etc.) can be packaged as microservices, such as inference microservices (e.g., NVIDIA NIM). These microservices can include containers (e.g., operating system (OS) level virtualization packages) that can include an application programming interface (API) layer, a server layer, a runtime layer, and / or a model "engine." For example, an inference microservice can include the container itself and the model (e.g., weights and biases). In some cases, such as when the machine learning model is small enough (e.g., has a sufficiently small number of parameters), the model can be included within the container itself. In other examples (e.g., when the model is large), the model can be hosted / stored in the cloud (e.g., in a data center) and / or hosted on-premises and / or at the edge (e.g., on a local server or computing device, but outside the container). In such embodiments, the model can be accessed via one or more APIs (e.g., REST APIs). Therefore, in some embodiments, the machine learning models described herein can be deployed as inference microservices to accelerate model deployment on any cloud, data center, or edge computing system while ensuring data security. For example, an inference microservice may include one or more APIs, pre-configured containers for simplified deployment, an optimized inference engine (e.g., execution software built using standardized AI model deployments, such as NVIDIA's Triton Inference Server, and / or one or more APIs for high-performance deep learning inference, which may include inference runtime and model optimizations that provide low latency and high throughput for production applications such as NVIDIA's TensorRT), and / or enterprise management data for telemetry (e.g., including identity, metrics, health checks, and / or monitoring). The machine learning models described herein may be included as part of a microservice along with an acceleration infrastructure capable of deployment with a single command and / or orchestrated and automatically scaled using a container orchestration system on the acceleration infrastructure (e.g., reaching data center scale on a single device). Therefore, the inference microservice may include a machine learning model (e.g., optimized for high-performance inference), inference runtime software for executing the machine learning model and providing output / response to inputs (e.g., user queries, prompts, etc.), and enterprise management software for providing health checks, identity verification, and / or other monitoring. In some embodiments, the inference microservice may include software for performing in-situ replacements and / or updates to the machine learning model. During replacement or update, the software performing the replacement / update may maintain user configurations for both the inference runtime software and the enterprise management software.

[0162] The systems and methods described herein may be used, but are not limited to, non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more advanced driver assistance systems (ADAS)), autonomous vehicles and machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, spacecraft, ships, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, underwater vehicles, unmanned aerial vehicles, and / or other vehicle types. Furthermore, the systems and methods described herein can be used for a variety of purposes, such as, but not limited to: machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and supervision, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twins, data center processing, conversational AI, optical transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, systems for performing generative AI operations, systems for implementing one or more language models (e.g., large language models (LLM)), cloud computing and / or any other suitable application.

[0163] The disclosed embodiments can be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aviation systems, medical systems, boating systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems containing one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems implementing one or more large language models (LLMs), one or more visual language models (VLMs), one or more multimodal language models, systems for performing optical transmission simulations, systems for performing collaborative content creation of 3D assets, systems implemented at least partially using cloud computing resources, and / or other types of systems.

[0164] Example autonomous vehicles

[0165] Figure 10AThis is an illustration of an example autonomous vehicle 1000 according to some embodiments of the present disclosure. The autonomous vehicle 1000 (which may be referred to herein as “vehicle 1000”) may include, but is not limited to, passenger vehicles such as cars, trucks, buses, first response vehicles, shuttle buses, electric or motorized bicycles, motorcycles, fire trucks, police vehicles, ambulances, boats, construction vehicles, underwater vessels, robotic vehicles, drones, aircraft, vehicles attached to trailers (e.g., semi-trailers for hauling goods), autonomous robots, humanoid robots, and / or other types of vehicles (e.g., driverless and / or vehicles that carry one or more passengers). Autonomous vehicles are typically described according to the levels of automation defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) in its "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201806, published June 15, 2018; Standard No. J3016-201609, published September 30, 2016; and previous and future versions of this standard). Vehicle 1000 may be able to perform one or more functions that meet Level 3-5 of the autonomous driving level. Vehicle 1000 may be able to perform one or more functions that meet Level 1-5 of the autonomous driving level. For example, depending on the embodiment, vehicle 1000 may be able to perform driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). The term “autonomy” as used herein can include any and / or all types of autonomy of the vehicle 1000 or other machines, such as full autonomy, high autonomy, conditional autonomy, partial autonomy, provision of auxiliary autonomy, semi-autonomy, primary autonomy or other specified autonomy.

[0166] Vehicle 1000 may include components such as chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. Vehicle 1000 may include a propulsion system 1050, such as an internal combustion engine, a hybrid power plant, an all-electric motor, and / or another type of propulsion system. Propulsion system 1050 may be connected to the drivetrain of vehicle 1000, which may include a transmission, to allow propulsion of vehicle 1000. Propulsion system 1050 may be controlled in response to receiving a signal from throttle / accelerator 1052.

[0167] A steering system 1054, which may include a steering wheel, can be used to steer the vehicle 1000 (e.g., along a desired path or route) when the propulsion system 1050 is operating (e.g., when the vehicle is in motion). The steering system 1054 may receive signals from the steering actuator 1056. For fully automatic (level 5) functionality, the steering wheel may be optional.

[0168] The brake sensor system 1046 can be used to operate the vehicle brakes in response to receiving a flag from the brake actuator 1048 and / or the brake sensor.

[0169] It may include one or more System-on-a-Chip (SoC) 1004 ( Figure 10C One or more controllers 1036, including one or more GPUs, may provide (e.g., indicating commands) signals to one or more components and / or systems of vehicle 1000. For example, one or more controllers may send signals to operate vehicle brakes via one or more brake actuators 1048, to operate steering system 1054 via one or more steering actuators 1056, and to operate propulsion system 1050 via one or more throttles / accelerators 1052. One or more controllers 1036 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operating commands (e.g., signals indicating commands) to allow autonomous and / or human-assisted driving class drivers to operate vehicle 1000. One or more controllers 1036 may include a first controller 1036 for autonomous driving functions, a second controller 1036 for functional safety functions, a third controller 1036 for artificial intelligence functions (e.g., computer vision), a fourth controller 1036 for infotainment functions, a fifth controller 1036 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 1036 may handle two or more of the above functions, two or more controllers 1036 may handle a single function, and / or any combination thereof.

[0170] One or more controllers 1036 may provide indications for controlling one or more components and / or systems of vehicle 1000 in response to sensor data (e.g., sensor inputs) received from one or more sensors. Sensor data may be received from, for example, but not limited to, Global Navigation Satellite System (“GNSS”) sensors 1058 (e.g., Global Positioning System sensors), RADAR sensors 1060, ultrasonic sensors 1062, LiDAR sensors 1064, inertial measurement unit (IMU) sensors 1066 (e.g., accelerometers, gyroscopes, magnetic compasses, magnetometers, etc.), microphones 1096, stereo cameras 1068, wide-angle cameras 1070 (e.g., fisheye cameras), infrared cameras 1072, surround cameras 1074 (e.g., 360-degree cameras), long-range and / or medium-range cameras 1098, speed sensors 1044 (e.g., for measuring the rate of vehicle 1000), vibration sensors 1042, steering sensors 1040, braking sensors (e.g., as part of braking sensor system 1046), and / or other sensor types.

[0171] One or more of the controllers 1036 may receive input (e.g., represented by input data) from the instrument panel 1032 of the vehicle 1000 and provide output (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 1034, an auditory sign, a speaker, and / or via other components of the vehicle 1000. These outputs may include information such as vehicle speed, rate, time, map data (e.g., [missing information]). Figure 10C Information such as high-definition (“HD”) map 1022, location data (e.g., the location of vehicle 1000 on the map), direction, and the location of other vehicles (e.g., occupying a grid), as well as information about objects and their states perceived by controller 1036, etc. For example, HMI display 1034 may display information about the existence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, leaving 34B in two miles, etc.).

[0172] Vehicle 1000 also includes a network interface 1024, which can communicate via one or more networks using one or more wireless antennas 1026 and / or a modem. For example, network interface 1024 may be able to communicate via Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multicarrier (“CDMA2000”), etc. One or more wireless antennas 1026 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using one or more local area networks such as Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, etc., and / or one or more low-power wide area networks (LPWAN) such as LoRaWAN, SigFox, etc.

[0173] Figure 10B For use in accordance with some embodiments of this disclosure Figure 10A This is an example of the camera position and field of view of an autonomous vehicle 1000. The camera and its respective field of view are an example embodiment and are not intended to be limiting. For example, additional and / or replaceable cameras may be included, and / or these cameras may be located at different positions on the vehicle 1000.

[0174] The camera type used for the camera may include, but is not limited to, a digital camera suitable for use with components and / or systems of vehicle 800. The camera may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. The camera type may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The camera may be able to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a red-white-white-white (RCCC) color filter array, a red-white-white-blue (RCCB) color filter array, a red-blue-green-white (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, a sharp-pixel camera, such as one having an RCCC, RCCB, and / or RBGC color filter array, may be used in efforts to improve light sensitivity.

[0175] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-function monocular camera can be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video). In some embodiments, image data from one or more cameras can be used to generate memories that enable long-term perception of vehicle 1000, as discussed herein.

[0176] One or more of the cameras can be mounted in mounting components such as custom-designed (3D-printed) parts to cut off stray light and reflections from inside the vehicle (e.g., reflections from the dashboard in the windshield mirror) that may interfere with the camera's image data capture capabilities. Regarding wing mirror mounting components, the wing mirror components can be custom-3D printed so that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras can be integrated into the wing mirror. For side-view cameras, one or more cameras can also be integrated into the four pillars at each corner of the cab.

[0177] A camera with a field of view that includes the environment in front of the vehicle 1000 (e.g., a front-facing camera) can be used for surround view to help identify forward paths and obstacles, and, with the assistance of one or more controllers 1036 and / or control SoCs, to provide information crucial for generating an occupancy grid and / or determining a preferred vehicle path. The front-facing camera can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. The front-facing camera can also be used in ADAS functions and systems, including lane departure warning (“LDW”), autonomous cruise control (“ACC”), and / or other functions such as traffic sign recognition.

[0178] A variety of cameras can be used in front-facing configurations, including, for example, monocular camera platforms that include complementary metal-oxide-semiconductor (“CMOS”) color imagers. Another example could be a wide-angle camera 1070, which can be used to perceive objects entering the field of view from the periphery (such as pedestrians, traffic at intersections, or bicycles). Although Figure 10B The middle image shows only one wide-angle camera, but any number (including zero) of wide-angle cameras 1070 can exist on vehicle 1000. Furthermore, any number of remote cameras 1098 (e.g., long-view stereo camera pairs) can be used for depth-based object detection, especially for objects for which neural networks have not yet been trained. Remote cameras 1098 can also be used for object detection and classification, as well as basic object tracking.

[0179] Any number of stereo cameras 1068 may also be included in a front-mounted configuration. In at least one embodiment, one or more stereo cameras 1068 may include an integrated control unit that includes a scalable processing unit that can provide a multi-core microprocessor and programmable logic (“FPGA”) with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. Such a unit can be used to generate a 3D map of the vehicle environment, including distance estimates for all points in the image. Alternative stereo cameras 1068 may include a compact stereo vision sensor that may include two camera lenses (one on each side) and an image processing chip that can measure the distance from the vehicle to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 1068 may be used in addition to those described herein or alternatively.

[0180] Cameras (e.g., side-view cameras) with a field of view including the side of the vehicle 1000 can be used for surround view, providing information for creating and updating occupancy grids and generating side-impact collision warnings. For example, surround camera 1074 (e.g., ... Figure 10B The four surround cameras 1074 shown can be mounted on the vehicle 1000. The surround cameras 1074 can include a wide-angle camera 1070, a fisheye camera, a 360-degree camera, and / or similar devices. Four examples are provided; the four fisheye cameras can be positioned at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround cameras 1074 (e.g., left, right, and rear) and can utilize one or more other cameras (e.g., forward-facing cameras) as a fourth surround-view camera.

[0181] A camera (e.g., a rear-view camera) having a field of view that includes the environment behind the vehicle 1000 can be used for assisted parking, surround view, rear collision warning, and creating and updating occupancy grids. A wide variety of cameras can be used, including but not limited to those also suitable as front cameras as described herein (e.g., long-range and / or mid-range camera 1098, stereo camera 1068, infrared camera 1072, etc.).

[0182] Figure 10C For use in accordance with some embodiments of this disclosure Figure 10AThe example autonomous vehicle 1000 is illustrated in the block diagram of an example system architecture. It should be understood that this arrangement, and other arrangements described herein, are merely illustrative. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used in addition to or in place of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities, which may be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by these entities can be implemented via hardware, firmware, and / or software. For example, the various functions can be implemented by a processor executing instructions stored in memory.

[0183] Figure 10C Each component, feature, and system in vehicle 1000 is illustrated as being connected via bus 1002. Bus 1002 may include a Controller Area Network (CAN) data interface (or, alternatively, referred to herein as the "CAN bus"). CAN may be a network within vehicle 1000 used to assist in controlling various features and functions of vehicle 1000, such as braking, acceleration, steering, windshield wipers, and other driving functions. The CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus can be read to find steering wheel angle, ground speed, engine speed per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.

[0184] Although bus 1002 is described herein as a CAN bus, this is not intended to be limiting. For example, FlexRay and / or Ethernet may be used in addition to or alternatively to a CAN bus. Furthermore, although bus 1002 is represented by a single line, this is not intended to be limiting. For example, any number of buses 1002 may exist, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 1002 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 1002 may be used for a collision avoidance function, and a second bus 1002 may be used for driving control. In any example, each bus 1002 may communicate with any component of vehicle 1000, and two or more buses 1002 may communicate with the same component. In some examples, each SoC 1004, each controller 1036, and / or each computer within the vehicle may have access to the same input data (e.g., input from sensors in the vehicle 1000) and may be connected to a common bus such as the CAN bus.

[0185] Vehicle 1000 may include one or more controllers 1036, such as those described herein. Figure 10A The controllers described herein. Controller 1036 can be used for a wide variety of functions. Controller 1036 can be coupled to any other different components and systems of vehicle 1000 and can be used for the control of vehicle 1000, artificial intelligence of vehicle 1000, infotainment and / or similar functions of vehicle 1000.

[0186] Vehicle 1000 may include one or more System-on-Chip (SoC) 1004. SoC 1004 may include CPU 1006, GPU 1008, processor 1010, cache 1012, accelerator 1014, data storage 1016, and / or other components and features not shown. SoC 1004 can be used to control vehicle 1000 in a wide variety of platforms and systems. For example, one or more SoCs 1004 may be combined with an HD map 1022 in a system (e.g., the system of vehicle 1000), the HD map being accessible from one or more servers (e.g., via a network interface 1024). Figure 10D One or more servers (1078) receive map refresh and / or updates.

[0187] CPU 1006 may include CPU clusters or CPU complexes (or, alternatively, referred to herein as "CCPLEX"). CPU 1006 may include multiple cores and / or L2 cache. For example, in some embodiments, CPU 1006 may include eight cores in a coherent multiprocessor configuration. In some embodiments, CPU 1006 may include four dual-core clusters, each cluster having a dedicated L2 cache (e.g., 2MB L2 cache). CPU 1006 (e.g., CCPLEX) may be configured to support simultaneous cluster operation, such that any combination of clusters of CPU 1006 can be active at any given time.

[0188] CPU 1006 can implement power management capabilities including one or more of the following features: automatic clock gating of hardware blocks when idle to save dynamic power; clock gating of each core when the core is not actively executing instructions due to the execution of WFI / WFE instructions; independent power gating of each core; independent clock gating of each core cluster when all cores are clock-gated or power-gated; and / or independent power gating of each core cluster when all cores are power-gated. CPU 1006 can further implement enhanced algorithms for managing power states, wherein allowed power states and desired wake-up times are specified, and the hardware / microcode determines the optimal power state to enter for the core, cluster, and CCPLEX. The processing core can support simplified power state entry sequences in software, with this work offloaded to the microcode.

[0189] GPU 1008 may include an integrated GPU (or, alternatively, referred to herein as an "iGPU"). GPU 1008 may be programmable and efficient for parallel workloads. In some examples, GPU 1008 may use an enhanced tensor instruction set. GPU 1008 may include one or more streaming microprocessors, wherein each streaming microprocessor may include an L1 cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In some embodiments, GPU 1008 may include at least eight streaming microprocessors. GPU 1008 may use a computation application programming interface (API). Furthermore, GPU 1008 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0190] In automotive and embedded applications, the GPU 1008 can be power-optimized for optimal performance. For example, the GPU 1008 can be fabricated on FinFETs. However, this is not intended to be limiting, and the GPU 1008 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can combine several mixed-precision processing cores divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Furthermore, the streaming microprocessor can include independent parallel integer and floating-point data paths to leverage the mixture of computation and addressing computations to provide efficient execution of workloads. Streaming microprocessors may include independent thread scheduling capabilities to allow for finer-grained synchronization and cooperation between parallel threads. Streaming microprocessors may include combined L1 data caches and shared memory units to improve performance while simplifying programming.

[0191] The GPU 1008 may include, in some examples, a High Bandwidth Memory (HBM) and / or a 16GB HBM2 memory subsystem providing a peak memory bandwidth of approximately 900GB / s. In some examples, in addition to HBM memory or alternatively, Synchronous Graphics Random Access Memory (SGRAM), such as Generation 5 Graphics Double Data Rate Synchronous Random Access Memory (GDDR5), may be used.

[0192] The GPU 1008 may include unified memory technology, which includes access counters to allow memory pages to be migrated more precisely to the processors that access them most frequently, thereby improving the efficiency of shared memory ranges between processors. In some examples, Address Translation Service (ATS) support can be used to allow the GPU 1008 to directly access the CPU 1006 page tables. In such examples, when the GPU 1008 Memory Management Unit (MMU) experiences a miss, the address translation request can be transferred to the CPU 1006. In response, the CPU 1006 can look up the virtual-physical mapping for the address in its page tables and transfer the translation back to the GPU 1008. Thus, unified memory technology can allow a single unified virtual address space for the memory of both the CPU 1006 and the GPU 1008, simplifying GPU 1008 programming and porting applications to the GPU 1008.

[0193] In addition, the GPU 1008 may include access counters that track how frequently the GPU 1008 accesses the memory of other processors. Access counters can help ensure that memory pages are moved to the physical memory of the processor that accesses those pages most frequently.

[0194] SoC 1004 may include any number of caches 1012, including those described herein. For example, cache 1012 may include an L3 cache available to both CPU 1006 and GPU 1008 (e.g., it is connected to both CPU 1006 and GPU 1008). Cache 1012 may include a write-back cache, which can track the state of rows, for example, using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, but a smaller cache size may also be used.

[0195] SoC 1004 may include an arithmetic logic unit (ALU) that can be utilized in processing of any of the various tasks or operations performed on vehicle 1000, such as processing a DNN. Furthermore, SoC 1004 may include a floating-point unit (FPU) (or other mathematical coprocessor or digital coprocessor type) for performing mathematical operations within the system. For example, SoC 1004 may include one or more FPUs integrated as execution units within CPU 1006 and / or GPU 1008.

[0196] SoC 1004 may include one or more accelerators 1014 (e.g., hardware accelerators, software accelerators, or combinations thereof). For example, SoC 1004 may include a hardware accelerator cluster, which may include optimized hardware accelerators and / or large on-chip memory. This large on-chip memory (e.g., 4MB SRAM) can enable the hardware accelerator cluster to accelerate neural networks and other computations. The hardware accelerator cluster can be used to secondary GPU 1008 and offload some tasks from GPU 1008 (e.g., freeing up more cycles of GPU 1008 to perform other tasks). As an example, accelerator 1014 can be used for targeted workloads (e.g., perception, convolutional neural networks (CNNs), etc.) that are stable enough to be easily controlled for acceleration. When used herein, the term "CNN" can include all types of CNNs, including region-based or region convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).

[0197] Accelerator 1014 (e.g., a hardware accelerator cluster) may include a Deep Learning Accelerator (DLA). The DLA may include one or more Tensor Processing Units (TPUs) that can be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured to perform image processing functions (e.g., for CNNs, RCNNs, etc.) and optimized for performing image processing functions. The DLA may be further optimized for a specific set of neural network types and floating-point operations and inference. The DLA is designed to provide higher performance per millimeter than a general-purpose GPU and significantly outperform CPUs. The TPU can perform several functions, including single-instance convolution functions, support for INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.

[0198] DLA can execute neural networks, especially CNNs, quickly and efficiently on processed or unprocessed data for any function across a wide variety of applications, such as, but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection and recognition using data from microphones; CNNs for face recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for safety and / or safety-related events.

[0199] The DLA can perform any function of the GPU 1008, and by using inference accelerators, for example, a designer can target either the DLA or the GPU 1008 for any function. For instance, a designer can focus the CNN processing and floating-point operations on the DLA and leave other functions to the GPU 1008 and / or other accelerators 1014.

[0200] Accelerator 1014 (e.g., a hardware accelerator cluster) may include a programmable vision accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA can provide a balance between performance and flexibility. For example, each PVA may include, for example, but not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.

[0201] RISC cores can interact with image sensors (such as the image sensor of any camera described herein), image label processors, and / or similar objects. Each of these RISC cores may include any amount of memory. Depending on the embodiment, the RISC core may use any of several protocols. In some examples, the RISC core may execute a real-time operating system (RTOS). RISC cores may be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, a RISC core may include an instruction cache and / or tightly coupled RAM.

[0202] DMA enables PVA components to access system memory independently of the CPU 1006. DMA can support any number of features to provide optimizations to the PVA, including but not limited to support for multidimensional addressing and / or circular addressing. In some examples, DMA can support addressing in up to six or more dimensions, which can include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.

[0203] A vector processor can be a programmable processor designed to efficiently and flexibly execute programming for computer vision algorithms and provide tag processing capabilities. In some examples, a PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may operate as the main processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core may include a digital tag processor, such as, for example, a Single Instruction Multiple Data (SIMD) or Very Long Instruction Word (VLIW) digital tag processor. The combination of SIMD and VLIW can enhance throughput and speed.

[0204] Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. Consequently, in some examples, each of the vector processors may be configured to execute independently of other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelization. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even different algorithms on a sequence of images or portions of an image. Among other things, any number of PVAs may be included in a cluster of hardware accelerators, and any number of vector processors may be included in each of these PVAs. Furthermore, the PVA may include additional error correction code (ECC) memory to enhance overall system security.

[0205] Accelerator 1014 (e.g., a hardware accelerator cluster) may include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for accelerator 1014. In some examples, on-chip memory may include at least 4MB of SRAM consisting of, for example, but not limited to, eight field-configurable memory blocks, accessible by both PVA and DLA. Each pair of memory blocks may include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. PVA and DLA may access memory via a backbone that provides high-speed memory access to PVA and DLA. The backbone may include (e.g., using an APB) an on-chip computer vision network that interconnects PVA and DLA to memory.

[0206] On-chip computer vision networks can include an interface that determines whether both the PVA and DLA provide a ready and valid flag before transmitting any control flags / addresses / data. Such an interface can provide separate phases and channels for transmitting control flags / addresses / data, as well as burst communication for continuous data transmission. This type of interface can conform to ISO 26262 or IEC 615010 standards, but other standards and protocols can also be used.

[0207] In some examples, SoC 1004 may include, for example, a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed August 10, 2018. This real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the location and extent of objects (e.g., within a world model) to generate real-time visualization simulations for RADAR sign interpretation, sound propagation synthesis and / or analysis, SONAR system simulation, general wave propagation simulation, comparison with LiDAR data for localization and / or other functional purposes, and / or for other uses. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing-related operations.

[0208] Accelerator 1014 (e.g., a cluster of hardware accelerators) has broad applications in autonomous driving. PVAs can be programmable vision accelerators used in critical processing stages of ADAS and autonomous vehicles. PVAs are well-suited to algorithmic domains requiring predictable processing, low power, and low latency. In other words, PVAs perform well in semi-dense or dense rule computation, even on small datasets requiring predictable runtimes with low latency and low power. Therefore, in the context of platforms for autonomous vehicles, PVAs are designed to run classical computer vision algorithms because they are efficient in object detection and integer arithmetic.

[0209] For example, according to one embodiment of this technology, PVA is used to perform computer stereo vision. In some examples, semi-global matching-based algorithms may be used, but this is not intended to be limiting. Many applications for Level 3-5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., from moving structures, pedestrian recognition, lane detection, etc.). PVA can perform computer stereo vision functions on input from two monocular cameras.

[0210] In some examples, PVA can be used to perform intensive optical flow, providing processed RADAR data from the raw RADAR data (e.g., using 4D Fast Fourier Transform). In other examples, PVA is used for time-of-flight depth processing, which, for example, involves processing raw time-of-flight data to provide processed time-of-flight data.

[0211] DLA can be used to run any type of network to enhance control and driving safety, including, for example, neural networks that output a confidence metric for each object detection. Such a confidence value can be interpreted as a probability or as providing a relative “weight” for each detection compared to other detections. This confidence value allows the system to make further decisions about which detections should be considered true positives rather than false positives. For example, the system can set a threshold for the confidence and only consider detections exceeding the threshold as true positives. In an Automatic Emergency Braking (AEB) system, false positives can cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run neural networks to regress the confidence values. The neural network can take at least some subset of parameters as its input, such as bounding box dimensions, ground plane estimates obtained (e.g. from another subsystem), outputs from inertial measurement unit (IMU) sensor 1066 related to the orientation and distance of vehicle 1000, 3D position estimates of objects obtained from the neural network and / or other sensors (e.g., LiDAR sensor 1064 or RADAR sensor 1060), etc.

[0212] SoC 1004 may include one or more data storage units 1016 (e.g., memory). The data storage unit 1016 may be on-chip memory of SoC 1004, which may store neural networks to be executed on the GPU and / or DLA. In some examples, for redundancy and security, the data storage unit 1016 may be large enough to store multiple instances of the neural network. The data storage unit 1016 may include L2 or L3 cache 1012. References to the data storage unit 1016 may include references to memory associated with the PVA, DLA, and / or other accelerators 1014 as described herein.

[0213] SoC 1004 may include one or more processors 1010 (e.g., embedded processors). Processor 1010 may include a startup and power management processor, which may be a dedicated processor and subsystem for handling startup power and management functions, as well as safety implementation. The startup and power management processor may be part of the SoC 1004 startup sequence and may provide runtime power management services. The startup power and management processor may provide clock and voltage programming, auxiliary system low-power state transitions, SoC 1004 thermal and temperature sensor management, and / or SoC 1004 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to the temperature, and SoC 1004 may use the ring oscillator to detect the temperature of CPU 1006, GPU 1008, and / or accelerator 1014. If it is determined that the temperature exceeds a threshold, the startup and power management processor may enter a temperature fault routine and place SoC 1004 into a lower power state and / or place vehicle 1000 into a driver-safe parking mode (e.g., safely stop vehicle 1000).

[0214] The processor 1010 may also include a set of embedded processors that can be used as an audio processing engine. The audio processing engine can be an audio subsystem that allows for full hardware support for multi-channel audio via multiple interfaces, as well as a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital flag processor and dedicated RAM.

[0215] The processor 1010 may also include an always-on-processor engine that can provide the necessary hardware features to support low-power sensor management and wake-up use cases. This always-on-processor engine may include a processor core, tightly coupled RAM, support for peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0216] The processor 1010 may also include a security cluster engine, which comprises a dedicated processor subsystem for handling security management for automotive applications. The security cluster engine may include two or more processor cores, tightly coupled RAM, support for peripheral devices (e.g., timers, interrupt controllers, etc.), and / or routing logic. In secure mode, the two or more cores may operate in lockstep mode and function as a single core with comparison logic that detects any differences between their operations.

[0217] The processor 1010 may also include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.

[0218] The processor 1010 may also include a high dynamic range flag processor, which may include an image flag processor, which is a hardware engine that is part of the camera processing pipeline.

[0219] Processor 1010 may include a video image compositer, which may be (e.g., implemented on a microprocessor) a processing block, implementing video post-processing functions required by the video playback application to generate the final image for the player window. The video image compositer may perform lens distortion correction on the wide-angle camera 1070, the surround camera 1074, and / or the in-cabin monitoring camera sensor. The in-cabin monitoring camera sensor is preferably monitored by a neural network running on another instance of an advanced SoC, configured to recognize in-cabin events and respond accordingly. The in-cabin system may perform lip reading to activate mobile phone services and make calls, dictate emails, change vehicle destinations, activate or change the vehicle's infotainment system and settings, or provide voice-activated web browsing. Some functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled in other situations.

[0220] Video image compositers can include enhanced temporal denoising for both spatial and temporal noise reduction. For example, in the case of motion in the video, denoising appropriately weights spatial information, reducing the weight of information provided by neighboring frames. In cases where the image or part of the image does not contain motion, the temporal denoising performed by the video image compositer can use information from previous images to reduce noise in the current image.

[0221] The video image compositer can also be configured to perform stereo correction on input stereo camera frames. When the operating system desktop is in use and the GPU 1008 does not need to continuously render new surfaces, the video image compositer can be further used for user interface components. Even when the GPU 1008 is powered on and active, performing 3D rendering, the video image compositer can be used to offload the GPU 1008 to improve performance and responsiveness.

[0222] SoC 1004 may also include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functions for receiving video and input from a camera. SoC 1004 may also include an input / output controller that can be software-controlled and can be used to receive I / O flags not submitted to a specific role.

[0223] SoC 1004 may also include a wide range of peripheral interfaces to allow communication with peripherals, audio codecs, power management and / or other devices. SoC 1004 can be used to process data from cameras and sensors (e.g., LiDAR sensor 1064, RADAR sensor 1060, etc., which can be connected via Gigabit Multimedia Serial Link and Ethernet), data from bus 1002 (e.g., vehicle 1000 speed, steering wheel position, etc.), and data from GNSS sensor 1058 (connected via Ethernet or CAN bus). SoC 1004 may also include a dedicated high-performance, high-capacity memory controller, which may include its own DMA engine and can be used to free up CPU 1006 from routine data management tasks.

[0224] The SoC 1004 can be an end-to-end platform with a flexible architecture spanning Automation Levels 3-5, providing a comprehensive functional safety architecture that leverages and efficiently utilizes computer vision and ADAS technologies for diversity and redundancy, along with deep learning tools to deliver a flexible and reliable driving software stack. The SoC 1004 can be faster, more reliable, and even more energy- and space-efficient than conventional systems. For example, when combined with the CPU 1006, GPU 1008, and data storage 1016, the accelerator 1014 can provide a fast and efficient platform for Level 3-5 autonomous vehicles.

[0225] Therefore, this technology offers capabilities and functionalities that cannot be achieved through conventional systems. For example, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages ​​such as C to execute a wide variety of processing algorithms across a diverse range of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for automotive ADAS applications and practical Level 3-5 autonomous vehicles.

[0226] In contrast to conventional systems, the techniques described in this paper, by providing CPU complexes, GPU complexes, and hardware accelerator clusters, allow multiple neural networks to be executed simultaneously and / or sequentially, and the results combined to achieve Level 3–5 autonomous driving capabilities. For example, a CNN executed on a DLA or dGPU (e.g., GPU 1020) could include text and word recognition, allowing a supercomputer to read and understand traffic signs, including those for which neural networks have not yet been specifically trained. The DLA could also include a neural network capable of recognizing, interpreting, and providing semantic understanding of the signs, and passing that semantic understanding to a path planning module running on the CPU complex.

[0227] As another example, multiple neural networks can operate simultaneously, as required for Level 3, 4, or 5 driving. For instance, a warning sign consisting of "Caution: Flashing lights indicate icy conditions," along with a light, can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a deployed first neural network (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" can be interpreted by a deployed second neural network that informs the vehicle's path planning software (preferably executing on a CPU complex) that icy conditions exist when the flashing lights are detected. The flashing lights can be identified by a deployed third neural network operating across multiple frames, informing the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks can operate simultaneously, for example, within a DLA and / or on a GPU 1008.

[0228] In some examples, the CNN used for facial recognition and owner identification can use data from camera sensors to identify the presence of an authorized driver and / or owner of vehicle 1000. A processing engine always on the sensors can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in safe mode, to disable the vehicle when the owner leaves. In this way, SoC 1004 provides security against theft and / or carjacking.

[0229] In another example, the CNN used for emergency vehicle detection and identification can use data from microphone 1096 to detect and identify emergency vehicle siren. In contrast to conventional systems that use a general classifier to detect siren and manually extract features, SoC 1004 uses a CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative shut-off rate of emergency vehicles (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the localized area in which the vehicle operates, as identified by GNSS sensor 1058. Thus, for example, when operating in the EU, the CNN will seek to detect EU siren, and when operating in the US, the CNN will seek to identify siren only in North America. Once an emergency vehicle is detected, with the assistance of ultrasonic sensor 1062, the control program can be used to execute emergency vehicle safety routines, causing the vehicle to slow down, pull over to the side of the road, stop, and / or idle until the emergency vehicle passes.

[0230] The vehicle may include a CPU 1018 (e.g., a discrete CPU or dCPU) that can be coupled to the SoC 1004 via a high-speed interconnect (e.g., PCIe). The CPU 1018 may include, for example, an X106 processor. The CPU 1018 can be used to perform any of a wide variety of functions, including, for example, arbitrating the results of potential inconsistencies between ADAS sensors and the SoC 1004, and / or monitoring the status and health of the controller 1036 and / or the infotainment SoC 1030.

[0231] Vehicle 1000 may include a GPU 1020 (e.g., a discrete GPU or dGPU) that can be coupled to SoC 1004 via a high-speed interconnect (e.g., NVIDIA's NVLINK). GPU 1020 may provide additional artificial intelligence capabilities, for example by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based at least in part on inputs (e.g., sensor data) from sensors of vehicle 1000.

[0232] Vehicle 1000 may also include a network interface 1024, which may include one or more wireless antennas 1026 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). Network interface 1024 can be used to enable wireless connectivity via the Internet to the cloud (e.g., with server 1078 and / or other network devices), with other vehicles, and / or with computing devices (e.g., passenger client devices). For communication with other vehicles, a direct link can be established between the two vehicles, and / or an indirect link can be established (e.g., across networks and via the Internet). A direct link can be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide vehicle 1000 with information about vehicles approaching vehicle 1000 (e.g., vehicles in front, to the side, and / or behind vehicle 1000). This functionality can be part of vehicle 1000's cooperative adaptive cruise control function.

[0233] Network interface 1024 may include a SoC that provides modulation and demodulation functions and enables controller 1036 to communicate via a wireless network. Network interface 1024 may include an RF front-end for up-conversion from baseband to RF and down-conversion from RF to baseband. Frequency conversion can be performed using known processes and / or using a superheterodyne process. In some examples, the RF front-end functionality may be provided by a separate chip. The network interface may include wireless functions for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0234] Vehicle 1000 may also include data storage 1028, which may include off-chip (e.g., off-chip SoC 1004) storage devices. Data storage 1028 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard disk, and / or other components and / or devices capable of storing at least one bit of data. In some embodiments, data storage 1028 includes a vector database and / or another type of data storage 208 that can be used to store memory 226 associated with vehicle 1000.

[0235] Vehicle 1000 may also include a GNSS sensor 1058. The GNSS sensor 1058 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) is used for auxiliary mapping, sensing, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 1058 can be used, including, for example, but not limited to, GPS using a USB connector with an Ethernet-to-serial (RS-232) bridge.

[0236] Vehicle 1000 may also include a RADAR sensor 1060. The RADAR sensor 1060 can be used by vehicle 1000 for remote vehicle detection even in dark and / or inclement weather conditions. The RADAR functional safety level may be ASIL B. The RADAR sensor 1060 can use CAN and / or bus 1002 (e.g., to transmit data generated by the RADAR sensor 1060) for control and access to object tracking data, and in some examples, Ethernet access for accessing raw data. A wide variety of RADAR sensor types can be used. For example, and without limitation, the RADAR sensor 1060 can be adapted for front, rear, and side RADAR use. In some examples, a pulse Doppler RADAR sensor is used.

[0237] The RADAR sensor 1060 can include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In some examples, the long-range RADAR can be used for adaptive cruise control functions. The long-range RADAR system can provide a wide field of view (e.g., within 250m) achieved through two or more independent scans. The RADAR sensor 1060 can help distinguish between stationary and moving objects and can be used by ADAS systems for emergency braking assist and forward collision warning. The long-range RADAR sensor can include a single-site multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the four central antennas can create a focused beam pattern designed to record the vehicle 1000's surroundings at higher rates with minimal traffic interference from adjacent lanes. The other two antennas can extend the field of view, enabling rapid detection of vehicles entering or leaving the vehicle 1000's lane.

[0238] As an example, a mid-range RADAR system can include a range of up to 1060m (front) or 80m (rear) and a field of view of up to 42 degrees (front) or 1050 degrees (rear). Short-range RADAR systems can include, but are not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor blind spots behind and beside the vehicle.

[0239] Short-range RADAR systems can be used in ADAS systems for blind spot detection and / or lane change assistance.

[0240] Vehicle 1000 may also include ultrasonic sensors 1062. Ultrasonic sensors 1062, which can be positioned at the front, rear, and / or sides of vehicle 1000, can be used for parking assistance and / or creating and updating occupancy grids. A wide variety of ultrasonic sensors 1062 can be used, and different ultrasonic sensors 1062 can be used for different detection ranges (e.g., 2.5m, 4m). Ultrasonic sensors 1062 can operate at functional safety level ASIL B.

[0241] Vehicle 1000 may include a LiDAR sensor 1064. The LiDAR sensor 1064 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LiDAR sensor 1064 may be of functional safety level ASIL B. In some examples, vehicle 1000 may include multiple LiDAR sensors 1064 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).

[0242] In some examples, the LiDAR sensor 1064 may be able to provide a list of objects and their distances within a 360-degree field of view. Commercially available LiDAR sensors 1064 may have an advertising range of, for example, approximately 1000m, with an accuracy of 2cm-3cm, and support for 1000Mbps Ethernet connectivity. In some examples, one or more non-protruding LiDAR sensors 1064 may be used. In such examples, the LiDAR sensor 1064 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of a vehicle 1000. In such examples, the LiDAR sensor 1064 may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for low-reflectivity objects, with a range of 200m. Front-mounted LiDAR sensors 1064 may be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0243] In some examples, LiDAR technologies such as 3D flash LiDAR can also be used. 3D flash LiDAR uses flashes of laser light as the emission source to illuminate the vehicle's surroundings up to approximately 200 meters. A flash LiDAR unit includes a receiver that records the laser pulse propagation time and reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LiDAR allows for the generation of highly accurate and distortion-free images of the surrounding environment using each laser flash. In some examples, four flash LiDAR sensors can be deployed, one on each side of the vehicle. Available 3D flash LiDAR systems include solid-state 3D staring array LiDAR cameras (e.g., non-scanning LiDAR devices) without moving parts other than fans. Flash LiDAR devices can use 5 nanosecond Class I (eye-safe) laser pulses per frame and can capture reflected laser light in the form of a 3D range point cloud and co-registered intensity data. By using flash LiDAR, and because flash LiDAR is a solid-state device with no moving parts, the LiDAR sensor 1064 is less susceptible to motion blur, vibration, and / or shock.

[0244] The vehicle may also include an IMU sensor 1066. In some examples, the IMU sensor 1066 may be located at the center of the rear axle of the vehicle 1000. The IMU sensor 1066 may include, for example, but not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 1066 may include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 1066 may include an accelerometer, a gyroscope, and a magnetometer.

[0245] In some embodiments, the IMU sensor 1066 can be implemented as a miniature, high-performance GPS-assisted inertial navigation system (GPS / INS) that combines a microelectromechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filter algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 1066 can enable the vehicle 1000 to estimate heading by directly observing and correlating velocity changes from GPS to the IMU sensor 1066 without input from a magnetic sensor. In some examples, the IMU sensor 1066 and the GNSS sensor 1058 can be combined into a single integrated unit.

[0246] The vehicle may include a microphone 1096 placed in and / or around the vehicle 1000. Among other things, the microphone 1096 may be used for emergency vehicle detection and identification.

[0247] The vehicle may also include any number of camera types, including stereo camera 1068, wide-angle camera 1070, infrared camera 1072, surround camera 1074, long-range and / or mid-range camera 1098, and / or other camera types. These cameras can be used to capture image data around the entire perimeter of the vehicle 1000. The types of cameras used depend on the embodiment and the requirements of the vehicle 1000, and any combination of camera types can be used to provide the necessary coverage around the vehicle 1000. Furthermore, the number of cameras may vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and without limitation, these cameras may support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described herein with respect to... Figure 10A and Figure 10B It was described in more detail.

[0248] Vehicle 1000 may also include vibration sensor 1042. Vibration sensor 1042 can measure vibrations of vehicle components such as axles. For example, changes in vibration can indicate changes in the road surface. In another example, when two or more vibration sensors 1042 are used, differences between vibrations can be used to determine friction or slippage on the road surface (e.g., when there is a vibration difference between the powered drive shaft and the free-rotating shaft).

[0249] Vehicle 1000 may include ADAS system 1038. In some examples, ADAS system 1038 may include SoC. ADAS system 1038 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC) and / or other features and functions.

[0250] The ACC system can use a RADAR sensor 1060, a LiDAR sensor 1064, and / or a camera. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to vehicles immediately in front of vehicle 1000 and automatically adjusts the vehicle speed to maintain a safe distance. Lateral ACC performs distance holding and, if necessary, advises vehicle 1000 to change lanes. Lateral ACC is associated with other ADAS applications such as LCA and CWS.

[0251] CACC uses information from other vehicles, which can be received indirectly from other vehicles via a wireless link or through a network connection (e.g., via the Internet) through network interface 1024 and / or wireless antenna 1026. Direct links can be provided by vehicle-to-vehicle (V2V) communication links, while indirect links can be infrastructure-to-vehicle (I2V) communication links. Typically, the V2V communication concept provides information about vehicles immediately ahead (e.g., vehicles immediately in front of vehicle 1000 and in the same lane), while the I2V communication concept provides information about traffic further ahead. A CACC system can include either or both I2V and V2V information sources. Given information about vehicles ahead of vehicle 1000, CACC can be more reliable, and it has the potential to improve traffic flow and reduce road congestion.

[0252] The Forward-Warping (FCW) system is designed to alert the driver to hazards, enabling the driver to take corrective action. The FCW system uses a front-facing camera and / or RADAR sensor 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as displays, speakers, and / or vibrating components. The FCW system can provide warnings in the form of, for example, audible, visual, haptic, and / or rapid braking pulses.

[0253] An AEB (Autonomous Emergency Braking) system detects an impending forward collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system can use a front-facing camera and / or RADAR sensor 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid a collision, and if the driver does not take corrective action, the AEB system can automatically apply the brakes to attempt to prevent or at least mitigate the effects of the predicted collision. The AEB system may include technologies such as dynamic brake support and / or collision approach braking.

[0254] The Lane Departure Warning (LDW) system provides visual, auditory, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle crosses lane markings. When the driver indicates intentional lane departure, the LDW system is deactivated by activating the turn sign. The LDW system can utilize a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as displays, speakers, and / or vibrating components.

[0255] The LKA system is a variation of the LDW system. If vehicle 1000 begins to leave the lane, the LKA system provides steering input or braking to correct vehicle 1000.

[0256] The BSW system detects and warns the driver of vehicles in the vehicle's blind spot. The BSW system can provide visual, auditory, and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses turn signs. The BSW system can utilize a rear-facing camera and / or RADAR sensor 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as displays, speakers, and / or vibrating components.

[0257] RCTW systems can provide visual, auditory, and / or tactile notifications when an object is detected outside the range of a rear-view camera while the vehicle is reversing. Some RCTW systems include AEB (Autonomous Emergency Braking) to ensure the application of the vehicle's brakes to avoid a collision. RCTW systems may use one or more rear-view RADAR sensors 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as displays, speakers, and / or vibrating components.

[0258] Conventional ADAS systems can be prone to false positives, which can be annoying and distracting for the driver, but typically not catastrophic, as ADAS systems alert the driver and allow them to determine whether a safe condition truly exists and take appropriate action. However, in an autonomous vehicle 1000, in the event of conflicting results, the vehicle 1000 itself must decide whether to heed the results from the main computer or auxiliary computer (e.g., the first controller 1036 or the second controller 1036). For example, in some embodiments, ADAS system 1038 may be a backup and / or auxiliary computer for providing perception information to a backup computer rationality module. The backup computer rationality monitor may run redundant and diverse software on hardware components to detect faults in perception and dynamic driving tasks. Outputs from ADAS system 1038 may be provided to a supervisory MCU. If the outputs from the main computer and the auxiliary computer conflict, the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.

[0259] In some examples, the master computer can be configured to provide a confidence score to the supervisory MCU, indicating the master computer's confidence level in the selected result. If the confidence score exceeds a threshold, the supervisory MCU can follow the master computer's direction regardless of whether the auxiliary computer provides conflicting or inconsistent results. If the confidence score does not meet the threshold and the master and auxiliary computers indicate different results (e.g., conflict), the supervisory MCU can arbitrate between these computers to determine the appropriate result.

[0260] The supervisory MCU can be configured to run a neural network trained and configured to determine, at least in part, the conditions under which the auxiliary computer provides a false alarm, based on outputs from both the host and auxiliary computers. Thus, the neural network in the supervisory MCU can learn when the output of the auxiliary computer can be trusted and when it cannot. For example, when the auxiliary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying a metallic object that is not actually dangerous, such as a drain grid or manhole cover that triggers an alarm. Similarly, when the auxiliary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to ignore the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In embodiments that include a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or GPU suitable for running the neural network using associated memory. In a preferred embodiment, the supervisory MCU may include a component of SoC 1004 and / or be included as a component of SoC 1004.

[0261] In other examples, ADAS system 1038 may include an auxiliary computer that performs ADAS functions using conventional computer vision rules. This allows the auxiliary computer to use classic computer vision rules (if-then), and the presence of neural networks in the supervising MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity make the entire system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functionality. For instance, if a software vulnerability or bug exists in the software running on the host computer and non-identical software code running on the auxiliary computer provides the same overall result, the supervising MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the host computer does not cause a substantial error.

[0262] In some examples, the output of the ADAS system 1038 can be fed to the perception block and / or the dynamic driving task block of the main computer. For example, if the ADAS system 1038 issues a forward collision warning because an object is immediately in front, the perception block can use this information when identifying the object. In other examples, the assistance computer can have its own neural network, which is trained and thus reduces the risk of false positives as described herein.

[0263] Vehicle 1000 may also include an infotainment SoC 1030 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more discrete components. The infotainment SoC 1030 may include a combination of hardware and software that can be used to provide vehicle 1000 with audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming media, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, WiFi, etc.) and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total coverage distance, brake fuel level, fuel level, door opening / closing, air filter information, etc.). For example, the infotainment SoC 1030 may include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, in-vehicle computer, in-vehicle entertainment, WiFi, steering wheel audio controls, hands-free voice controls, head-up display (HUD), HMI display 1034, telematics device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems) and / or other components. The infotainment SoC 1030 may further be used to provide information (e.g., visual and / or auditory) to users of the vehicle, such as information from the ADAS system 1038, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0264] The infotainment SoC 1030 may include GPU functionality. The infotainment SoC 1030 can communicate with other devices, systems, and / or components of the vehicle 1000 via bus 1002 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 1030 may be coupled to a supervisory MCU, allowing the GPU of the infotainment system to perform some autonomous driving functions in the event of a failure of the main controller 1036 (e.g., the primary and / or backup computer of the vehicle 1000). In such an example, the infotainment SoC 1030 may place the vehicle 1000 into a driver-safe parking mode as described herein.

[0265] Vehicle 1000 may also include instrument panel 1032 (e.g., digital instrument cluster, electronic instrument cluster, digital instrument panel, etc.). Instrument panel 1032 may include a controller and / or supercomputer (e.g., a discrete controller or supercomputer). Instrument panel 1032 may include a set of instruments such as speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, seat belt warning light, parking brake warning light, engine malfunction indicator, airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between infotainment SoC 1030 and instrument panel 1032. In other words, instrument panel 1032 may be included as part of infotainment SoC 1030, or vice versa.

[0266] Figure 10D For cloud-based servers and according to some embodiments of this disclosure Figure 10A This is a system diagram illustrating communication between example autonomous vehicles 1000. System 1076 may include server 1078, network 1090, and vehicles including vehicle 1000. Server 1078 may include multiple GPUs 1084(A)-1084(H) (collectively referred to herein as GPU 1084), PCIe switches 1082(A)-1082(D) (collectively referred to herein as PCIe switch 1082), and / or CPUs 1080(A)-1080(B) (collectively referred to herein as CPU 1080). GPU 1084, CPU 1080, and PCIe switches may be interconnected with high-speed interconnects and / or PCIe connections 1086, such as, but not limited to, the NVLink interface 1088 developed by NVIDIA. In some examples, GPU 1084 is connected via NVLink and / or NVSwitch SoC, and GPU 1084 and PCIe switch 1082 are connected via PCIe interconnect. Although the diagram illustrates eight GPUs 1084, two CPUs 1080, and two PCIe switches, it is not intended to be limiting. Depending on the embodiment, each of the servers 1078 may include any number of GPUs 1084, CPUs 1080, and / or PCIe switches. For example, each of the servers 1078 may include eight, sixteen, thirty-two, and / or more GPUs 1084.

[0267] Server 1078 can receive image data from vehicles via network 1090, representing images of unexpected or altered road conditions, such as recently initiated roadworks. Server 1078 can also transmit neural network 1092, updated neural network 1092, and / or map information 1094, including information about traffic and road conditions, to vehicles via network 1090. Updates to map information 1094 may include updates to HD map 1022, such as information about construction sites, potholes, bends, floods, or other obstacles. In some examples, neural network 1092, updated neural network 1092, and / or map information 1094 may have been generated from new training and / or data received from any number of vehicles in the environment, and / or based on experience gained from training performed at a data center (e.g., using server 1078 and / or other servers).

[0268] Server 1078 can be used to train machine learning models (e.g., neural networks) based on training data. Training data can be generated by the vehicle and / or generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., where the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., where the neural network does not require supervised learning). Training can be performed according to any one or more categories of machine learning techniques, including but not limited to categories such as: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, joint learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, it can be used by the vehicle (e.g., transmitted to the vehicle via network 1090), and / or the machine learning model can be used by server 1078 to remotely monitor the vehicle.

[0269] In some examples, server 1078 can receive data from a vehicle and apply that data to a state-of-the-art real-time neural network for real-time intelligent reasoning. For example, server 1078 can convert a stream of sensor input from the vehicle into fragments, generate descriptive text for these fragments, and store representations of these descriptive texts and fragments as memories in a data store. Server 1078 can also, or alternatively, use LLM, VLM, and / or other types of machine learning models to selectively retrieve question-related memories, use the retrieved memories to generate answers to the question, and / or use the retrieved memories to perform other question-related tasks. Server 1078 may include a deep learning supercomputer powered by GPU 1084 and / or a dedicated AI computer, such as the DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 1078 may include a deep learning infrastructure in a data center using only CPU power.

[0270] The deep learning infrastructure of server 1078 may be capable of rapid real-time inference and can be used to assess and verify the health status of the processor, software, and / or associated hardware in vehicle 1000. For example, the deep learning infrastructure may receive periodic updates from vehicle 1000, such as image sequences and / or objects located in those image sequences that vehicle 1000 has already located (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure may run its own neural network to identify objects and compare them with objects identified by vehicle 1000. If the results do not match and the infrastructure concludes that the AI ​​in vehicle 1000 has malfunctioned, then server 1078 may transmit a flag to vehicle 1000, instructing vehicle 1000's fail-safe computer to take control, notify passengers, and complete a safe stopping operation.

[0271] For inference, server 1078 may include GPU 1084 and one or more programmable inference accelerators (such as NVIDIA's TensorRT). The combination of a GPU-powered server and inference acceleration enables real-time response. In other examples, such as where performance is less critical, CPU, FPGA, and other processor-powered servers can be used for inference.

[0272] In summary, the disclosed technology provides long-term perception for robots (and other machine types) operating in dynamic environments. During the memory-building phase, the robot generates a stream of video frames, location information, temporal information, and / or other types of sensor data. This data is aggregated into “segments” corresponding to discrete time intervals (e.g., every 3 seconds). A visual language model (VLM) and / or another type of machine learning model is used to generate descriptive text for each segment, and this descriptive text is converted into an embedded representation (e.g., using an embedding model). The embedded representation, location information, temporal information, and / or other information generated and / or collected during the memory-building phase are stored as the robot’s corresponding “memory” in a vector database residing on the robot and / or in remote locations accessible to the robot.

[0273] During the subsequent query phase, a Large Language Model (LLM) and / or other types of machine learning models translate the user's question to the robot into a set of queries against a vector database. These queries retrieve relevant memories from the vector database, and the retrieved memories are added to the context used by the LLM. This process is iteratively repeated using a new set of queries generated from the updated context until the LLM determines that the updated context is sufficient to answer the question. The LLM then uses the information from the updated context to generate and output the answer to the question. This answer may include text related to the question, location information, time information, and / or time duration information. Some or all of this information may also be used to generate navigation goals for the robot, trajectories for the robot, commands to move the robot toward the navigation goals, and / or other outputs that cause the robot to perform one or more actions related to the question.

[0274] One technical advantage of the disclosed technique compared to prior methods is the continuous conversion of sensor data collected by the robot (or other machine type) into efficient memory representations, which can be used to retrieve objects, scenes, and / or dynamic events perceived by the robot. Therefore, the disclosed technique improves performance on tasks involving long time spans compared to conventional methods with short memory durations. Another technical advantage of the disclosed technique is its ability to rapidly retrieve and reason about memories matching queries. Consequently, the disclosed technique allows the robot to generate answers to any set of spatiotemporal questions and perform related tasks in a timely and feasible manner. Yet another technical advantage of the disclosed technique is its ability to detect, identify, and recall “open set” objects during query processing. Therefore, compared to conventional methods, the disclosed technique allows the robot to answer a wider range of questions and / or perform a broader range of tasks related to perception and / or retrieval, which specify “closed set” objects for mapping and / or perception.

[0275] 1. In some embodiments, a method includes: converting one or more sensing inputs obtained using one or more sensors of a machine into a plurality of segments; for each segment included in the plurality of segments: generating explanatory text for the segment via executing a machine learning model; and storing a representation of the explanatory text in association with the segment in a data repository; and having the machine perform one or more actions based at least on one or more queries to the data repository.

[0276] 2. The method of claim 1, wherein storing the representation of the explanatory text in association with the fragment comprises: converting the explanatory text into an embedding corresponding to the representation of the explanatory text via executing a second machine learning model; and storing the embedding and the fragment in a vector database corresponding to the data repository.

[0277] 3. The method of any one of Clauses 1-2, wherein performing the one or more actions comprises: matching a first query included in the one or more queries with one or more fragments in the data repository; generating a second query included in the one or more queries based at least on the one or more fragments; and determining the one or more actions based at least on one or more additional fragments in the data repository that match the second query.

[0278] 4. The method of any one of Clauses 1-3, wherein generating the second query comprises: inputting the following into a second machine learning model: (i) a context including information from the one or more fragments; and (ii) a hint for generating the second query based at least on the context.

[0279] 5. The method of any one of Clauses 1-4 further includes: generating the one or more queries by executing a second machine learning model, based at least on a question from the user.

[0280] 6. The method of any one of Clauses 1-5, wherein the one or more actions include at least one of: outputting an answer to the one or more queries or navigating to a location associated with the one or more queries.

[0281] 7. The method as described in any one of Clauses 1-6, wherein the one or more queries include at least one of the following: location, time, or description.

[0282] 8. The method of any one of clauses 1-7, wherein each of the plurality of segments spans a time interval.

[0283] 9. The method of any one of Clauses 1-8, wherein the machine learning model comprises a visual language model (VLM).

[0284] 10. The method of any one of Clauses 1-9, wherein the segment comprises at least one of the following: one or more locations of the machine, one or more images captured by one or more cameras included in the one or more sensors, or one or more timestamps.

[0285] 11. In some embodiments, at least one processor includes: processing circuitry configured to cause operations to be performed, the operations including: for an individual segment of a plurality of segments of sensor data: generating descriptive text for the individual segment via executing a machine learning model; and storing a representation of the descriptive text and time and location information thereon in a data repository; receiving one or more requests; generating one or more responses to the one or more requests based at least on querying the data repository; and using one or more output devices of a robot to elicit a visual or auditory presentation of the one or more responses.

[0286] 12. At least one processor as described in Clause 11, wherein storing the representation of the descriptive text comprises: converting the descriptive text into an embedding corresponding to the representation of the descriptive text via executing a second machine learning model; and storing the embedding in a vector database corresponding to the data repository.

[0287] 13. At least one processor as described in any one of Clauses 11-12, wherein the at least one processor is included in the robot, on a locally deployed computing system of the robot, or in a remotely located data center of the robot.

[0288] 14. At least one processor as described in any one of clauses 11-13, wherein generating the one or more responses comprises: using the data repository to determine a context associated with the one or more requests; and determining one of time or location related to the context, wherein the one or more responses are generated based at least on the context, the time, or the location.

[0289] 15. At least one processor as described in any one of Clauses 11-14, wherein the one or more responses comprise at least one of the following: text information, location, time, duration, or binary answer.

[0290] 16. At least one processor as described in any one of Clauses 11-15, wherein the descriptive text describes perceptual information corresponding to the static and dynamic aspects of a scene associated with and included in the sequence of video frames in the individual segment.

[0291] 17. At least one processor as described in any one of Clauses 11-16, wherein the machine learning model is a visual language model (VLM) or a multimodal language model (MMLM), and one or more of the following: the one or more responses are generated using a second machine learning model different from the machine learning model; or the one or more responses are generated using the machine learning model.

[0292] 18. At least one processor as described in any one of Clauses 11-17, wherein said at least one processor comprises at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more analog operations; a system for performing one or more digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system implemented using a robot; a system for performing one or more conversational AI operations; a system for performing one or more generative AI operations; a system implementing one or more large language models (LLMs); a system implementing one or more visual language models (VLMs); a system implementing one or more multimodal language models; a system for generating synthetic data; a system comprising one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

[0293] 19. In some embodiments, a robot includes: one or more graphics processing units (GPUs); one or more central processing units (CPUs); one or more hardware accelerators; one or more sensors; and a data repository, wherein the robot is configured to perform one or more operations based at least on one or more descriptive memories stored in the data repository, wherein the one or more descriptive memories are determined using sensor data acquired using the one or more sensors over one or more time intervals, and wherein the one or more descriptive memories are stored as including at least descriptive text, time, and associated location.

[0294] 20. The robot as described in Clause 19, wherein the robot comprises at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulated operations; a system for performing one or more digital twin operations; a system for performing optical transmission simulation; a system for performing collaborative content creation of 3D assets; a system for performing one or more deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; a system implemented using a robot; a system for performing one or more conversational AI operations; a system for performing one or more generative AI operations; a system implementing one or more large language models (LLMs); a system implementing one or more visual language models (VLMs); a system implementing one or more multimodal language models; a system for generating synthetic data; a system comprising one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources.

[0295] This disclosure can be described in the general context of machine-usable instructions or computer code, including computer-executable instructions such as program modules, which are executed by a computer or other machine such as a personal digital assistant or other handheld device. Typically, a program module, including routines, programs, objects, components, data structures, etc., refers to code that performs a specific task or implements a specific abstract data type. This disclosure can be practiced in a wide variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. This disclosure can also be practiced in distributed computing environments in which tasks are performed by remote processing devices linked via a communication network.

[0296] As used herein, the phrase "and / or" relating to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, "element A, element B, and / or element C" can include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or element A, B, and C. Furthermore, "at least one of element A or element B" can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0297] This document describes in detail the subject matter of this disclosure to meet statutory requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the discloser has envisioned that the claimed subject matter may be embodied in other ways to include steps different from or similar combinations of steps described herein in conjunction with other current or future techniques. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of the method employed, these terms should not be construed as suggesting any particular order among or between the various steps disclosed herein, unless the order of the steps is explicitly described.

Claims

1. A method comprising: Convert one or more sensor inputs obtained using one or more sensors of the machine into multiple segments; For each of the plurality of segments: Descriptive text for the fragment is generated by executing a machine learning model; and The description text is stored in the data repository in association with the fragment; as well as The machine performs one or more actions based on at least one or more queries to the data repository.

2. The method of claim 1, wherein storing the representation of the explanatory text associated with the fragment comprises: The descriptive text is converted into an embedding corresponding to the representation of the descriptive text by executing a second machine learning model; as well as The embeddings and fragments are stored in a vector database corresponding to the data repository.

3. The method of claim 1, wherein performing the one or more actions comprises: The first query included in the one or more queries will be matched with one or more fragments in the data repository; A second query, which is included in the one or more queries, is generated based on at least one or more of the fragments. as well as The one or more actions are determined based on at least one or more additional fragments in the data repository that match the second query.

4. The method of claim 3, wherein generating the second query comprises: The following should be input into the second machine learning model: (i) context including information from the one or more segments; (ii) a hint for generating the second query based at least on the context.

5. The method of claim 1, further comprising: The one or more queries are generated by executing a second machine learning model, based at least on the user's question.

6. The method of claim 1, wherein the one or more actions include at least one of: outputting an answer to the one or more queries or navigating to a location associated with the one or more queries.

7. The method of claim 1, wherein the one or more queries include at least one of the following: location, time, or description.

8. The method of claim 1, wherein each of the plurality of segments spans a time interval.

9. The method of claim 1, wherein the machine learning model includes a visual language model (VLM).

10. The method of claim 1, wherein the fragment comprises at least one of the following: one or more locations of the machine, one or more images captured by one or more cameras included in the one or more sensors, or one or more timestamps.

11. At least one processor, comprising: Processing circuitry for causing an operation to be performed, the operation including: For an individual segment within a multiple segments of sensor data: Descriptive text for the individual segments is generated by executing a machine learning model; and The descriptive text, along with its time and location information, is stored in a data repository. Receive one or more requests; At least based on querying the data repository, one or more responses to the one or more requests are generated; and Use one or more output devices of the robot to elicit a visual or auditory presentation in response to the one or more of them.

12. The at least one processor of claim 11, wherein storing the representation of the descriptive text comprises: The descriptive text is converted into an embedding corresponding to the representation of the descriptive text by executing a second machine learning model; as well as The embedding is stored in a vector database corresponding to the data repository.

13. The at least one processor of claim 11, wherein the at least one processor is included in the robot, on a locally deployed computing system communicatively coupled to the robot, or in a data center communicatively coupled to a remote location of the robot.

14. The at least one processor of claim 11, wherein generating the one or more responses comprises: Use the data repository to determine the context associated with the one or more requests; as well as Determine either the time or the location relevant to the context. The one or more responses are generated based at least on the context, the time, or the location.

15. The at least one processor of claim 11, wherein the one or more responses include at least one of the following: text information, location, time, duration, or binary answer.

16. The at least one processor of claim 11, wherein the descriptive text describes perceptual information corresponding to the static and dynamic aspects of a scene associated with and included in the sequence of video frames in the individual segments.

17. The at least one processor of claim 11, wherein the machine learning model is a visual language model (VLM) or a multimodal language model (MMLM), and one of the following: The one or more responses are generated using a second machine learning model that is different from the machine learning model; or The one or more responses are generated using the machine learning model.

18. The at least one processor as claimed in claim 11, wherein the at least one processor is included in at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system for performing one or more simulation operations; A system for performing one or more digital twin operations; A system for performing optical transmission simulation; A system for collaborative content creation of 3D assets; A system for performing one or more deep learning operations; Systems implemented using edge devices; A system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; Systems implemented using robots; A system for performing one or more conversational AI operations; A system for performing one or more generative AI operations; A system that implements one or more large language model LLMs; A system that implements one or more Visual Language Models (VLMs); A system that implements one or more multimodal language models; A system for generating synthetic data; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.

19. A robot comprising: One or more graphics processing units (GPUs); One or more central processing units (CPUs); One or more hardware accelerators; One or more sensors; as well as Data repository The robot is used to perform one or more operations based at least on one or more descriptive memories stored in the data repository, wherein the one or more descriptive memories are determined using sensor data obtained using the one or more sensors over one or more time intervals, and wherein the one or more descriptive memories are stored as including at least descriptive text, time, and associated location.

20. The robot of claim 19, wherein the robot is included in at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system for performing one or more simulation operations; A system for performing one or more digital twin operations; A system for performing optical transmission simulation; A system for collaborative content creation of 3D assets; A system for performing one or more deep learning operations; Systems implemented using edge devices; A system for generating or presenting at least one of virtual reality content, augmented reality content, or mixed reality content; Systems implemented using robots; A system for performing one or more conversational AI operations; A system for performing one or more generative AI operations; A system that implements one or more large language model LLMs; A system that implements one or more Visual Language Models (VLMs); A system for implementing one or more multimodal language models; a system for generating synthetic data; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or a system that is at least partially implemented using cloud computing resources.

Citation Information

Patent Citations

  • Method for programmable timeouts of tree traversal mechanisms in hardware

    US10885698B2