AUTOMATIC HAZARD DETECTION AND NOTIFICATION USING MULTIMODAL MODELS
Multimodal machine learning models automatically detect and notify road hazards, addressing user input-related issues in conventional systems, enhancing safety and accuracy in hazard reporting.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- NVIDIA CORP
- Filing Date
- 2025-11-18
- Publication Date
- 2026-05-21
AI Technical Summary
Conventional systems for reporting road hazards require user input, leading to potential distractions, inaccuracies, and incomplete information, which can result in unsafe driving decisions.
Utilizing multimodal machine learning models to automatically detect and notify road hazards using sensor data from vehicles, reducing the need for user input and enhancing accuracy.
This approach increases safety by minimizing driver distractions and improving the accuracy of hazard reporting, ensuring timely and reliable information is provided to drivers and other users.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND
[0001] Determining information about road hazards, such as vehicle accidents blocking roads, the presence of emergency vehicles on and / or near roads, the presence of animals on and / or near roads, road construction, and the like, is important for both human driving and the autonomous and semi-autonomous functionality of machines. Thus, driving applications allow users to report information describing road hazards, such as the types of hazards, their locations, and important details (how many vehicles are involved in an accident, waiting times, etc.). Furthermore, these driving applications then make the reported information available to other users, who use it to perform various driving tasks.Human drivers can, for example, reduce speeds when approaching road hazards and / or change routes to avoid road hazards.
[0002] Conventional systems that provide these driving applications, however, require user input, such as reporting, verifying, and / or updating information, which can increase the risk of hazards caused by distracted drivers. For example, a driver reporting a road hazard using a driving application may be at least partially distracted when entering the information into a user device. Furthermore, conventional systems may receive inaccurate and / or incomplete information—for example, from users accidentally entering incorrect information and / or from malicious users intentionally reporting false information—and / or may fail to receive information about unreported road hazards.This false, unreported and / or incomplete information can also lead drivers to take unnecessary action for road hazards that do not exist, and / or lead drivers who rely too heavily on the driving application to refrain from taking necessary action for unreported road hazards. SUMMARY
[0003] The invention is defined by the claims. To illustrate the invention, aspects and embodiments that may or may not be within the scope of protection of the claims are described herein.
[0004] This section describes techniques for automatic hazard detection and notification using multimodal models. The systems and methods described here can use one or more machine learning models (one or more models), such as one or more vision language models (and / or any other type of model), to determine information about hazards in an environment. For example, sensor data obtained using one or more sensors on a machine can be processed, along with data representing one or more prompts associated with hazard identification, using the one or more models. Based on at least this processing, the one or more models can generate output data representing information associated with a hazard, such as...a type of hazard, a location of the hazard and / or any other details related to the hazard.
[0005] Embodiments of the present disclosure relate to automatic hazard detection and notification using multimodal models. The systems and methods described herein can use one or more machine learning models, such as one or more vision language models (VLMs), one or more multimodal language models (MMLMs), and / or any other type of model to determine information associated with hazards present in an environment. Sensor data obtained using one or more sensors of a machine, such as image data, audio data, LiDAR data, ultrasonic data, radar data, input data, and / or the like, can be processed using the one or more models. In some examples, additional data is processed using the one or more models, such as...Data representing one or more prompts associated with hazard identification. Based at least on processing, the one or more models can generate output data representing information associated with a hazard, such as the type of hazard, the location of the hazard, and / or any other details related to the hazard. Systems and procedures described herein can then perform tasks using this information, such as providing this information to one or more application systems that notify other users of road hazards.
[0006] In contrast to conventional systems, the systems of this disclosure, in some embodiments, can use one or more models to automatically determine information associated with hazards and / or to report the information to other users. Thus, the systems of this disclosure may require no or only minimal user input to identify and / or report hazards. This can increase safety for users and / or pedestrians compared to conventional systems, for example, by reducing driver distractions associated with hazard notification.Furthermore, this can increase the accuracy associated with hazard reporting compared to conventional systems, as one or more models can be configured to detect any number of hazards, generate information associated with the hazards, verify that the information is accurate using one or more techniques, and / or report the accurate information.
[0007] Further features of the disclosure are characterized by the independent and dependent claims.
[0008] Any feature of one aspect of the disclosure can be applied in any suitable combination to other aspects of the disclosure. In particular, procedural aspects can be applied to apparatus or system aspects, and vice versa.
[0009] Furthermore, features implemented in hardware can be implemented in software and vice versa. Any reference to software and hardware features herein should be interpreted accordingly.
[0010] Each system or device feature described herein can also be provided as a process feature, and vice versa. System and / or device aspects that are functionally described (including means plus functional features) can alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and allocated working memory.
[0011] It is also understood that certain combinations of the various features described and defined in each aspect of the revelation can be implemented and / or provided and / or used independently of one another.
[0012] The disclosure also provides computer programs and computer program products comprising software code designed to perform one of the methods described herein when executed on a data processing device and / or to embody one of the device and system features described herein, including one or all component steps of a method.
[0013] The disclosure also includes a computer or computing system (including networked or distributed systems) with an operating system that supports a computer program for carrying out the procedures described herein and / or for embodying the device or system features described herein.
[0014] The revelation also provides a computer-readable medium on which one or more of the aforementioned computer programs are stored.
[0015] The revelation also provides a signal that carries one or more of the aforementioned computer programs.
[0016] The disclosure extends to methods and / or devices and / or systems as described herein with reference to the accompanying drawings.
[0017] Aspects and embodiments of the disclosure will now be described purely by way of example with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The systems and procedures presented here for automatic hazard detection and notification using multimodal models are described in detail below with reference to the accompanying drawings. These show: Fig. 1A An exemplary data flow diagram for a process for automatically determining and / or providing information associated with hazards, according to some embodiments of the present disclosure; Fig. 1B a data flow diagram of a process for training one or more machine learning models to determine information associated with hazards, according to some embodiments of the present disclosure; Fig. 2 an example of a machine that navigates in an environment containing multiple hazards, according to some embodiments of the present disclosure; Fig. 3 an example of using one or more machine learning models to determine information associated with hazards, according to some embodiments of the present disclosure; Fig. 4A an example of processing input data using one or more iterations to determine information associated with a hazard, according to some embodiments of the present disclosure; Fig. 4B an example of processing input data using one or more iterations associated with one or more prompts, according to some embodiments of the present disclosure; Fig. 5 an example of the use of language to determine information associated with a hazard, according to some embodiments of the present disclosure; Fig. 6A-6B is an example of a user interface associated with an application that provides information related to hazards, according to some embodiments of the present disclosure; Fig. 7 An example of one or more systems configured to perform one or more of the processes described herein, according to some embodiments of the present disclosure. Fig. 8 a flowchart showing a method for automatically determining and providing information associated with a hazard, according to some embodiments of the present disclosure. Fig. 9 a flowchart showing a procedure for determining how information associated with a hazard should be reported, according to some embodiments of the present disclosure; Fig. 10A a block diagram of an exemplary generative language model system suitable for use in implementing at least some embodiments of the present disclosure; Fig. 10B a block diagram of an exemplary generative language model containing a transformer-encoder-decoder suitable for use in implementing at least some embodiments of the present disclosure; Fig. 10C a block diagram of an exemplary generative language model containing a decoder-transformer-only architecture suitable for use in implementing at least some embodiments of the present disclosure; Fig. 11A an illustration of an exemplary autonomous vehicle, according to some embodiments of the present disclosure; Fig. 11B is an example of camera locations and fields of view for the exemplary autonomous vehicle from Fig. 11A, according to some embodiments of the present disclosure; Fig. 11C a block diagram of an exemplary system architecture for the exemplary autonomous vehicle from Fig. 11A, according to some embodiments of the present disclosure; Fig. 11D a system diagram for the communication between one or more cloud-based servers and the example autonomous vehicle from Fig. 11A, according to some embodiments of the present disclosure; Fig. 12 a block diagram of an exemplary computing device suitable for use in the implementation of some embodiments of the present disclosure; and Fig. 13 a block diagram of an exemplary data center suitable for use in the implementation of some embodiments of the present disclosure. DETAILED DESCRIPTION
[0019] Systems and methods are disclosed with respect to automatic hazard detection and notification using multimodal models. Although the present disclosure relates to an exemplary autonomous or semi-autonomous vehicle or an exemplary autonomous or semi-autonomous machine 1100 (here alternatively referred to as "Vehicle 1100", "Ego-Vehicle 1100", "Ego-Machine 1100" or "Machine 1100"), which is exemplary with respect to Fig. The fact that the systems and procedures described herein can be described in sections 11A-11D is not intended to be restrictive. For example, the systems and procedures described herein can be used without restriction by non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, steered and unsteered robots or robotic platforms, warehouse vehicles, all-terrain vehicles, vehicles coupled to one or more trailers, hydrofoils, boats, shuttle vehicles, emergency vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, underwater vehicles, drones, and / or other types of vehicles.Furthermore, although the present disclosure may be described in relation to the identification and reporting of information associated with hazards, this is not intended to be limiting, and the systems and methods described herein may be used in augmented reality, virtual reality, mixed reality, robotics, security and surveillance, autonomous or semi-autonomous machine applications and / or any other technology fields where the identification and reporting of information associated with hazards may be used.
[0020] One or more systems can, for example, obtain sensor data using one or more sensors from a machine, such as a vehicle, a semi-autonomous vehicle, an autonomous vehicle, and / or other types of robots navigating their environment. As described here, the sensor data can, without limitation, include image data obtained using one or more image sensors, LiDAR data obtained using one or more LiDAR sensors, RADAR data obtained using one or more RADAR sensors, location data obtained using one or more location sensors, audio data obtained using one or more microphones, and input data obtained using one or more input sensors (e.g., a touch-sensitive display, etc.).) are obtained, and / or contain any other type of sensor data obtained using any other type of sensor. The one or more systems can then process at least some of the sensor data to determine whether a hazard may be present in the environment, and / or to determine information associated with a hazard that is present in the environment.
[0021] As described here, a hazard can, without limitation, include: a roadway that is blocked (e.g., by an object such as a pedestrian, animal, vehicle, and / or any other type of object, and / or due to a collision), an object located near the roadway, the presence of emergency services and / or personnel, road construction, a hazardous driving condition (e.g., a gravel road, a pothole, etc.), dangerous drivers, and / or any other type of hazard that may alter the way a driver, semi-autonomous vehicle, and / or autonomous vehicle navigates. Additionally, information associated with a hazard can, without limitation, include the type of hazard, the location of the hazard, a waiting state created by the hazard (e.g., whether a lane is blocked, a time frame for detouring around the hazard, etc.).), all potential updates regarding the hazard (e.g., whether the hazards have already been reported), a number of objects (e.g., vehicles, emergency personnel, etc.) associated with the hazard, one or more types of objects associated with the hazard, and / or any other information describing the hazard.
[0022] If, in the first example, a hazard involves a collision between two vehicles, then one or more systems can determine information describing the nature of the hazard, a collision, the number of vehicles involved, whether the vehicles are blocking one or more lanes, and / or any other information associated with the collision. If, in the second example, a hazard involves a police officer conducting a traffic stop, then one or more systems can determine information describing the nature of the hazard as a traffic stop, the number of police vehicles and / or officers involved, the side of the roadway where the traffic stop is taking place, and / or any other information associated with the traffic stop.However, in a third example, if a hazard involves an obstacle in a lane, one or more systems can determine information describing that the nature of the hazard is a blocked lane, the type of object causing the obstacle, the lane being blocked, and / or any other information associated with the obstacle.
[0023] As described here, the one or more systems can use one or more techniques to determine the information associated with hazards using the sensor data. In some examples, the one or more systems can, for instance, input at least some of the sensor data into one or more machine learning models (the one or more hazard models), such as one or more speech models, one or more vision-speech models, one or more text-to-speech models, one or more speech-to-text models, and / or any other type of model. The one or more systems can, for example, input at least image data into the one or more hazard models, representing one or more images depicting the environment that at least partially surrounds the machine.Additionally, in some examples, the one or more systems input additional data into the one or more hazard models, such as prompt data representing one or more prompts associated with hazard identification. As described here, in some examples, a prompt may contain a general prompt to identify any hazards located near the machine. Additionally or alternatively, in some examples, a prompt may contain a specific prompt to identify a particular type of hazard (e.g., a collision) located near the machine. In any of the examples, based at least on the processing of the input data, the one or more hazard models can generate and / or output data representing information associated with one or more hazards that may be located near the machine.
[0024] In some examples, one or more systems may be configured to process the input data using one or more iterations to identify the information associated with a hazard. In a first example, the one or more systems may initially input the input data (e.g., sensor data, etc.) into the one or more hazard models along with a first prompt, where the first prompt is associated with the general identification of hazards. Then, if the output data from the one or more hazard models indicates that a hazard is present in the environment, the one or more systems may input at least a portion of the input data, the output data, and a second, more specific prompt into the one or more hazard models.The second prompt can, for example, be assigned to identifying a type of hazard and / or the location of the hazard. Thus, one or more hazard models can then generate additional output data providing further details about the hazard, such as the type of hazard and / or its location. This process can then be repeated for any number of iterations, so that one or more systems receive additional information associated with the hazard.
[0025] In a second example, one or more systems can initially input the data into one or more hazard models along with a first prompt, where the first prompt is associated with a first type of hazard. For example, the first prompt might be associated with identifying collisions in the vicinity, such as by containing "Determine if there is a collision in the vicinity." The one or more systems can also input the data into one or more hazard models along with a second prompt, where the second prompt is associated with a second type of hazard. For example, the second prompt might be associated with identifying emergency vehicles in the vicinity, such as by containing "Identify all emergency vehicles located in the vicinity."The one or more systems can then use all the outputs from the one or more hazard models to determine whether a hazard is present in the environment and / or to determine the information associated with the hazard. For example, an initial output might indicate that there is no collision in the environment, while a second output might indicate that an emergency vehicle is present in the environment.
[0026] In some examples, the one or more hazard models can be fine-tuned to identify hazards and / or one or more types of hazards. For instance, the one or more hazard models can be trained to identify collisions, emergency vehicles, and / or any other type of hazard in the environment. In some examples, the one or more hazard models may include a general model used to perform various tasks beyond simply detecting hazards. In such examples, the one or more systems can use prompts to instruct the one or more hazard models to perform the tasks described here, such as identifying hazards, generating information associated with the hazards, and / or reporting hazards.
[0027] In some examples, the one or more systems can use one or more additional models (e.g., the one or more reporting models), such as one or more language models (and / or any other type of model), to determine how to report the information associated with the hazards. For example, the one or more systems can input the output data from the one or more hazard models into the one or more reporting models. Based at least on the processing of the output data, the one or more reporting models can then generate output data representing one or more techniques associated with reporting the information.The output data can, for example, be a request to verify the information, a request for additional information related to a hazard, a request to provide the information, and / or any other type of request. The one or more systems can then perform one or more operations using the output data.
[0028] If the output data from one or more reporting models in a first example represents a request to verify the information, the one or more systems can provide content that fulfills this request. As described here, the content can include, without limitation, audio representing speech, visual content representing text and / or graphics, and / or any other type of content. The one or more systems can then receive input from the user, such as audio representing speech, input representing a selection, input representing text, and / or any other type of input, and use the input to determine whether the information needs to be verified. For example, if the input indicates that the information is correct, the one or more systems can verify the information associated with the hazard.However, if the input indicates that the information is incorrect, one or more systems may be unable to verify the information associated with the hazard and / or may request additional information from the user.
[0029] If the output data of one or more reporting models in a second example represents a request for additional information associated with the hazard, the one or more systems can re-provide the content that fulfills the request. Additionally, the one or more systems can receive an input representing the additional information associated with the hazard. This input could, for example, contain audio data representing language that describes the hazard, such as the type of hazard and / or its location. The one or more systems can then use this input to update the information associated with the hazard.In some examples, updating the information may involve using one or more hazard models, one or more notification models, and / or any other type of model to process the input in order to generate output data representing the updated information.
[0030] In some examples, the one or more systems can use outputs from one or more other processing components of the machine, such as one or more systems, one or more classifiers, one or more machine learning models, one or more neural networks, one or more modules, one or more software applications, one or more hardware processors, and / or the like, to determine information associated with hazards. If, in a first example, the output from the one or more hazard models indicates that a hazard contains an emergency vehicle, then the one or more systems can process at least some of the sensor data (e.g., audio data, etc.) using one or more processing components configured to determine types of emergency vehicles.In this way, one or more systems can use one or more hazard models to determine initial information associated with the hazard, such as the presence of an emergency vehicle in the vicinity, and one or more processing components can determine additional information, such as the type of emergency vehicle. In a second example, if the output from one or more hazard models indicates that a hazard contains an obstacle, one or more systems can process at least some of the sensor data (e.g., image data, radar data, LiDAR data, etc.) using one or more processing components (e.g., one or more perception systems) configured to determine information associated with objects.In this way, one or more systems can use one or more hazard models to determine initial information associated with the hazard, such as that the hazard contains an obstacle and / or a type of object causing the obstacle, and one or more processing components to determine additional information, such as the location of the object in the environment.
[0031] The one or more systems can then perform one or more processes using the information associated with the hazards. In some examples, the one or more systems can, for instance, send data representing at least part of the information to one or more additional systems (the one or more application systems). As described here, the one or more application systems can be assigned to provide information associated with the hazards to users. For example, the one or more application systems can generate, manage, and / or update an application that provides user interfaces displaying information about hazards present in environments.For example, in the case of a hazard, a user interface can display at least one type of hazard, the hazard's location in the environment, and / or any other details relevant to the hazard. Thus, one or more application systems can use the received data to update one or more user interfaces to include at least some of the information associated with the hazards. In this way, other users of the application are able to use the one or more user interfaces to identify information associated with the hazards and / or take precautions based on the information, such as changing driving actions and / or taking new routes.
[0032] Additionally or alternatively, in some examples, the one or more systems can use information associated with hazards to determine one or more operations for the machine to perform. In a first example, the one or more systems can cause the machine to output content corresponding to a hazard, such as audio content stating the information associated with the hazard, visual content displaying the information associated with the hazard, and / or any other type of content. In a second example, if the machine is semi-autonomous and / or autonomous, the one or more systems can determine one or more trajectories for the machine to navigate to avoid the hazard and / or navigate safely around it. The one or more systems can then cause the machine to execute at least one trajectory.However, in a third example, the machine can notify a user to contact emergency services and / or automatically contact emergency services based on the information. In such examples, the machine can also send at least some of the information to the emergency services so that they can determine the hazard. While these are just a few example operations that can be performed using the information associated with the hazards, one or more systems may cause the machine to perform additional and / or alternative operations in other examples.
[0033] While the examples here describe identifying hazards, determining information associated with hazards, and / or providing the information associated with hazards (e.g., to update an application), other examples may use similar processes to report information for other types of objects and / or features in environments.
[0034] In some examples, models (e.g., machine learning models, deep neural networks, language models, LLMs, VLMs, multimodal language models, perceptual models, tracking models, fusion models, transformer models, diffusion models, encoder-only models, decoder-only models, encoder-decoder models, neural rendering field (NERF) models, etc.) can be packaged as a microservice, such as an inference microservice (e.g., NVIDIA's NIMs), which may contain a container (e.g., an operating system (OS)-level virtualization package) that can contain an application programming interface (API) layer, a server layer, a runtime layer, and / or a model "engine." The inference microservice might contain, for example, the container itself and one or more models (e.g., weights and biases). In some cases, such as...If the one or more machine learning models are small enough (i.e., have a sufficiently small number of parameters), they can be contained within the container itself. In other examples—such as when the one or more models are large—they can be hosted / stored in the cloud (e.g., in a data center) and / or on-premises and / or at the edge (e.g., on a local server or computing device, but outside the container). In such implementations, the one or more models can be accessed via one or more APIs, such as REST APIs.Therefore, in some embodiments, one or more of the machine learning models described herein can be deployed as an inference microservice to accelerate the deployment of one or more models in any cloud, data center, or edge computing system while ensuring data security. For example, the inference microservice may include one or more APIs, a pre-configured container for simplified deployment, an optimized inference engine (e.g., built using standardized software for deploying and running AI models, such as NVIDIA's Triton inference server), and / or one or more APIs for high-performance deep learning inference, which may include inference runtime and model optimizations that provide low latency and high throughput for production applications—such as…TensorRT from NVIDIA) and / or enterprise management data for telemetry (e.g., including identity, metrics, health checks, and / or monitoring).
[0035] The one or more machine learning models described herein can be included as part of the microservice, along with an accelerated infrastructure capable of being deployed with a single command and / or orchestrated and automatically scaled using a container orchestration system on an accelerated infrastructure (e.g., from a single device to the size of a data center). Therefore, the inference microservice can include the one or more machine learning models (e.g., optimized for high-performance inference), inference runtime software to execute the one or more machine learning models and provide outputs / responses to inputs (e.g., user queries, prompts, etc.), and enterprise management software to provide health checks, identity verification, and / or other monitoring.In some embodiments, the inference microservice may include software to perform an on-premises exchange and / or update of one or more machine learning models. During the exchange or update, the software performing the exchange / update may retain the user configurations of the inference runtime software and the enterprise management software.
[0036] Additionally, in some embodiments, the systems and methods described here can be performed in a simulation environment (e.g., NVIDIA's DriveSIM, ISAAC GYM, and / or ISAAC SIM) using simulated data (e.g., simulated sensor data from a virtual or simulated machine). Simulated sensor data and / or map data (simulated or real) can be used, for example, to perform various operations within the simulation environment, such as generating the simulation data and / or operating a machine. These simulated operations can be used to test the performance of the underlying algorithms, systems, and / or processes before deployment in the real world. In some cases, simulation can be used to generate synthetic training data, such as training data containing hazards, landmarks, features, objects, etc.to generate the synthetic training data (in addition to or alternatively from real data) so that it can then be processed to perform one or more of the operations described here.
[0037] In each example, such as when a simulation environment is used for testing, verification, training, etc., the simulation environment and / or the associated training data can be rendered or otherwise generated using one or more light transport algorithms, such as ray tracing and / or path tracing algorithms. In some embodiments, the simulation environment and / or one or more objects, features, or components thereof can be generated or managed within a three-dimensional (3D) content collaboration platform (e.g., NVIDIA's OMNIVERSE) for industrial digitization, generative physical AI, and / or other use cases, applications, or services. The content collaboration platform or system can, for example, be a system for using or developing Universal Scene Descriptor (USD) data (e.g., OpenUSD) for managing objects, features, scenes, etc.in a simulated environment, digital environment, etc. The platform can include real-world physics simulation, such as using NVIDIA's PhysX SDK, to simulate real-world physics and physical interactions with simulations hosted by the platform. The platform can integrate OpenUSD, along with ray tracing / path tracing / light transport simulation (e.g., NVIDIA's RTX rendering technologies), into software tools and simulation workflows to build, train, deploy, or test AI systems, such as systems for testing, validating, and training (e.g., machine learning models, neural networks, etc.) and / or other tasks related to automotive, robotics, machinery, or other applications.
[0038] In some embodiments, the teleoperation or remote control of a vehicle or other machine can be performed using a remote control or teleoperation system. The systems and methods described here can be used, for example, to identify information associated with hazards in the environment. This information can then be included in a visualization or representation of the environment to assist a remote operator in guiding an autonomous or semi-autonomous machine through an environment or in providing waypoints or other instructions for control or navigation. For example, information relating to hazards can be provided to a remote operator (e.g., via a visualization) to assist them in making navigation, planning, and / or control decisions for the (at least partially) remotely controlled vehicle or machine.
[0039] The systems and procedures described here can be used without restriction by non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more adaptive driver assistance systems (ADAS)), autonomous vehicles or machines, steered and non-steered robots or robotic platforms, warehouse vehicles, all-terrain vehicles, vehicles coupled to one or more trailers, hydrofoils, boats, shuttle vehicles, emergency vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, underwater vehicles, drones, and / or other vehicle types.Furthermore, the systems and methods described here can be used for a variety of purposes, including but not limited to machine control, machine locomotion, machine propulsion, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and monitoring, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twinning, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or other suitable applications.
[0040] The disclosed embodiments can comprise a variety of different systems, such as automotive systems (e.g.,a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine, systems implemented with a robot, aviation systems, media systems, boat systems, intelligent area surveillance systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using an edge device, systems implementing large language models (LLMs), systems implementing one or more vision language models (VLMs), systems implementing one or more multimodal language models, systems using or employing one or more interference microservices, systems implementing one or more machine learning models in a service or microservice together with an operating system-level virtualization package (e.g.,integrate or deploy systems that include one or more virtual machines (VMs), systems that perform operations to generate synthetic data, systems that are at least partially implemented in a data center, systems that perform operations using conversational AI, systems that perform light transport simulations, systems that perform collaborative content creation for 3D assets, systems that perform operations using generative AI, systems that are at least partially implemented using cloud computing resources, and / or other types of systems.
[0041] With reference to Fig. 1A, illustrated Fig. 1A An exemplary data flow diagram for a process 100 for automatically determining and / or providing information associated with hazards according to some embodiments of the present disclosure; It is pointed out that this and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, arrays, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as discrete or distributed components, or in conjunction with other components, in any suitable combination and at any suitable location. Various functions described herein that are performed by entities may be executed by hardware, firmware, and / or software.Various functions can be performed, for example, by a processor that executes instructions stored in main memory. In some embodiments, the systems, methods, and processes described here can be implemented using similar components, features, and / or functions to those of the exemplary autonomous vehicle 1100 from [Company Name]. Fig. 11A-11D, the exemplary calculating device 1200 of Fig. 12 and / or the exemplary data center 1300 of Fig. 13 will be executed.
[0042] The process 100 can, for example, include one or more sensors 102 that generate sensor data 104. As described here, the sensor data 104 can, without limitation, include image data obtained using one or more image sensors, LiDAR data obtained using one or more LiDAR sensors, RADAR data obtained using one or more RADAR sensors, location data obtained using one or more location sensors, audio data obtained using one or more microphones, input data obtained using one or more input sensors (e.g., a touch-sensitive display, etc.), and / or any other type of sensor data obtained using any other type of sensor.In some examples, the one or more sensors 102 may be included as part of and / or associated with a machine. For example, the one or more sensors 102 may be included as part of a vehicle, a semi-autonomous vehicle, and / or an autonomous vehicle (e.g., an exemplary autonomous vehicle 1102) that navigates in an environment. Additionally or alternatively, in some examples, the one or more sensors 102 may be included as part of one or more user devices 106. For example, the one or more sensors 102 may be part of a user device 106 that is associated with an occupant of the machine.
[0043] As described here, sensor data 104 can represent one or more hazards located in the environment. In some examples, a hazard, without limitation, can include: a roadway that is blocked (e.g., by an object such as a pedestrian, animal, vehicle, and / or any other type of object, and / or due to a collision), an object located near the roadway, the presence of emergency services and / or personnel, road construction, a hazardous driving condition (e.g., a gravel road, a pothole, etc.), dangerous drivers, and / or any other type of hazard that can alter the way a driver navigates a semi-autonomous and / or autonomous vehicle. In one example, a hazard might be an object such as...A hazard can include a pedestrian, an animal, a vehicle, and / or any other type of object located within at least one lane of a roadway. A second example is an emergency vehicle, such as a police car, an ambulance, and / or the like, located on or near a roadway. A third example is a pothole located on the roadway.
[0044] Fig. Figure 2 illustrates, for example, an example of a machine 202 navigating in an environment 204 containing multiple hazards, according to some embodiments of the present disclosure. As shown, the environment 204 may contain at least a first road 206 with two lanes 208(1)-(2), a second road 210 with two lanes 212(1)-(2), and a third road 214 with two lanes 216(1)-(2), wherein the machine 202 navigates on the first lane 208(1) of the first road 206. During navigation, the machine 202 may generate sensor data (which may contain and / or be similar to the sensor data 104) representing the multiple hazards.For example, a first hazard may be assigned to an emergency vehicle 218 located on the first road 206, a second hazard may be assigned to a collision 220 between two vehicles on the second road 210, and a third hazard may be assigned to an obstacle 222 located on the third road 214.
[0045] Referring again to the example of Fig. 1A, the process 100 may include inputting at least some of the sensor data 104 into one or more machine learning models 108 (one or more models 108), such as one or more speech models, one or more vision speech models, one or more text-to-speech models, one or more speech-to-text models, and / or any other type of model. For example, if the one or more models 108 include the one or more vision speech models, then the process 100 may include inputting at least image data representing one or more images of the environment into the one or more models 108. As shown, the process 100 may further include inputting additional data into the one or more models 108, such as prompt data 110 representing one or more prompts configured to instruct the one or more models 108 to perform one or more tasks.The one or more prompts may be assigned to guide the one or more Model 108s to identify hazards in the environment and / or to determine information associated with the identified hazards.
[0046] In some examples, a prompt might contain a general prompt that directs the one or more Models 108 to perform a general task associated with hazard identification. For a first example, a prompt might direct the one or more Models 108 to identify one or more hazards using at least the Sensor Data 104. A prompt might include, for example, "See one or more hazards in the environment" or "See a collision, emergency vehicle, objects, or animals in the environment." In a second example, a prompt might direct the one or more Models 108 to generate information associated with an identified hazard. A prompt might include, for example, "Please provide information about each identified hazard in the environment."
[0047] In some examples, a prompt may additionally or alternatively contain a specific prompt that directs one or more Models 108 to perform a particular task associated with hazard identification. In a first example, a prompt may direct one or more Models 108 to identify a specific type of hazard, such as emergency vehicles located on a roadway. For example, a prompt might include: "See a vehicle being stopped on the road," "Are there any objects on the road," or "See an emergency vehicle in the vicinity." In a second example, a prompt may direct one or more Models 108 to generate specific information associated with a hazard, such as...Information describing a location associated with a hazard, as represented by one or more sensor displays of sensor data 104. A prompt might include, for example, "Can you tell us the location of the vehicle on the road?" or "Can you determine the type of emergency vehicle on the road?"
[0048] Additionally or alternatively, in some examples the prompts may contain multiple prompts that are entered into one or more Models 108 in a row to instruct one or more Models 108 to perform different tasks, which is described in more detail here.
[0049] The process 100 may include one or more models 108 processing the sensor data 104 and / or the prompt data 110 and, based on at least this processing, generating and / or outputting data that represents hazard information 112 associated with one or more potential hazards located in the environment. In some examples, information 112 associated with a hazard may, without limitation, include a type of hazard, a location of the hazard, a waiting state created by the hazard (e.g., whether a lane is blocked, a time frame for detouring around the hazard, etc.), any potential updates regarding the hazard (e.g., whether the hazards have already been reported), and a number of objects (e.g., vehicles, emergency personnel, etc.).A hazard type can describe one or more types of objects associated with the hazard, and / or any other information that describes the hazard. Furthermore, a hazard type can include an obstacle located on a roadway, a collision, an emergency vehicle, construction work, a hazardous driving condition, dangerous drivers, and / or any other type of hazard. Additionally, a hazard location can be a general location in the environment, such as the x-coordinate location, the y-coordinate location, and / or the z-coordinate location; a roadway; a lane; a shoulder of the roadway; and / or the like; or a specific location associated with a sensor representation (e.g., an image), such as a boundary shape (e.g., a boundary frame, etc.).), which indicates part of the sensor display that represents the hazard, and / or contains any other type of location information.
[0050] Fig. Figure 3 illustrates, for example, an example of the use of the one or more models 108 to determine information associated with hazards, according to some embodiments of the present disclosure. As shown, sensor data 302, obtained using one or more sensors of the machine 202, together with prompt data 304 representing one or more prompts, can be input into the one or more models 108. The one or more models 108 can then process the data and, based at least on the processing, generate output data 306(1)-(3) associated with the hazards in the environment 204. For example, and as shown, the first output data 306(1) can represent information associated with the first hazard, such as that a first type 308(1) of the first hazard contains the emergency vehicle 218 and a first location 310(1) associated with the first hazard.The first location can specify the location in some examples of the first hazard in the vicinity 204 (e.g. coordinates, the first street 206, the first lane 208(1) etc.) and / or a location in a sensor representation that represents the first hazard.
[0051] The second output data 306(2) can additionally represent information associated with the second hazard, such as that a second type 308(2) of the second hazard contains the collision 220 and a second location 310(2) associated with a second hazard. In some examples, the second location can specify the location of the second hazard in the environment 204 (e.g., coordinates, the second road 210, the second lane 212(2), etc.) and / or a location in a sensor representation that represents the second hazard. Furthermore, the third output data 306(3) can represent information associated with the third hazard, such as that a third type 308(3) of the third hazard contains the obstacle 222 and a third location 310(3) is associated with a third hazard. In some examples, the third location can be the location of the third hazard in the vicinity 204 (e.g. coordinates, the third street 214, the third lane 216(1), etc.).) and / or specify a location in a sensor display that represents the third hazard.
[0052] Referring again to the example of Fig. In some examples, process 100 may involve processing the input data using one or more iterations to identify the hazard information 112 associated with a hazard. Initially, in a first example, the sensor data 104 may be input into one or more models 108 along with a first prompt, where the first prompt is associated with the general identification of hazards. If the hazard information 112 generated using one or more models 108 then indicates that a hazard is present in the environment, the sensor data 104 (and / or additional sensor data 104), data representing the hazard information 112, and a second, more specific prompt may be input into one or more models 108. The second prompt may, for example, be associated with identifying a type of hazard and / or a location of the hazard.Thus, one or more models 108 can then generate additional hazard information 112 that provides further details about the hazard, such as the type of hazard and / or the location of the hazard. This process 100 can then be repeated for any number of iterations, so that one or more models 108 continue to generate additional hazard information 112 associated with the hazard.
[0053] Fig. Figure 4A illustrates, for example, an instance of processing input data using one or more iterations to determine information associated with a hazard, according to some embodiments of the present disclosure. As shown, initial sensor data 404(1), obtained using one or more sensors of the machine 202 together with a first prompt 406(1), can be input into the one or more models 108 during a first iteration 402. In some examples, the first prompt 406(1) can contain a general prompt that instructs the one or more models 108 to determine whether the initial sensor data 404(1) represent one or more hazards present in the environment, such as by including "Determine if there are any hazards represented by the images".Based at least on the processing of the data, one or more models 108 can generate initial output data 408(1) indicating whether one or more of the hazards are located in the environment. The initial output data 408(1) can, for example, indicate that the first hazard associated with the emergency vehicle 218 is located in the environment 204.
[0054] Next, during a second iteration, 410 second sensor data 404(2) (which may contain or differ from the first sensor data 404(1)), a second prompt 406(2), and the first output data 408(1) can be input into the one or more models 108. In some examples, the second prompt 406(2) may contain a specific prompt that directs the one or more models 108 to perform a specific task, such as identifying the type of hazard and / or the location of the hazard. Thus, based on processing the data, the one or more models 108 can generate second output data 408(2) that specifies the type of first hazard, such as the emergency vehicle 218, and / or the location associated with the first hazard.In some examples, this process can be repeated for one or more additional iterations, with each iteration using a different prompt to guide the one or more models 108 to determine additional information associated with the hazards.
[0055] Referring again to the example of Fig. In some examples, process 100 may involve inputting multiple prompts into one or more models 108 to perform multiple tasks. For example, process 100 may initially involve inputting sensor data 104 along with a first prompt into one or more models 108, where the first prompt is associated with a first type of hazard. The first prompt may be associated with identifying collisions in the environment, such as by including "Determine if there is a collision in the environment." Additionally, one or more models 108 may generate initial hazard information 112 associated with the first type of hazard. Process 100 may also involve inputting sensor data 104 along with a second prompt into one or more models 108, where the second prompt is associated with a second type of hazard.The second prompt can, for example, be assigned to identifying emergency vehicles in the vicinity, such as by including "Identify all emergency vehicles located in the vicinity". Additionally, one or more models can generate second hazard information (112) assigned to the second type of hazard.
[0056] Fig. Figure 4B illustrates, for example, an example of processing input data using one or more iterations associated with one or more prompts, according to some embodiments of the present disclosure. As shown, during a first iteration 412, sensor data 414, obtained using one or more sensors of the machine 202 together with a first prompt 416(1), can be input into one or more models 108. In some examples, the first prompt 416(1) can be associated with performing a specific task, such as identifying a particular type of hazard. For example, the first prompt 416(1) can be associated with identifying emergency vehicles in the vicinity 204.Thus, based on the processing of the data, one or more models 108 can generate initial output data 418(1) representing information related to the first hazard assigned to the emergency vehicle 218.
[0057] Additionally, the sensor data 414, obtained using the one or more sensors of the machine 202 together with a second prompt 416(2), can be input into the one or more models 108 during a second iteration 412. In some examples, the second prompt 416(2) can be assigned to performing a specific task, such as identifying a particular type of hazard. For example, the second prompt 416(2) can be assigned to identifying collisions in the environment 204. Thus, based on processing the data, the one or more models 108 can generate second output data 418(2) representing information related to the second hazard associated with the collision 220.In this way, different prompts can be used to instruct one or more models 104 to perform various tasks associated with the reporting of hazards in the environment 202.
[0058] Referring again to the example of Fig. 1A The process 100 may involve the use of one or more additional models 108 in some examples, such as one or more language models (and / or any other type of model) to determine how to report the hazard information 112 associated with the hazards. The one or more additional models 108 may, for example, process data representing at least part of the hazard information 112. Based at least on the processing of the data, the one or more additional models 108 may generate output data representing one or more techniques associated with reporting the hazard information 112.The output data can, for example, represent a request to verify hazard information 112, a request for additional hazard information 112 associated with a hazard, a request to provide the hazard information 112, and / or any other type of request. The process 100 can then include performing one or more operations based on at least the reporting, which is determined using one or more additional models 108.
[0059] For more details illustrated Fig. Figure 5 provides an example of using language to determine information associated with a hazard, according to some embodiments of the present disclosure. As shown, sensor data 504, obtained using one or more sensors of the machine 202 together with one or more prompts 506, can be input into the one or more models 108 during a first instance. Based at least on the processing of the data, the one or more models 108 can generate output data 508 representing information associated with a hazard. The output data 508 can, for example, represent text describing the hazard. One or more language models 510 (one or more text-to-speech models, one or more speech-to-text models, one or more large language models, etc.) can then process the output data 508 to determine how to report the information.For example, one or more language models 510 can generate request data 512 that represent a request to confirm the information associated with the hazard, a request to provide additional information associated with the hazard, and / or any other type of request.
[0060] Machine 202 can then provide the request to one or more users of Machine 202. For example, Machine 202 can output audio content representing the request, display visual content containing text that corresponds to the request, send the request data 512 to a user device for display to one or more users, and / or use any other technique.
[0061] Next, and in a second instance 514, the machine 202 can use one or more sensors (e.g., one or more microphones) to receive audio data 516 representing a user language that meets the request. The user language can, for example, verify the information, such as "The information associated with the hazard is correct," specify the requested information, such as "The first hazard contains an ambulance located in the first lane 208(1) of the first road 206," and / or provide any other details about the hazard. The one or more models 108 and / or the one or more language models 510 can then process the audio data 516 to generate additional output data 518 associated with the hazard.If, in a first example, the language indicates that the information is correct, then the additional output data 518 can represent the information determined by the one or more models 108, as represented by the output data 508. If, in a second example, the language indicates additional information associated with the hazard, then the additional output data 518 can represent the additional information indicated by the language and / or the information represented by the output data 508.
[0062] The processes that are in the example of Fig. The processes described in section 5 can be repeated in some examples for additional iterations. For instance, the one or more models 108 and / or the one or more language models 510 can process the additional output data 518 to determine whether further information associated with the hazard should be requested. If the one or more models 108 and / or the one or more language models 510 determine that further information should be requested, a further request for the additional information can be provided. Additionally, the one or more models 108 and / or the one or more language models 510 can process additional audio data representing additional speech that specifies this further information.In other words, these processes can continue to repeat themselves until one or more models 108 and / or one or more language models 510 determine that sufficient information related to the hazard is available.
[0063] Referring again to the example of Fig. 1A Process 100 can include processing at least some of the sensor data 104 using one or more processing components 114 associated with the machine to determine additional hazard information 116 associated with hazards. As described here, a processing component 114 can, without restriction, include a machine learning model, a neural network, a classifier, an algorithm, a module, software, hardware, and / or any other type of processing component. For example, a processing component 114 can include a perception system, a location system, a trajectory system, a tracking system, and / or any other type of system associated with the machine.
[0064] For a first example of determining additional hazard information 116, a processing component 114 can be configured to identify various types of sensors, such as police sirens, ambulance sirens, and / or any other type of emergency vehicle siren. Thus, the processing component 114 can process at least some of the sensor data 104, such as audio data representing the sound of a siren, to generate hazard information 116 that represents the type of siren and / or the type of emergency vehicle associated with that type of siren. The hazard information 116 can then (1) be processed by one or more models 108 to assist in further guiding one or more models 108 to generate hazard information 112, and / or (2) be added to the hazard information 112 to provide further details about the hazard.
[0065] In a second example, a processing component 114 can be configured to process sensor data 104, such as image data, radar data, LiDAR data, and / or any other type of sensor data, to determine information associated with objects in the environment. The processing component 114 can, for example, include and / or be associated with a machine's perception system. Thus, based on processing at least some of the sensor data 104, the processing component 114 can generate hazard information 116 that represents information associated with one or more objects corresponding to a hazard, such as one or more types of the one or more objects involved in the hazard, one or more locations of the one or more objects, and / or any other information associated with the one or more objects.The hazard information 116 can then (1) be processed by the one or more models 108 to assist in further guiding the one or more models 108 to generate the hazard information 112, and / or (2) be added to the hazard information 112 to provide further details about the hazard.
[0066] In a third example, a processing component 114 can be configured to analyze sensor data 104, such as location data, image data, radar data, LiDAR data, and / or any other type of sensor data, to determine pose information associated with the machine. The processing component 114 can, for example, include and / or be associated with a location system of the machine. Thus, based at least on the processing of the sensor data 104, the processing component 114 can generate hazard information 116 that represents the pose of the machine in the environment, such as its location (e.g., the x-coordinate location, the y-coordinate location, and / or the z-coordinate location) and / or its orientation (e.g., roll, pitch, and / or yaw).The hazard information 116 can then (1) be processed by the one or more models 108 to assist in guiding the one or more models 108 to generate the hazard information 112, and / or (2) be added to the hazard information 112 to provide further details about the hazard.
[0067] Process 100 may then include performing one or more operations using hazard information 112 (and / or hazard information 116). For example, as shown, process 100 may include providing hazard information 112 (and / or hazard information 116) to one or more application systems 118 that generate, manage, update, and / or provide an application to users, where the application may be represented by application data 120. For example, as described here, the application may provide at least information about hazards present in an environment, with the information being provided by users of the application, emergency personnel, city officials, authorized personnel, and / or any other entity.In some examples, information associated with a hazard can, without limitation, describe a type of hazard, a location of the hazard, a waiting state created by the hazard (e.g., whether a lane is blocked, a time period to detour around the hazard, etc.), any potential updates regarding the hazard (e.g., whether the hazards have already been reported), a number of objects (e.g., vehicles, emergency personnel, etc.) associated with the hazard, one or more types of objects associated with the hazard, and / or the like.
[0068] Fig. Figures 6A-6B illustrate, for example, a user interface 602 associated with an application that provides information related to hazards, according to some embodiments of the present disclosure. As illustrated by the example of Fig. As shown in Figure 6A, the user interface 602 can contain at least information 604(1)-(3) of the hazards associated with the environment 204. For example, the first piece of information 604(1) can indicate the first hazard associated with the emergency vehicle 218 located on the first road 206, the second piece of information 604(2) can indicate the second hazard associated with the collision 220 between two vehicles on the second road 210, and the third piece of information 604(3) can indicate the third hazard associated with the obstacle 222 located on the third road 214. As shown, the indicators 604(1)-(3) are located at approximately similar locations on a map of the environment 204, since the hazards are located in the environment 204.
[0069] The user interface 602 may further contain information 606(1)-(3), each associated with a hazard. The first piece of information 606(1), associated with the first hazard, may, for example, indicate that the emergency vehicle 218 contains an ambulance, that the emergency vehicle 218 is on the first road 206 and / or the first lane 208(1), and / or any other information associated with the first hazard. Furthermore, the second piece of information 606(2), associated with the second hazard, may indicate that the second hazard involves a collision, that the collision blocks the second road 210 and / or the second lane 212(2), that two vehicles are involved in the collision, and / or any other information associated with the second hazard.Furthermore, the third information 606(3) associated with the third hazard can specify a type of obstacle 222, wherein the obstacle 222 blocks the third road 214 and / or the first lane 216(1), and / or any other information associated with the third hazard. Thus, by performing one or more of the processes described herein, the one or more application systems 118 are able to use the hazard information determined using the one or more models 108 (and / or any other technique described herein) to update the application associated with the user interface 602.
[0070] For example, and as illustrated by the example of Fig. As shown in Figure 6B, the machine 202 and / or one or more other machines can provide updated hazard information to one or more application systems 118. The updated hazard information may, for example, indicate that the initial information 606(1) associated with the first hazard is correct. Thus, the one or more application servers 118 can cause the user interface 602 to retain the initial information 606(1) for the first hazard. However, the updated hazard information may also include updated information 608 associated with the second hazard. The updated information 608 may, for example, indicate that the vehicles have moved at least away from the second road 210, so that the vehicles are no longer obstructing traffic.Thus, one or more application servers can cause user interface 602 to update the second piece of information 606(2) to include the updated information 608. Furthermore, the updated hazard information can indicate that the third hazard is no longer present in environment 204. Therefore, one or more application servers can cause user interface 602 to update itself by indicating that the third hazard no longer exists in environment 204.
[0071] Referring again to the example of Fig. 1A Other processes can be carried out in relation to the hazard information 112, the hazard information 116, and / or the additional hazard information provided by the application. For example, the machine can provide one or more users of the machine with content associated with the hazard information 112 (and / or the hazard information 116 and / or the additional hazard information). In some examples, the content can include, without limitation: audio content representing speech indicating the hazard information 112 (and / or the hazard information 116 and / or the additional hazard information), visual content representing the hazard information 112 (and / or the hazard information 116 and / or the additional hazard information), one or more warnings indicating the hazards, and / or any other type of content.
[0072] Furthermore, in some examples, the machine can determine one or more operations to be performed based on at least hazard information 112 (and / or hazard information 116 and / or additional hazard information). In one example, the machine can determine to reduce speed when navigating near a hazard represented by hazard information 112. In a second example, the machine can determine a new route to navigate to avoid a hazard represented by hazard information 112. Finally, in a third example, the machine can notify a user to contact emergency services and / or automatically contact emergency services based on hazard information 112.In such examples, the machine can also send at least part of the hazard information 112 to the emergency services so that they can determine the hazard. While these are just a few examples of operations that the machine can perform using hazard information 112, hazard information 116, and / or additional hazard information, the machine can perform additional and / or alternative operations in other examples.
[0073] In some examples, process 100 can continue to repeat itself while the machine continues to generate sensor data 104 using one or more sensors 102. In some examples, process 100 can repeat itself based on the occurrence of one or more events. In a first example, process 100 can repeat itself after certain time intervals, such as every second (and / or at another time interval). In a second example, process 100 can repeat itself based on the machine determining that there is a potential hazard in the environment, such as based on application data 120 indicating the potential hazard (e.g., reported by another user) and / or the machine user indicating the potential hazard.Process 100 can be performed, for example, to verify and / or update information associated with a reported hazard. In a third example, Process 100 can be repeated based on a sudden change in the driving conditions associated with the machine, such as the machine suddenly changing direction (e.g., turning, etc.) and / or speed (e.g., decelerating).
[0074] As described here, one or more models 108 can be trained in some examples to determine the hazard information 112 that is associated with hazards. Fig. Figure 1B illustrates, for example, a data flow diagram of a process 122 for training one or more machine learning models 124 (the one or more models 124 which may contain and / or be similar to the one or more models 108) to determine information associated with hazards, according to some embodiments of the present disclosure. As shown, the one or more models 124 may be trained using training data, such as sensor data 126 (which may be similar to the sensor data 104) and prompt data 128 (which may be similar to the prompt data 110). The training data may be synthetically generated (e.g., generated from computer models or renderings), actually generated (e.g., designed and generated from real data), and / or a combination thereof.
[0075] The one or more models 124 can be trained using the training data and corresponding ground truth data 130. As shown, in some examples, the ground truth data 130 may contain hazard information 132 that indicates whether instances of the sensor data 126 represent hazards, types of hazards represented by instances of the sensor data 126, and / or details associated with hazards. For example, if the sensor data 126 contains image data representing images, then individual images and / or groups of images may contain corresponding ground truth data 130 representing hazard information 132.
[0076] To train one or more models 124, as shown, one or more training engines 134 can use one or more loss functions to measure the loss (e.g., error) in output data 136 compared to the ground truth data 130. In some examples, any type of loss function can be used. Furthermore, in some examples, different outputs can have different loss functions. For example, hazard information indicating whether instances of sensor data 126 represent hazards can have a first loss function, while hazard information indicating types of hazards can have a second loss function, and / or so. In such examples, the loss functions can be combined to form a total loss (where one or more losses can be weighted), and the total loss can be used to train the one or more models 124 (e.g.,(to update the parameters of these). In any given example, backward computations can be performed to recursively calculate gradients of one or more loss functions with respect to training parameters. In some examples, weights and / or biases of one or more models can be used to compute these gradients.
[0077] Fig. Figure 7 illustrates an example of one or more systems 702 configured to perform one or more of the processes described herein, according to some embodiments of the present disclosure. In some examples, at least a part of the one or more systems 702 may be contained within a machine, such as an exemplary autonomous vehicle 1100 (and / or any other type of machine). In some examples, at least a part of the one or more systems 702 may be located remotely from the machine and communicate with it. In such examples, the one or more systems 702 may receive sensor data 104 from the machine and / or send hazard information 112 (and / or hazard information 116) back to the machine.
[0078] As shown, the one or more systems 902 can contain one or more processors 704 (which may contain and / or be similar to one or more CPUs 1108, one or more GPUs 1110, a processor 1112, one or more CPUs 1120, one or more GPUs 1120, one or more CPUs 1206 and / or one or more GPUs 1208), one or more communication interfaces 706 (which may contain and / or be similar to a network interface 1124 and / or one or more communication interfaces 1210), and a memory 708 (which may contain and / or be similar to a memory 1204). The memory 708 can also store the one or more models 108, the prompt data 110, and / or the one or more processing components 114.Additionally, the one or more processors 704 can execute the one or more models 108 and / or the one or more processing components 114 to perform one or more of the processes described herein.
[0079] With reference to Fig. 8 and Fig. Each block of Methods 800 and 900 described herein comprises a computational process that can be performed using any combination of hardware, firmware, and / or software. Various functions can be performed, for example, by a processor executing instructions stored in main memory. Methods 800 and 900 can also be embodied as computer-usable instructions stored on computer storage media. Methods 800 and 900 can be provided to another product by a standalone application, a service or hosted service (standalone or in combination with another hosted service), or a plug-in, to name a few. Additionally, these Methods 800 and 900 are illustrated by way of example with respect to Fig. 1 and Fig. 7 described. However, these procedures 800 and 900 can additionally or alternatively be executed by any system or combination of systems, including, but not limited to, the systems described herein.
[0080] Fig. Figure 8 illustrates a flowchart that defines a method 800 for automatically determining and providing information associated with a hazard, according to some embodiments of the present disclosure. The method 800 may include, in block B802, determining, based on at least initial sensor data obtained using one or more sensors, a location associated with a machine in the environment. The one or more systems 702 (e.g., the one or more processing components 114) may, for example, use the initial sensor data 104 to determine the location associated with the machine. As described here, in some examples, the initial sensor data 104 is generated using the one or more sensors 102 of the machine.Additionally or alternatively, in some examples the first sensor data 104 are generated using one or more sensors 102 of a user device 106.
[0081] Method 800, in block B804, can include determining, using one or more vision language models and based on at least secondary sensor data obtained from one or more secondary sensors of the machine, information associated with a hazard located in the environment. The one or more systems 702 can input the secondary sensor data 104, for example, into the one or more models 108. As described here, the secondary sensor data 104 can include at least image data representing one or more images of the environment. In some examples, the one or more systems 702 can input additional data into the one or more models 108, such as prompt data 110 representing one or more prompts. The one or more models 108 can then process the data and generate the hazard information 112 associated with the hazard.In some examples, one or more systems 702 may use one or more additional models 108, such as one or more language models, to verify the hazard information 112 and / or to determine whether to request additional hazard information associated with the hazard.
[0082] Procedure 800 may, in block B806, include sending data to update one or more applications to one or more systems to provide the information associated with the hazard at approximately the location. The one or more systems 702 may, for example, send the data, which represents at least the hazard information 112, to the one or more application systems 118. The one or more application systems 118 may then use the data to update the one or more applications so that the one or more applications provide at least the hazard information 112 associated with the hazard.
[0083] Fig. Figure 9 illustrates a flowchart showing a method 900 for determining how information associated with a hazard is reported, according to some embodiments of the present disclosure. The method 900 may, in block B902, include determining, using one or more vision language models and based on at least sensor data obtained using one or more sensors of a machine, initial output data representing information associated with a hazard located in the environment. The one or more systems 702 may input the sensor data 104 into the one or more models 108. As described here, the sensor data 104 may include at least image data representing one or more images of the environment. In some examples, the one or more systems 702 may input additional data into the one or more models 108, such as...The prompt data 110, which represents one or more prompts. The one or more models 108 can then process the input data and generate the first output data, which represents the hazard information 112 associated with the hazard.
[0084] Procedure 900 in block B904 can determine, using one or more language models and based at least on the first output, a second output that represents a type of message associated with the information. For example, the one or more systems 702 can input the first output into one or more additional models 108. The one or more language models can process the first output data to generate the second output data, which represents the type of message. As described here, in some examples, the type of message can include, without limitation: requesting that a user verify the hazard information 112, requesting that the user provide additional hazard information 116 associated with the hazard, providing the hazard information 112 to the one or more application systems 118, and / or performing any other type of message.
[0085] Procedure 900 can include in block B906 the performance of one or more operations based on at least the message type. The one or more systems 702 can perform the one or more operations, for example, based on the message type from the one or more additional models 108. If, in a first example, the message type includes verifying hazard information 112, then the one or more systems 702 can generate and / or provide content that fulfills the hazard information verification requirement 112. If, in a second example, the message type includes requesting additional hazard information 116, then the one or more systems 702 can generate and / or provide content that fulfills the additional hazard information requirement 116.However, if the type of message in a third example includes the provision of hazard information 112, then one or more systems 702 can send data representing hazard information 112 to one or more application systems 118. EXEMPLARY LANGUAGE MODELS
[0086] In at least some embodiments, language models such as large language models (LLMs), vision language models (VLMs), multimodal language models (MMLMs), and / or other types of generative artificial intelligence (AI) can be implemented. These models may be capable of understanding, summarizing, translating, and / or otherwise generating text (e.g., natural language text, code, etc.), images, videos, computer-aided design (CAD) assets, Omniverse and / or Metaverse file information (e.g., in USD format, such as OpenUSD), and / or the like, based on the context provided in input prompts or queries. These language models may be considered "large" in embodiments based on the fact that the models are trained on massive datasets and have architectures with a large number of learning network parameters (weights and biases), such as millions or billions of parameters.The LLMs / VLMs / MMLMs / etc. can be implemented to summarize text data, analyze data and derive insights from it (e.g., text, image, video, etc.) and generate new text / image / video / etc. in user-specified styles, tones, and / or formats. The LLMs / VLMs / MMLMs / etc. of this disclosure can, in embodiments, be used exclusively for text processing, while in other embodiments, multimodal LLMs can be implemented to accept, understand, and / or generate text and / or other types of content such as images, audio, 2D and / or 3D data (e.g., in USD formats), and / or video. For example, vision language models (VLMs) or, more generally, multimodal language models (MMLMs) can be implemented to process input data types such as image, video, audio, text, 3D design (e.g.,CAD) and / or others to accept and / or to generate or output data types such as image, video, audio, text, 3D design and / or others.
[0087] Different types of architectures for LLMs / VLMs / MMLMs, etc., can be implemented in various embodiments. For example, different architectures can be implemented that use different techniques for understanding and generating outputs, such as text, audio, video, image, 2D and / or 3D design or asset data, etc. In some embodiments, architectures for LLMs / VLMs / MMLMs, etc., can be used, such as recurrent neural networks (RNNs) or long-short-term memory (LSTM) networks, while in other embodiments, transformer architectures, such as those based on self-attention and / or cross-attention mechanisms (e.g., between context data and text data), are used to understand and recognize relationships between words or tokens and / or context data (e.g., other text, video, image, design data, USD, etc.).One or more generative processing pipelines containing LLMs / VLMs / MMLMs / etc. may also contain one or more diffusion blocks (e.g., denoisers). The LLMs / VLMs / MMLMs / etc. of this disclosure may contain one or more encoder and / or decoder blocks. For example, discriminative or encoder-only models such as BERT (Bidirectional Encoder Representations from Transformers) may be implemented for tasks involving language understanding, such as classification, sentiment analysis, question answering, and named entity recognition. As another example, generative or decoder-only models such as GPT (Generative Pretrained Transformer) may be used for tasks involving language and content generation, such as... B. Text completion, story generation, and dialogue generation. LLMs / VLMs / MMLMs / etc.Architectures containing both encoder and decoder components, such as T5 (Text-to-Text Transformer), can be implemented to understand and generate content, for example, for translations and summaries. These examples are not intended to be restrictive, and any architecture type, including but not limited to those described herein, can be implemented depending on the specific implementation and the one or more tasks performed using the LLMs / VLMs / MMLMs / etc.
[0088] In various embodiments, LLMs / VLMs / MMLMs / etc. can be trained using unsupervised learning, where an LLM / VLM / MMLM / etc. learns patterns from large amounts of unlabeled text / audio / video / image / design / USD / etc. data. Due to the extensive training, the models in some embodiments may not require task-specific or domain-specific training. LLMs / VLMs / MMLMs / etc. that have undergone extensive pre-training with enormous amounts of unlabeled data can be referred to as baseline models and may be suitable for a variety of tasks, such as answering questions, summarizing, filling in missing information, translating, and generating image / video / design / USD / data. Some LLMs / VLMs / MMLMs / etc.They can be tailored for a specific use case using techniques such as prompt tuning, fine-tuning, retrieval augmented generation (RAG), adding adapters (e.g., custom neural networks and / or layers of neural networks that tune or adapt prompts or tokens to align the language model with a particular task or domain), and / or using other fine-tuning or tailoring techniques that optimize the models for use in specific tasks and / or domains.
[0089] In some embodiments, the LLMs / VLMs / MMLMs / etc. of the present disclosure can be implemented using various model alignment techniques. For example, in some embodiments, guardrails can be implemented to identify impermissible or unwanted inputs (e.g., prompts) and / or outputs of the models. The system can use the guardrails and / or other model alignment techniques to either prevent a specific unwanted input from being processed using the LLMs / VLMs / MMLMs / etc., and / or to prevent the output or presentation (e.g., display, audio output, etc.) of information generated using the LLMs / VLMs / MMLMs / etc. In some embodiments, one or more additional models—or layers thereof—can be implemented to identify problems with the inputs and / or outputs of the models.For example, these "security models" can be trained to identify inputs and / or outputs that are "safe" or otherwise acceptable or desirable, and / or that are "unsafe" or otherwise undesirable for the particular application / implementation. As a result, the LLMs / VLMs / MMLMs / etc. of this disclosure are less likely to output speech / text / audio / video / design data / USD data / etc. that is offensive, vulgar, inappropriate, unsafe, non-technical, and / or otherwise undesirable for the particular application / implementation.
[0090] In some embodiments, the LLMs / VLMs / etc. may be configured or able to access or use one or more plug-ins, application programming interfaces (APIs), databases, data stores, repositories, etc. For example, for certain tasks or operations for which it is not ideally suited, the model may have instructions (e.g., as a result of training and / or based on instructions in a given prompt) to access one or more plug-ins (e.g., third-party plug-ins) to obtain assistance in processing the current input. In such an example, where at least part of a prompt is related to restaurants or the weather, the model may access one or more restaurant or weather plug-ins (e.g., via one or more APIs) to retrieve the relevant information.As another example where at least part of a response requires a mathematical calculation, the model can access one or more mathematical plugins or APIs to assist in solving the one or more problems and then use the plugin's and / or API's response in the model's output. This process can be repeated, for example recursively, for any number of iterations and using any number of plugins and / or APIs until a response to the input prompt can be generated that addresses each question / request / requirement / process / operation, etc. Therefore, the one or more models can rely not only on their own knowledge gained from training on one or more large datasets but also on the expertise or optimized nature of one or more external resources, such as APIs, plugins, and / or the like.
[0091] In some embodiments, multiple language models (e.g., LLMs / VLMs / MMLMs / etc.), multiple instances of the same language model, and / or multiple prompts served to the same language model or instance of the same language model can be implemented, executed, or accessed (e.g., using one or more plug-ins, user interfaces, APIs, databases, data stores, repositories, etc.) to provide output in response to the same query or to separate parts of a query. In at least one embodiment, multiple language models, such as language models with different architectures or language models trained on different (e.g., updated) datasets, can be served with the same input query and the same prompt (e.g., a set of constraints, conditioners, etc.).In one or more embodiments, the language models can be different versions of the same basic model. In one or more embodiments, at least one language model can be instantiated as multiple agents—for example, more than one prompt can be provided to constrain, direct, or otherwise influence the style, content, or character of the output provided. In one or more non-constraintive embodiments, the same language model can be prompted to provide output corresponding to a different role, perspective, character, or knowledge base, as defined by a provided prompt.
[0092] In each of these embodiments, the output of two or more (e.g., each) language models, two or more versions of at least one language model, two or more instantiated agents of at least one language model, and / or two further prompts provided to at least one language model can be further processed, e.g., aggregated, compared, or filtered, or used to determine (and provide) a consensus response. In one or more embodiments, the output of a language model—or a version, instance, or agent—can be provided as input to another language model for further processing and / or validation. In one or more embodiments, a language model can be instructed to generate or otherwise obtain output with respect to input source material, the output being associated with the input source material.Such an assignment might involve, for example, generating a label or a portion of text that is embedded (e.g., as metadata) in input source text or image. In one or more embodiments, an output from a language model can be used to determine the validity of input source material for further processing or insertion into a dataset. For example, a language model can be used to evaluate the presence (or absence) of a target word in a portion of text or an object in an image, annotating the text or image to indicate such presence (or absence). Alternatively, the determination from the language model can be used to determine whether the source material should be included in a maintained dataset, for example, without restriction.
[0093] Fig. Figure 10A is a block diagram of an exemplary generative language model system 1000, suitable for use in implementing at least some embodiments of the present disclosure. In the Fig. In the illustrated example 10A, the generative language model system 1000 includes a Retrieval Extended Generation (RAG) component 1092, an input processor 1005, a tokenizer 1010, an embedding component 1020, plug-ins / APIs 1095, and a generative language model (LM) 1030 (which may contain an LLM, a VLM, a multimodal LM, etc.).
[0094] Generally speaking, the input processor 1005 can receive an input 1001 containing text and / or other types of input data (e.g., audio data, video data, image data, sensor data (e.g., LiDAR, RADAR, ultrasound, etc.), 3D design data, CAD data, universal scene describer (USD) data such as OpenUSD, etc.), depending on the architecture of the generative LM 1030 (e.g., LLM / VLM / MMLM / etc.). In some embodiments, the input 1001 contains plain text in the form of one or more sentences, paragraphs, and / or documents. Additionally or alternatively, the input 1001 can contain numerical sequences, pre-computed embeddings (e.g., word or sentence embeddings), and / or structured data (e.g., in tabular formats, JSON, or XML). In some implementations where the generative LM 1030 is capable of processing multimodal inputs, the input 1001 can be text with image data, audio data, video data, design data, USD data and / or other types of input data, such as...B., but without restriction, combine (or omit) the methods described herein. Using raw input text as an example, the Input Processor 1005 can prepare raw input text in various ways. For example, the Input Processor 1005 can perform different types of text filtering to remove noise (e.g., special characters, punctuation, HTML markup, stop words, parts of one or more images, parts of audio, etc.) from relevant text content. In an example that includes stop words (frequent words that tend to convey little semantic meaning), the Input Processor 1005 can remove stop words to reduce noise and focus the generative LM 1030 on more meaningful content.The 1005 input processor can apply text normalization, for example by converting all characters to lowercase, removing accents, and / or handling special cases such as contractions or abbreviations to ensure consistency. These are just a few examples, and other types of input processing can also be applied.
[0095] In some embodiments, a RAG component 1092 (which may contain one or more RAG models and / or be implemented using the generative LM 1030 itself) can be used to retrieve additional information to be used as part of the input 1001 or the prompt. RAG can be used to enhance the input to the LLM / VLM / MMLM / etc. with external knowledge, making the answers to specific questions, queries, or requirements more relevant, such as in a case where specialized knowledge is needed. The RAG component 1092 can retrieve this additional information (e.g., grounding information such as grounding text / image / video / audio / USD / CAD / etc.) from one or more external sources, which can then be fed into the LLM / VLM / MMLM / etc. along with the prompt to improve the accuracy of the model's responses or outputs.
[0096] For example, in some embodiments, the input 1001 can be generated using the query or input into the model (e.g., a question, a request, etc.) in addition to the data retrieved using the RAG component 1092. In some embodiments, the input processor 1005 can analyze the input 1001 and communicate with the RAG component 1092 (or the RAG component 1092 can be part of the input processor 1005 in some embodiments) to identify relevant text and / or other data to be provided to the generative LM 1030 as additional context or information sources from which the reaction, response, or output 1090 is generally to be identified.For example, if the input indicates that the user is interested in a desired tire pressure for a specific make and model of vehicle, the RAG component 1092—for example, using a RAG model that performs a vector search in an embedding space—can retrieve the tire pressure information or the corresponding text from a digital (embedded) version of the user manual for that specific vehicle make and model. Similarly, if a user revisits a chatbot in connection with a specific product offering or service, the RAG component 1092 can retrieve a previously stored conversation history—or at least a summary thereof—and include the previous conversation history, along with the current question / request, as part of the input 1001 in the generative LM 1030.
[0097] The RAG component 1092 can employ various RAG techniques. For example, naive RAG can be used when documents are indexed, split into pieces, and applied to an embedding model to generate embeddings that correspond to the pieces. A user request can also be applied to the embedding model and / or another embedding model of the RAG component 1092, and the piece embeddings can be compared with the request's embeddings to identify the most similar embeddings that can be fed to the generative LM 1030 to generate output.
[0098] In some embodiments, more advanced RAG techniques can be used. For example, the pieces can undergo pre-fetching processes (e.g., forwarding, rewriting, metadata analysis, extension, etc.) before being passed to the embedding model. Furthermore, post-fetching processes (e.g., re-ranking, prompt compression, etc.) can be performed on the outputs of the embedding model before the final embeddings are generated and used as a comparison to an input query.
[0099] Another example is modular RAG techniques, such as those similar to naive and / or extended RAG, but which may also include features like hybrid search, recursive retrieval and query engines, step-back approaches, subqueries, and hypothetical document embedding.
[0100] As another example, Graph-RAG can use knowledge graphs as a source of contextual or factual information. Graph-RAG can be implemented using a graph database as a source of contextual information, which is then sent to the LLM / VLM / MMLM / etc. Instead of providing the model with (or in addition to) snippets of data extracted from larger documents, which can result in a lack of context, factual accuracy, linguistic precision, etc., Graph-RAG can also provide structured entity information to the LLM / VLM / MMLM / etc. by combining the structured text description of the entity with its many properties and relationships, allowing the model deeper insights. When implementing Graph-RAG, the systems and procedures described herein use a graph as a content store, extract relevant snippets from documents, and request them from the LLM / VLM / MMLM / etc.to respond using this. In such embodiments, the knowledge graph can contain relevant text content and metadata about the knowledge graph and be integrated into a vector database. In some embodiments, the graph RAG can use a graph as a domain expert, extracting descriptions of concepts and entities relevant to a query / prompt and passing them to the model as semantic context. These descriptions can include relationships between the concepts. In other examples, the graph can be used as a database, where part of a query / prompt can be allocated to a graph query, the graph query can be executed, and the LLM / VLM / MMLM / etc. can summarize the results.In such an example, the graph can store relevant factual information, and a query (natural language query) to a graph query tool (NL-to-Graph-query tool) and an entity join can be used. In some implementations, graph RAG (e.g., using a graph database) can be combined with standard RAG (e.g., vector database) and / or other RAG types to benefit from multiple approaches.
[0101] In all embodiments, the RAG component 1092 can implement a plug-in, an API, a user interface, and / or other functionality to perform RAG. For example, a graph RAG plug-in can be used by the LLM / VLM / MMLM / etc. to query the knowledge graph to extract relevant information for feeding into the model, and a standard or vector RAG plug-in can be used to query a vector database. For example, the graph database can interact with a plug-in's REST interface, thus decoupling the graph database from the vector database and / or the embedding models.
[0102] The Tokenizer 1010 can segment (e.g., processed) text data into smaller units (tokens) for subsequent analysis and processing. Depending on the implementation, the tokens can represent individual words, partial words, characters, parts of audio / video / images, etc. Word-based tokenization divides the text into individual words, with each word treated as a separate token. Partial word tokenization breaks words down into smaller meaning-bearing units (e.g., prefixes, suffixes, stems), enabling the generative LM 1030 to understand morphological variations and more effectively handle words not in the vocabulary. Character-based tokenization represents each character as a separate token, allowing the generative LM 1030 to process text at a fine-grained level.The choice of tokenization strategy can depend on factors such as the language being processed, the task at hand, and / or the characteristics of the training dataset. Therefore, Tokenizer 1010 can convert the (e.g., processed) text into a structured format according to the tokenization scheme implemented in the respective implementation.
[0103] The embedding component 1020 can use any known embedding technique to convert discrete tokens into (e.g., dense, continuous vector) representations of semantic meaning. For example, the embedding component 1020 can use pre-trained word embeddings (e.g., Word2Vec, GloVe, or FastText), one-hot coding, term frequency-inverse document frequency (TF-IDF) coding, one or more neural network embedding layers, and / or other techniques.
[0104] In some implementations where the input 1001 contains image data / video data / etc., the input processor 1001 can resize the data to a standard size compatible with the format of a corresponding input channel, and / or normalize pixel values to a common range (e.g., 0 to 1) to ensure a uniform representation, and the embedding component 1020 can encode the image data using any known technique (e.g., using one or more convolutional neural networks (CNNs) to extract visual features).In some implementations where the input 1001 contains audio data, the input processor 1001 can resample an audio file to a uniform sampling rate for uniform processing, and the embedding component 1020 can use any known technique to extract and encode audio features, for example, in the form of a spectrogram (e.g., a Mel spectrogram). In some implementations where the input 1001 contains video data, the input processor 1001 can extract frames or apply resizing to extracted frames, and the embedding component 1020 can extract features such as optical flow or video embeddings and / or encode temporal information or sequences of frames. In some implementations where the input 1001 contains multimodal data, the embedding component 1020 can create representations of the different data types (e.g.,fusing text, image, audio, USD, video, design, etc.) using techniques such as early fusion (concatenation), late fusion (sequential processing), attention-based fusion (e.g., self-attention, cross-attention), etc.
[0105] The generative LM 1030 and / or other components of the generative LM system 1000 can use different types of neural network architectures, depending on the implementation. For example, transformer-based architectures, such as those used in models like GPT, can be implemented, incorporating self-attention mechanisms that weigh the importance of different words or tokens in the input sequence, and / or feedforward networks that process the output of the self-attention layers, applying nonlinear transformations to the input representations and extracting higher-level features. Some non-restrictive example architectures include transformers (e.g.,Encoder-decoder, decoder-only, multimodal), RNNs, LSTMs, fusion models, diffusion models, crossmodal embedding models that learn shared embedding spaces, graph neural networks (GNNs), hybrid architectures that combine different types of architectures, adversarial networks such as generative adversarial networks (GANs) or adversarial autoencoders (AAEs) for joint distributional learning, and others. Therefore, depending on the implementation and architecture, the embedding component 1020 can apply a coded representation of the input 1001 to the generative LM 1030, and the generative LM 1030 can process the coded representation of the input 1001 to generate an output 1090 that may contain response text and / or other types of data.
[0106] As described herein, the generative LM 1030, in some embodiments, may be configured to access or use plug-ins / APIs 1095—or have the capability to access or use them (which may include one or more plug-ins, application programming interfaces (APIs), databases, data stores, repositories, etc.). For example, for certain tasks or operations for which the generative LM 1030 is not ideally suited, the model may have instructions (e.g., as a result of training and / or based on instructions in a given prompt, such as those retrieved using the RAG component 1092) to access one or more plug-ins / APIs 1095 (e.g., third-party plug-ins) to obtain assistance in processing the current input.In such an example, where at least part of a prompt is related to restaurants or weather, the model can access one or more restaurant or weather plugins (e.g., via one or more APIs), send at least part of the prompt related to the specific API 1095 to the plugin / API 1095, the plugin / API 1095 can process the information and return a response to the generative LM 1030, and the generative LM 1030 can use the response to generate the output 1090. This process can be repeated for any number of iterations and using any number of plugins / APIs 1095, e.g., recursively, until an output 1090 can be generated that addresses every question / request / request / process / operation / etc. from the input 1001.Therefore, one or more models can rely not only on their own knowledge from training with one or more large datasets and / or from data retrieved using the RAG component 1092, but also on the expertise or optimized nature of one or more external resources, such as the plug-ins / APIs 1095.
[0107] Fig. Figure 10B is a block diagram of an example implementation where the generative LM 1030 contains a transformer-encoder-decoder. For example, suppose an input text such as "Who discovered gravity?" is tokenized (e.g., by the Tokenizer 810 from...) Fig. 10A) into tokens such as words, and each token is encoded into a corresponding embedding (e.g., of size 512) (e.g., by the embedding component 1020 of Fig. 910A). Since these token embeddings typically do not represent the token's position in the input sequence, any known technique can be used to add positional encoding to each token embedding to encode the sequential relationships and context of the tokens in the input sequence. Therefore, the (e.g., resulting) embeddings can be applied to one or more encoders 1035 of the generative LM 1030.
[0108] In an exemplary implementation, the one or more encoders 1035 form an encoder stack, with each encoder containing a self-attention layer and a feedforward network. In an exemplary transformer architecture, each token (e.g., word) flows through a separate path. Therefore, each encoder can accept a sequence of vectors, with each vector passing through the self-attention layer, then through the feedforward network, and then up to the next encoder in the stack. Any known self-attention technique can be used.For example, to calculate a self-attention score for each token (word), a query vector, a key vector, and a value vector can be created for each token. A self-attention score can be calculated for pairs of tokens by taking the dot product of the query vector with the corresponding key vectors, normalizing the resulting numerical values, multiplying them by the corresponding value vectors, and summing the weighted value vectors. The encoder can employ multi-head attention, where the attention mechanism is applied multiple times in parallel with different learned weight matrices. Any number of encoders can be cascaded to generate a context vector that encodes the input. An attention projection layer 1040 can convert the context vector into attention vectors (keys and values) for one or more decoders 1045.
[0109] In an exemplary implementation, the one or more decoders 1045 form a decoder stack, with each decoder containing a self-attention layer, an encoder-decoder self-attention layer that uses the encoder's attention vectors (keys and values) to focus on relevant parts of the input sequence, and a feedforward network. As with the one or more encoders 1035, in an exemplary transformer architecture, each token (e.g., word) flows through a separate path in the one or more decoders 1045. During a first iteration, the one or more decoders 1045, a classifier 1050, and a generation mechanism 1055 can generate an initial token, and the generation mechanism 1055 can apply the generated token as input during a second iteration. The process can repeat in a loop, successively passing tokens (e.g., words ...words) are generated and added to the output of the previous pass, and the token embeddings of the compound sequence with positional encodings are applied as an input to one or more decoders 1045 during a subsequent pass, generating one token at a time (known as autoregression) until a symbol or token is predicted that represents the end of the response. Within each decoder, the self-attention layer is typically restricted to focusing only on previous positions in the output sequence by applying a masking technique (e.g., setting future positions to negative infinity) before the softmax operation. In an exemplary implementation, the encoder-decoder attention layer works similarly to the (e.g.,Multi-head) self-attention in the one or more encoders 1035, except that it creates its queries from the underlying layer and takes the keys and values (e.g. matrix) from the output of the one or more encoders 1035.
[0110] Therefore, the one or more decoders 1045 can output a decoded (e.g., vector) representation of the input applied during a given iteration. The classifier 1050 can include a multi-class classifier comprising one or more layers of a neural network that project the decoded (e.g., vector) representation into an appropriate dimensionality (e.g., one dimension for each supported word or token in the output vocabulary) and a softmax operation that converts logits into probabilities. Therefore, the generation mechanism 1055 can select or sample a word or token based on an appropriate predicted probability (e.g., selecting the word with the highest predicted probability) and append it to the output of a previous iteration, generating each word or token sequentially.The generation mechanism 1055 can repeat the process, triggering successive decoder inputs and corresponding predictions until a symbol or token is selected or sampled that represents the end of the response, after which the generation mechanism 1055 can output the generated response.
[0111] Fig. Figure 10C is a block diagram of an exemplary implementation where the generative LM 1030 incorporates a decoder-transformer-only architecture. For example, one or more 1060 decoders can be used. Fig. 10C similar to one or more 1045 decoders. Fig. 10B work, with the exception that each of the one or more decoders 1060 of Fig. 10C omits the encoder-decoder self-attention layer (since there is no encoder in this implementation). Therefore, the one or more Decoders 1060 can form a decoder stack, with each decoder containing a self-attention layer and a feedforward network. Furthermore, instead of encoding the input sequence, a symbol or token representing the end of the input sequence (or the beginning of the output sequence) can be appended to the input sequence, and the resulting sequence (e.g., corresponding embeddings with positional encodings) can be applied to the one or more Decoders 1060. As with the one or more Decoders 1045 of Fig. In 10B, each token (e.g., word) can flow through a separate path in the one or more decoders 1060, and the one or more decoders 1060, a classifier 1065, and a generation mechanism 1070 can use autoregression to sequentially generate one token after another until a symbol or token is predicted that represents the end of the response. The classifier 1065 and the generation mechanism 1070 can be used similarly to the classifier 1050 and the generation mechanism 1055 of Fig. 10B operates wherein the generation mechanism 1070 selects or samples each successive output token based on a corresponding predicted probability and appends it to the output of a previous pass, with each token being generated sequentially until a symbol or token is selected or sampled that represents the end of the response. This and other architectures described herein are intended only as examples, and other suitable architectures may be implemented within the scope of protection of this disclosure. EXEMPLARY AUTONOMOUS VEHICLE
[0112] Fig. Figure 11A is an illustration of an exemplary autonomous vehicle 1100, according to some embodiments of the present disclosure. The autonomous vehicle 1100 (here alternatively referred to as "vehicle 1100") may, without limitation, include: a passenger vehicle, such as a car, truck, bus, emergency service vehicle, shuttle, electric or motorized bicycle, motorcycle, fire engine, police vehicle, ambulance, boat, construction vehicle, underwater vehicle, robotic vehicle, drone, aircraft, a vehicle coupled to a trailer (e.g., a semi-trailer truck used for transporting cargo), and / or another type of vehicle (e.g., one that is unmanned and / or carries one or more passengers).Autonomous vehicles are generally described in terms of automation levels defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) standard "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201806, published June 15, 2018, Standard No. J3016-201609, published September 30, 2016, and earlier and future versions of this standard). The Vehicle 1100 may exhibit functionality corresponding to one or more of the Levels 3 through 5 of autonomous driving levels.The Vehicle 1100 can exhibit functionality corresponding to one or more of the Levels 1 to 5 of autonomous driving. For example, depending on its configuration, the Vehicle 1100 may be capable of driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). The term "autonomous," as used here, may encompass any and / or all types of autonomy for the Vehicle 1100 or any other machine, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, assistive autonomy, semi-autonomous, primary autonomous, or any other designation.
[0113] The vehicle 1100 can include components such as a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. The vehicle 1100 can include a propulsion system 1150, such as an internal combustion engine, a hybrid electric power plant, a pure electric motor, and / or another type of propulsion. The propulsion system 1150 can be connected to a drivetrain of the vehicle 1100, which may include a transmission to enable the propulsion of the vehicle 1100. The propulsion system 1150 can be controlled in response to signals received from the throttle valve or accelerator device 1152.
[0114] A steering system 1154, which may include a steering wheel, can be used to steer the vehicle 1100 (e.g., along a desired path or route) when the drive system 1150 is in operation (e.g., when the vehicle is in motion). The steering system 1154 can receive signals from a steering actuator 1156. The steering wheel is optional for full automation (level 5).
[0115] The brake sensor system 1146 can be used to actuate the vehicle brakes in response to receiving signals from the brake actuators 1148 and / or the brake sensors.
[0116] The one or more controllers 1136, the one or more systems-on-chips (SoCs) 1104 ( Fig. 11C) and / or GPUs, can provide signals (e.g., representing instructions) to one or more components and / or systems of the vehicle 1100. For example, the one or more controllers can send signals to actuate the vehicle brakes via one or more brake actuators 1148, to actuate the steering system 1154 via one or more steering actuators 1156, and to actuate the propulsion system 1150 via one or more throttle / accelerator devices 1152. The one or more controllers 1136 can include one or more built-in (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and issue operating commands (e.g., signals representing commands) to enable autonomous driving and / or to assist a human driver in driving the vehicle 1100.The one or more controllers 1136 can include a first controller 1136 for autonomous driving functions, a second controller 1136 for functional safety functions, a third controller 1136 for artificial intelligence functions (e.g., computer vision), a fourth controller 1136 for infotainment functions, a fifth controller 1136 for emergency redundancy, and / or other controllers. In some examples, a single controller 1136 can perform two or more of the above-mentioned functionalities, two or more controllers 1136 can perform a single functionality, and / or any combination thereof.
[0117] The one or more controllers 1136 can provide the signals for controlling one or more components and / or systems of the vehicle 1100 in response to sensor data received from one or more sensors (e.g. sensor inputs). The sensor data can be received, for example, without restriction, from one or more of the following: Global Navigation Satellite Systems (GNSS) sensor(s) 1158 (e.g., Global Positioning System sensor(s)), RADAR sensor(s) 1160, Ultrasonic sensor(s) 1162, LIDAR sensor(s) 1164, Inertial Measurement Unit (IMU) sensor(s) 1166 (e.g., accelerometer(s), gyroscope(s), magnetic compass(s), magnetometer(s), etc.), microphone(s) 1196, stereo camera(s) 1168, wide-angle camera(s) 1170 (e.g., fisheye cameras), infrared camera(s) 1172, ambient camera(s) 1174 (e.g.,360-degree cameras), long-range and / or medium-range camera(s) 1198, speed sensor(s) 1144 (e.g. for measuring the speed of the vehicle 1100), vibration sensor(s) 1142, steering sensor(s) 1140, brake sensor(s) (e.g. as part of the brake sensor system 1146), and / or other sensor types.
[0118] One or more of the controllers 1136 can receive inputs (e.g., in the form of input data) from an instrument cluster 1132 of the vehicle 1100 and provide outputs (e.g., in the form of output data, display data, etc.) via a human-machine interface (HMI) display 1134, an acoustic alarm, a loudspeaker, and / or via other components of the vehicle 1100. The outputs can include information such as vehicle speed, engine speed, time, map data (e.g., the high-definition (HD) map 1122). Fig. 11C), location data (e.g., the location of vehicle 1100, e.g., on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and the status of objects as perceived by the one or more controllers 1136, etc. For example, the HMI display 1134 can show information about the presence of one or more objects (e.g., a road sign, a warning sign, a changing traffic light, etc.) and / or information about driving maneuvers that the vehicle has performed, is currently performing, or will perform (e.g., changing lanes now, taking exit 34B in two miles, etc.).
[0119] The vehicle 1100 also includes a network interface 1124, which can use one or more wireless antennas 1126 and / or modems for communication over one or more networks. The network interface 1124 can be suitable, for example, for communication via Long-Term Evolution (LTE), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Global System for Mobile Communication (GSM), IMT-CDMA Multi-Carrier (CDMA2000), etc. The one or more wireless antennas 1126 can also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc.and / or low power wide area networks (LPWANs), such as LoRaWAN, SigFox, etc.
[0120] Fig. 11B is an example of camera locations and fields of view for the exemplary autonomous vehicle 1100. Fig. 11A, according to some embodiments of the present disclosure. The cameras and respective fields of view are an exemplary embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be located at different locations on the vehicle 1100.
[0121] The camera types may include, but are not limited to, digital cameras designed for use with the components and / or systems of the 1100 vehicle. The one or more cameras may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. Depending on the configuration, the camera types may be capable of any frame rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc. The cameras may use roller shutters, global shutters, another type of shutter, or a combination thereof.In some examples, the color filter array may include a red-clear-clear-clear (RCCC) color filter array, a red-clear-clear-blue (RCCB) color filter array, a red-blue-green (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor color filter array (RGGB), a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, cameras with clear pixels, such as cameras with an RCCC, RCCB, and / or RBGC color filter array, may be used to increase light sensitivity.
[0122] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For instance, a multi-function monocular camera can be installed to provide features including lane departure warning, traffic sign recognition, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).
[0123] One or more cameras can be mounted in a bracket, such as a specially designed (three-dimensional ("3D") printed bracket, to eliminate stray light and reflections from inside the vehicle (e.g., reflections of the dashboard in the windshield) that could interfere with the camera's image acquisition. Regarding the mounting of exterior mirrors, the mirrors can be individually 3D printed so that the camera mounting plate is shaped to fit the mirror. In some cases, the one or more cameras can be integrated into the exterior mirror itself. For side cameras, the one or more cameras can also be integrated into the four pillars at each corner of the cabin.
[0124] Cameras with a field of view that includes portions of the environment in front of the vehicle (e.g., forward-facing cameras) can be used for surround view to help identify forward paths and obstacles and, with the aid of one or more controllers and / or control SoCs, to provide information critical for creating an occupancy grid and / or determining preferred vehicle paths. Forward-facing cameras can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. Forward-facing cameras can also be used for ADAS functions and systems that include lane departure warnings (LDW), autonomous cruise control (ACC), and / or other functions such as traffic sign recognition.
[0125] A variety of cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform containing a complementary metal oxide semiconductor (CMOS) color imager. Another example is the 1170 wide-angle camera, which can be used to capture objects moving into view from the periphery (e.g., pedestrians, crossing vehicles, or bicycles). Although in Fig. While Figure 11B illustrates only one wide-angle camera, the vehicle 1100 can have any number (including zero) of wide-angle cameras 1170. Furthermore, any number of long-range cameras 1198 (e.g., a pair of long-range stereo cameras) can be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. The one or more long-range cameras 1198 can also be used for object detection and classification, as well as basic object tracking.
[0126] Any number of stereo cameras 1168 can also be included in a forward-facing configuration. In at least one embodiment, one or more of the stereo cameras 1168 can include an integrated control unit comprising a scalable processing unit that can provide programmable logic (“FPGA”) and a multicore microprocessor with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. Such a unit can be used to create a 3D map of the vehicle's surroundings that includes a distance estimate for all points in the image. Alternatively, one or more stereo cameras 1168 can include a compact stereo vision sensor that may contain two camera lenses (one left and one right) and an image processing chip that measures the distance between the vehicle and the target object and processes the generated information (e.g.,Metadata) can be used to activate the autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 1168 can be used in addition to or as an alternative to those described here.
[0127] Cameras with a field of view that includes sections of the environment to the sides of the vehicle 1100 (e.g., side cameras) can be used for the surround view and provide information that is used to create and update the occupancy grid and to generate side-impact collision warnings. For example, one or more surround cameras 1174 (e.g., four surround cameras 1174, as in Fig. (11B illustrated) are positioned on the vehicle 1100. The one or more surround-view cameras 1174 can include one or more wide-angle cameras 1170, one or more fisheye cameras, one or more 360-degree cameras, and / or the like. For example, four fisheye cameras can be mounted at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround-view cameras 1174 (e.g., left, right, and rear) and one or more other cameras (e.g., a forward-facing camera) as a fourth surround-view camera.
[0128] Cameras with a field of view that includes sections of the area behind the vehicle 1100 (e.g., reversing cameras) can be used for parking assistance, surround view, rear-impact warnings, and creating and updating the occupancy grid. A variety of cameras can be used, including cameras that are also suitable as one or more forward-facing cameras (e.g., one or more long-range and / or medium-range cameras 1198, one or more stereo cameras 1168, one or more infrared cameras 1172, etc.), as described herein.
[0129] Fig. 11C is a block diagram of an exemplary system architecture for the exemplary autonomous vehicle 1100 from Fig. 11A, according to some embodiments of the present disclosure. It should be noted that these and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, arrays, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as single or distributed components, or in conjunction with other components, in any suitable combination and at any suitable location. Various functions described herein, which are performed by entities, may be executed by hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in memory.
[0130] Each of the components, features and systems of the 1100 vehicle in Fig. 11C is illustrated as being connected via bus 1102. Bus 1102 may contain a Controller Area Network (CAN) data interface (here alternatively referred to as a "CAN bus"). A CAN can be a network within the vehicle 1100 that serves to support the control of various features and functions of the vehicle 1100, such as the operation of brakes, acceleration, braking, steering, windshield wipers, etc. A CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). The CAN bus can be read to determine the steering wheel angle, vehicle speed, engine speed (rpm), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.
[0131] Although bus 1102 is described here as a CAN bus, this should not be interpreted as a limitation. For example, FlexRay and / or Ethernet can be used in addition to or as an alternative to the CAN bus. Furthermore, while a single line is used to represent bus 1102, this is not intended as a restriction. For instance, there can be any number of buses 1102, which may contain one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using a different protocol. In some examples, two or more buses 1102 can be used to perform different functions and / or for redundancy. For example, a first bus 1102 can be used for collision avoidance functionality, and a second bus 1102 can be used for actuation control.In each example, each bus 1102 can communicate with one of the vehicle's components 1100, and two or more buses 1102 can communicate with the same components. In some examples, each SoC 1104, each controller 1136, and / or each computer within the vehicle can have access to the same input data (e.g., inputs from vehicle 1100 sensors) and be connected to a common bus, such as the CAN bus.
[0132] The vehicle 1100 can contain one or more controllers 1136, as described herein with reference to Fig. 11A. The one or more controllers 1136 can be used for a variety of functions. The one or more controllers 1136 can be coupled with one or more of the various other components and systems of the vehicle 1100 and can be used for controlling the vehicle 1100, for the artificial intelligence of the vehicle 1100, for infotainment for the vehicle 1100 and / or the like.
[0133] The vehicle 1100 can contain one or more systems-on-a-chip (SoC) 1104. The SoC 1104 can contain one or more CPUs 1106, one or more GPUs 1108, one or more processors 1110, one or more caches 1112, one or more accelerators 1114, one or more data storage devices 1116, and / or other components and features not illustrated. The one or more SoCs 1104 can be used to control the vehicle 1100 in a variety of platforms and systems. For example, the one or more SoCs 1104 in a system (e.g., the system of the vehicle 1100) can be combined with an HD card 1122, which is accessed via a network interface 1124 by one or more servers (e.g., the one or more servers 1178). Fig. 11D) Receive map refreshes and / or updates.
[0134] The one or more CPUs 1106 can contain a CPU cluster or CPU complex (hereinafter referred to as "CCPLEX"). The one or more CPUs 1106 can contain multiple cores and / or L2 caches. In some embodiments, the one or more CPUs 1106 can, for example, contain eight cores in a coherent multiprocessor configuration. In some embodiments, the one or more CPUs 1106 can contain four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2 MB L2 cache). The one or more CPUs 1106 (e.g., the CCPLEX) can be configured to support the concurrent operation of clusters, so that any combination of clusters of the one or more CPUs 1106 can be active at any given time.
[0135] The one or more CPUs 1106 can implement power management functions that include one or more of the following features: individual hardware blocks can be automatically clocked when idle to dynamically save power; each core clock can be controlled when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core can be independently power-controlled; each core cluster can be independently clock-controlled when all cores are clock-controlled or power-controlled; and / or each core cluster can be independently power-controlled when all cores are power-controlled.The one or more CPUs 1106 can also implement an improved power state management algorithm where permissible power states and expected wake-up times are defined, and the hardware / microcode determines the best power state to input for the core, cluster, and CCPLEX. The processing cores can support simplified sequences for inputting the power state to software, offloading the work to the microcode.
[0136] The one or more GPUs 1108 can include an integrated GPU (referred to herein alternatively as an "iGPU"). The one or more GPUs 1108 can be programmable and can be efficient for parallel workloads. The one or more GPUs 1108 can use an extended Tensor instruction set in some examples. The one or more GPUs 1108 can include one or more streaming microprocessors, each of which can contain an L1 cache (e.g., an L1 cache of at least 96 KB), and two or more of the streaming microprocessors can share an L2 cache (e.g., an L2 cache of 512 KB). In some embodiments, the one or more GPUs 1108 can contain at least eight streaming microprocessors. The one or more GPUs 1108 can use one or more application programming interfaces (APIs) for computation.Furthermore, the one or more GPUs 1108 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0137] The one or more GPUs 1108 can be power-optimized for best performance in automotive and embedded applications. The one or more GPUs 1108 can be manufactured, for example, on a FinFET field-effect transistor. However, this is not a limitation, and the one or more GPUs 1108 can also be manufactured using other semiconductor manufacturing processes. Each streaming microprocessor can contain an array of mixed-precision processing cores, divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks.In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA TENSOR COREs for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64 KB register file. Furthermore, the streaming microprocessors can include independent parallel integer and floating-point data paths to enable efficient execution of workloads with a mix of computations and addressing operations. The streaming microprocessors can include an independent thread scheduling function to enable fine-grained synchronization and cooperation between parallel threads. The streaming microprocessors can include a combined L1 data cache and a shared memory unit to improve performance while simplifying programming.
[0138] The one or more GPUs 1108 can include high-bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem to provide a peak memory bandwidth of approximately 900 GB / second in some examples. In some examples, synchronous graphics random access memory (SGRAM), such as double-data-rate type five synchronous graphics random access memory (GDDR5), can be used in addition to or as an alternative to HBM memory.
[0139] The one or more GPUs 1108 can incorporate a unified memory technology that includes access counters to enable more accurate migration of memory pages to the processor that accesses them most frequently, thereby improving the efficiency of memory areas shared by processors. In some examples, support for Address Translation Services (ATS) can be used so that the one or more GPUs 1108 can directly access the page tables of the one or more CPUs 1106. In such examples, if the Memory Management Unit (MMU) of the one or more GPUs 1108 fails, an address translation request can be sent to the one or more CPUs 1106.In response, the one or more CPUs 1106 can search their page tables for the virtual-physical mapping for the address and send the translation back to the one or more GPUs 1108. This unified memory technology thus enables a single, unified virtual address space for the memory of both the one or more CPUs 1106 and the one or more GPUs 1108, thereby simplifying the programming of the one or more GPUs 1108 and the porting of applications to the one or more GPUs 1108.
[0140] Additionally, the one or more GPUs 1108 can contain an access counter that tracks the frequency of accesses by the one or more GPUs 1108 to the memory of other processors. This access counter can help ensure that memory pages are moved to the physical memory of the processor that accesses them most frequently.
[0141] The one or more SoCs 1104 can contain any number of caches 1112, including those described here. The one or more caches 1112 can, for example, contain an L3 cache that is available to both the one or more CPUs 1106 and the one or more GPUs 1108 (e.g., one that is connected to both the one or more CPUs 1106 and the one or more GPUs 1108). The one or more caches 1112 can contain a write-back cache that can track the states of the rows, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache can be 4 MB or larger, depending on the implementation, although smaller cache sizes can also be used.
[0142] The one or more SoCs 1104 can contain one or more arithmetic logic units (ALUs) that can be used to perform processing related to one of the many tasks or operations of the vehicle 1100—such as DNN processing. Additionally, the one or more SoCs 1104 can contain one or more floating-point units (FPUs)—or other mathematical or numerical coprocessors—for performing mathematical operations within the system. For example, the one or more SoCs 1104 can contain one or more FPUs integrated as execution units into one or more CPUs 1106 and / or one or more GPUs 1108.
[0143] The one or more SoCs 1104 can contain one or more accelerators 1114 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the one or more SoCs 1104 can contain a hardware acceleration cluster, which may include optimized hardware accelerators and / or a large amount of on-chip memory. The large on-chip memory (e.g., 4 MB SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used to complement the one or more GPUs 1108 and offload some of the tasks from the one or more GPUs 1108 (e.g., to free up more cycles of the one or more GPUs 1108 for other tasks). The one or more accelerators 1114 can, for example, be used for specific workloads (e.g.,Perception, convolutional neural networks (CNNs), etc., are used that are stable enough to be suitable for acceleration. The term "CNN" as used here can include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).
[0144] The one or more Accelerators 1114 (e.g., the Hardware Acceleration Cluster) can include a Deep Learning Accelerator (DLA). The one or more DLAs can include one or more Tensor Processing Units (TPUs) configured to provide an additional ten trillion operations per second for deep learning applications and inference. The TPUs can be accelerators configured and optimized to perform image processing functions (e.g., for CNNs, RCNNs, etc.). The one or more DLAs can also be optimized for a specific set of neural network types and floating-point operations, as well as for inference. The design of the one or more DLAs can deliver more performance per millimeter than a general-purpose GPU and far surpasses the performance of a CPU.The one or more TPUs can perform multiple functions, including a convolution function for a single instance that supports, for example, INT8, INT16 and FP16 data types for both features and weights, as well as post-processor functions.
[0145] One or more DLAs can quickly and efficiently run neural networks, especially CNNs, on processed or unprocessed data for a variety of functions, including, but not limited to: a CNN for object identification and detection using camera sensor data; a CNN for distance estimation using camera sensor data; a CNN for emergency vehicle detection and identification using microphone data; a CNN for facial recognition and vehicle owner identification using camera sensor data; and / or a CNN for security and / or protection-related events.
[0146] The one or more DLAs can execute any function of the one or more GPUs 1108, and by using an inference accelerator, a developer can, for example, allocate either the one or more DLAs or the one or more GPUs 1108 to each function. For example, the developer can concentrate the processing of CNNs and floating-point operations on the one or more DLAs and leave other functions to the one or more GPUs 1108 and / or other accelerators 1114.
[0147] The one or more Accelerators 1114 (e.g., the Hardware Acceleration Cluster) can contain a Programmable Vision Accelerator (PVA), which can also be referred to here as a Computer Vision Accelerator. The one or more PVAs can be designed and configured to accelerate computer vision algorithms for Advanced Driver Assistance Systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The one or more PVAs can offer a balance between performance and flexibility. For example, each PVA can contain any number of Reduced Instruction Set Computer (RISC) cores, Direct Memory Access (DMA) cores, and / or any number of vector processors, without limitation.
[0148] The RISC cores can interact with image sensors (e.g., the image sensors of one of the cameras described here), image signal processors, and / or the like. Each RISC core can contain any amount of memory. Depending on the implementation, the RISC cores can use any number of protocols. In some examples, the RISC cores can run a real-time operating system (RTOS). The RISC cores can be implemented with one or more integrated circuits, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC cores can include an instruction cache and / or tightly coupled RAM.
[0149] The DMA can allow components of the PVA(s) to access the system's memory independently of the one or more CPUs. The DMA can support any number of features that serve to optimize the PVA, including, but not limited to, support for multidimensional and / or circular addressing. In some examples, the DMA can support up to six or more dimensions of addressing, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.
[0150] Vector processors can be programmable processors designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, the PVA may contain a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, one or more DMA machines (e.g., two DMA machines), and / or other peripheral devices. The vector processing subsystem may operate as the primary processing machine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or working memory (e.g., VMEM).A VPU core can contain a digital signal processor, such as a single instruction, multiple data (SIMD) or a very long instruction word (VLIW). The combination of SIMD and VLIW can increase throughput and speed.
[0151] Each vector processor can contain an instruction cache and can be coupled to dedicated memory. Therefore, in some examples, each vector processor can be configured to operate independently of the others. In other examples, the vector processors contained in a particular PVA can be configured to use data parallelism. For example, in some embodiments, the multiple vector processors contained in a single PVA can execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors contained in a particular PVA can simultaneously execute different computer vision algorithms on the same image, or even different algorithms on successive images or sections of an image.Among other things, any number of PVAs can be contained in the hardware acceleration cluster, and any number of vector processors can be contained in each of the PVAs. Furthermore, one or more PVAs can contain additional memory for error-correcting code (ECC) to increase the overall security of the system.
[0152] The one or more Accelerators 1114 (e.g., the hardware acceleration cluster) can include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for the Accelerators 1114. In some examples, the on-chip memory can include at least 4 MB of SRAM, consisting, for example, and without limitation, of eight field-configurable memory blocks accessible to both the PVA and the DLA. Each pair of memory blocks can include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory can be used. The PVA and the DLA can access the memory via a backbone, enabling high-speed memory access for both the PVA and the DLA.The backbone can include an on-chip computer vision network that connects the PVA and DLA to the main memory (e.g., using the APB).
[0153] The on-chip computer vision network can include an interface that, prior to the transmission of control signals / addresses / data, ensures that both the PVA and the DLA provide ready-to-use and valid signals. Such an interface can provide separate phases and channels for the transmission of control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface can conform to ISO 26262 or IEC 61508 standards, although other standards and protocols can also be used.
[0154] In some examples, one or more SoCs 1104 can include a real-time ray tracing hardware accelerator as described in US patent application No. 16 / 101,232, filed on August 10, 2018. The real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the positions and extents of objects (e.g., within a world model) for generating real-time visualization simulations, for radar signal interpretation, for sound propagation synthesis and / or analysis, for simulating SONAR systems, for general wave propagation simulation, for comparison with lidar data for localization purposes, and / or for other functions and / or purposes. In some embodiments, one or more Tree Traversal Units (TTUs) can be used to perform one or more operations related to ray tracing.
[0155] The single or multiple Accelerator 1114 (e.g., the hardware accelerator cluster) have a wide range of applications for autonomous driving. The PVA can be a programmable vision accelerator used for critical processing steps in ADAS and autonomous vehicles. The PVA's capabilities are well-suited to algorithmic domains requiring predictable processing with low power consumption and low latency. In other words, the PVA is well-suited for semi-dense or dense regular computations, even with small datasets, that demand predictable runtimes with low latency and low power consumption. Therefore, in the context of autonomous vehicle platforms, PVAs are designed to execute classic computer vision algorithms, as they are efficient at object detection and operate with integer mathematics.
[0156] According to one embodiment of the technology, the PVA is used, for example, to perform computer stereovision. In some examples, a semi-global matching-based algorithm can be used, although this is not intended as a limitation. Many applications for Level 3-5 autonomous driving require spontaneous motion estimation or stereo matching (e.g., structure of motion, pedestrian detection, lane detection, etc.). The PVA can perform computer stereovision on input from two monocular cameras.
[0157] In some examples, the PVA can be used to perform dense optical flow processing. This involves processing raw radar data (e.g., using a 4D Fast Fourier Transform) to provide processed radar data. In other examples, the PVA is used for time-of-flight depth processing, for example, by processing raw time-of-flight data to deliver processed time-of-flight data.
[0158] The DLA can be used to power any type of network to improve control and driving safety; this includes, for example, a neural network that outputs a confidence score for each object detection. Such a confidence score can be interpreted as a probability or as providing a relative "weighting" of each detection compared to other detections. This confidence score allows the system to make further decisions about which detections should be considered true positives and not false positives. For example, the system can set a confidence threshold and consider only those detections that exceed the threshold as true positives.In an automatic emergency braking (AEB) system, false positive detections would cause the vehicle to automatically initiate emergency braking, which is obviously undesirable. Therefore, only the safest detections should be considered as triggers for AEB. The DLA can employ a neural network for confidence regression. The neural network can use as input at least a subset of parameters, such as the dimensions of the boundary frame, the ground plane estimate (obtained, for example, from another subsystem), the output of the inertial measurement unit (IMU) sensor 1166 correlated with the vehicle's orientation 1100, distance, and 3D position estimates of the object obtained from the neural network and / or other sensors (e.g., one or more LiDAR sensors 1164 or one or more radar sensors 1160).
[0159] The one or more SoCs 1104 can contain the one or more data stores 1116 (e.g., main memory). The one or more data stores 1116 can be on-chip main memory on the one or more SoCs 1104, where neural networks can be stored to run on the GPU and / or the DLA. In some examples, the one or more data stores 1116 can be large enough to store multiple instances of neural networks for redundancy and security. The one or more data stores 1112 can include one or more L2 or L3 caches 1112. The reference to the one or more data stores 1116 can include a reference to the main memory allocated to the PVA, the DLA, and / or one or more other accelerators 1114, as described here.
[0160] The one or more SoCs 1104 can contain one or more processors 1110 (e.g., embedded processors). The one or more processors 1110 can contain a boot and power management processor, which can be a dedicated processor and subsystem to handle boot power and management functions and the associated security enforcement. The boot and power management processor can be part of the boot sequence of the one or more SoCs 1104 and can provide runtime power management services. The boot and power management processor can provide clock and voltage programming, support for system transitions to a low-power state, management of the thermals and temperature sensors of the one or more SoCs 1104, and / or management of the one or more SoCs 1104 power states.Each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to the temperature, and the one or more SoCs 1104 can use the ring oscillators to detect the temperatures of the one or more CPUs 1106, the one or more GPUs 1108, and / or the one or more accelerators 1114. If it is determined that the temperatures exceed a threshold, the boot and power management processor can enter a temperature fault routine and put the one or more SoCs 1104 into a reduced-power state and / or put the vehicle 1100 into a chauffeur-to-safe-stop mode (e.g., bring the vehicle 1100 to a safe stop).
[0161] The one or more 1110 processors can also include a number of embedded processors that can serve as an audio processing engine. The audio processing engine can be an audio subsystem that provides full hardware support for multi-channel audio across multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor and dedicated RAM.
[0162] The one or more 1110 processors can also include an always-on processor machine, which provides the necessary hardware functions to support low-power sensor management and wake-up from use cases. The always-on processor machine can include a processor core, tightly coupled RAM, supporting peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0163] The one or more 1110 processors can also include a security cluster machine, which contains a dedicated processor subsystem for the security management of automotive applications. The security cluster machine can include two or more processor cores, tightly coupled RAM, supporting peripherals (such as timers, an interrupt controller, etc.), and / or routing logic. In a security mode, the two or more cores can operate in lockstep mode, functioning as a single core with comparison logic that detects any differences between their operations.
[0164] The one or more 1110 processors can also contain a real-time camera machine, which may include a dedicated processor subsystem for managing the real-time camera.
[0165] The one or more 1110 processors may also include a high dynamic range signal processor, which may include an image signal processor, which is a hardware machine that is part of the camera processing pipeline.
[0166] The one or more 1110 processors can include a video image compositor, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to generate the final image for the player window. The video image compositor can perform lens distortion correction on the one or more 1170 wide-angle cameras, the one or more 1174 ambient lighting cameras, and / or on the sensors of the in-cabin surveillance camera. The in-cabin surveillance camera sensor is preferably monitored by a neural network running on a separate instance of the extended SoC and configured to detect events in the cabin and respond accordingly.A system in the cabin can lip-read to activate mobile service and make a call, dictate emails, change the destination, activate or change the infotainment system and vehicle settings, or enable voice-controlled internet browsing. Certain functions are only available to the driver when the vehicle is operating in autonomous mode and are otherwise deactivated.
[0167] The video image compositor can include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, if there is motion in a video, the noise reduction weights the spatial information accordingly and reduces the impact of information provided by adjacent frames. If a frame or portion of a frame does not contain motion, the temporal noise reduction performed by the video image compositor can use information from the previous frame to reduce noise in the current frame.
[0168] The video image compositor can also be configured to perform stereo equalization of the input stereo lens images. Furthermore, the video image compositor can be used for user interface design when the operating system desktop is in use and the one or more GPUs 1108 do not need to constantly render new surfaces. Even when the one or more GPUs 1108 are powered on and actively performing 3D rendering, the video image compositor can be used to offload the workload from the GPUs 1108, thus improving performance and responsiveness.
[0169] The one or more 1104 SoCs can also include a serial camera interface with a Mobile Industry Processor Interface (MIPI) for receiving video and camera input, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functions. The one or more 1104 SoCs can also include one or more input / output controllers, one or more of which can be software-controlled and used for receiving I / O signals that are not assigned to a specific role.
[0170] The one or more 1104 SoCs can also include a wide range of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. The one or more 1104 SoCs can be used to process data from cameras (e.g., via Gigabit Multimedia Serial Link and Ethernet), sensors (e.g., one or more 1164 LiDAR sensors, one or more 1160 radar sensors, etc., which can be connected via Ethernet), data from the 1102 bus (e.g., vehicle speed 1100, steering wheel position, etc.), and data from one or more 1158 GNSS sensors (e.g., connected via Ethernet or CAN bus).Furthermore, the one or more SoCs 1104 can contain dedicated high-performance mass storage controllers, which can contain their own DMA machines and can be used to offload routine data management tasks from the one or more CPUs 1106.
[0171] The single or multiple 1104 SoCs can form an end-to-end platform with a flexible architecture spanning automation levels 3-5, thereby providing a comprehensive functional safety architecture that supports and efficiently utilizes computer vision and ADAS techniques for diversity and redundancy, and provides a platform for a flexible, reliable driving software stack along with deep learning tools. The single or multiple 1104 SoCs can be faster, more reliable, and even more energy- and space-efficient than conventional systems. For example, the single or multiple 1114 accelerators, in combination with the single or multiple 1106 CPUs, the single or multiple 1108 GPUs, and the single or multiple 1116 data stores, can form a fast, efficient platform for level 3-5 autonomous vehicles.
[0172] This technology thus offers capabilities and functions that cannot be achieved by conventional systems. For example, computer vision algorithms can be run on CPUs that can be configured using a high-level programming language, such as C, to execute a variety of processing algorithms on a wide range of visual data. However, CPUs are often unable to meet the performance requirements of many computer vision applications, such as execution time and power consumption. In particular, many CPUs are unable to execute complex object detection algorithms in real time, which is a requirement for in-vehicle ADAS applications and a prerequisite for practical Level 3-5 autonomous vehicles.
[0173] Unlike conventional systems, the technology described herein, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, enables the simultaneous and / or sequential execution of multiple neural networks and the combination of their results to enable Level 3-5 autonomous driving functionality. For example, a CNN running on the DLA or the dGPU (e.g., one or more GPUs 1120) can include text and word recognition, allowing the supercomputer to read and understand traffic signs, including signs for which the neural network has not been specifically trained. The DLA can further include a neural network capable of identifying and interpreting the sign, providing a semantic understanding, and passing this semantic understanding to the path planning modules running on the CPU complex.
[0174] Another example is that multiple neural networks can run simultaneously, as required for driving at levels 3, 4, or 5. For instance, a warning sign reading "Caution: Flashing lights indicate black ice" accompanied by an electric light can be interpreted independently or jointly by several neural networks. The sign itself can be identified as a traffic sign by a first neural network (e.g., a trained one), while the text "Flashing lights indicate black ice" can be interpreted by a second neural network, which then informs the vehicle's path planning software (preferably running on the CPU) that the presence of black ice indicates the presence of flashing lights.The turn signal can be identified across multiple images by a third neural network, which informs the vehicle's path planning software about the presence (or absence) of turn signals. All three neural networks can run simultaneously, e.g., within the DLA and / or on one or more GPUs 1108.
[0175] In some examples, a CNN for facial recognition and vehicle owner identification can use data from camera sensors to identify the presence of an authorized driver and / or owner of the vehicle 1100. The always-on sensor processing unit can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and to disable the vehicle in security mode when the owner leaves. In this way, one or more SoCs 1104 provide security against theft and / or carjacking.
[0176] In another example, a CNN for emergency vehicle detection and identification can use data from microphones 1196 to detect and identify emergency vehicle sirens. Unlike conventional systems that use general classifiers to detect sirens and manually extract features, the one or more SoCs 1104 use the CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to detect the relative approach speed of the emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area in which the vehicle is operating, as identified by one or more GNSS sensors 1158.For example, the CNN will attempt to detect European sirens when operating in Europe, and when operating in the United States, the CNN will attempt to identify only North American sirens. Once an emergency vehicle is detected, a controller can be used to execute an emergency vehicle safety routine, slowing the vehicle down, pulling over to the side of the road, parking the vehicle, and / or letting the vehicle idle, using the 1162 ultrasonic sensors, until one or more emergency vehicles pass.
[0177] The vehicle may contain one or more CPUs 1118 (e.g., one or more discrete CPUs or one or more dCPUs) that may be coupled to the one or more SoCs 1104 via a high-speed connection (e.g., PCIe). The CPUs 1118 may, for example, contain an x86 processor. The CPUs 1118 may be used, for example, to perform a variety of functions, including reconciling potentially inconsistent results between ADAS sensors and the one or more SoCs 1104 and / or monitoring the status and health of the one or more Controllers 1136 and / or the Infotainment SoC 1130.
[0178] The Vehicle 1100 can contain one or more GPUs 1120 (e.g., one or more discrete GPUs or one or more dGPUs) that can be coupled to the one or more SoCs 1104 via a high-speed connection (e.g., NVIDIA's NVLINK). The one or more GPUs 1120 can provide additional artificial intelligence capabilities, such as running redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs (e.g., sensor data) from sensors in the Vehicle 1100.
[0179] The vehicle 1100 may also include the network interface 1124, which may contain one or more wireless antennas 1126 (e.g., one or more wireless antennas for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). The network interface 1124 can be used to establish a wireless connection via the internet to the cloud (e.g., to one or more servers 1178 and / or other network devices), to other vehicles, and / or to computing devices (e.g., client devices of passengers). To communicate with other vehicles, a direct connection between the two vehicles and / or an indirect connection (e.g., via networks and the internet) can be established. Direct connections can be established via vehicle-to-vehicle communication.Vehicle-to-vehicle communication can provide vehicle 1100 with information about vehicles in its vicinity (e.g., vehicles in front of, beside, and / or behind vehicle 1100). This functionality can be part of a cooperative adaptive cruise control function of vehicle 1100.
[0180] The network interface 1124 can include a system-on-a-chip (SoC) that provides modulation and demodulation functions, enabling one or more controllers 1136 to communicate over wireless networks. The network interface 1124 can include a high-frequency (RF) front end for upconversion from baseband to RF and downconversion from RF to baseband. The frequency conversions can be performed using known methods and / or superheterodyne methods. In some examples, the RF front-end functionality can be provided by a separate chip. The network interface can include wireless functionality for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0181] The vehicle 1100 may further include one or more data storage devices 1128, which may be located outside the chip (e.g., outside the SoCs 1104). The one or more data storage devices 1128 may contain one or more memory elements, including RAM, SRAM, DRAM, VRAM, flash, hard disks, and / or other components and / or devices capable of storing at least one bit of data.
[0182] The 1100 vehicle can also include one or more 1158 GNSS sensors. The one or more 1158 GNSS sensors (e.g., GPS, supported GPS sensors, differential GPS (DGPS) sensors, etc.) assist with mapping, perception, grid generation, and / or path planning. Any number of 1158 GNSS sensors can be used, including, for example, a single GPS unit that uses a USB connection with an Ethernet-to-serial (RS-232) bridge.
[0183] The vehicle 1100 can also include one or more RADAR sensors 1160. The one or more RADAR sensors 1160 can be used by the vehicle 1100 to detect vehicles at long range, even in darkness and / or adverse weather conditions. The functional safety level of the RADAR can be ASIL B. The one or more RADAR sensors 1160 can use the CAN bus and / or the 1102 bus (e.g., for transmitting the data generated by the one or more RADAR sensors 1160) for control and access to object tracking data, with some examples using Ethernet for access to the raw data. A variety of RADAR sensor types can be used. The one or more RADAR sensors 1160 can be suitable for front, rear, and side RADAR applications without restriction. In some examples, one or more pulse-Doppler RADAR sensors are used.
[0184] The single or multiple RADAR 1160 sensors can incorporate various configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range with side coverage, etc. In some examples, long-range RADAR can be used for adaptive cruise control. Long-range RADAR systems can provide a wide field of view, achieved through two or more independent scans, for example, at a range of 250 m. The single or multiple RADAR 1160 sensors can assist in distinguishing between stationary and moving objects and can be used by ADAS systems for emergency braking and forward collision warning. Long-range RADAR sensors can incorporate a monostatic multimodal RADAR with multiple (e.g., six or more) fixed RADAR antennas and a high-speed CAN and FlexRay interface.In an example with six antennas, the four central antennas can generate a focused beam pattern designed to detect the surroundings of vehicle 1100 at higher speeds with minimal interference from traffic in adjacent lanes. The other two antennas can expand the field of view, enabling the rapid detection of vehicles entering or exiting vehicle 1100's lane.
[0185] Medium-range radar systems can, for example, have a range of up to 1160 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 1150 degrees (rear). Short-range radar systems can include, among other things, radar sensors designed for installation at both ends of the rear bumper. When such a radar sensor system is installed at both ends of the rear bumper, it can generate two beams that continuously monitor the blind spot behind and to the sides of the vehicle.
[0186] Short-range radar systems can be used in an ADAS system for blind spot detection and / or as a lane change assistant.
[0187] The vehicle 1100 can also include one or more ultrasonic sensors 1162. The one or more ultrasonic sensors 1162, which can be mounted on the front, rear, and / or sides of the vehicle 1100, can be used for parking assistance and / or for creating and updating an occupancy grid. A variety of ultrasonic sensors 1162 can be used, and different ultrasonic sensors 1162 can be used for different detection ranges (e.g., 2.5 m, 4 m). The one or more ultrasonic sensors 1162 can operate with functional safety levels of ASIL B.
[0188] The vehicle 1100 can contain one or more LiDAR sensors 1164. The one or more LiDAR sensors 1164 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The one or more LiDAR sensors 1164 can meet the functional safety level ASIL B. In some examples, the vehicle 1100 can contain multiple LiDAR sensors 1164 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to deliver data to a Gigabit Ethernet switch).
[0189] In some examples, one or more LiDAR sensors 1164 may be able to provide a list of objects and their distances for a 360-degree field of view. Commercially available LiDAR sensors 1164 may, for example, have a specified range of approximately 1100 m, with an accuracy of 2 cm to 3 cm and support for an 1100 Mbit / s Ethernet connection. In some examples, one or more non-protruding LiDAR sensors 1164 may be used. In such examples, the one or more LiDAR sensors 1164 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of the vehicle 1100. In such examples, one or more LIDAR sensors 1164 can provide a horizontal field of view of up to 120 degrees and a vertical field of view of up to 35 degrees, with a range of 200 m, even with objects of low reflectivity.The one or more front-mounted LIDAR sensors 1164 can be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0190] In some examples, LiDAR technologies, such as 3D flash LiDAR, can also be used. 3D flash LiDAR uses a laser pulse as a transmission source to illuminate the vehicle's surroundings up to approximately 200 m. A flash LiDAR unit contains a sensor that records the travel time of the laser pulse and the reflected light at each pixel, which in turn corresponds to the distance between the vehicle and the objects. Flash LiDAR can enable the generation of highly accurate and distortion-free images of the surroundings with each laser pulse. In some examples, four flash LiDAR sensors can be used, one on each side of the vehicle. Available 3D flash LiDAR systems include a solid-state 3D focal plane array LiDAR camera that contains no moving parts other than a fan (e.g., a non-scanning LiDAR device).The flash LIDAR device can use a 5-nanosecond pulse of a Class I (eye-safe) laser per frame and capture the reflected laser light in the form of 3D distance point clouds and co-registered intensity data. By using flash LIDAR, and because flash LIDAR is a solid-state device with no moving parts, the single or multiple LIDAR sensors can be less susceptible to motion blur, vibration, and / or shock.
[0191] The vehicle may also contain one or more IMU sensors 1166. In some examples, the one or more IMU sensors 1166 may be located in the center of the rear axle of the vehicle 1100. The one or more IMU sensors 1166 may, for example, and without limitation, contain one or more accelerometers, one or more magnetometers, one or more gyroscopes, one or more magnetic compasses, and / or other sensor types. In some examples, such as six-axis applications, the one or more IMU sensors 1166 may contain accelerometers and gyroscopes, while in nine-axis applications, the one or more IMU sensors 1166 may contain accelerometers, gyroscopes, and magnetometers.
[0192] In some embodiments, the one or more IMU sensors 1166 can be implemented as a miniaturized, high-performance GPS-aided inertial navigation system (GPS / INS) that combines inertial sensors of a microelectromechanical system (MEMS), a highly sensitive GPS receiver, and advanced Kalman filter algorithms to provide estimates of position, velocity, and orientation. Thus, in some examples, the one or more IMU sensors 1166 can enable the vehicle 1100 to estimate its course without requiring input from a magnetic sensor by directly observing and correlating velocity changes from the GPS with the one or more IMU sensors 1166.In some examples, one or more IMU sensors 1166 and one or more GNSS sensors 1158 can be combined in a single integrated unit.
[0193] The vehicle may contain one or more microphones 1196, which are mounted in and / or around the vehicle 1100. The one or more microphones 1196 may be used, among other things, for the detection and identification of emergency vehicles.
[0194] The vehicle may also include any number of camera types, including one or more stereo cameras 1168, one or more wide-angle cameras 1170, one or more infrared cameras 1172, one or more surround-view cameras 1174, one or more long-range and / or medium-range cameras 1198, and / or other camera types. The cameras can be used to capture image data around the entire periphery of the vehicle 1100. The types of cameras used depend on the embodiment and requirements of the vehicle 1100, and any combination of camera types can be used to ensure the necessary coverage around the vehicle 1100. Furthermore, the number of cameras can vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or any other number of cameras.The cameras can, for example and without limitation, support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the one or more cameras is described here with reference to... Fig. 11A and Fig. 11B described in more detail.
[0195] The vehicle 1100 may also contain one or more vibration sensors 1142. The one or more vibration sensors 1142 can measure vibrations of vehicle components, such as one or more axles. For example, changes in vibrations may indicate a change in the road surface. In another example, if two or more vibration sensors 1142 are used, the differences between the vibrations can be used to determine friction or slippage on the road surface (e.g., if the difference in vibration is between a driven axle and a freely rotating axle).
[0196] The vehicle 1100 may include an ADAS system 1138. In some examples, the ADAS system 1138 may include a SoC. The ADAS system 1138 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning systems (CWS), lane centering (LC), and / or other features and functions.
[0197] The ACC systems can use one or more radar sensors, one or more lidar sensors, and / or one or more cameras. The ACC systems can include longitudinal and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front and automatically adjusts the vehicle speed to maintain a safe distance from vehicles ahead. Lateral ACC maintains the distance and advises the vehicle to change lanes if necessary. Lateral ACC interacts with other ADAS applications, such as LCA and CWS.
[0198] The CACC uses information from other vehicles, which can be received via the network interface 1124 and / or the one or more wireless antennas 1126 from other vehicles via a wireless connection or indirectly via a network connection (e.g., via the internet). Direct connections can be provided via a vehicle-to-vehicle (V2V) communication link, while indirect connections can be an infrastructure-to-vehicle (I2V) communication link. Generally, the V2V communication concept provides information about the vehicles immediately ahead (e.g., vehicles directly in front of the vehicle 1100 and in the same lane), while the I2V communication concept provides information about traffic further ahead. CACC systems can incorporate both I2V and V2V information sources.Given the information about the vehicles ahead of vehicle 1100, the CACC can be more reliable and has the potential to improve traffic flow and reduce congestion on the road.
[0199] FCW systems are designed to warn the driver of a hazard, allowing them to take corrective action. FCW systems use a forward-facing camera and / or one or more RADAR 1160 sensors coupled with a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to the driver feedback system, such as a display, speaker, and / or vibrating component. FCW systems can provide a warning in the form of an audible signal, a visual warning, a vibration, and / or a rapid braking pulse.
[0200] AEB systems detect an impending forward collision with another vehicle or object and can automatically apply the brakes if the driver does not take corrective action within a specific time or distance parameter. AEB systems can use one or more forward-facing cameras and / or one or more radar sensors coupled with a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first warns the driver so they can take corrective action to avoid the collision; if the driver does not take corrective action, the AEB system can automatically apply the brakes to prevent or at least mitigate the effects of the predicted collision. AEB systems may incorporate techniques such as dynamic brake assist and / or emergency braking for an impending collision.
[0201] Lane Departure Warning (LDW) systems provide visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver if the vehicle crosses lane markings. An LDW system will not activate if the driver indicates an intentional lane departure by using a turn signal. LDW systems may utilize forward-facing cameras coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to the feedback signal for the driver, such as a display, speaker, and / or vibrating component.
[0202] LKA systems are a variant of LDW systems. LKA systems provide steering or braking inputs to correct the vehicle 1100 if the vehicle 1100 begins to leave its lane.
[0203] Blind Spot Warning (BSW) systems detect and warn the driver of vehicles in the car's blind spot. BSW systems can provide visual, audible, and / or tactile warnings to indicate that merging into or changing lanes is unsafe. The system can issue an additional warning if the driver activates a turn signal. BSW systems can utilize one or more rear-facing cameras and / or radar sensors coupled with a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to the driver for feedback, such as a display, speaker, and / or vibrating component.
[0204] RCTW systems can provide visual, audible, and / or tactile alerts when an object is detected outside the reversing camera's field of view while the vehicle is reversing. Some RCTW systems incorporate AEB (Automatic Emergency Braking) to ensure the vehicle's brakes are applied to prevent a collision. RCTW systems can utilize one or more rear-facing radar sensors coupled with a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to the driver for feedback, such as a display, speaker, and / or vibrating component.
[0205] Conventional ADAS systems can produce false positives, which, while annoying and distracting for the driver, are generally not catastrophic because the ADAS systems warn the driver and give them the opportunity to decide whether a safety issue truly exists and to act accordingly. However, in an autonomous vehicle 1100, the vehicle 1100 itself must decide, in the event of conflicting results, whether to follow the result from a primary computer or a secondary computer (e.g., a first controller 1136 or a second controller 1136). In some embodiments, the ADAS system 1138 can, for example, be a backup and / or secondary computer that provides information about perception to a rationality module of the backup computer.The backup computer rationality monitor can run redundant, diverse software on hardware components to detect errors in perception and dynamic driving tasks. The outputs of the ADAS system 1138 can be provided to a monitoring MCU. If the outputs of the primary and secondary computers conflict, the monitoring MCU must determine how to resolve the conflict to ensure safe operation.
[0206] In some examples, the primary computer can be configured to provide the monitoring MCU with a confidence score indicating its confidence in the chosen outcome. If the confidence score exceeds a threshold, the monitoring MCU can follow the primary computer's instruction, regardless of whether the secondary computer returns a conflicting or inconsistent result. If the confidence score does not reach the threshold and the primary and secondary computers display different results (e.g., conflicting results), the monitoring MCU can mediate between the computers to determine the appropriate outcome.
[0207] The monitoring MCU can be configured to run one or more neural networks that are trained and configured to determine, based on the output of the primary and secondary computers, the conditions under which the secondary computer will trigger false alarms. This allows the one or more neural networks in the monitoring MCU to learn when the output of the secondary computer can be trusted and when it cannot. For example, if the secondary computer is a radar-based FCW system, a neural network in the monitoring MCU can learn when the FCW system identifies metallic objects that do not actually pose a threat, such as a drain grate or manhole cover, triggering an alarm.Similarly, if the secondary computer is a camera-based lane departure warning (LDW) system, a neural network in the supervising MCU can learn to override the LDW system when cyclists or pedestrians are present and leaving the lane is indeed the safest maneuver. In embodiments containing one or more neural networks running on the supervising MCU, the supervising MCU can include at least one DLA or GPU suitable for executing the one or more neural networks with associated memory. In preferred embodiments, the supervising MCU can include and / or be contained as a component of the one or more SoCs 1104.
[0208] In other examples, the ADAS system 1138 can include a secondary computer that executes the ADAS functionality according to the classical rules of computer vision. Thus, the secondary computer can use classical computer vision rules (if-then), and the presence of one or more neural networks in the monitoring MCU can improve reliability, safety, and performance. For example, the diverse implementation and intentional non-identity make the overall system more fault-tolerant, especially against errors caused by software (or software-hardware interfaces).For example, if a software bug or error occurs in the software on the primary computer and the non-identical software code on the secondary computer produces the same overall result, the monitoring MCU can have greater confidence that the overall result is correct and that the bug in the software or hardware on the primary computer does not cause a significant error.
[0209] In some examples, the output of the ADAS system 1138 can be fed into the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, if the ADAS system 1138 displays a forward collision warning due to an object immediately in front of the vehicle, the perception block can use this information in object identification. In other examples, the secondary computer may have its own trained neural network, thus reducing the risk of false positives, as described herein.
[0210] The Vehicle 1100 may also include the Infotainment SoC 1130 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system may not actually be an SoC and may contain two or more discrete components. The Infotainment SoC 1130 may include a combination of hardware and software that can be used to provide the Vehicle 1100 with audio (e.g., music, a personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation systems, rear parking sensors, a radio data system, vehicle-related information such as fuel level, total distance traveled, brake fluid level, oil level, door open / close status, air filter information, etc.).The Infotainment SoC 1130 can include, for example, radios, record players, navigation systems, video players, USB and Bluetooth connectivity, car computers, in-car entertainment, Wi-Fi, steering wheel audio controls, hands-free systems, a head-up display (HUD), an HMI display 1134, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, functions, and / or systems), and / or other components. The Infotainment SoC 1130 can also be used to provide information (e.g., visual and / or audible) to one or more vehicle users, such as information from the ADAS system 1138, autonomous driving information such as planned vehicle maneuvers, road layouts, environmental information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0211] The infotainment SoC 1130 can include GPU functionality. The infotainment SoC 1130 can communicate with other devices, systems, and / or components of the vehicle 1100 via the bus 1102 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 1130 can be coupled with a monitoring MCU so that the infotainment system's GPU can perform some self-driving functions if one or more primary controllers 1136 (e.g., the vehicle 1100's primary and / or backup computers) fail. In such an example, the infotainment SoC 1130 can put the vehicle 1100 into a chauffeur-to-safe-stop mode, as described here.
[0212] The vehicle 1100 may also include an instrument cluster 1132 (e.g., a digital instrument panel, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 1132 may include a controller and / or supercomputer (e.g., a discrete controller or supercomputer). The instrument cluster 1132 may contain a number of instruments, such as a speedometer, fuel gauge, oil pressure gauge, tachometer, odometer, turn signals, shift position indicator, seat belt warning light(s), parking brake warning light(s), engine malfunction light(s), airbag system (SRS) information, lighting controls, safety system controls, navigation information, etc. In some examples, information from the infotainment SoC 1130 and the instrument cluster 1132 may be displayed and / or shared. In other words, the Instrument Cluster 1132 can be included as part of the Infotainment SoC 1130, or vice versa.
[0213] Fig. Figure 11D is a system diagram for the communication between the one or more cloud-based servers and the exemplary autonomous vehicle 1100. Fig. 11A, according to some embodiments of the present disclosure. The system 1176 may include the one or more servers 1178, the one or more networks 1190, and the vehicles, including the vehicle 1100. The server(s) 1178 may include multiple GPUs 1184(A)-1184(H) (here collectively referred to as GPUs 1184), PCIe switches 1182(A)-1182(H) (here collectively referred to as PCIe switches 1182), and / or CPUs 1180(A)-1180(B) (here collectively referred to as CPUs 1180). The GPUs 1184, the CPUs 1180, and the PCIe switches can be interconnected via high-speed connections, such as, but without limitation, NVIDIA's NVLink interfaces 1188 and / or PCIe connections 1186. In some examples, the GPUs 1184 are connected via NVLink and / or NVSwitch SoCs, and the GPUs 1184 and the PCIe switches 1182 are connected via PCIe connections.Although eight GPUs 1184, two CPUs 1180, and two PCIe switches are illustrated, this should not be interpreted as a limitation. Depending on the configuration, each Server 1178 can contain any number of GPUs 1184, CPUs 1180, and / or PCIe switches. For example, one or more Server 1178s can each contain eight, sixteen, thirty-two, and / or more GPUs 1184.
[0214] The one or more servers 1178 can receive image data from the vehicles via the one or more networks 1190. This image data is representative of images showing unexpected or changed road conditions, such as recently started roadworks. The one or more servers 1178 can transmit neural networks 1192, updated neural networks 1192, and / or map information 1194 to the vehicles via the one or more networks 1190. This map information contains information about traffic and road conditions. The map information updates 1194 can include updates for the HD map 1122, such as information about construction sites, potholes, detours, flooding, and / or other obstacles.In some examples, the neural networks 1192, the updated neural networks 1192 and / or the map information 1194 may result from new training and / or experience represented in the data received from any number of vehicles in the environment, and / or may be based on training performed in a data center (e.g. using one or more servers 1178 and / or other servers).
[0215] One or more Server 1178 systems can be used to train machine learning models (e.g., neural networks) based on training data. The training data can be generated by the vehicles and / or in a simulation (e.g., using a game machine). In some examples, the training data is tagged (e.g., if the neural network benefits from supervised learning) and / or subjected to other preprocessing, while in other examples, the training data is not tagged and / or preprocessed (e.g., if the neural network does not require supervised learning).Training can be performed using one or more classes of machine learning techniques, including, but not limited to, classes such as: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and cluster analysis), multilinear subspace learning, diverse learning, representational learning (including substitute dictionary learning), rule-based machine learning, anomaly detection, and all variants or combinations thereof. Once the machine learning models are trained, they can be used by the vehicles (e.g., transmitted to the vehicles via one or more networks 1190) and / or used by one or more servers 1178 for remote monitoring of the vehicles.
[0216] In some examples, one or more Server 1178 units can receive data from the vehicles and apply that data to advanced neural networks in real time for intelligent, real-time inference. The one or more Server 1178 units can include deep learning supercomputers and / or dedicated AI computers powered by GPUs 1184, such as NVIDIA's DGX and DGX Station machines. However, in some examples, the one or more Server 1178 units can include a deep learning infrastructure that uses only CPU-powered data centers.
[0217] The deep learning infrastructure of one or more servers 1178 can perform fast, real-time inference and can use this capability to assess and verify the state of the processors, software, and / or associated hardware in the vehicle 1100. For example, the deep learning infrastructure can receive periodic updates from the vehicle 1100, such as a sequence of images and / or objects that the vehicle 1100 has located within that sequence of images (e.g., via computer vision and / or other machine learning object classification techniques).The deep learning infrastructure can run its own neural network to identify the objects and compare them with the objects identified by vehicle 1100. If the results do not match and the infrastructure concludes that the AI in vehicle 1100 is not working correctly, one or more servers 1178 can send a signal to vehicle 1100, instructing a fail-safe computer in vehicle 1100 to take control, notify the passengers, and perform a safe parking maneuver.
[0218] For inference, one or more Server 1178 systems can include GPUs 1184 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of GPU-driven servers and inference accelerators can enable real-time responsiveness. In other scenarios, such as when performance is less critical, servers powered by CPUs, FPGAs, and other processors can be used for inference. EXAMPLE CALCULATION DEVICE
[0219] Fig. Figure 12 is a block diagram of an exemplary computing device 1200 suitable for use in implementing some embodiments of the present disclosure. The computing device 1200 may include a connection system 1202 that directly or indirectly couples the following devices: main memory 1204, one or more central processing units (CPUs) 1206, one or more graphics processing units (GPUs) 1208, a communication interface 1210, input / output (I / O) ports 1212, input / output components 1214, a power supply 1216, one or more presentation components 1218 (e.g., display(s)), and one or more logic units 1220. In at least one embodiment, the one or more computing devices 1200 may comprise one or more virtual machines (VMs), and / or each of the components thereof may comprise virtual components (e.g., virtual hardware components).As non-restrictive examples, one or more of the GPUs 1208 can comprise one or more vGPUs, one or more of the CPUs 1206 can comprise one or more vCPUs, and / or one or more of the logic units 1220 can comprise one or more virtual logic units. Thus, a computing device 1200 can contain discrete components (e.g., a complete GPU allocated to the computing device 1200), virtual components (e.g., a portion of a GPU allocated to the computing device 1200), or a combination thereof.
[0220] Although the various blocks of Fig. Where components 12 are shown as connected via the connection system 1202, this is not intended as a limitation and is for clarity only. In some embodiments, for example, a presentation component 1218, such as a display device, can be considered an I / O component 1214 (e.g., if the display is a touchscreen). As another example, the CPUs 1206 and / or GPUs 1208 can contain memory (e.g., the memory 1204 can represent a storage device in addition to the memory of the GPUs 1208, the CPUs 1206, and / or other components). In other words, the computing device of Fig. Section 12 is for illustrative purposes only. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "game console", "electronic control unit (ECU)", "virtual reality system" and / or other device or system types, as all are within the scope of protection of the computing device of Fig. 12 are being considered.
[0221] The 1202 interconnection system can represent one or more connections or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The 1202 interconnection system can include one or more bus or connection types, such as an Industry Standard Architecture (ISA) bus, an Extended ISA bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI Express (PCIe) bus, and / or another type of bus or connection. In some embodiments, there are direct connections between components. For example, the 1206 CPU can be directly connected to the 1204 main memory. Furthermore, the 1206 CPU can be directly connected to the 1208 GPU.In a direct or point-to-point connection between components, the 1202 connection system can include a PCIe link to establish the connection. In these examples, a PCI bus does not need to be included in the 1200 computing device.
[0222] The 1204 main memory can contain a variety of computer-readable media. Computer-readable media can be any available media that the 1200 computing device can access. Computer-readable media can include both volatile and non-volatile media, as well as removable and non-removable media. For example, and without limitation, computer-readable media can include computer storage media and communication media.
[0223] Computer storage media can include both volatile and non-volatile media, and / or removable and non-removable media, implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, and / or other data types. For example, main memory can store computer-readable instructions (e.g., representing one or more programs and / or one or more program elements, such as an operating system). Computer storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other storage technologies; CD-ROM, Digital Versatile Discs (DVDs), or other optical disk storage; magnetic cartridges, magnetic tapes, magnetic disk storage, or other magnetic storage devices; or any other medium that can be used to store the desired information and that the computing device can access.As used here, computer storage media do not inherently contain signals.
[0224] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave or other transport mechanism, and include any media for transmitting information. The term "modulated data signal" can refer to a signal in which one or more of its properties are set or modified to encode information within the signal. Computer storage media can include, but are not limited to, wired media, such as a wired network or a direct-wired connection, and wireless media, such as acoustic, RF, infrared, and other wireless media. Combinations of the above should also be included in the scope of protection of the computer-readable media.
[0225] The one or more CPUs 1206 can be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1200 to perform one or more of the procedures and / or processes described herein. The one or more CPUs 1206 can each contain one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing a multitude of software threads simultaneously. The one or more CPUs 1206 can contain any type of processor and can contain different types of processors depending on the type of computing device 1200 implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).Depending on the type of computing device 1200, the processor can be, for example, an Advanced RISC Machine (ARM) processor implemented with Reduced Instruction Set Computing (RISC), or an x86 processor implemented with Complex Instruction Set Computing (CISC). The computing device 1200 can contain one or more CPUs 1206, in addition to one or more microprocessors or additional coprocessors, such as mathematical coprocessors.
[0226] In addition to or as an alternative to the one or more CPUs 1206, the one or more GPUs 1208 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1200 to perform one or more of the procedures and / or processes described herein. One or more of the GPUs 1208 can be an integrated GPU (e.g., with one or more of the CPUs 1206) and / or one or more of the GPUs 1208 can be a discrete GPU. In embodiments, one or more of the GPUs 1208 can be a coprocessor of one or more of the CPUs 1206. The one or more GPUs 1208 can be used by the computing device 1200 to render graphics (e.g., 3D graphics) or to perform general-purpose computing. The one or more GPUs 1208 can be used, for example, for general-purpose computing on GPUs (GPGPU).The one or more GPUs 1208 can contain hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The one or more GPUs 1208 can generate pixel data for output images in response to rendering commands (e.g., rendering commands from the one or more CPUs 1206 received via a host interface). The one or more GPUs 1208 can include graphics memory, such as display memory, for storing pixel data or other suitable data, such as GPGPU data. The display memory can be included as part of the 1204 main memory. The one or more GPUs 1208 can contain two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using NVLINK) or connect them via a switch (e.g., using NVSwitch).When combined, each GPU can generate 1208 pixel data or GPGPU data for different sections of an output or for different outputs (e.g., a first GPU for a first image and a second GPU for a simulated image). Each GPU can have its own dedicated memory or share memory with other GPUs.
[0227] In addition to or as an alternative to the one or more CPUs 1206 and / or the one or more GPUs 1208, the one or more logic units 1220 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1200 to perform one or more of the methods and / or processes described herein. In embodiments, the one or more CPUs 1206, the GPUs 1208, and / or the one or more logic units 1220 may discretely or jointly execute any combination of the methods, processes, and / or sections thereof. One or more of the logic units 1220 may be part of and / or integrated into one or more of the CPUs 1206 and / or one or more of the GPUs 1208, and / or one or more of the logic units 1220 may be discrete components or otherwise separate from the CPUs 1206 and / or the GPUs 1208.In embodiments, one or more of the logic units 1220 can be a co-processor of one or more of the CPUs 1206 and / or one or more of the GPUs 1208.
[0228] Examples of one or more logic units 1220 contain one or more processing cores and / or components thereof, such as data processing units (DPUs), tensor cores (TCs), tensor processing units (TPUs), pixel visual cores (PVCs), vision processing units (VPUs), graphics processing clusters (GPCs), texture processing clusters (TPCs), streaming multiprocessors (SMs), tree traversal units (TTUs), artificial intelligence accelerators (AIAs), deep learning accelerators (DLAs), arithmetic logic units (ALUs), and application-specific integrated circuits. (Application-Specific Integrated Circuits, ASICs), Floating Point Units (FPUs),Input / output (I / O) elements, peripheral component interconnect (PCI) or PCI Express (PCIe) elements, and / or similar.
[0229] The 1210 communication interface can include one or more receivers, transmitters, and / or transceivers that enable the 1200 computing device to communicate with other computing devices over an electronic network, including wired and / or wireless communication. The 1210 communication interface can include components and functions that enable communication over a variety of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., Ethernet or InfiniBand communication), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.In one or more embodiments, the one or more logic units 1220 and / or the communication interface 1210 may contain one or more data processing units (DPUs) to transfer data received via a network and / or via the connection system 1202 directly to one or more GPUs 1208 (e.g., a memory thereof).
[0230] The I / O ports 1212 enable the computing device 1200 to be logically coupled with other devices, including the I / O components 1214, one or more presentation components 1218, and / or other components, some of which may be built into (e.g., integrated with) the computing device 1200. Illustrative I / O components 1214 include a microphone, mouse, keyboard, joystick, gamepad, game controller, satellite dish, scanner, printer, wireless device, etc. The I / O components 1214 can provide a natural user interface (NUI) that processes air gestures, speech, or other physiological inputs generated by a user. In some cases, the inputs can be transmitted to a suitable network element for further processing.A NUI can implement any combination of speech capture, stylus capture, face capture, biometric capture, gesture capture (both on-screen and off-screen), air gestures, head and eye tracking, and touch capture (as further described below) associated with a display of the Computing Device 1200. The Computing Device 1200 can include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture capture and recognition. Additionally, the Computing Device 1200 can include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit (IMU)) that enable motion detection. In some examples, the output from the accelerometers or gyroscopes can be used by the Computing Device 1200 to render immersive augmented reality or virtual reality.
[0231] The power supply 1216 can include a hardwired power supply, a battery power supply, or a combination thereof. The power supply 1216 can power the computing device 1200 to enable the operation of the computing device 1200's components.
[0232] The one or more presentation components 1218 can include a display (e.g., a monitor, a touchscreen, a television screen, a head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The one or more presentation components 1218 can receive data from other components (e.g., the one or more GPUs 1208, the one or more CPUs 1206, DPUs, etc.) and output the data (e.g., as an image, video, sound, etc.). EXEMPLARY DATA CENTER
[0233] Fig. Figure 13 illustrates an exemplary data center 1300 that can be used in at least one embodiment of the present disclosure. The data center 1300 can include an infrastructure layer 1310, a framework layer 1320, a software layer 1330, and / or an application layer 1340.
[0234] As in Fig. As shown in Figure 13, the infrastructure layer 1310 of the data center can contain a resource orchestrator 1312, grouped compute resources 1314 and node compute resources (“node RRs”) 1316(1)-1316(N), where “N” is any positive integer. In at least one embodiment, the node RRs 1316(1)-1316(N) may include, but are not limited to, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic solid storage), storage devices (e.g., solid-state or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power supply modules and / or cooling modules, etc.In some embodiments, one or more Node RRs among Node RRs 1316(1)-1316(N) may correspond to a server that has one or more of the computing resources mentioned above. Furthermore, in some embodiments, Node RRs 1316(1)-1316(N) may contain one or more virtual components, such as vGPUs, vCPUs, and / or the like, and / or one or more of Node RRs 1316(1)-1316(N) may correspond to a virtual machine (VM).
[0235] In at least one embodiment, the grouped compute resources 1314 can contain separate groupings of node RRs 1316, which are housed in one or more racks (not shown) or in many racks in data centers at different geographic locations (also not shown). Separate groupings of node RRs 1316 within grouped compute resources 1314 can contain grouped compute, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, multiple node RRs 1316, including the CPUs, GPUs, DPUs, and / or other processors, can be grouped in one or more racks to provide compute resources to support one or more workloads.The one or more racks can also contain any number of power supply modules, cooling modules and / or network switches in any combination.
[0236] The resource orchestrator 1312 can configure or otherwise control one or more node RRs 1316(1)-1316(N) and / or grouped compute resources 1314. In at least one embodiment, the resource orchestrator 1312 can include an entity for managing the software design infrastructure (SDI) for the data center 1300. The resource orchestrator 1312 can include hardware, software, or a combination thereof.
[0237] In at least one embodiment, as in Fig. As shown in Figure 13, the framework layer 1320 can contain a job scheduler 1333, a configuration manager 1334, a resource manager 1336, and / or a distributed file system 1338. The framework layer 1320 can contain a framework that supports the software 1332 of the software layer 1330 and / or one or more applications 1342 of the application layer 1340. The software 1332 or the one or more applications 1342 can each contain web-based service software or applications such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 1320 can be a type of free and open-source software web application framework, such as Apache Spark™ (hereinafter "Spark"), which can utilize a distributed file system 1338 for processing large amounts of data (e.g., "Big Data"), without being limited to it.In at least one embodiment, the job scheduler 1333 can include a Spark driver to facilitate the scheduling of workloads supported by different layers of the data center 1300. The configuration manager 1334 can be able to configure different layers, such as the software layer 1330 and the framework layer 1320, which contains Spark and the distributed file system 1338, to support the processing of large amounts of data. The resource manager 1336 can be able to manage clustered or grouped compute resources allocated or assigned to support the distributed file system 1338 and the job scheduler 1333. In at least one embodiment, the clustered or grouped compute resources can include the grouped compute resource 1314 on the infrastructure layer 1310 of the data center.The resource manager 1336 can coordinate with the resource orchestrator 1312 to manage these allocated or assigned computing resources.
[0238] In at least one embodiment, the software 1330 contained in software layer 1332 may include software used by at least sections of the node RRs 1316(1)-1316(N), the grouped compute resources 1314, and / or the distributed file system 1338 of framework layer 1320. One or more types of software may include, among others, web page search software, email virus scanning software, database software, and streaming video content software.
[0239] In at least one embodiment, the applications 1342 contained in the application layer 1340 may include one or more types of applications used by at least sections of the node RRs 1316(1)-1316(N), the grouped compute resources 1314, and / or the distributed file system 1338 of the framework layer 1320. One or more types of applications may include, but are not limited to, any number of genome applications, cognitive computations, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.
[0240] In at least one embodiment, a configuration manager 1334, a resource manager 1336, and a resource orchestrator 1312 can implement any number and type of self-modifying actions based on any set and type of data acquired in any technically feasible manner. Self-modifying actions can relieve a data center operator of data center 1300 of potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly functioning sections of a data center.
[0241] The Data Center 1300 may contain tools, services, software, or other resources to train one or more machine learning models or to predict or infer information using one or more machine learning models according to one or more embodiments described herein. For example, one or more machine learning models may be trained by calculating weighting parameters according to a neural network architecture, using software and / or computing resources described above with reference to the Data Center 1300.In at least one embodiment, trained or deployed machine learning models corresponding to one or more neural networks can be used to infer or predict information using the resources described above with reference to the Computing Center 1300 by using weighting parameters calculated by one or more training techniques such as, but not limited to, those described herein.
[0242] In at least one embodiment, the data center can use 1300 CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or equivalent virtual computing resources) to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above can be configured as a service to allow users to train or infer information, such as image capture, speech capture, or other artificial intelligence services. EXEMPLARY NETWORK ENVIRONMENTS
[0243] Network environments suitable for implementing embodiments of the disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may run on one or more instances of the one or more computing devices. Fig. 12 are implemented - e.g., each device may contain similar components, features, and / or functionality to one or more computing devices 1200. If backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices may also be included as part of a data center 1300, an example of which is given herein with reference to Fig. 13 is described in more detail.
[0244] The components of a network environment can communicate with each other over one or more networks, which can be wired, wireless, or both. The network can contain multiple networks or a network of networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the internet and / or a public switched telephone network (PSTN), and / or one or more private networks. If the network includes a wireless telecommunications network, components such as a base station, a communications tower, or even access points (as well as other components) can provide wireless connectivity.
[0245] Compatible network environments can contain one or more peer-to-peer network environments—in which case a server cannot be included in a network environment—and one or more client-server network environments—in which case one or more servers can be included in a network environment. In peer-to-peer network environments, the functionality described here can be implemented on any number of client devices with reference to one or more servers.
[0246] In at least one embodiment, a network environment can include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. A framework layer can include a framework to support software of a software layer and / or one or more applications of an application layer. The software or the one or more applications may each include web-based service software or applications. In embodiments, one or more of the client devices can use the web-based service software or applications (e.g.,by accessing the service software and / or applications via one or more application programming interfaces (APIs). The framework layer can be a type of free and open-source software web application framework that uses, for example, a distributed file system for processing large amounts of data (e.g., "Big Data"), but is not limited to this.
[0247] A cloud-based network environment can provide cloud computing and / or cloud storage, performing any combination (or parts thereof) of the computing and / or data storage functions described herein. Each of these different functions can be distributed across multiple locations of central or core servers (e.g., one or more data centers, which may be distributed across a state, region, country, the globe, etc.). If a connection to a user (e.g., a client device) is relatively close to one or more edge servers, one or more core servers can offload at least some functionality to the one or more edge servers. A cloud-based network environment can be private (e.g., restricted to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0248] The one or more client devices can have at least some of the components, features, and functions of the one or more mentioned here in relation to Fig.The 12 exemplary computing devices described may include 1200. By way of example, and not as a limitation, a client device may be a personal computer (PC), a laptop, a mobile device, a smartphone, a tablet computer, a smartwatch, a portable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or global positioning device, a video player, a video camera, a surveillance device or surveillance system, a vehicle, a boat, a hydrofoil, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or gaming system, an entertainment system, a vehicle computer system, an embedded system controller, a remote control, a device, a consumer electronics device, a workstation, an edge device,any combination of these described devices or any other suitable device may be embodied.
[0249] The revelation can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions such as program modules, which are executed by a computer or other machine, such as a personal data assistant or other handheld device. Generally, program modules contain routines, programs, objects, components, data structures, etc., and refer to code that performs specific tasks or implements certain abstract data types. The revelation can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc.The revelation can also be practiced in distributed computing environments, where tasks are performed by remote processing devices that are connected to each other via a network for communication.
[0250] As used herein, any mention of "and / or" in relation to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, "element A, element B and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Furthermore, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Additionally, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0251] The subject matter of this disclosure is specifically described herein to satisfy legal requirements. However, the description itself is not intended to limit the scope of protection afforded by this disclosure. Rather, the inventors have considered that the claimed subject matter may also be embodied in other ways to include various steps or combinations of steps similar to those described in this document, in conjunction with other present or future technologies. Although the terms “step” and / or “block” may be used herein to denote various elements of the methods employed, these terms should not be interpreted as implying any particular sequence among or between the various steps disclosed herein, except where the sequence of each step is expressly described. EXAMPLE PARAGRAPHS A: Method comprising: Determining, based on at least sensor data obtained using one or more first sensors of a machine, a location associated with the machine in an environment; Determining, based on at least one or more vision language models processing at least image data obtained using one or more second sensors of the machine, information associated with a road hazard located in the environment; and sending data to one or more systems to update one or more applications to indicate at least some of the information associated with the road hazard approximately at that location. B: Procedure according to paragraph A, further comprising: determining, based on at least one or more language models processing second data representing the information, a request to verify the information; providing content associated with the request to verify the information; and receiving input data indicating that the information associated with the road hazard has been verified, the sending of the data to update the one or more applications being based at least on the road hazard being verified. C: Method according to either paragraph A or paragraph B, further comprising: determining, based on at least one or more language models processing second data representing the information, a request for second information associated with the road hazard; providing content associated with the request for the second information and receiving input data specifying the second information associated with the road hazard, the data further serving to update the one or more applications to specify at least part of the second information. D: Method according to any of paragraphs AC, wherein the determination of the information associated with the road hazard is further based on at least one or more vision language models processing second data comprising at least one of the following: a first prompt requesting information associated with road hazards that may be in the environment; or a second prompt requesting second information associated with one or more road hazards that may be in the environment, wherein the one or more road hazards include at least the road hazard. E: Method according to any of paragraphs AD, wherein: determining the information associated with the road hazard is further based on at least one or more vision language models processing second data representing a first prompt associated with the road hazard; and the method further comprises determining, based on at least one or more vision language models processing at least the image data and third data representing a second prompt associated with a second road hazard, second information associated with the second road hazard that may be in the vicinity. F: Method according to one of claims AE, wherein determining the information associated with the road hazard comprises: determining, based on at least one or more vision language models processing the image data and second data representing a first prompt, initial information indicating that the road hazard is in the vicinity; and determining, based on at least one or more vision language models processing the image data and third data representing the initial information and a second prompt, the information associated with the road hazard. G: Method according to any of paragraphs AF, further comprising: Determining, based on at least second sensor data obtained using one or more third sensors of the machine, second information relating to the road hazard, wherein the data further serve to update the one or more applications in order to indicate at least part of the second information relating to the road hazard. H: Method according to one of paragraphs AG, wherein the information associated with the road hazard describes at least one of the following: a type of hazard; a location of the hazard in the surroundings; or a location of the hazard as represented by one or more images provided by the image data. I: Procedure according to any of paragraphs AH, wherein the road hazard includes at least one of the following: an obstruction located near a roadway; an emergency service taking place near the roadway; a construction activity associated with the roadway; an animal located near the roadway; a pedestrian located near the roadway; or a hazardous road condition associated with the roadway. J: System, comprising: one or more processors, for: receiving sensor data obtained using one or more sensors of a machine navigating in an environment; determining, using one or more machine learning models and based on at least the sensor data and one or more prompts associated with one or more hazards, information corresponding to a hazard located in the environment; and sending, to one or more systems, data associated with the information corresponding to the hazard, the data being used to update an application to indicate the information corresponding to the hazard. K: System according to paragraph J, wherein the one or more processors further serve to: Determine, based on at least second sensor data obtained using one or more second sensors, a location associated with the hazard, wherein the data further associates the information. L: System either according to paragraph J or according to paragraph K, wherein the one or more processors further serve to: determine, using the one or more machine learning models and based on at least the information, a request to verify the information; provide content associated with the request to verify the information; and receive input data indicating that the information associated with the hazard has been verified, the data being sent based on at least the hazard being verified. M: System according to any of paragraphs JL, wherein the one or more processors further serve to: determine, using the one or more machine learning models and based on at least the information, a request for a second piece of information associated with the hazard; provide content associated with the request for the second piece of information; and receive input data specifying the second piece of information associated with the hazard, wherein the data further associates the second piece of information. N: System according to any of paragraphs JM, wherein the one or more processors further serve to: determine, using the one or more machine learning models and based on at least the information, to provide the application with the information to update, wherein the data is sent based on at least the determination to provide the information. O: System according to any of paragraphs JN, wherein the one or more processors further serve to: Determine, using one or more machine learning models and based at least on sensor data and a prompt associated with one or more second hazards, second information corresponding to a second hazard potentially present in the environment. P: System according to one of paragraphs JO, wherein the determination of the information associated with the hazard comprises: determining, using the one or more machine learning models and based at least on the sensor data and a first prompt of the one or more prompts, initial information indicating that the hazard is in the environment; and determining, using the one or more machine learning models and based at least on the sensor data, the initial information and a second prompt of the one or more prompts, the information associated with the hazard. Q: System according to one of paragraphs JP, wherein the one or more processors further serve to: Determine, based on at least second sensor data obtained using one or more second sensors of the machine, second information associated with the hazard, wherein the data further associates the second information. R: System according to any of paragraphs JQ, wherein the system comprises at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system that provides one or more cloud gaming applications; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using one or more large language models (LLMs);A system for performing operations using one or more vision language models; a system for performing operations using one or more multimodal language models; a system for performing one or more operations with conversational AI; a system for generating synthetic data; a system for presenting at least one type of virtual reality, augmented reality, or mixed reality content; systems that implement one or more multimodal language models; systems that use or employ one or more inference microservices; systems that integrate or employ one or more machine learning models in a service or microservice along with an operating system (OS)-level virtualization package (e.g., a container); a system that includes one or more virtual machines (VMs); a system that is at least partially deployed in a data center;or a system that is implemented at least partially using cloud computing resources. S: Autonomous or semi-autonomous machine comprising: one or more central processing units (CPUs); one or more graphics processing units (GPUs); one or more hardware accelerators and one or more external sensors with fields of view or sensory fields located outside the autonomous or semi-autonomous machine, wherein the autonomous or semi-autonomous machine sends data representing information associated with a hazard to one or more systems, wherein the information is determined based on at least one or more language models processing one or more prompts associated with the hazard and sensor data obtained using the one or more sensors. T: Machine according to paragraph S, wherein the machine includes or comprises at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system that provides one or more cloud gaming applications; a system for performing one or more deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing one or more generative AI operations; a system for performing operations using one or more large language models (LLMs);A system for performing operations using one or more vision language models; a system for performing operations using one or more multimodal language models; a system for performing one or more operations with conversational AI; a system for generating synthetic data; a system for presenting at least one type of virtual reality, augmented reality, or mixed reality content; systems that implement one or more multimodal language models; systems that use or employ one or more inference microservices; systems that integrate or employ one or more machine learning models in a service or microservice along with an operating system (OS)-level virtualization package (e.g., a container); a system that includes one or more virtual machines (VMs); a system that is at least partially deployed in a data center;or a system that is implemented at least partially using cloud computing resources.
[0252] It is understood that the aspects and embodiments described above are purely exemplary and that modifications of details may be made within the scope of protection of the claims.
[0253] Each device, each method and each feature disclosed in the description, and (where applicable) the claims and drawings, may be provided independently or in any suitable combination.
[0254] Reference numerals appearing in the claims are for illustrative purposes only and do not restrict the scope of protection of the claims. QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] US 16 / 101,232
[0154] Cited non-patent literature
[0000] Standard No. J3016-201806, published on June 15, 2018
[0112] Standard No. J3016-201609, published on 30 September 2016
[0112]
Claims
[1] Procedure, encompassing: Determine, based on at least sensor data obtained using one or more of a machine's primary sensors, a location associated with the machine in an environment; Determine, based on at least one or more vision language models that process at least image data obtained using one or more secondary sensors of the machine, information associated with a road hazard located in the environment; and Sending data to one or more systems to update one or more applications to indicate at least some of the information related to the road hazard at approximately the location. [2] The method of claim 1, further comprising: Determine, based on at least one or more language models that process second data representing the information, a requirement to verify the information; Providing content that is related to the requirement to verify the information; and Receiving input data indicating that the information associated with the road hazard has been verified, where the sending of data to update one or more applications is based at least on the road hazard that is verified. [3] Method according to claim 1 or 2, further comprising: Determine, based on at least one or more language models that process second data representing the information, a requirement for second information related to the road hazard; Providing content that corresponds to the request for the second piece of information; and Receiving input data that specifies the second set of information associated with the road hazard, where the data also serve to update one or more applications in order to provide at least part of the second information. [4] A method according to any of the preceding claims, wherein the determination of the information associated with the road hazard is further based on at least one or more vision language models that process second data which represent at least one of the following: an initial prompt requesting information related to road hazards that may be present in the vicinity; or a second prompt to request second pieces of information related to one or more road hazards potentially located in the vicinity, wherein the one or more road hazards include at least the road hazard. [5] Method according to any one of the preceding claims, wherein: Determining the information associated with the road hazard is further based on at least one or more vision language models that process second data representing a first prompt associated with the road hazard; and The method further includes determining, based on at least one or more vision language models that process at least the image data and third data that represent a second prompt associated with a second road hazard, second information associated with the second road hazard that may be in the vicinity. [6] A method according to any of the preceding claims, wherein determining the information related to the road hazard comprises: Determine, based on at least one or more vision language models that process the image data and second data that provide a first prompt of initial information indicating that the road hazard is in the vicinity; and Determine, based on at least one or more vision language models that process the image data and third data representing the initial information and a second prompt, the information associated with the road hazard. [7] Method according to any one of the preceding claims, further comprising: Determine, based on at least second sensor data obtained using one or more third sensors of the machine, second information related to road hazards, where the data also serve to update one or more applications in order to provide at least part of the second set of information related to the road hazard. [8] A method according to any of the preceding claims, wherein the information associated with the road hazard describes at least one of the following: a type of danger; a location of the danger in the vicinity; or a location of the hazard as represented by one or more images depicted by the image data. [9] Method according to any of the preceding claims, wherein the road hazard comprises at least one of the following: an obstacle located near a driving surface; an emergency service located near the driving surface; a construction activity that is associated with the roadway; an animal that is located near the driving surface; a pedestrian who is near the roadway; or a dangerous road condition that is associated with the driving surface. [10] System, encompassing: one or more processors, for example: Receiving sensor data obtained using one or more sensors of a machine navigating an environment; Determine, using one or more machine learning models and based on at least sensor data and one or more prompts associated with one or more hazards, information corresponding to a hazard located in the environment; and Sending, to one or more systems, data that is associated with the information corresponding to the hazard, wherein the data is used to update an application to specify the information corresponding to the hazard. [11] System according to claim 10, wherein the one or more processors further serve to: Determine, based on at least second sensor data obtained using one or more second sensors, a location associated with the hazard, where the data are further assigned to the information. [12] System according to any one of claims 10 to 11, wherein the one or more processors further serve to: Determine, using one or more machine learning models and based on at least the information, a requirement to verify the information; Providing content that is related to the request to verify the information; and Receiving input data indicating that the information associated with the hazard has been verified, the data is sent based at least on the risk that is verified. [13] System according to any one of claims 10 to 12, wherein the one or more processors further serve to: Determine, using one or more machine learning models and based at least on the information, a requirement for a second piece of information that is associated with the hazard; Providing content that corresponds to the request for the second piece of information, and Receiving input data that specifies the second set of information associated with the hazard, where the data are further assigned to the second set of information. [14] System according to any one of claims 10 to 13, wherein the one or more processors further serve to: Determine, using one or more machine learning models and based on at least the information, to provide the application with the information to update, where the data is sent based at least on determining the provision of the information. [15] System according to any one of claims 10 to 14, wherein the one or more processors further serve to: Determine, using one or more machine learning models and based at least on the sensor data and a prompt associated with one or more second hazards, second information corresponding to a second hazard that is potentially in the environment. [16] System according to any one of claims 10 to 15, wherein the determination of the information associated with the hazard comprises: Determine, using one or more machine learning models and based at least on the sensor data and an initial prompt from one or more prompts, initial information indicating that the hazard is in the vicinity; and Determine, using one or more machine learning models and based at least on the sensor data, the initial information and a second prompt of one or more prompts, the information associated with the hazard. [17] System according to any one of claims 10 to 16, wherein the one or more processors further serve to: Determine, based on at least second sensor data obtained using one or more second sensors of the machine, second information that is associated with the hazard, where the data are further assigned to the second set of information. [18] System according to any one of claims 10 to 17, wherein the system comprises at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulations; a system for conducting collaborative content creation for 3D assets; a system that provides one or more cloud gaming applications; a system for performing one or more deep learning operations; a system that is implemented using an edge device; a system that is implemented using a robot; a system for performing one or more operations using generative AI; a system for performing operations using one or more large language models (LLMs); a system for performing operations using one or more vision language models; a system for performing operations using one or more multimodal language models; a system for performing one or more operations using conversational AI; a system for generating synthetic data; a system for presenting at least one type of virtual reality content, augmented reality content, or mixed reality content; Systems that implement one or more multimodal language models; Systems that use or employ one or more inference microservices; Systems that integrate or use one or more machine learning models in a service or microservice together with an operating system (OS) level virtualization package (e.g., a container); a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources. [19] Autonomous or semi-autonomous machine, comprising: one or more central processing units (CPUs); one or more graphics processing units (GPUs); one or more hardware accelerators; and one or more external sensors with fields of view or sensory fields outside the autonomous or semi-autonomous machine, wherein the autonomous or semi-autonomous machine sends data representing information associated with a hazard to one or more systems, wherein the information is determined based on at least one or more language models processing one or more prompts associated with the hazard and sensor data obtained using the one or more sensors. [20] Machine according to claim 19, wherein the machine contains or comprises at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing one or more simulation operations; a system for performing one or more digital twin operations; a system for performing light transport simulations; a system for conducting collaborative content creation for 3D assets; a system that provides one or more cloud gaming applications; a system for performing one or more deep learning operations; a system that is implemented using an edge device; a system that is implemented using a robot; a system for performing one or more operations using generative AI; a system for performing operations using one or more large language models (LLMs); a system for performing operations using one or more vision language models; a system for performing operations using one or more multimodal language models; a system for performing one or more operations using conversational AI; a system for generating synthetic data; a system for presenting at least one type of virtual reality content, augmented reality content, or mixed reality content; Systems that implement one or more multimodal language models; Systems that use or employ one or more inference microservices; Systems that integrate or use one or more machine learning models in a service or microservice together with an operating system (OS) level virtualization package (e.g., a container); a system that includes one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is implemented at least partially using cloud computing resources.