Techniques for adaptive driving using language models

By receiving scene descriptions and traffic rule texts, and using a language model to generate adaptive driving instructions, the problem of traditional autonomous vehicles complying with traffic rules in different geographical locations has been solved, achieving safe and efficient autonomous driving.

CN121399012APending Publication Date: 2026-01-23NVIDIA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480039295.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-10-31
Filing Date
2024-11-12
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Traditional autonomous vehicles struggle to safely comply with local regulations when traffic rules change in different geographical locations, leading to potential safety issues and reduced transport efficiency.

Method used

By receiving scenario descriptions, driving plans, and traffic rule texts, the system uses a language model to generate adaptive driving instructions, ensuring that the vehicle complies with local traffic rules.

Benefits of technology

This enables autonomous vehicles to safely comply with traffic rules in different geographical locations, improving driving safety and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121399012A_ABST
    Figure CN121399012A_ABST
Patent Text Reader

Abstract

One embodiment of a method for controlling a vehicle includes receiving a first text containing a description of a scene and a first plan for driving the vehicle; based on the description of the scene and the first plan, extracting at least one part of a traffic rule set; generating a first cue requesting a driving instruction and including a description of the scene, a first plan, and at least a portion of a set of traffic rules; processing the first cue via a first trained language model to generate a second plan for driving the vehicle; and generating a driving instruction based on the second plan.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 548,344, filed November 13, 2023, entitled “Adaptive Policy Selection for Autonomous Machine Operation Using Generative Pre-trained Transformer Models,” and U.S. Patent Application No. 18 / 933,354, filed October 31, 2024, entitled “Techniques for Adaptive Driving Using Language Models.” The subject matter of these related applications is incorporated herein by reference. background Technical Field

[0003] Embodiments of this disclosure generally relate to the fields of computer science, machine learning and artificial intelligence (AI) and autonomous vehicles, and more specifically, to techniques for adaptive driving using language models. Background Technology

[0004] Machine learning can be used to discover trends, patterns, relationships, and / or other attributes associated with large, complex, interconnected, and / or multidimensional datasets. To gather insights from large datasets, input-output pairs from the data can be used to train artificial neural networks, regression models, support vector machines, decision trees, Naive Bayes classifiers, and / or other types of machine learning models. In turn, the trained machine learning models can be used to guide decisions and / or perform tasks related to the data and / or other similar data.

[0005] One type of task that machine learning models can be trained to perform is controlling vehicles, such as autonomous and semi-autonomous vehicles. Typically, when using a machine learning model to control a vehicle, the model is trained to obey a specific set of traffic rules. However, traffic rules can vary significantly in different geographical locations where the vehicle might ultimately be traveling. For example, some locations may require vehicles to travel on the left side of the road, while others may require them to travel on the right. As another example, in some locations, turning right at a red light is permitted, but in others it is not.

[0006] When a vehicle is controlled by a machine learning model trained to obey a specific set of traffic rules, the trained model may end up controlling the vehicle incorrectly in a way that violates local traffic rules when local traffic rules differ from that specific set of rules. Such erroneous vehicle control can lead to dangerous driving scenarios that cause safety problems. Therefore, traditional autonomous vehicles are typically restricted to operating within geofenced areas where traffic rules remain unchanged. Restricting the operation of traditional autonomous vehicles to geofenced areas reduces their utility when transporting passengers to and from locations outside the geofenced area.

[0007] As mentioned earlier, what is needed in the field is a more efficient technology for using machine learning models to control vehicles. Summary of the Invention

[0008] One embodiment of this disclosure proposes a computer-implemented method for controlling a vehicle. The method includes: receiving first text, the first text comprising a description of a scenario and a first plan for driving the vehicle; and extracting at least a portion of a traffic rule set based on the scenario description and the first plan. The method further includes generating a first prompt, the first prompt requesting driving instructions and comprising the scenario description, the first plan, and at least a portion of the traffic rule set. The method further includes processing the first prompt via a first trained language model to generate a second plan for driving the vehicle. Furthermore, the method includes generating driving instructions based on the second plan.

[0009] Other embodiments of this disclosure include, but are not limited to, one or more computer-readable media containing instructions for performing one or more aspects of the disclosed technology, and one or more computing systems for performing one or more aspects of the disclosed technology.

[0010] Compared to existing technologies, at least one technical advantage of the disclosed technology is that it enables vehicles controlled using trained machine learning models to adapt to traffic rules in different geographical locations. Therefore, when the disclosed technology is implemented to control a vehicle using a trained machine learning model, the vehicle can drive in a manner that complies with local traffic rules and is safer than what is typically achievable using traditional machine learning models. Alternatively, when the disclosed technology is implemented to respond to user input, it enables users to drive more safely and comply with local traffic rules. These technical advantages represent one or more technical improvements superior to existing methods. Attached Figure Description

[0011] To gain a more detailed understanding of the features of the various embodiments described above, the inventive concept briefly outlined above can be described in more specific terms with reference to the various embodiments, some of which are illustrated in the accompanying drawings. However, it should be noted that the drawings only illustrate typical embodiments of the inventive concept and should not be considered as any limitation on the scope of the invention; other equally effective embodiments exist.

[0012] Figure 1 A block diagram of a computing system configured to implement one or more aspects of various embodiments is shown;

[0013] Figure 2A Illustrations of exemplary autonomous vehicles according to various embodiments are shown;

[0014] Figure 2B Various embodiments are shown. Figure 2A Exemplary camera positions and fields of view for an exemplary autonomous vehicle;

[0015] Figure 2C According to various embodiments Figure 2A A block diagram of an exemplary system architecture for an exemplary autonomous vehicle;

[0016] Figure 2D It is a cloud-based server according to various embodiments and Figure 2A The system diagram in the example shows communication between autonomous vehicles;

[0017] Figure 3 According to various embodiments Figure 1 A more detailed illustration of the language model driving assistance system in the image;

[0018] Figure 4 It is included according to various embodiments Figure 1 Language modeling in autonomous vehicle (AV) driver assistance systems;

[0019] Figure 5 It is included according to various embodiments Figure 1 The application of language models in navigation systems for driver assistance systems;

[0020] Figure 6 According to various embodiments Figure 5 How can a navigation application in a navigation application generate an example plan?

[0021] Figure 7 This is a flowchart of method steps for providing driving assistance to autonomous vehicles or users according to various embodiments; and

[0022] Figure 8 This is a flowchart of method steps for extracting portions of local traffic rules according to various embodiments. Detailed Implementation

[0023] Numerous specific details are set forth in the following description to provide a more comprehensive understanding of the various embodiments. However, those skilled in the art will understand that the inventive concept can be practiced even without one or more of these specific details.

[0024] General Overview

[0025] Embodiments of this disclosure provide techniques for generating driving instructions for autonomous vehicles or users. In some embodiments, a language model-based driving assistance system receives text as input describing a scenario, a plan for driving the vehicle, the current situation, and traffic rules describing the geographical location of the vehicle. Given this input, the language model-based driving assistance system extracts one or more portions of the traffic rules relating to the scenario description, plan, and / or situation. For example, in some embodiments, the language model-based driving assistance system may prompt a trained language model to extract one or more keywords from the scenario description, plan, and / or current situation, which contain common traffic-related phrases. In this case, the language model-based driving assistance system may also search for one or more keywords in the traffic rules and then extract the portions (e.g., paragraphs) of the traffic rules containing those keywords. The language model-based driving assistance system further prompts the trained language model to generate an updated plan for driving the vehicle, which takes into account that portion of the traffic rules. The updated plan may be output to the user as a driving instruction via a display device and / or a speaker. Alternatively, the updated plan may be used to update an automatically generated motion plan, which can then be applied to control the vehicle.

[0026] Technologies used to generate driving commands for autonomous vehicles or users have numerous real-world applications. For example, these technologies can be used to control autonomous or semi-autonomous vehicles in real-world or virtual environments. As another example, these technologies can be used to output driving commands to users via display and / or audio devices.

[0027] The examples above are not intended to be limiting. As those skilled in the art will understand, the techniques described herein for generating driving instructions can generally be applied to any suitable application scenario.

[0028] System Overview

[0029] Figure 1A block diagram of a computing system 100 is shown, configured to implement one or more aspects of various embodiments. In some embodiments, the computing system 100 may include any type of computing system, including but not limited to server machines, server platforms, desktop computers, laptop computers, handheld / mobile devices, digital kiosks, in-vehicle infotainment systems, and / or wearable devices. In some embodiments, the computing system 100 is a server machine operating in a data center or cloud computing environment that provides scalable computing resources as a service over a network.

[0030] In some embodiments, the computing system 100 includes, but is not limited to, a processor 112 and a memory 114, which are coupled to the parallel processing subsystem 112 via a memory bridge 105 and a communication path 106. The language model driving assistance system 130 (hereinafter referred to as...) Figures 3 to 8 (To be described in more detail) is stored in memory 114 and executed on processor 112 of computing system 100. Although this document is primarily described with respect to language model driving assistance system 130, the techniques disclosed herein can also be implemented, in whole or in part, in other software and / or hardware, such as in parallel processing subsystem 112. Memory bridge 105 is also coupled to I / O (input / output) bridge 107 via communication path 106, which in turn is coupled to switch 116.

[0031] In some embodiments, I / O bridge 107 is configured to receive user input from optional input device 108, such as a keyboard, mouse, touchscreen, sensor data analysis (e.g., evaluating gestures, voice, or other information about the field of view or perception of one or more sensors for one or more purposes), and forward the input to processor 112 for processing. In some embodiments, computing system 100 may be a server machine in a cloud computing environment. In these embodiments, computing system 100 may not include input device 108, but may receive equivalent input information in the form of messages transmitted over a network and received via network adapter 118 by receiving commands (e.g., in response to one or more inputs from a remote computing device). In some embodiments, switch 116 is configured to provide connectivity between I / O bridge 107 and other components of computing system 100 (e.g., network adapter 118 and various add-on cards 120 and 121).

[0032] In some embodiments, I / O bridge 107 is coupled to system disk 114, which may be configured to store content, applications, and data for use by processor 112 and parallel processing subsystem 112. In some embodiments, system disk 114 provides non-volatile storage for applications and data and may include fixed or removable hard disk drives, flash memory devices, and CD-ROMs (optical disc read-only memory), DVD-ROMs (digital versatile optical discs), Blu-ray discs, HD-DVDs (high-definition DVDs), or other magnetic, optical, or solid-state storage devices. In some embodiments, other components such as universal serial bus or other port connections, compact disc drives, digital versatile optical disc drives, film recording devices, etc., may also be connected to I / O bridge 107.

[0033] In some embodiments, memory bridge 105 may be a northbridge chip, and I / O bridge 107 may be a southbridge chip. Furthermore, communication paths 106 and 113, as well as other communication paths within computing system 100, may be implemented using any technically suitable protocol, including but not limited to AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol known in the art.

[0034] In some embodiments, the parallel processing subsystem 112 includes a graphics subsystem that passes pixels to an optional display device 110, which may be any conventional cathode ray tube, liquid crystal display, light-emitting diode display, and / or the like. In these embodiments, the parallel processing subsystem 112 may include circuitry optimized for graphics and video processing, including, for example, video output circuitry. Such circuitry may be incorporated into one or more parallel processing units (PPUs) (also referred to herein as parallel processors) included within the parallel processing subsystem 112.

[0035] In some embodiments, the parallel processing subsystem 112 includes circuitry optimized for general and / or computational processing (e.g., optimized circuitry). Similarly, such circuitry may be incorporated into one or more PPUs included within the parallel processing subsystem 112, which are configured to perform such general and / or computational operations. In other embodiments, one or more PPUs included within the parallel processing subsystem 112 may be configured to perform graphics processing, general processing, and / or computational processing operations. The system memory 114 includes at least one device driver configured to manage the processing operations of one or more PPUs within the parallel processing subsystem 112.

[0036] In some embodiments, the parallel processing subsystem 112 can be connected with Figure 1One or more other components can be integrated to form a single system. For example, the parallel processing subsystem 112 can be integrated with the processor 112 and other interconnect circuitry on a single chip to form a system-on-a-chip (SoC).

[0037] In some embodiments, processor 112 includes the main processor of computing system 100 for controlling and coordinating the operation of other system components. In some embodiments, processor 112 issues commands to control the operation of the PPU. In some embodiments, communication path 113 is a PCI Express link, in which a dedicated channel is allocated to each PPU. Other communication paths may also be used. Advantageously, the PPU implements a highly parallel processing architecture, and the PPU can be equipped with any number of local parallel processing memories (PP memories).

[0038] It should be understood that the system shown herein is illustrative only, and various changes and modifications are possible. The connection topology (including the number and arrangement of bridges, the number of processors 112, and the number of parallel processing subsystems 112) can be modified as needed. For example, in some embodiments, system memory 114 may be directly connected to processor 112 instead of via memory bridge 105, and other devices may communicate with system memory 114 via memory bridge 105 and processor 112. In other embodiments, parallel processing subsystem 112 may be connected to I / O bridge 107 or directly to processor 112 instead of via memory bridge 105. In still other embodiments, I / O bridge 107 and memory bridge 105 may be integrated into a single chip rather than existing as one or more discrete devices. In some embodiments, Figure 1 One or more of the components shown may be absent. For example, switch 116 may be omitted, and network adapter 118 and add-on cards 120, 121 may be directly connected to I / O bridge 107. Finally, in some embodiments, Figure 1 One or more components shown can be implemented as virtualized resources in a virtual computing environment (e.g., a cloud computing environment). Specifically, in some embodiments, the parallel processing subsystem 112 can be implemented as a virtualized parallel processing subsystem. For example, the parallel processing subsystem 112 can be implemented as a virtual graphics processing unit (vGPU) that renders graphics on a virtual machine (VM) running on a server machine, whose GPU and other physical resources are shared among one or more VMs.

[0039] In some embodiments, the computing system 100 described herein can be used with Figure 2A-2DThe components, features, and / or functions of the example autonomous vehicle 200 described herein are similar to those components, features, and / or functions used in its execution. The systems and methods described herein can be used, but are not limited to, non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more adaptive driver assistance systems (ADAS)), driver- or driverless robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, spacecraft, ships, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, engineering vehicles, submarines, drones, and / or other vehicle types. Furthermore, the systems and methods described herein can be used for a variety of purposes, as examples and not limitations, for machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, safety and supervision, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twins, data center processing, conversational AI, optical transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or any other suitable applications.

[0040] The disclosed embodiments may be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aerial systems, medical systems, rowing systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems including one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, systems for performing optical transmission simulations, systems for performing collaborative content creation of 3D assets, systems implemented at least partially using cloud computing resources, and / or other types of systems.

[0041] Example autonomous vehicles

[0042] Figure 2AThis is an illustration of an exemplary autonomous vehicle 200 according to various embodiments. The autonomous vehicle 200 (which may, alternatively, be referred to herein as "vehicle 200") may include, but is not limited to, passenger vehicles such as cars, trucks, buses, ambulances, shuttles, electric or motorized bicycles, motorcycles, fire trucks, police cars, ambulances, boats, engineering vehicles, submarines, robotic vehicles, drones, aircraft, vehicles coupled to trailers (e.g., semi-trailer trucks for transporting goods) and / or other types of vehicles (e.g., driverless and / or capable of accommodating one or more passengers). Autonomous vehicles are typically described according to the levels of automation defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) in its "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201506, published June 15, 2018; Standard No. J3016-201609, published September 30, 2016; and previous and future versions of this standard). Vehicle 200 is capable of performing one or more functions that meet Level 3-5 of the autonomous driving level. Vehicle 200 is capable of performing one or more functions that meet Level 1-5 of the automated driving level. For example, depending on the embodiment, vehicle 200 is capable of performing driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). The term “autonomy” as used herein may include any and / or all types of autonomy for 200 or other machines, such as full autonomy, high autonomy, conditional autonomy, partial autonomy, providing auxiliary autonomy, semi-autonomy, primary autonomy, or other names.

[0043] Vehicle 200 may include components such as chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. Vehicle 200 may include a propulsion system 250, such as an internal combustion engine, a hybrid power plant, an all-electric motor, and / or another type of propulsion system. Propulsion system 250 may be connected to the drivetrain of vehicle 200, which may include a transmission, to enable propulsion of vehicle 200. Propulsion system 250 may be controlled in response to receiving a signal from throttle / accelerator 252.

[0044] A steering system 254, which may include a steering wheel, can be used to steer the vehicle 200 (e.g., along a desired path or route) when the propulsion system 250 is operating (e.g., when the vehicle is in motion). The steering system 254 may receive signals from the steering actuator 256. For fully automatic (level 5) functionality, the steering wheel may be optional.

[0045] The brake sensor system 246 can be used to operate the vehicle brakes in response to receiving signals from the brake actuator 148 and / or the brake sensor.

[0046] It may include one or more CPUs, System-on-a-Chip (SoC) 204 ( Figure 2C One or more controllers 236, including and / or one or more GPUs, may provide signals (e.g., signals representing commands) to one or more components and / or systems of vehicle 200. For example, one or more controllers may send signals to operate vehicle brakes via one or more brake actuators 248, to operate steering system 254 via one or more steering actuators 256, and / or to operate propulsion system 250 via one or more throttles / accelerators 252. One or more controllers 236 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operating commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving vehicle 200. One or more controllers 236 may include a first controller 236 for autonomous driving functions, a second controller 236 for functional safety functions, a third controller 236 for artificial intelligence functions (e.g., computer vision), a fourth controller 236 for infotainment functions, a fifth controller 236 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 236 can handle two or more of the functions described above, and two or more controllers 236 can handle a single function, and / or any combination thereof.

[0047] One or more controllers 236 may provide signals for controlling one or more components and / or systems of vehicle 200 in response to sensor data (e.g., sensor inputs) received from one or more sensors. Sensor data may be received from, for example, but not limited to, Global Navigation Satellite System (“GNSS”) sensors 258 (e.g., Global Positioning System sensors), RADAR sensors 260, ultrasonic sensors 262, LIDAR sensors 264, inertial measurement unit (IMU) sensors 266 (e.g., accelerometers, gyroscopes, magnetic compasses, magnetometers, etc.), microphones 296, stereo cameras 268, wide-angle cameras 270 (e.g., fisheye cameras), infrared cameras 272, surround cameras 274 (e.g., 360-degree cameras), long-range and / or medium-range cameras 298, speed sensors 244 (e.g., for measuring the rate of vehicle 100), vibration sensors 242, steering sensors 240, braking sensors (e.g., as part of braking sensor system 246), and / or other sensor types.

[0048] One or more of the controllers 236 may receive inputs (e.g., represented by input data) from the instrument cluster 232 of the vehicle 200 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 234, an auditory signaling device, a speaker, and / or via other components of the vehicle 200. These outputs may include information such as vehicle speed, rate, time, map data (e.g., [missing information]). Figure 2C Information such as high-definition (“HD”) maps 222, location data (e.g., the location of vehicle 200 on the map), orientation, and the location of other vehicles (e.g., occupying grids), as well as information about objects and their states perceived by controller 236, etc. For example, HMI display 234 may display information about the existence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, leaving 34B in two miles, etc.).

[0049] Vehicle 200 further includes a network interface 224, which can communicate via one or more networks using one or more wireless antennas 226 and / or a modem. For example, network interface 224 may be able to communicate via Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multicarrier (“CDMA2000”), etc. One or more wireless antennas 226 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using one or more local area networks such as Bluetooth, Bluetooth Low Energy (“LE”), Z-Wave, ZigBee, etc., and / or one or more low-power wide area networks (“LPWAN”) such as LoRaWAN, SigFox, etc.

[0050] Figure 2B The illustration shows the use of various embodiments for Figure 2A The exemplary camera positions and fields of view of the exemplary autonomous vehicle 200 are shown below. The cameras and their respective fields of view are an example embodiment and are not intended to be limiting. For example, additional and / or replaceable cameras may be included, and / or these cameras may be located at different positions on the vehicle 200.

[0051] The camera type used for the camera may include, but is not limited to, a digital camera suitable for use with components and / or systems of vehicle 200. The camera may operate at Automotive Safety Integrity Level (ASIL) B and / or another ASIL. The camera type may have any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The camera may be able to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a red-white-white-white (RCCC) color filter array, a red-white-white-blue (RCCB) color filter array, a red-blue-green-white (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, a sharp-pixel camera, such as a camera with RCCC, RCCB, and / or RBGC color filter arrays, may be used in efforts to improve light sensitivity.

[0052] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-function monocular camera can be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).

[0053] One or more of the cameras can be mounted in mounting components such as custom-designed (3D-printed) components to cut off stray light and reflections from inside the vehicle (e.g., reflections from the dashboard reflected in the windshield mirror) that may interfere with the camera's image data capture capabilities. Regarding the wing mirror mounting components, the wing mirror components can be custom-3D printed so that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras can be integrated into the wing mirror. For side-view cameras, one or more cameras can also be integrated into the four pillars at each corner of the cab.

[0054] A camera with a field of view that includes the environment in front of the vehicle 200 (e.g., a front-facing camera) can be used for surround view to help identify forward paths and obstacles, and, with the assistance of one or more controllers 236 and / or control SoCs, to provide information crucial for generating an occupancy grid and / or determining a preferred vehicle path. The front-facing camera can be used to perform many of the same ADAS functions as LiDAR, including emergency braking, pedestrian detection, and collision avoidance. The front-facing camera can also be used in ADAS functions and systems, including Lane Departure Warning (“LDW”), Autonomous Cruise Control (ACC), and / or other functions such as traffic sign recognition.

[0055] A variety of cameras can be used in front-facing configurations, including, for example, monocular camera platforms including complementary metal-oxide-semiconductor (“CMOS”) color imagers. Another example could be a wide-angle camera 270, which can be used to perceive objects entering the field of view from the periphery (such as pedestrians, traffic at intersections, or bicycles). Although Figure 2B The image shows only one wide-angle camera, but any number (including zero) of wide-angle cameras 270 can be present on vehicle 200. Furthermore, any number of remote cameras 298 (e.g., a pair of remote-view stereo cameras) can be used for depth-based object detection, especially for objects for which neural networks have not yet been trained. Remote cameras 298 can also be used for object detection and classification, as well as basic object tracking.

[0056] Any number of stereo cameras 268 may also be included in the front-mounted configuration. In some embodiments, one or more of the stereo cameras 268 may include an integrated control unit comprising a scalable processing unit that can provide a multi-core microprocessor and programmable logic (“FPGA”) with an integrated controller area network (“CAN”) or Ethernet interface on a single chip. Such a unit can be used to generate a 3D map of the vehicle environment, including distance estimates for all points in the image. Alternative stereo cameras 268 may include a compact stereo vision sensor that may include two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from the vehicle to a target object and activate autonomous emergency braking and lane departure warning functions using the generated information (e.g., metadata). Other types of stereo cameras 268 may be used as a supplement to or replacement for those stereo cameras described herein.

[0057] Cameras with a field of view including the side portion of the environment of vehicle 200 (e.g., side-view cameras) can be used for surround view, providing information for creating and updating occupancy grids and generating side-impact collision warnings. For example, surround camera 274 (e.g., ... Figure 2B The four surround cameras 274 shown can be mounted on the vehicle 200. The surround cameras 274 can include wide-angle cameras 270, fisheye cameras, 360-degree cameras, and / or the like. For example, four fisheye cameras can be positioned at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle can use three surround cameras 274 (e.g., left, right, and rear) and can utilize one or more other cameras (e.g., forward-facing cameras) as a fourth surround-view camera.

[0058] A camera with a field of view that includes the environment behind the vehicle 200 (e.g., a rear-view camera) can be used for parking assistance, surround view, rear collision warning, and creating and updating occupancy grids. A wide variety of cameras can be used, including but not limited to those also suitable as front-facing cameras as described herein (e.g., long-range and / or mid-range camera 298, stereo camera 268, infrared camera 272, etc.).

[0059] Figure 2C According to various embodiments Figure 2AThe example system architecture block diagram of the example autonomous vehicle 200 is shown below. It should be understood that this and other arrangements described herein are merely illustrative examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groups, etc.) may be used as supplements to or replacements of these arrangements and elements, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities can be implemented via hardware, firmware, and / or software. For example, the various functions can be implemented by a processor executing instructions stored in memory.

[0060] Figure 2C Each component, feature, and system in vehicle 200 is illustrated as being connected via bus 202. Bus 202 may include a Controller Area Network (CAN) data interface (or, alternatively, referred to herein as the "CAN bus"). CAN may be a network within vehicle 200 used to assist in the control of various features and functions of vehicle 200, such as the actuation of brakes, acceleration, braking, steering, windshield wipers, etc. The CAN bus can be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus can be read to find steering wheel angle, ground speed, engine speed per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.

[0061] Although bus 202 is described herein as a CAN bus, this is not intended to be limiting. For example, FlexRay and / or Ethernet may be used in addition to or alternatively to a CAN bus. Furthermore, although bus 202 is represented by a single line, this is not intended to be limiting. For example, any number of buses 202 may exist, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 202 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 202 may be used for a collision avoidance function, and a second bus 202 may be used for drive control. In any example, each bus 202 may communicate with any component of vehicle 200, and two or more buses 202 may communicate with the same component. In some examples, each SoC 204, each controller 236, and / or each computer within the vehicle may have access to the same input data (e.g., input from sensors of vehicle 200) and may be connected to a common bus such as a CAN bus.

[0062] Vehicle 200 may include one or more controllers 236, such as those described herein. Figure 2A The controllers described herein. Controller 236 can be used for a wide variety of functions. Controller 236 can be coupled to any other different components and systems of vehicle 200 and can be used for the control of vehicle 200, artificial intelligence of vehicle 200, infotainment and / or the like for vehicle 200.

[0063] Vehicle 200 may include one or more system-on-a-chip (SoC) 204. SoC 204 may include CPU 206, GPU 208, processor 210, cache 212, accelerator 214, data storage 216, and / or other components and functions not shown. In some embodiments, the components included in vehicle 200 (e.g., CPU 210 and data storage 216) may be combined with the above description. Figure 1 The corresponding components (e.g., processor 142 and memory 144) included in the described computing system 100 are the same or similar. In a wide variety of platforms and systems, SoC 204 can be used to control vehicle 200. For example, (one or more) SoC 204 can be combined with a high-definition (HD) map 222 in a system (e.g., the system of vehicle 200), the HD map 222 being accessible via network interface 224 from one or more servers (e.g., [server name missing]). Figure 2D The server in question (278) retrieves map refresh and / or updates.

[0064] CPU 206 may include a CPU cluster or CPU complex (or, alternatively, referred to herein as "CCPLEX"). CPU 206 may include multiple cores and / or L2 cache. For example, in some implementations, CPU 206 may include eight cores in a coherent multiprocessor configuration. In some implementations, CPU 206 may include four dual-core clusters, each with a dedicated L2 cache (e.g., 2MB L2 cache). CPU 206 (e.g., CCPLEX) may be configured to support simultaneous cluster operation, such that any combination of clusters of CPU 206 can be active at any given time.

[0065] CPU 206 can implement power management capabilities including one or more of the following features: automatic clock gating of hardware blocks when idle to conserve dynamic power; clock gating of each core when the core is not actively executing instructions due to the execution of WFI / WFE instructions; independent power gating of each core; independent clock gating of each core cluster when all cores are clock-gated or power-gated; and / or independent power gating of each core cluster when all cores are power-gated. CPU 206 can further implement enhanced algorithms for managing power states, wherein allowed power states and desired wake-up times are specified, and the hardware / microcode determines the optimal power state to enter for the core, cluster, and CCPLEX. The processing core can support simplified power state entry sequences in software, with this work offloaded to the microcode.

[0066] GPU 208 may include an integrated GPU (or, alternatively, referred to herein as an "iGPU"). GPU 208 may be programmable and efficient for parallel workloads. In some examples, GPU 208 may use an enhanced tensor instruction set. GPU 208 may include one or more streaming microprocessors, wherein each streaming microprocessor may include an L1 cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In some embodiments, GPU 208 may include at least eight streaming microprocessors. GPU 208 may use a computation application programming interface (API). Furthermore, GPU 208 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0067] In automotive and embedded applications, the GPU 208 can be power-optimized for optimal performance. For example, the GPU 208 can be fabricated on FinFETs. However, this is not intended to be limiting, and the GPU 208 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can combine several mixed-precision processing cores divided into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor cores for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, dispatch units, and / or a 64KB register file. Furthermore, the streaming microprocessor can include independent parallel integer and floating-point data paths to provide efficient execution of workloads by leveraging the mixture of computation and addressing computation. The streaming microprocessor can include independent thread scheduling capabilities to allow for finer-grained synchronization and cooperation between parallel threads. Streaming microprocessors can include a combination of L1 data cache and shared memory units to improve performance while simplifying programming.

[0068] The GPU 208 may include, in some examples, a High Bandwidth Memory (HBM) and / or a 16GB HBM2 memory subsystem providing a peak memory bandwidth of approximately 900GB / s. In some examples, in addition to HBM memory or alternatively, Synchronous Graphics Random Access Memory (SGRAM), such as Generation 5 Graphics Double Data Rate Synchronous Random Access Memory (GDDR5), may be used.

[0069] GPU 208 may include unified memory technology, which includes access counters to allow memory pages to be migrated more precisely to the processors that access them most frequently, thereby improving the efficiency of shared memory ranges between processors. In some examples, Address Translation Service (ATS) support can be used to allow GPU 208 to directly access CPU 206 page tables. In such examples, when GPU 208 Memory Management Unit (MMU) experiences a miss, the address translation request can be transferred to CPU 206. In response, CPU 206 can look up the virtual-physical mapping for the address in its page tables and transfer the translation back to GPU 208. In this way, unified memory technology can allow a single unified virtual address space for the memory of both CPU 206 and GPU 208, thereby simplifying GPU 208 programming and porting applications to GPU 208.

[0070] In addition, GPU 208 may include access counters that track how frequently GPU 208 accesses the memory of other processors. Access counters can help ensure that memory pages are moved to the physical memory of the processor that accesses those pages most frequently.

[0071] SoC 204 may include any number of caches 212, including those described herein. For example, cache 212 may include an L3 cache available to both CPU 206 and GPU 208 (e.g., it is connected to both CPU 206 and GPU 208). Cache 212 may include a write-back cache, which can track the state of rows, for example, using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4 MB or more, but a smaller cache size may also be used.

[0072] SoC 204 may include one or more arithmetic logic units (ALUs) that can be used to perform processing of any of a variety of tasks or operations related to vehicle 200, such as processing a DNN. Furthermore, SoC 204 may include a floating-point unit (FPU) or other mathematical coprocessor or digital coprocessor type for performing mathematical operations within the system. For example, SoC 204 may include one or more FPUs integrated as execution units within CPU 206 and / or GPU 208.

[0073] SoC 204 may include one or more accelerators 214 (e.g., hardware accelerators, software accelerators, or combinations thereof). For example, SoC 204 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. This large on-chip memory (e.g., 4MB SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used to supplement GPU 208 and offload some tasks from GPU 208 (e.g., freeing up more cycles of GPU 208 to perform other tasks). As an example, accelerator 214 can be used for targeted workloads (e.g., perceptrons, convolutional neural networks (CNNs), etc.) that are stable enough to be easily controlled for acceleration. When used herein, the term "CNN" can include all types of CNNs, including region-based or region convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).

[0074] Accelerator 214 (e.g., a hardware acceleration cluster) may include a Deep Learning Accelerator (DLA). The DLA may include one or more Tensor Processing Units (TPUs) that can be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured to perform image processing functions (e.g., for CNNs, RCNNs, etc.) and optimized for performing image processing functions. The DLA may be further optimized for a specific set of neural network types and floating-point operations and inference. The DLA is designed to provide higher performance per millimeter than a general-purpose GPU and significantly outperform CPUs. The TPU can perform several functions, including single-instance convolution functions, support for INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.

[0075] DLA can execute neural networks, especially CNNs, quickly and efficiently on processed or unprocessed data for any function across a wide variety of applications, such as, but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection and recognition using data from microphones; CNNs for face recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for safety and / or safety-related events.

[0076] The DLA can perform any function of the GPU 208, and by using inference accelerators, for example, a designer can make the DLA or GPU 208 target any function. For example, a designer can focus the CNN processing and floating-point operations on the DLA and leave other functions to the GPU 208 and / or other accelerators 214.

[0077] Accelerator 214 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. A PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. A PVA can provide a balance between performance and flexibility. For example, each PVA may include, for example, but not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.

[0078] RISC cores can interact with image sensors (such as the image sensor of any camera described herein), image signal processors, and / or the like. Each of these RISC cores may include any amount of memory. Depending on the embodiment, the RISC core may use any of several protocols. In some examples, the RISC core may execute a real-time operating system (RTOS). RISC cores may be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, a RISC core may include an instruction cache and / or tightly coupled RAM.

[0079] DMA enables PVA components to access system memory independently of the CPU 206. DMA can support any number of features used to provide optimizations to the PVA, including but not limited to support for multidimensional addressing and / or circular addressing. In some examples, DMA can support addressing in up to six or more dimensions, which can include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.

[0080] A vector processor can be a programmable processor designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, a PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may operate as the main processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core may include a digital signal processor, such as, for example, a Single Instruction Multiple Data (SIMD) or Very Long Instruction Word (VLIW) digital signal processor. The combination of SIMD and VLIW can enhance throughput and speed.

[0081] Each of the vector processors may include an instruction cache and may be coupled to dedicated memory. Consequently, in some examples, each of the vector processors may be configured to execute independently of other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelism. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even different algorithms on a sequence of images or portions of an image. Among other things, any number of PVAs may be included in a hardware-accelerated cluster, and any number of vector processors may be included in each of these PVAs. Furthermore, the PVA may include additional error-correcting code (ECC) memory to enhance overall system security.

[0082] Accelerator 214 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for accelerator 214. In some examples, on-chip memory may include at least 4MB of SRAM consisting of, for example, but not limited to, eight field-configurable memory blocks, accessible by both the PVA and DLA. Each pair of memory blocks may include an Advanced Peripheral Bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and DLA may access memory via a backbone that provides high-speed memory access to the PVA and DLA. The backbone may include (e.g., using an APB) an on-chip computer vision network that interconnects the PVA and DLA to memory.

[0083] On-chip computer vision networks can include interfaces that ensure both the PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. Such interfaces can provide separate phases and channels for transmitting control signals / addresses / data, as well as burst communication for continuous data transmission. This type of interface can conform to ISO 26262 or IEC 61508 standards, but other standards and protocols can also be used.

[0084] In some examples, SoC 204 may include, for example, a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed August 10, 2018. This real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the location and extent of objects (e.g., within a world model) to generate real-time visualization simulations for RADAR signal interpretation, sound propagation synthesis and / or analysis, SONAR system simulation, general wave propagation simulation, comparison with LiDAR data for localization and / or other functional purposes, and / or for other uses. In some embodiments, one or more Tree Traversal Units (TTUs) may be used to perform one or more ray tracing-related operations.

[0085] Accelerators 214 (e.g., hardware accelerator clusters) have broad applications in autonomous driving. PVAs can be programmable vision accelerators used in critical processing stages of ADAS and autonomous vehicles. The capabilities of PVAs are a good match for algorithmic domains requiring predictable processing, low power, and low latency. In other words, PVAs perform well in semi-dense or dense rule computation, even on small datasets requiring predictable runtimes with low latency and low power. Therefore, in the context of platforms for autonomous vehicles, PVAs are designed to run classical computer vision algorithms because they are efficient in object detection and integer mathematical operations.

[0086] For example, according to one embodiment of this technology, PVA is used to perform computer stereo vision. In some examples, semi-global matching-based algorithms may be used, but this is not intended to be limiting. Many applications for Level 3–5 autonomous driving require instantaneous motion estimation / stereo matching (e.g., from moving structures, pedestrian recognition, lane detection, etc.). PVA can perform computer stereo vision functions on input from two monocular cameras.

[0087] In some examples, PVA can be used to perform intensive optical flow, processing raw RADAR data (e.g., using 4D Fast Fourier Transform) to provide processed RADAR. In other examples, PVA is used for time-of-flight depth processing, which, for example, involves processing raw time-of-flight data to provide processed time-of-flight data.

[0088] DLA can be used to run any type of network to enhance control and driving safety, including, for example, neural networks that output a confidence metric for each object detection. Such a confidence value can be interpreted as a probability or as providing a relative “weight” for each detection compared to other detections. This confidence value allows the system to make further decisions about which detections should be considered true positives rather than false positives. For example, the system can set a threshold for the confidence and only consider detections exceeding the threshold as true positives. In an Automatic Emergency Braking (AEB) system, false positives can cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run a neural network to regress the confidence value. This neural network can take at least a subset of parameters as input, such as bounding box dimensions, ground plane estimates (e.g., from another subsystem), inertial measurement unit (IMU) sensor 266 outputs related to vehicle orientation and distance, 3D position estimates of objects obtained from the neural network and / or other sensors (e.g., LiDAR sensor 264 or RADAR sensor 260), etc.

[0089] SoC 204 may include one or more data stores 216 (e.g., memory). Data stores 216 may be on-chip memory of SoC 204, which may store neural networks to be executed on the GPU and / or DLA. In some examples, for redundancy and security, data stores 216 may be large enough to store multiple instances of the neural network. Data stores 212 may include L2 or L3 caches 212. References to data stores 216 may include references to memory associated with PVA, DLA, and / or other accelerators 214 as described herein.

[0090] SoC 204 may include one or more processors 210 (e.g., embedded processors). Processor 210 may include a startup and power management processor, which may be a dedicated processor and subsystem for handling startup power and management functions, as well as safety implementation. The startup and power management processor may be part of the SoC 204 startup sequence and may provide runtime power management services. The startup power and management processor may provide clock and voltage programming, auxiliary system low-power state transitions, SoC 204 thermal and temperature sensor management, and / or SoC 204 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to the temperature, and SoC 204 may use the ring oscillator to detect the temperature of CPU 206, GPU 208, and / or accelerator 214. If it is determined that the temperature exceeds a threshold, the startup and power management processor may enter a temperature fault routine and place SoC 204 into a lower power state and / or place vehicle 200 into a driver-safe parking mode (e.g., safely stop vehicle 200).

[0091] Processor 210 may further include a set of embedded processors that can be used as an audio processing engine. The audio processing engine can be an audio subsystem that allows for full hardware support for multi-channel audio via multiple interfaces, as well as a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor and dedicated RAM.

[0092] The processor 210 may further include an always-on-processor engine that can provide the necessary hardware features to support low-power sensor management and wake-up use cases. This always-on-processor engine may include a processor core, tightly coupled RAM, support for peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0093] Processor 210 may further include a security cluster engine, which includes a dedicated processor subsystem for handling security management of automotive applications. The security cluster engine may include two or more processor cores, tightly coupled RAM, support for peripheral devices (e.g., timers, interrupt controllers, etc.), and / or routing logic. In secure mode, the two or more cores may operate in lockstep mode and function as a single core with comparison logic that detects any differences between their operations.

[0094] The processor 210 may further include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.

[0095] The processor 210 may further include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.

[0096] Processor 210 may include a video image compositer, which may be (e.g., implemented on a microprocessor) a processing block, implementing video post-processing functions required by the video playback application to generate the final image for the player window. The video image compositer may perform lens distortion correction on the wide-angle camera 270, the surround camera 274, and / or the in-cabin monitoring camera sensor. The in-cabin monitoring camera sensor is preferably monitored by a neural network running on another instance of an advanced SoC, configured to recognize in-cabin events and respond accordingly. The in-cabin system may perform lip reading to activate mobile phone services and make calls, dictate emails, change vehicle destinations, activate or change the vehicle's infotainment system and settings, or provide voice-activated web browsing. Some functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled in other situations.

[0097] Video image compositers can include enhanced temporal denoising for both spatial and temporal noise reduction. For example, in the case of motion in the video, denoising appropriately weights spatial information, reducing the weight of information provided by neighboring frames. In cases where the image or part of the image does not contain motion, the temporal denoising performed by the video image compositer can use information from previous images to reduce noise in the current image.

[0098] The video image compositer can also be configured to perform stereo correction on input stereo camera frames. When the operating system desktop is in use and the GPU 208 does not need to continuously render new surfaces, the video image compositer can be further used for user interface components. Even when the GPU 208 is powered on and active, performing 3D rendering, the video image compositer can be used to offload the GPU 208 to improve performance and responsiveness.

[0099] SoC 204 may further include a Mobile Industry Processor Interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functions for receiving video and input from a camera. SoC 204 may further include an input / output controller that can be software-controlled and can be used to receive I / O signals not assigned to a specific role.

[0100] SoC 204 may further include a wide range of peripheral interfaces to enable communication with peripherals, audio codecs, power management and / or other devices. SoC 204 can be used to process data from cameras and sensors (e.g., LIDAR sensor 264, RADAR sensor 260, etc., which can be connected via Gigabit Multimedia Serial Link and Ethernet), data from bus 202 (e.g., vehicle 200 speed, steering wheel position, etc.), and data from GNSS sensor 258 (connected via Ethernet or CAN bus). SoC 204 may further include a dedicated high-performance, high-capacity memory controller, which may include its own DMA engine, and which can be used to free up CPU 206 from routine data management tasks.

[0101] SoC 204 can be an end-to-end platform with a flexible architecture spanning Levels 3-5 of automation, providing a comprehensive functional safety architecture that leverages and efficiently utilizes computer vision and ADAS technologies for diversity and redundancy, along with deep learning tools, to deliver a flexible and reliable driving software stack. SoC 204 can be faster, more reliable, and even more energy- and space-efficient than conventional systems. For example, when combined with CPU 206, GPU 208, and data storage 216, accelerator 214 can provide a fast and efficient platform for Level 3-5 autonomous vehicles.

[0102] Therefore, this technology offers capabilities and functionalities that cannot be achieved through conventional systems. For example, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages ​​such as C to execute a wide variety of processing algorithms across a diverse range of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for automotive ADAS applications and practical Level 3-5 autonomous vehicles.

[0103] In contrast to conventional systems, the techniques described in this paper, by providing CPU complexes, GPU complexes, and hardware acceleration clusters, allow multiple neural networks to be executed simultaneously and / or sequentially, and the results to be combined to achieve Level 3–5 autonomous driving capabilities. For example, a CNN executed on a DLA or dGPU (e.g., GPU 220) could include text and word recognition, allowing a supercomputer to read and understand traffic signs, including those for which neural networks have not yet been specifically trained. The DLA could further include a neural network capable of recognizing, interpreting, and providing semantic understanding of the signs, and passing that semantic understanding to a path planning module running on a CPU complex.

[0104] As another example, multiple neural networks can operate simultaneously, as required for Level 3, 4, or 5 driving. For instance, a warning sign consisting of "Caution: Flashing lights indicate icy conditions" along with a light can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a deployed first neural network (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" can be interpreted by a deployed second neural network that informs the vehicle's path planning software (preferably executing on a CPU complex) that icy conditions exist when the flashing lights are detected. The flashing lights can be identified by a deployed third neural network operating across multiple frames, informing the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks can operate simultaneously, for example, within a DLA and / or on a GPU 208.

[0105] In some examples, the CNN used for facial recognition and owner identification can use data from camera sensors to identify the presence of an authorized driver and / or owner of vehicle 200. A processing engine always on the sensors can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in safe mode, to disable the vehicle when the owner leaves. In this way, SoC 204 provides security against theft and / or carjacking.

[0106] In another example, the CNN used for emergency vehicle detection and identification can use data from microphone 296 to detect and identify emergency vehicle siren. In contrast to conventional systems that use a general classifier to detect siren and manually extract features, SoC 204 uses a CNN to classify environmental and urban sounds as well as visual data. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative shut-off rate of emergency vehicles (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the localized area in which the vehicle operates, as identified by GNSS sensor 258. Thus, for example, when operating in Europe, the CNN will seek to detect European siren, and when operating in the United States, the CNN will seek to identify siren only in North America. Once an emergency vehicle is detected, with the assistance of ultrasonic sensor 262, the control program can be used to execute emergency vehicle safety routines, causing the vehicle to slow down, pull over to the side of the road, stop, and / or idle until the emergency vehicle passes.

[0107] The vehicle may include a CPU 218 (e.g., a discrete CPU or dCPU) that can be coupled to the SoC 204 via a high-speed interconnect (e.g., PCIe). The CPU 218 may include, for example, an x86 processor. The CPU 218 can be used to perform any of a wide variety of functions, including, for example, arbitrating the results of potential inconsistencies between ADAS sensors and the SoC 204, and / or monitoring the status and health of the controller 236 and / or the infotainment SoC 230.

[0108] Vehicle 200 may include GPU 220 (e.g., a discrete GPU or dGPU) that can be coupled to SoC 204 via a high-speed interconnect (e.g., NVIDIA's NVLINK). GPU 220 may provide additional artificial intelligence capabilities, for example by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs from sensors of vehicle 200 (e.g., sensor data).

[0109] Vehicle 200 may further include a network interface 224, which may include one or more wireless antennas 226 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). Network interface 224 can be used to enable wireless connectivity via the Internet to the cloud (e.g., with server 278 and / or other network devices), with other vehicles, and / or with computing devices (e.g., a passenger's client device). For communication with other vehicles, a direct link can be established between the two vehicles, and / or an indirect link can be established (e.g., across networks and via the Internet). A direct link can be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide vehicle 200 with information about vehicles approaching vehicle 200 (e.g., vehicles in front, to the side, and / or behind vehicle 200). This functionality can be part of vehicle 200's cooperative adaptive cruise control function.

[0110] Network interface 224 may include a SoC that provides modulation and demodulation functions and enables controller 236 to communicate via a wireless network. Network interface 224 may include an RF front-end for up-conversion from baseband to RF and down-conversion from RF to baseband. Frequency conversion can be performed using known processes and / or using a superheterodyne process. In some examples, the RF front-end functionality may be provided by a separate chip. The network interface may include wireless functions for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0111] Vehicle 200 may further include data storage 228, which may include off-chip (e.g., off-chip SoC 204) storage devices. Data storage 228 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard disk, and / or other components and / or devices capable of storing at least one bit of data.

[0112] Vehicle 200 may further include GNSS sensor 258. GNSS sensor 258 (e.g., GPS, assisted GPS sensor, differential GPD (DGPS) sensor, etc.) is used for assisted mapping, sensing, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 258 can be used, including, for example, but not limited to, GPS using a USB connector with an Ethernet-to-serial (RS-232) bridge.

[0113] Vehicle 200 may further include a RADAR sensor 260. The RADAR sensor 260 can be used by vehicle 200 for remote vehicle detection even in dark and / or inclement weather conditions. The RADAR functional safety level may be ASIL B. The RADAR sensor 260 can use CAN and / or bus 202 (e.g., to transmit data generated by the RADAR sensor 260) for control and access to object tracking data, and in some examples, Ethernet access for accessing raw data. A wide variety of RADAR sensor types can be used. For example, and without limitation, the RADAR sensor 260 can be adapted for front, rear, and side RADAR use. In some examples, a pulse Doppler RADAR sensor is used.

[0114] RADAR sensor 260 can include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In some examples, long-range RADAR can be used for adaptive cruise control functions. A long-range RADAR system can provide a wide field of view (e.g., within 250m) achieved through two or more independent scans. RADAR sensor 260 can help distinguish between stationary and moving objects and can be used by ADAS systems for emergency braking assist and forward collision warning. Long-range RADAR sensors can include a single-site multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the four central antennas can create a focused beam pattern designed to record the vehicle 200's surroundings at higher rates with minimal traffic interference from adjacent lanes. The other two antennas can extend the field of view, enabling rapid detection of vehicles entering or leaving the vehicle 200's lane.

[0115] As an example, a mid-range RADAR system can include a range of up to 260m (front) or 80m (rear) and a field of view of up to 42 degrees (front) or 250 degrees (rear). A short-range RADAR system can include, but is not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two beams that continuously monitor blind spots behind and beside the vehicle.

[0116] Short-range RADAR systems can be used in ADAS systems for blind spot detection and / or lane change assistance.

[0117] Vehicle 200 may further include ultrasonic sensors 262. Ultrasonic sensors 262, which may be positioned at the front, rear, and / or sides of vehicle 200, can be used for parking assistance and / or creating and updating occupancy grids. A wide variety of ultrasonic sensors 262 can be used, and different ultrasonic sensors 262 can be used for different detection ranges (e.g., 2.5m, 4m). Ultrasonic sensors 262 can operate at functional safety level ASIL B.

[0118] Vehicle 200 may include a LIDAR sensor 264. The LIDAR sensor 264 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 264 may be of functional safety level ASIL B. In some examples, vehicle 200 may include multiple LIDAR sensors 264 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).

[0119] In some examples, the LIDAR sensor 264 may be able to provide a list of objects and their distances within a 360-degree field of view. Commercially available LIDAR sensors 264 may have an advertising range of, for example, approximately 200m, with an accuracy of 2cm-3cm, and support for 200Mbps Ethernet connectivity. In some examples, one or more non-protruding LIDAR sensors 264 may be used. In such examples, the LIDAR sensor 264 may be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of a vehicle 200. In such examples, the LIDAR sensor 264 may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for low-reflectivity objects, with a range of 200m. Front-mounted LIDAR sensors 264 may be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0120] In some examples, LiDAR technologies such as 3D flash LiDAR can also be used. 3D flash LiDAR uses a flash of laser light as the emission source to illuminate the environment around a vehicle up to approximately 200 meters in distance. A flash LiDAR unit includes a receiver that records the laser pulse propagation time and reflected light onto each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LiDAR allows for the generation of highly accurate and distortion-free images of the surrounding environment using each laser flash. In some examples, four flash LiDAR sensors can be deployed, one on each side of the vehicle. Available 3D flash LiDAR systems include solid-state 3D staring array LiDAR cameras (e.g., non-browsing LiDAR devices) without moving parts other than a fan. Flash LiDAR devices can use 5 nanosecond Class I (eye-safe) laser pulses per frame and can capture reflected laser light in the form of a 3D range point cloud and co-registered intensity data. By using a flash LiDAR, and because a flash LiDAR is a solid-state device with no moving parts, the LiDAR sensor 264 is less susceptible to motion blur, vibration, and / or shock.

[0121] The vehicle may further include an IMU sensor 266. In some examples, the IMU sensor 266 may be located at the center of the rear axle of the vehicle 200. The IMU sensor 266 may include, for example, but not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 266 may include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 266 may include an accelerometer, a gyroscope, and a magnetometer.

[0122] In some embodiments, the IMU sensor 266 can be implemented as a miniature, high-performance GPS-assisted inertial navigation system (GPS / INS) that combines a microelectromechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filter algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 266 can enable the vehicle 200 to estimate heading by directly observing and correlating velocity changes from GPS to the IMU sensor 266 without input from a magnetic sensor. In some examples, the IMU sensor 266 and the GNSS sensor 258 can be combined into a single integrated unit.

[0123] The vehicle may include a microphone 296 placed in and / or around the vehicle 200. Among other things, the microphone 296 may be used for emergency vehicle detection and identification.

[0124] The vehicle may further include any number of camera types, including stereo camera 268, wide-angle camera 270, infrared camera 272, surround camera 274, long-range and / or mid-range camera 298, and / or other camera types. These cameras can be used to capture image data around the entire perimeter of the vehicle 200. The camera types used depend on the embodiment and the requirements of the vehicle 200, and any combination of camera types can be used to provide the necessary coverage around the vehicle 200. Furthermore, the number of cameras may vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and without limitation, these cameras may support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras is described herein with respect to... Figure 2A and Figure 2B It was described in more detail.

[0125] Vehicle 200 may further include vibration sensor 242. Vibration sensor 242 can measure vibrations of vehicle components such as axles. For example, changes in vibration can indicate changes in the road surface. In another example, when two or more vibration sensors 242 are used, differences between vibrations can be used to determine friction or slippage on the road surface (e.g., when there is a vibration difference between a power drive shaft and a free-rotating shaft).

[0126] Vehicle 200 may include ADAS system 238. In some examples, ADAS system 238 may include SoC. ADAS system 238 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC) and / or other features and functions.

[0127] The ACC system can use a RADAR sensor 260, a LIDAR sensor 264, and / or a camera. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to vehicles immediately in front of vehicle 200 and automatically adjusts the vehicle speed to maintain a safe distance. Lateral ACC performs distance holding and, if necessary, advises vehicle 200 to change lanes. Lateral ACC is associated with other ADAS applications such as LCA and CWS.

[0128] CACC uses information from other vehicles, which can be received indirectly from other vehicles via a wireless link or through a network connection (e.g., via the Internet) through network interface 224 and / or wireless antenna 226. Direct links can be provided by vehicle-to-vehicle (V2V) communication links, while indirect links can be infrastructure-to-vehicle (I2V) communication links. Typically, the V2V communication concept provides information about vehicles immediately ahead (e.g., vehicles immediately in front of vehicle 200 and in the same lane), while the I2V communication concept provides information about traffic further ahead. A CACC system can include either or both of these I2V and V2V information sources. Given information about vehicles ahead of vehicle 200, CACC can be more reliable, and it has the potential to improve traffic flow and reduce road congestion.

[0129] The Forward-Looking Warning (FCW) system is designed to alert the driver to hazards, enabling the driver to take corrective action. The FCW system uses a front-facing camera and / or RADAR sensor 260 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components. The FCW system can provide warnings in the form of, for example, audible, visual, haptic, and / or rapid braking pulses.

[0130] An AEB (Autonomous Emergency Braking) system detects an impending forward collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system can use a front-facing camera and / or RADAR sensor 260 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid a collision, and if the driver does not take corrective action, the AEB system can automatically apply the brakes to attempt to prevent or at least mitigate the effects of the predicted collision. The AEB system may include technologies such as dynamic brake support and / or collision approach braking.

[0131] The Lane Departure Warning (LDW) system provides visual, auditory, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle 200 crosses a lane marking. When the driver indicates intentional lane departure, the LDW system is deactivated by activating a turn signal. The LDW system can utilize a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.

[0132] The LKA system is a variant of the LDW system. If vehicle 200 begins to leave the lane, the LKA system provides corrective steering input or braking to vehicle 200.

[0133] The BSW system detects and warns the driver of vehicles in the vehicle's blind spot. The BSW system can provide visual, auditory, and / or tactile alerts to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses turn signals. The BSW system can use one or more rear-facing cameras and / or one or more RADAR sensors 260, coupled to a dedicated processor, DSP, FPGA, and / or ASIC (which is electrically coupled to driver feedback, such as a display, speaker, and / or vibration components).

[0134] RCTW systems can provide visual, auditory, and / or tactile notifications when an object is detected outside the range of a rear-view camera while the vehicle is reversing. Some RCTW systems include AEB (Autonomous Emergency Braking) to ensure the application of the vehicle's brakes to avoid a collision. RCTW systems may use one or more rear-view RADAR sensors 260 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibrating components.

[0135] Conventional ADAS systems can be prone to false positives, which can be frustrating and distracting for the driver, but typically not catastrophic, as they alert the driver and allow them to determine whether a safe condition truly exists and take appropriate action. However, in autonomous vehicle 200, in the event of conflicting results, vehicle 200 itself must decide whether to heed the results from the main computer or auxiliary computer (e.g., first controller 236 or second controller 236). For example, in some embodiments, ADAS system 238 may be a backup and / or auxiliary computer for providing perception information to a backup computer rationality module. The backup computer rationality monitor may run redundant and diverse software on hardware components to detect faults in perception and dynamic driving tasks. Output from ADAS system 238 may be provided to a supervisory MCU. If the outputs from the main computer and auxiliary computer conflict, the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.

[0136] In some examples, the master computer can be configured to provide a confidence score to the supervisory MCU, indicating the master computer's confidence level in the selected result. If the confidence score exceeds a threshold, the supervisory MCU can follow the master computer's direction regardless of whether the auxiliary computer provides conflicting or inconsistent results. If the confidence score does not meet the threshold and the master and auxiliary computers indicate different results (e.g., conflict), the supervisory MCU can arbitrate between these computers to determine the appropriate result.

[0137] The supervisory MCU can be configured to run a neural network trained and configured to determine the conditions under which the auxiliary computer provides a false alarm based on outputs from both the host and auxiliary computers. Thus, the neural network in the supervisory MCU can learn when the output of the auxiliary computer can be trusted and when it cannot. For example, when the auxiliary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW system is identifying a metallic object that is not actually dangerous, such as a drain grid or manhole cover that triggers an alarm. Similarly, when the auxiliary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to ignore the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In embodiments that include a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or GPU suitable for running the neural network using associated memory. In a preferred embodiment, the supervisory MCU may include components of SoC 104 and / or be included as components of SoC 204.

[0138] In other examples, ADAS system 238 may include an auxiliary computer that performs ADAS functions using conventional computer vision rules. This allows the auxiliary computer to use classic computer vision rules (if-then), and the presence of neural networks in the supervising MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity make the entire system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functionality. For instance, if a software vulnerability or bug exists in the software running on the host computer and non-identical software code running on the auxiliary computer provides the same overall result, the supervising MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the host computer does not cause a substantial error.

[0139] In some examples, the output of ADAS system 238 can be fed to the perception block and / or the dynamic driving task block of the main computer. For example, if ADAS system 238 issues a forward collision warning because an object is immediately in front, the perception block can use this information when recognizing the object. In other examples, the assistance computer can have its own neural network, which is trained and thus reduces the risk of false positives as described herein.

[0140] Vehicle 200 may further include an infotainment SoC 230 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more discrete components. The infotainment SoC 230 may include a combination of hardware and software that can be used to provide vehicle 200 with audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming media, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.) and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total coverage distance, brake fuel level, fuel level, door opening / closing, air filter information, etc.). For example, the infotainment SoC 230 may include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, in-vehicle computer, in-vehicle entertainment, Wi-Fi, steering wheel audio controls, hands-free voice controls, head-up display (HUD), HMI display 234, telematics device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems) and / or other components. The infotainment SoC 230 may further be used to provide information (e.g., visual and / or auditory) to the vehicle's users, such as information from ADAS system 238, autonomous driving information such as planned vehicle maneuvers, trajectory, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0141] The infotainment SoC 230 may include GPU functionality. The infotainment SoC 230 can communicate with other devices, systems, and / or components of the vehicle 200 via bus 202 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 230 may be coupled to a supervisory MCU, allowing the GPU of the infotainment system to perform autonomous driving functions in the event of a failure of the main controller 236 (e.g., the primary and / or backup computer of the vehicle 200). In such an example, the infotainment SoC 230 may place the vehicle 200 into a driver-safe parking mode as described herein.

[0142] Vehicle 200 may further include instrument cluster 232 (e.g., digital instrument panel, electronic instrument cluster, digital instrument panel, etc.). Instrument cluster 232 may include a controller and / or a supercomputer (e.g., a discrete controller or supercomputer). Instrument cluster 232 may include a set of instruments such as speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, seatbelt warning light, parking brake warning light, engine malfunction indicator, airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between infotainment SoC 230 and instrument cluster 232. In other words, instrument cluster 232 may be included as part of infotainment SoC 230, or vice versa.

[0143] Figure 2D For cloud-based servers and according to various embodiments Figure 2A This is a schematic diagram of a system for communication between exemplary autonomous vehicles 200. System 276 may include server 278, network 290, and vehicles including vehicle 200. Server 278 may include multiple GPUs 284(A)-284(H) (collectively referred to herein as GPU 284), PCIe switches 282(A)-282(H) (collectively referred to herein as PCIe switch 282), and / or CPUs 280(A)-280(B) (collectively referred to herein as CPU 280). GPUs 284, CPUs 280, and PCIe switches may be interconnected with high-speed interconnects and / or PCIe connections 286, such as, but not limited to, NVLink interface 288 developed by NVIDIA. In some examples, GPUs 284 are connected via NVLink and / or NVSwitch SoCs, and GPUs 284 and PCIe switches 282 are connected via PCIe interconnects. Although eight GPUs 284, two CPUs 280, and two PCIe switches are shown in the diagram, this is not intended to be limiting. Depending on the embodiment, each of the servers 278 may include any number of GPUs 284, CPUs 280, and / or PCIe switches. For example, each of the servers 278 may include eight, sixteen, thirty-two, and / or more GPUs 284.

[0144] Server 278 can receive image data from vehicles via network 290, representing images of unexpected or changed road conditions such as recently commenced roadworks. Server 278 can also transmit neural network 292, updated neural network 292, and / or map information 294, including information about traffic and road conditions, to vehicles via network 290. Updates to map information 294 may include updates to HD map 222, such as information about construction sites, potholes, bends, floods, or other obstacles. In some examples, neural network 292, updated neural network 292, and / or map information 294 may have been generated from new training and / or data received from any number of vehicles in the environment, and / or based on experience gained from training performed at a data center (e.g., using server 278 and / or other servers).

[0145] Server 278 can be used to train machine learning models (e.g., neural networks) based on training data. Training data can be generated by the vehicle and / or generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., where the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., where the neural network does not require supervised learning). Training can be performed according to any one or more classes of machine learning techniques, including but not limited to: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, joint learning, transfer learning, feature learning (including principal component analysis and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including alternative dictionary learning), rule-based machine learning, anomaly detection, and any variations or combinations thereof. Once the machine learning model is trained, it can be used by the vehicle (e.g., transmitted to the vehicle via network 290), and / or the machine learning model can be used by server 278 to remotely monitor the vehicle.

[0146] In some examples, server 278 can receive data from the vehicle and apply that data to a state-of-the-art real-time neural network for real-time intelligent inference. Server 278 may include a deep learning supercomputer powered by GPU 284 and / or a dedicated AI computer, such as the DGX and DGX station machines developed by NVIDIA. However, in some examples, server 278 may include a deep learning infrastructure in a data center that uses only CPU power.

[0147] The deep learning infrastructure of server 278 may be capable of rapid real-time inference and can be used to assess and verify the health status of the processor, software, and / or associated hardware in vehicle 200. For example, the deep learning infrastructure may receive periodic updates from vehicle 200, such as image sequences and / or objects located in those image sequences that vehicle 200 has already located (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure may run its own neural network to identify objects and compare them with objects identified by vehicle 200. If the results do not match and the infrastructure concludes that the AI ​​in vehicle 200 has malfunctioned, then server 278 may transmit a signal to vehicle 200 instructing the vehicle 200's fail-safe computer to take control, notify passengers, and complete a safe stopping operation.

[0148] For inference, server 278 may include GPU 284 and one or more programmable inference accelerators (such as NVIDIA's TensorRT 3). The combination of a GPU-powered server and inference acceleration enables real-time response. In other examples, such as where performance is less critical, CPU, FPGA, and other processor-powered servers can be used for inference.

[0149] Using language models to control autonomous vehicles

[0150] Figure 3 According to various embodiments Figure 1 A more detailed description of the language model driving assistance system 130 is provided below. As shown, the language model driving assistance system 130 includes a traffic rule extractor 310, a planner 330, and a language model 320. Although the language model 320 is shown as being included in the language model driving assistance system 130, in some embodiments, the language model 320 may be independent of the language model driving assistance system 130. For example, in some embodiments, the language model 320 may execute in a cloud computing environment and be accessed by the language model driving assistance system 130 via an application programming interface (API).

[0151] During operation, the language model driving assistance system 130 receives a scenario description 302, a motion plan (“plan”) 304, a situation 306, and local traffic rules 316. The scenario description 302 is a description of the current environment from the vehicle's perspective. The motion plan 304 is a nominal execution plan for driving the vehicle. The situation 306 is the situation the vehicle is currently encountering. In some embodiments, under normal circumstances, such as when no specific situation is input to the language model driving assistance system 130, the situation defaults to “normal state”. In other cases, any other situation can be received, such as “someone honks at me,” “a car suddenly appears,” “a pedestrian suddenly appears,” “it’s raining heavily,” “there’s trash on the road,” “the road is a little wet,” “it suddenly starts raining,” etc. In some embodiments, the scenario description 302, motion plan 304, and situation 306 can be received in text format. For example, the scenario description 302, motion plan 304, and / or situation 306 can be input by a user in text form via a keyboard or touchscreen. As another example, the scene description 302, motion plan 304, and / or condition 306 can be generated from the user's speech using any technically feasible speech-to-text technology. As yet another example, one or more of the scene description, motion plan, or condition can be generated using one or more trained machine learning models, such as a trained visual language model that generates text of the scene description and / or condition based on one or more images (or other sensor data) of the environment surrounding the vehicle, or a language model trained to generate a high-level semantic description of the motion plan (e.g., GPT-Driver), and so on.

[0152] Traffic rule extractor 310 extracts portions of local traffic rules 316 based on scene description 302, motion plan 304, and situation 306. Local traffic rules 316 can be obtained in any technically feasible manner, such as based on user input specifying the current geographic location, automatically obtained based on the vehicle's Global Positioning System (GPS) coordinates, etc. Although, by way of reference, local traffic rules 316 are shown as being input into the language model driving assistance system 130, in some embodiments, local traffic rules may also be included in the language model driving assistance system 130. Extracting portions of local traffic rules 316 can be beneficial because passing all local traffic rules 316 to the language model 320 may be redundant, and in some cases, the language model 320 does not allow this. Furthermore, irrelevant information from local traffic rules 316 can sometimes affect the performance of the language model 320. In some embodiments, the language model driving assistance system 130 may extract portions of local traffic rules 316 by: (1) prompting the language model 320 to obtain common traffic-related phrases within the scenario description 302, motion plan 304, and situation 306; and (2) extracting paragraphs from the local traffic rules 316 containing these common traffic-related phrases. Illustratively, the traffic rule extractor 310 includes a prompt generator that generates a prompt 318. The prompt 318 requests common traffic-related phrases from the scenario description 302, motion plan 304, and situation 306, which are also included in the prompt 318. Given the prompt 318, the language model 320 outputs keywords 322 containing the common traffic-related phrases. In some embodiments, any technically feasible language model 320 may be used, such as a trained Large Language Model (LLM). In some embodiments, the language 320 may be a pre-trained model used in a zero-shot manner. In some other embodiments, the language model 320 can be specifically fine-tuned using additional training data to generate driving instructions.

[0153] Traffic rule extractor 310 also includes a paragraph extractor for extracting one or more paragraphs 324 containing the keyword 322 from local traffic rule 316. In some embodiments, paragraph extractor 314 may search for keyword 322 in local traffic rule 316. Paragraph extractor 314 may then extract the paragraphs 324 from local traffic rule 316 containing the keyword. Paragraph extractor 314 may also segment local traffic rule 316 into paragraphs before searching for keyword 322 and extracting paragraphs 324. While this document has primarily described in conjunction with extracting paragraphs of local traffic rules as a reference example, in other embodiments, any suitable portion of the local traffic rule (e.g., sentences) may also be extracted. Furthermore, in other embodiments, traffic rule extractor 310 may extract portions of local traffic rule 316 in any technically feasible manner (e.g., using embedded search).

[0154] Given the extracted paragraph 324, scene description 302, motion plan 304, and situation 306, the planner 330 prompts the language model 320 to generate an updated plan 336, which can then be output by the language model driving assistance system 130. Illustratively, the planner 330 includes a prompt generator 332. The prompt generator 332 generates a prompt 334 that requests driving instructions and includes the scene description, motion plan, situation, and extracted paragraph 324. The planner 332 inputs the prompt 334 into the language model 320, which outputs the updated plan 336. The updated plan 336 contains driving instructions specific to the scene description 302, motion plan 304, and situation 306, and adheres to the traffic rules in the extracted paragraph 324.

[0155] The language model driving assistance device 130 outputs an updated motion plan 336 to provide driving assistance to the user or other software. In some embodiments, the updated motion plan 336 can be output to the user as a driving command via a display device and / or a speaker. In this case, the driving command can be converted into an audio signal (e.g., via text-to-speech technology) and output via a speaker; and / or, the driving command can be output in text form via a display device. In other embodiments, the updated motion plan 336 can be output to an AV motion planner, which uses the updated motion plan 336 to generate a planned motion for the vehicle. In this case, the planned motion can be transmitted to a controller (not shown) that controls the vehicle's steering, accelerator, and / or braking to achieve the planned motion, as described below. Figure 4 To describe in more detail.

[0156] Figure 4The illustrations show the inclusion of various embodiments. Figure 1 The autonomous vehicle (AV) application 400 of the language model driving assistance system 130 is shown in the figure. As shown, the AV application 400 includes an AV motion planner 410 and the language model driving assistance system 130.

[0157] During operation, the AV application 400 receives sensor data 402 and vehicle status information 404. In some embodiments, sensor data 402 may include image data, light detection and ranging (LIDAR) data, and / or radar (RADAR) data. In some embodiments, vehicle status information 404 may include controller local area network (CAN) bus data, which may indicate the vehicle's speed, acceleration, steering angle, braking, position, etc. The AV motion planner 410 processes the sensor data 402 and vehicle status information 404 to generate a motion plan ("plan") 412, which contains text describing the planned motion of the vehicle and the corresponding future trajectory (e.g., [(x1,y1),...,(xN,yN)]). In some embodiments, any technically feasible AV motion planner can be used, including motion planners that accept input text and output text containing the planned motion and future trajectory. For example, in some embodiments, the AV motion planner 410 may be a GPT-Driver, DriveGPT4, MotionLM, etc.

[0158] AV application 400 uses language model driving assistance system 130 to process motion plan 412 and local traffic rules 416 received as input, thereby generating guidance 414 for AV motion planner 410. In some embodiments, language model driving assistance system 130 performs the above combination. Figure 3The process involves extracting portions of local traffic rules 416 containing keywords related to motion plan 412. In some embodiments, the extracted portions of local traffic rules 416 may also be related to scenarios and conditions (not shown). The language model driving assistance system 130 prompts a language model (e.g., language model 320) to generate driving instructions (“instructions”) 414. Driving instructions 414 are then fed back to the AV motion planner 410, which uses instructions 414 along with sensor data 404, vehicle state information 404, and / or motion plan 412 to generate an updated plan 406. The updated plan 406 may include text describing the planned motion and the corresponding future trajectory (e.g., [(x1',y1'),…,(xN',yN')]), which modifies the previous motion plan 412 and associated trajectory to fit driving instructions 414. The updated motion plan 406 and / or associated motion trajectory can then be transmitted to a controller (not shown) that controls the vehicle’s steering, accelerator and / or braking to achieve the planned motion and / or associated trajectory.

[0159] Figure 5 The illustrations show the inclusion of various embodiments. Figure 1 The navigation application 500 is a language model-based driver assistance system 130. As shown in the figure, the navigation application 500 includes a visual language model 504 and the language model-based driver assistance system 130. The visual language model 504 is a trained machine learning model that combines natural language processing and computer vision to understand and generate text about images. During operation, the navigation application 500 uses the visual language model 504 to process image data 502 of the vehicle's surroundings to generate a scene description 506. For example, the navigation application 500 can prompt the visual language model 504 to describe the scene in the image data 502. The scene description 506 is combined with the above. Figure 3 The scene described is similar to 302 and includes a textual description of the environment from the vehicle's perspective.

[0160] The navigation application 500 also uses the language model driving assistance system 130 to process the scene description 506, as well as the plan 508, situation 510, and local traffic rules 512 received in the form of text input (e.g., input via keyboard or touchscreen, or converted from audio signals using speed-to-text technology), in order to generate an updated plan 514. Alternatively, in some embodiments, the situation 510 may also be generated using a visual language model 504. In some embodiments, the language model driving assistance system 130 performs the above combination. Figure 3The process described above extracts portions of the local traffic rule 512 containing keywords relevant to the scenario description 506, plan 608, and situation 510. The language model driving assistance system 130 then prompts the language model (e.g., language model 320) to generate an updated plan 514 that takes into account the local traffic rule 412. Once generated, the navigation application 500 can output the updated plan 514 in any technically feasible manner (e.g., via text instructions displayed on a display device and / or audio instructions output via one or more speaker devices).

[0161] Figure 6 Various embodiments are shown. Figure 5 The navigation application 500 can generate an exemplary plan. As shown, given an image 602 of the environment from the vehicle's perspective (which may be captured by a camera mounted on the vehicle, for example, as part of a video), the navigation application 500 inputs image 602 into a visual language model 504 to generate a scene description 604: "Sparse traffic ahead on the bridge, leading to a modern city skyline under a clear sky." The navigation application 500 inputs scene description 604, a driver's manual 610 containing traffic regulations for the current location, a movement plan received from the user that is "already in the left lane and going straight," and a situation received from the user that "a car suddenly appears on the right," into a driver assistance language model 130. Given such inputs, the language model driver assistance system 130 performs the above combination. Figure 3 and Figure 5 The described processing is used to (1) extract the portion of the driver's manual 610 containing keywords related to scenario description 604, plan 606, and situation 608; and (2) prompt the language model to generate an updated plan 610, "Slow down, check your rearview mirror, and give way to oncoming traffic on your right," which takes into account the traffic rules in the extracted portion of the driver's manual. The navigation application 500 can then output the updated plan 610 in any technically feasible manner, such as text instructions displayed via a display device and / or audio instructions output via one or more speaker devices.

[0162] Figure 7 This is a flowchart of method steps for providing driving assistance to an autonomous vehicle or user according to various embodiments. Although combined... Figures 1 to 5 These method steps have been described, but those skilled in the art will understand that any system configured to perform these method steps in any order falls within the scope of this disclosure.

[0163] As shown in the figure, method 700 begins at step 702, in which the language model driving assistance system 130 receives a scene description, a motion plan, and a status. The scene description is a description of the current environment from the vehicle's perspective. The motion plan is the vehicle's nominal execution plan. The status is the situation the vehicle is encountering. In some embodiments, the scene description, motion plan, and status may be received in text format. For example, in some embodiments, the scene description, motion plan, and status may be input as text by a user. As another example, in some embodiments, the scene description, motion plan, and status may be text generated from the user's speech via any technically feasible speech-to-text technology. As yet another example, in some embodiments, one or more of the scene description, motion plan, and status may be generated using a machine learning model. For example, in some embodiments, the scene description and / or status may be generated by a visual language model provided with one or more images of the vehicle's surroundings as input, the motion plan may be generated by a language model specifically trained to generate motion plans and corresponding trajectories in text format, and so on, as described above. Figures 5 to 6 As stated above.

[0164] In step 704, the language model driving assistance system 130 extracts one or more portions of local traffic rules based on the scene description, motion plan, and situation. In some embodiments, the language model driving assistance system 130 may extract one or more portions of local traffic rules by: generating a prompt requesting common traffic-related phrases, which includes a scene description, motion plan, and situation; processing the prompt using a trained language model to identify one or more keywords; and extracting one or more paragraphs containing the one or more keywords from the local traffic rules, as described below. Figure 8 Described in more detail. In some other embodiments, the language model driving assistance system 130 can extract (one or more) portions of local traffic rules in any technically feasible manner (e.g., using embedding search).

[0165] In step 706, the language model driving assistance system 130 generates a prompt requesting a driving instruction. In addition to the driving instruction request, the prompt includes a scenario description, a motion plan, a status, and (one or more) portions of traffic rules extracted from local traffic rules in step 704.

[0166] In step 708, the language model driving assistance system 130 processes the cues using a trained language model to generate an updated motion plan. In some embodiments, the language model driving assistance system 130 may input the cues generated in step 806 into the trained language model. Given such input, the trained language model can output driving instructions, which the language model driving assistance system 130 can use as an updated motion plan. In some embodiments, the trained language model may be a pre-trained LLM that is not retrained. In some embodiments, additional training data may be used to fine-tune the trained language model to generate driving instructions.

[0167] In step 710, the language model driving assistance system 130 generates driving instructions based on the updated motion plan. In some embodiments, any suitable driving assistance can be generated in any technically feasible manner. For example, in some embodiments, driving instructions can be generated and output as text and / or as audio, with the text output to the user via a display device and the audio output to the user via one or more speaker devices. As another example, in some embodiments, driving instructions can be generated as text used by other software and / or hardware, such as an AV motion planner that uses the driving instructions to generate and thus control the planned motion and / or trajectory of the vehicle.

[0168] Figure 8 This is a flowchart of method steps for extracting a portion of local traffic rules in step 704 of method 700, according to various embodiments. Although these method steps are combined... Figures 1 to 5 The methods described herein are for informational purposes only, but those skilled in the art will understand that any system configured to perform these method steps in any order falls within the scope of this disclosure.

[0169] As shown in the figure, in step 802, the language model driving assistance system 130 generates a prompt requesting common traffic-related phrases. This prompt requests common traffic-related phrases from the scene description, movement plan, and situation received in step 702, which may also be included in the prompt.

[0170] In step 804, the language model driving assistance system 130 processes the prompt using a trained language model to determine one or more keywords. In some embodiments, the language model driving assistance system 130 may input the prompt generated in step 802 into the trained language model. Given such input, the trained language model outputs one or more keywords containing common traffic-related phrases requested by the prompt.

[0171] In step 806, the language model driving assistance system 130 extracts one or more paragraphs containing the keyword from local traffic rules. In some embodiments, the language model driving assistance system 130 may search for the keyword in local traffic rules. In this case, the language model driving assistance system 130 may extract the paragraphs containing the keyword from the local traffic rules.

[0172] In summary, techniques for generating driving instructions for autonomous vehicles or users are disclosed. In some embodiments, a language model-based driving assistance system receives text as input describing a scenario, a plan for driving the vehicle, the current situation, and traffic rules describing the geographical location of the vehicle. Given such input, the language model-based driving assistance system extracts one or more portions of the traffic rules that relate to the scenario description, plan, and / or situation. For example, in some embodiments, the language model-based driving assistance system may prompt a trained language model to extract one or more keywords containing common traffic-related phrases from the scenario description, plan, and / or current situation. In this case, the language model-based driving assistance system may also search for the keyword(s) in the traffic rules and then extract the portions (e.g., paragraphs) of those traffic rules that contain the keyword(s). The language model-based driving assistance system also prompts a trained language model to generate an updated plan for driving the vehicle that takes into account the portion(s) of the traffic rules. The updated plan may be output to the user as a driving instruction via a display device and / or a speaker. Alternatively, the updated plan may be used to update an automatically generated motion plan, which may then be applied to control the vehicle.

[0173] Compared to existing technologies, at least one technical advantage of the disclosed technology is that it enables vehicles controlled using trained machine learning models to adapt to traffic rules in different geographical locations. Therefore, when implemented to control a vehicle using a trained machine learning model, the disclosed technology allows for safer driving by adhering to local traffic rules compared to what is typically achievable using traditional machine learning models. Alternatively, when implemented to respond to user input, the disclosed technology enables users to drive more safely and comply with local traffic rules. These technical advantages represent one or more technical improvements over existing methods.

[0174] 1. In some embodiments, a computer-implemented method for controlling a vehicle includes: receiving first text, the first text comprising a description of a scenario and a first plan for driving the vehicle; extracting at least a portion of a traffic rule set based on the description of the scenario and the first plan; generating a first prompt, the first prompt requesting driving instructions and including the description of the scenario, the first plan, and the at least a portion of the traffic rule set; processing the first prompt via a first trained language model to generate a second plan for driving the vehicle; and generating driving instructions based on the second plan.

[0175] 2. The computer-implemented method according to Clause 1 further includes: receiving a second text indicating a situation, wherein the extraction of at least a portion of the traffic rule set is also based on the situation, and wherein the prompt further includes the situation.

[0176] 3. The computer-implemented method according to Clause 1 or 2, wherein extracting the at least portion of the traffic rule set comprises: generating a second prompt, the second prompt requesting a traffic phrase and including the description of the scenario and the first plan; processing the second prompt via the first trained language model to extract one or more keywords from the description of the scenario and the first plan; and extracting one or more paragraphs from the traffic rule set, the one or more paragraphs including at least a first keyword contained in the one or more keywords.

[0177] 4. A computer-implemented method according to any one of clauses 1-3, wherein the driving instructions are transmitted to a trained planning model, and further comprising: generating one or more trajectories of the vehicle from the trained planning model and based on the driving instructions; and causing one or more operations for controlling the vehicle to be performed based on the one or more trajectories.

[0178] 5. The computer-implemented method according to any one of the clauses 1-4 further includes transmitting the driving instructions to the driver via a speaker device.

[0179] 6. The computer-implemented method according to any one of the clauses 1-5 further includes displaying the driving instructions to the driver via a display device.

[0180] 7. The computer-implemented method according to any one of the clauses 1-6 further includes processing image data associated with the vehicle using a visual language model to generate the description of the scene.

[0181] 8. The computer-implemented method according to any one of the provisions 1-7 further includes generating the first plan via a second trained language model based on second text indicating at least one of sensor data or status information associated with the vehicle.

[0182] 9. The computer-implemented method according to any one of the clauses 1-8, wherein the set of traffic rules is included in a driver's manual.

[0183] 10. The computer-implemented method according to any one of the clauses 1-9, wherein the first trained language model includes a trained large language model.

[0184] 11. In some embodiments, one or more non-transitory computer-readable media storing instructions that, when executed by at least one processor, cause the at least one processor to perform the following steps: receiving first text, the first text comprising a description of a scenario and a first plan for driving a vehicle; extracting at least a portion of a traffic rule set based on the description of the scenario and the first plan; generating a first prompt, the first prompt requesting driving instructions and including the description of the scenario, the first plan, and the at least a portion of the traffic rule set; processing the first prompt via a first trained language model to generate a second plan for driving the vehicle; and generating driving instructions based on the second plan.

[0185] 12. One or more non-transitory computer-readable media as described in Clause 11, wherein, when executed by the at least one processor, the instructions further cause the at least one processor to perform the step of receiving a second text indicating a situation, wherein the extraction of the at least a portion of the traffic rule set is also based on the situation, and wherein the prompt further includes the situation.

[0186] 13. One or more non-transitory computer-readable media as described in Clause 11 or 12, wherein extracting said at least a portion of the traffic rule set comprises: generating a second prompt, the second prompt requesting a traffic phrase and including the description of the scenario and the first plan; processing the second prompt via the first trained language model to extract one or more keywords from the description of the scenario and the first plan; and extracting one or more paragraphs from the traffic rule set, said one or more paragraphs including at least a first keyword contained in said one or more keywords.

[0187] 14. One or more non-transitory computer-readable media according to any one of clauses 11-13, wherein the driving instructions are transmitted to a trained planning model, and when executed by the at least one processor, the instructions further cause the at least one processor to perform the following steps: generating one or more trajectories of the vehicle from the trained planning model and based on the driving instructions; and causing one or more operations for controlling the vehicle to be performed based on the one or more trajectories.

[0188] 15. One or more non-transitory computer-readable media according to any one of clauses 11-14, wherein, when executed by the at least one processor, the instructions further cause the at least one processor to perform the step of: transmitting the driving instructions to the driver via an audio signal output from a speaker device.

[0189] 16. One or more non-transitory computer-readable media according to any one of clauses 11-15, wherein, when executed by the at least one processor, the instructions further cause the at least one processor to perform the step of: transmitting the driving instructions to the driver via text displayed by a display device.

[0190] 17. One or more non-transitory computer-readable media according to any one of clauses 11-16, wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of: processing image data associated with the vehicle using a visual language model to generate the description of the scene.

[0191] 18. One or more non-transitory computer-readable media according to any one of clauses 11-17, wherein the instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of: generating the first plan based on second text indicating at least one of sensor data or status information associated with the vehicle via a second trained language model.

[0192] 19. One or more non-transitory computer-readable media as described in any of clauses 11-18, wherein the first text and the first plan are received from a user or at least one of one or more trained machine learning models.

[0193] 20. In some embodiments, a system includes: one or more memories storing instructions; and one or more processors coupled to the memories and configured, upon executing the instructions, to: receive first text comprising a description of a scenario and a first plan for driving a vehicle; extract at least a portion of a set of traffic rules based on the description of the scenario and the first plan; generate a first prompt requesting driving instructions and including the description of the scenario, the first plan, and the at least a portion of the set of traffic rules; process the first prompt via a first trained language model to generate a second plan for driving the vehicle; and generate driving instructions based on the second plan.

[0194] Any element of any claim and / or any combination of any element described in this application, in whatever manner, is within the scope of this disclosure and its intended protection.

[0195] The descriptions of various embodiments are presented for illustrative purposes and are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations can be readily made by those skilled in the art without departing from the scope and spirit of the described embodiments.

[0196] Various aspects of this embodiment may be embodied as a system, method, or computer program product. Therefore, various aspects of this disclosure may take the form of a completely hardware embodiment, a completely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which are generally collectively referred to herein as a “module” or a “system.” Furthermore, various aspects of this disclosure may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.

[0197] Any combination of one or more computer-readable media may be used. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, such as, but not limited to, any one of the foregoing, or a suitable combination of any of the foregoing. More specific examples (not an exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer floppy disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or a suitable combination of any of the foregoing. In this document, a computer-readable storage medium may be any tangible medium capable of containing or storing a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0198] Various aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams illustrating methods, apparatus (systems), and computer program products of various embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine. When these instructions are executed by the processor of the computer or other programmable data processing apparatus, the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams are implemented. Such processors can be, but are not limited to, general-purpose processors, special-purpose processors, application-specific processors, or field-programmable gate arrays.

[0199] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, code segment, or portion of code, containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the order in which the functions marked in the blocks are executed may differ from the order shown in the figures. For example, two blocks shown consecutively in the figures may actually execute substantially simultaneously, or, depending on the functions involved, these blocks may sometimes execute in reverse order. It should also be noted that each block in the block diagrams and / or flowcharts, as well as combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware system or a combination of dedicated hardware and computer instructions that performs the specified function or action.

[0200] While the above description relates to embodiments of this disclosure, other embodiments of this disclosure may be designed without departing from the basic scope of this disclosure, the scope of which is defined by the following claims.

Claims

1. A computer-implemented method for controlling a vehicle, the method comprising: receiving first text that includes a description of a scenario and a first plan for driving a vehicle; extracting at least a portion of a traffic rule set based on the description of the scenario and the first plan; generating a first prompt that requests driving instructions and includes the description of the scenario, the first plan, and the at least a portion of the traffic rule set; processing the first prompt via a first trained language model to generate a second plan for driving the vehicle; and generating driving instructions based on the second plan. extracting the at least a portion of the traffic rule set is further based on a condition, and wherein the prompt further includes the condition.

2. The computer-implemented method of claim 1, further comprising receiving a second text indicating a condition, wherein, extracting the at least a portion of the traffic rule set comprises:

3. The computer-implemented method of claim 1, wherein, generating a second prompt that requests traffic phrases and includes the description of the scenario and the first plan; processing the second prompt via the first trained language model to extract one or more keywords from the description of the scenario and the first plan; and extracting one or more passages from the traffic rule set that include at least a first keyword included in the one or more keywords. the driving instructions are communicated to a trained planning model, and further include:

4. The computer-implemented method of claim 1, wherein, generating, by the trained planning model and based on the driving instructions, one or more trajectories of the vehicle; and causing one or more operations for controlling the vehicle to be performed based on the one or more trajectories.

5. The computer-implemented method of claim 1, further comprising communicating the driving instructions to a driver via a speaker device.

6. The computer-implemented method of claim 1, further comprising displaying the driving instructions to a driver via a display device.

7. The computer-implemented method of claim 1, further comprising processing image data associated with the vehicle using a visual language model to generate the description of the scenario.

8. The computer-implemented method of claim 1, further comprising generating, via a second trained language model, the first plan based on second text that indicates at least one of sensor data or state information associated with the vehicle. the traffic rule set is included in a driver’s manual.

9. The computer-implemented method of claim 1, wherein, the first trained language model includes a trained large language model.

10. The computer-implemented method of claim 1, wherein, 11. One or more non-transitory computer-readable media storing instructions that, when executed by at least one processor, cause the at least one processor to perform the steps of: receiving first text that includes a description of a scenario and a first plan for driving a vehicle; extracting at least a portion of a traffic rule set based on the description of the scenario and the first plan; generating a first prompt that requests driving instructions and includes the description of the scenario, the first plan, and the at least a portion of the traffic rule set; ​ processing the first prompt via a first trained language model to generate a second plan for driving the vehicle; and generating driving instructions based on the second plan.

12. The one or more non-transitory computer-readable media of claim 11, wherein, The instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of receiving second text indicative of a situation, wherein the extracting the at least a portion of the traffic rules set is further based on the situation, and wherein the prompt further includes the situation.

13. The one or more non-transitory computer-readable media of claim 11, wherein, The extracting the at least a portion of the traffic rules set includes: generating a second prompt requesting a traffic phrase and including the description of the scenario and the first plan; processing the second prompt via the first trained language model to extract one or more keywords from the description of the scenario and the first plan; and extracting one or more passages from the traffic rules set that include at least a first keyword contained in the one or more keywords.

14. The one or more non-transitory computer-readable media of claim 11, wherein, The driving instructions are communicated to a trained planning model, and the instructions, when executed by the at least one processor, further cause the at least one processor to perform the steps of: generating, by the trained planning model and based on the driving instructions, one or more trajectories of the vehicle; and causing one or more operations for controlling the vehicle to be performed based on the one or more trajectories.

15. The one or more non-transitory computer-readable media of claim 11, wherein, The instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of communicating the driving instructions to a driver via an audio signal output by a speaker device.

16. The one or more non-transitory computer-readable media of claim 11, wherein, The instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of communicating the driving instructions to a driver via text displayed by a display device.

17. The one or more non-transitory computer-readable media of claim 11, wherein, The instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of processing image data associated with the vehicle using a visual language model to generate the description of the scenario.

18. The one or more non-transitory computer-readable media of claim 11, wherein, The instructions, when executed by the at least one processor, further cause the at least one processor to perform the step of generating, via a second trained language model, the first plan based on second text indicative of at least one of sensor data or state information associated with the vehicle.

19. The one or more non-transitory computer-readable media of claim 11, wherein, The first text and the first plan are received from at least one of a user or one or more trained machine learning models.

20. A system comprising: one or more memories storing instructions; and one or more processors coupled with the one or more memories and, when executing the instructions, configured to: receive first text containing a description of a scenario and a first plan for driving a vehicle; extract at least a portion of a traffic rules set based on the description of the scenario and the first plan; generating a first prompt requesting driving instructions and including the description of the scene, the first plan, and the at least a portion of the traffic rules set; processing the first prompt via a first trained language model to generate a second plan for driving the vehicle; and generating driving instructions based on the second plan.

Citation Information

Patent Citations

  • Method for programmable timeouts of tree traversal mechanisms in hardware

    US10885698B2