Machine operation assistance using language model to enhance operator monitoring

By using large language models to provide natural language responses in vehicles, the driving risk problem caused by ignoring useful information in the prior art is solved, and the effect of simplifying the driving experience and reducing cognitive burden is achieved.

CN119928899APending Publication Date: 2025-05-06NVIDIA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411543605.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-01
Filing Date
2024-10-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing vehicle technology ignores potential useful information that may simplify the driving experience and reduce the operator's cognitive burden during driving, resulting in increased distractions, unsafe driving conditions and dangerous obstacles, thereby increasing safety risks and accidents.

Method used

By using large language models (LLMs) to provide context-specific natural language responses, it provides operators with rich information, such as natural language texts of traffic signs, weather data, traffic data, etc., to help operators better understand environmental conditions and reduce cognitive burdens.

Benefits of technology

This approach can simplify the driving experience, reduce the operator's cognitive burden, reduce the risk of distraction and unsafe driving, and thus reduce accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119928899A_ABST
    Figure CN119928899A_ABST
Patent Text Reader

Abstract

The invention discloses machine operation assistance using a language model to enhance operator monitoring. Various embodiments of the present disclosure relate to operator assistance based on operator monitoring. For example, during long-distance driving, the driver may be drowsy or may be unalerted due to other reasons. Thus, particular embodiments have the ability to initiate a conversation with the driver based on driver interest and / or detecting whether the driver is drowsy. In illustrative examples, a driver monitoring system (DMS) camera of a vehicle may employ components that acquire pixel-level information that displays nodding, hand sagging, and the like. Based on image pattern features in the image data, a particular embodiment generates a score representing a level of alertness. A representation of the alertness level may be provided as input to a machine learning model so that the model may generate appropriate natural language or other responses, such as starting a conversation with personalized questions and answers, sending control signals to horn, etc.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is related to U.S. patent application No. 18 / 499,913 filed on November 1, 2023, entitled "Machine Operation Assistance Using Language Model-Augmented Perception." Background Art

[0003] Vehicles may be equipped with technology that makes driving safer or more convenient. For example, a car virtual assistant can understand and respond to simple voice commands. This allows the driver to control various functions of the car without taking their hands off the steering wheel or their eyes off the road. In addition, the driver can use the car virtual assistant to control the car's entertainment system. This includes playing music from various sources such as streaming services, radio, or personal devices, as well as adjusting the volume and switching between tracks. Other features of vehicle technology include providing navigation assistance, phone integration that allows the driver to make calls or text messages hands-free, and smart home integration, so that the driver can control home smart devices from the car.

[0004] One of the main drawbacks of these and other conventional technologies is that driving typically requires a high level of concentration to perceive and respond to stimuli in the environment, which often requires a considerable cognitive load to safely operate the vehicle. Distractions, unsafe driving conditions, and / or hazardous obstacles can create safety risks and accidents. However, by limiting the information that conventional technologies provide to the operator, these technologies ignore potentially useful information that could simplify the driving experience and reduce the cognitive load on the operator, thereby potentially limiting or even interfering with the operator's ability to safely operate the vehicle in the environment. Summary of the invention

[0005] Embodiments of the present disclosure relate to operator (e.g., driver) assistance using one or more language models (e.g., large language models (LLMs)). For example, some embodiments relate to providing context-specific information to an operator as part of a natural language conversation or other natural language output. In an illustrative example, a particular embodiment generates a natural language utterance (e.g., "Be warned, there are a lot of deer in the area") based on extracting natural language text from a nearby traffic sign (e.g., a sign that reads "Deer Xing"). Additionally or alternatively, some embodiments relate to interacting with an operator based on operator monitoring (e.g., via a relevant conversation). For example, during a long drive, a driver may become drowsy or may not be alert. Therefore, particular embodiments have the ability to interact with the driver (e.g., start or continue a conversation) based on driver interest and / or in response to detecting the possibility that the driver is drowsy.

[0006] Compared to the above conventional systems, such as the systems described above, various embodiments provide the operator with rich information, such as natural language responses generated by a language model. Such natural language responses can simplify the driving or operating experience and reduce the cognitive burden on the operator. This allows the operator to safely guide the ego-machine through the environment. In this way, distractions, unsafe driving conditions, and / or dangerous obstacles can be reduced, resulting in fewer safety risks and accidents.

[0007] In various example embodiments, such natural language responses may include, for example, but not limited to, natural language text summaries of event data (e.g., weather data, traffic data, radio broadcast data) based on the geographic location of the machine, a generated natural language sentence indicating whether a parking space is available, a generated natural language sentence indicating whether the operator saw or missed a traffic sign, a natural language text summary of information about the destination and / or route of travel, a natural language response based on the operator's alertness level, a natural language response based on the operator's interests, and / or a natural language response to an operator's utterance.

[0008] In order to generate a natural language response, the language model may ingest or receive various inputs (or portions of inputs). For example, the input may be or include various prompts that have been prompt engineered, prompt adjusted, and / or fine-tuned to elicit appropriate natural language responses using a large language model (LLM). For example, when the natural language response relates to the operator's alertness level, the prompt provided to the language model may include personalized information (e.g., an indication that the driver likes sports team X), a detected (e.g., calculated, inferred, etc.) representation of the operator's alertness level (e.g., "the operator is very sleepy"), a natural language instruction (e.g., "start a conversation with the driver that is in line with the driver's interests"), a one-shot or few-shot prompt example of a representative input and / or output (e.g., an example conversation initiated at the KSS level), and / or "send a control signal to honk the horn". In response to the language model ingesting such a prompt, the language model may output a natural language response, such as "I just honked the horn because you were falling asleep. Can you name some players who have played for Team X and entered the Hall of Fame?" BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The present system and method for subcutaneous authentication are described in detail below with reference to the accompanying drawings, in which:

[0010] Figure 1 is a data flow diagram illustrating an example natural language extraction pipeline according to some embodiments of the present disclosure;

[0011] Figure 2 is a data flow diagram illustrating an example alertness level detector pipeline according to some embodiments of the present disclosure;

[0012] Figure 3 is a block diagram of a large language model that uses specific inputs to generate specific natural language responses according to some embodiments;

[0013] Figure 4A shows image data of a stop sign and a parking space 408 according to some embodiments;

[0014] Figure 4B According to some embodiments, Figure 4A Traffic signs and / or Figure 4A A natural language response of natural language characters extracted from the parking space;

[0015] Figure 5A showing a present machine operator in a drowsy state and a responsive natural language response according to some embodiments;

[0016] Figure 5B According to some embodiments Figure 5A The operator of Figure 5A The response of the audio frequency response and the model response of the operator's response 5;

[0017] Figure 6 is a flow chart of an example process for generating a natural language response based on extracting natural language characters represented in image data according to some embodiments;

[0018] Figure 7 is a flow chart of an example process for extracting natural language characters from one or more objects according to some embodiments;

[0019] Figure 8 is a flow chart of an example process for generating a natural language response based on a calculated alertness level of an operator of a host machine according to some embodiments;

[0020] Fig. 9 is a flow chart of an example process for generating natural language characters based on providing different image data sets to the LLM according to some embodiments;

[0021] Fig. 10A is an illustration of an example autonomous vehicle according to some embodiments of the present disclosure;

[0022] Fig. 10B According to some embodiments of the present disclosure Fig. 10A Examples of autonomous vehicle camera positions and fields of view;

[0023] Fig. 10C According to some embodiments of the present disclosure Fig. 10A a block diagram of an example system architecture for an example autonomous vehicle;

[0024] Fig. 10D According to some embodiments of the present disclosure, a method for Fig. 10A System diagram of an example of communication between autonomous vehicles;

[0025] Fig.11 is a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and

[0026] Fig.12 is a block diagram of an example data center suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0027] Embodiments of the present disclosure relate to operator (e.g., driver) assistance using one or more language models (e.g., large language models (LLMs)). Although the present disclosure may be related to an example autonomous or semi-autonomous vehicle or machine 1000 (alternatively referred to herein as “vehicle 1000” or “machine 1000”), the example Figures 10A-10D For example, the systems and methods described herein may be used by non-autonomous vehicles or machines, semi-autonomous vehicles or machines (e.g., in one or more advanced driver assistance systems (ADAS)), autonomous vehicles or machines, manned and unmanned robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, aircraft, ships, space shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, engineering vehicles, trains, underwater vehicles, remotely controlled vehicles (e.g., drones), and / or other vehicle types, but are not limited thereto. In addition, although the present disclosure may be described with respect to a model that generates a natural language response based on extracting natural language information from an object and / or detecting an operator's alertness level, this is not limiting, and the systems and methods described herein may be used in augmented reality, virtual reality, mixed reality, robotics, security and surveillance, autonomous or semi-autonomous machine applications, and / or any other technology field where identity authentication may be used.

[0028] Various embodiments of the present disclosure relate to providing context-specific information to a driver or other operator as a natural language response (e.g., one or more characters, words, etc.). For example, systems and methods are disclosed for extracting natural language characters (e.g., English phrases) from image data representing an object (e.g., a traffic sign, a billboard, a convenience store) and generating a responsive natural language output (e.g., "You cannot park at this location because it is not yet 6 p.m."). In an illustrative example, one or more sensors (e.g., one or more cameras) mounted on a machine (e.g., a vehicle) may capture image data including representations of traffic signs in an environment. The sensor may detect one or more regions of the image data in which the traffic sign is detected via object detection or segmentation. In some cases, a speed sensor (e.g., an inductive sensor) in a vehicle may detect that the vehicle is below and / or exceeds a threshold speed. In response to the vehicle being below and / or exceeding a speed threshold and / or the detection / segmentation of a traffic sign, a particular embodiment extracts natural language characters within one or more regions of the image data representing the detected traffic sign. For example, optical character recognition (OCR) may be used to encode elements of the image data into machine-readable natural language text (e.g., via pattern matching or feature extraction methods). Such natural language text may be, for example, "No parking allowed between 8 a.m. and 6 p.m.," derived from a parking traffic sign. Thus, the OCR function or any other extraction of natural language characters may be triggered in response to any suitable event, such as detecting (e.g., via object detection) an object and / or determining whether the machine is under and / or over a speed threshold.

[0029] Continuing with the example, in response to detecting natural language characters from a traffic sign, particular embodiments provide such natural language characters as input to one or more machine learning models, such that the machine learning models generate other natural language characters associated with the environment as output (e.g., a natural language response indicating whether the current context permits parking in a detected parking space associated with the detected traffic sign) based at least on the detection of the natural language characters represented in the traffic sign. For example, the machine learning model may include a large language model (LLM), wherein a representation of the detected natural language characters may be included in a prompt of the LLM. The prompt typically includes a natural language instruction (e.g., a query) and / or other natural language information, which causes the LLM to return a responsive natural language output. For example, using the above example, a prompt may include one or more of the following: a zero-shot, one-shot, or few-shot example of representative input and / or output, entity data (e.g., a label) describing a particular entity in the natural language characters contained in the traffic sign (e.g., performed via named entity recognition (NER)), a hierarchical data structure (e.g., a waterfall model) representing multiple features of the environment (e.g., each road and its connecting roads and traffic signs), the time of day when the element representing the stop sign was detected, an indication of whether the parking space is available (e.g., determined via object detection or segmentation of visual data representing the parking space), and / or a query (e.g., "Can I park in this parking space?"). The LLM may then ingest the prompt and responsively generate an output of natural language characters and / or words based on the confidence interval. Such output may be generated based on the LLM having been prompt-adjusted, prompt-engineered, and / or fine-tuned to generate such output.

[0030] A representation of the LLM output (e.g., a visual indication of whether the detected traffic sign allows parking in the then-detected parking space) may then be presented on a sound device or display device (e.g., an infotainment display) visible to an operator or occupant of the vehicle. The LLM output (e.g., the generated natural language characters) may initially be in a text format. Thus, for example, particular embodiments may process the natural language characters using a text-to-speech module in order to present the natural language characters as audio data to a virtual assistant device built into the vehicle. For example, using the stop sign illustration above, the audio output of the sound device may be an utterance representing the generated natural language characters stating: "I'm sorry, but it looks like parking is not allowed at this location at this time. Although there is no car in the parking space, it is now 2 p.m., which is earlier than the designated time (6 p.m.) when parking is permitted according to the stop sign."

[0031] Some systems and methods of the present disclosure also involve attracting the driver by presenting one or more generated natural language characters in response to detecting the driver's alertness level or other events. For example, a driver monitoring system (DMS) camera in the steering column of a vehicle can employ an eye tracking module that emits infrared light that reflects in the driver's eyes. This reflection can be represented in image data picked up by a camera in the vehicle, where a specific embodiment can, for example, detect pupil position, eye gaze, etc. This image data can alternatively or additionally include other information of any other part of the driver, such as pixel-level information showing whether the driver's eyes are open or closed, such as nodding, hand drooping, or the driver's eyes. Based on image pattern features in the image data, a specific embodiment generates a score representing a first alertness level. For example, a model (e.g., a convolutional neural network) can be fine-tuned to classify different alertness levels of the driver based on different positions of the driver indicated in the image. For example, the model can adjust the weights of the model during training to indicate that when the driver closes his eyes, this is the lowest alert level. This can be based on labeling several training images of the driver with their eyes closed as "minimum alert level 1" or other similar classifications.

[0032] In some embodiments, such a classification or score is mapped (e.g., via a hand-coded data structure) to any suitable alertness scoring system (e.g., represented in natural language or some other coded input) so that it can be provided as input to one or more LLMs or other models. For example, some embodiments map image data classifications to the Karolinska Sleepiness Scale (“KSS”) via a data structure, which contains (e.g., textual) representations of all nine alertness levels—“(1) Extremely alert,” “(2) Very alert,” “(3) Alert,” “(4) Fairly alert,” “(5) Neither alert nor sleepy,” “(6) Some signs of sleepiness,” “(7) Sleepy, but no effort to stay awake,” “(8) Sleepy, some effort to stay awake,” and “(9) Sleepy, very much trying to stay awake, resisting sleep.” Some embodiments convert the calculated alert level into a corresponding (e.g., a natural language phrase or other) representation of the calculated alert level and provide the representation as input (e.g., at least a portion of the input) to a machine learning model so that the machine learning model can generate corresponding natural language characters (e.g., a natural language response indicating that the operator's KSS level is 7 and / or certain personalized natural language responses, such as personalized questions and answers, to attempt to lower the operator's KSS level).

[0033] In some embodiments, these generated natural language characters are processed using a text-to-speech module or provided to a text-to-speech module such that presentation of such output of natural language characters includes synthesizing audio data into a phrase utterance that initiates a conversation with the driver based at least on a score corresponding to an alertness level. For example, in response to detecting that the operator's alertness level is classified as a specified threshold or below / above a specified threshold (e.g., at KSS level 9 alertness), the phrase utterance may be "Warning: You are falling asleep. Would you like directions to the nearest hotel or rest area?" Some embodiments may (e.g., substantially simultaneously) send a control signal to a volume control in the vehicle to increase the volume of such utterance (e.g., at a predetermined decibel level) in order to interact with the driver based on the detected alertness level. Such control signals may additionally or alternatively cause other tangible actions, such as honking a horn, slowing down and / or stopping the car, etc.

[0034] Some embodiments cause presentation of a natural language output predicted by one or more machine learning models based on detecting whether the driver is not seeing a traffic sign. Such detection can be based on analyzing image data representing a traffic sign (e.g., image data captured by an object detection camera) for image data representing one or more portions of the driver (e.g., image data captured by a DMS camera). For example, as described above, an eye tracking module can be used to detect the driver's eye gaze in an image data frame associated with a particular time slice representing the time when the eye gaze was detected to determine the direction the driver was looking at at a particular timestamp. For example, an eye gaze may indicate that the driver is looking straight ahead at a particular time, but the detected traffic sign may only be visible outside the field of view associated with the detected gaze at that time (e.g., through a passenger-side window).

[0035] In addition, to facilitate detecting whether the driver (or other operator) noticed the detected traffic sign, a particular embodiment may programmatically call, for example, a routine that accesses a waterfall data structure and / or a geographic location indicator indicating the location of the vehicle in a particular time slice, the waterfall data structure indicating one or more (e.g., all) features in the driving environment and their locations (e.g., all connected road names, traffic signs, traffic lights, building structures, their orientations in space, etc.). In this way, the eye gaze direction and timestamp may be associated with and / or intersected with the orientation of the detected traffic sign (e.g., as indicated by the waterfall structure, the geographic location indicator, and the corresponding timestamp). Thus, using the above example, since the driver's eyes are detected looking forward in the relevant frame, the vehicle is determined to be near the traffic sign at the time slice (e.g., within a certain specified threshold distance), and / or the waterfall data structure indicates that the traffic sign detected at the time slice is located on the right side of the road or on the right side of the vehicle, so a particular embodiment may generate a score (e.g., a confidence level) representing the likelihood that the driver did not see the detected traffic sign. In response, particular embodiments may provide a second indicator of a second score (e.g., a natural language phrase) as at least part of the input to the machine learning model, and the resulting action (e.g., an output presenting corresponding natural language characters) may include a phrase representing a response to the driver not seeing the traffic sign. For example, a score indicating the likelihood that the driver did not see the traffic sign may initially be a binary classification (e.g., "0") or other value indicating that the driver likely did not see the traffic sign. Particular embodiments may then responsively map such a score to a corresponding natural language input, such as "The driver did not see the stop sign," via a hand-coded data structure (e.g., a hash map). Such a phrase may be placed in a prompt such that the LLM may make adjustments to the prompt to output a phrase, such as "You just missed the stop sign. Please be careful. I'll let you know before you get to the next stop sign."

[0036] Some embodiments may cause one or more machine learning models to output natural language characters based on a geolocation indicator. The geolocation indicator may represent a location of the machine detected in an environment or otherwise obtained. For example, in some embodiments, the geolocation indicator represents geographic coordinates derived from a GPS module coupled to a vehicle or user device of a driver or other operator, which captures or otherwise updates the geographic coordinates of the device in a cyclic manner (e.g., every Nth interval, such as every 2 seconds). In another example, the geolocation indicator represents or includes a natural language identifier representing a city, state, or other area. In some of these embodiments, such identifiers may be accessed in any suitable manner, such as object detection of real-world objects (e.g., a "Welcome to Arizona" sign), and then responsively extracting natural language characters contained in image data in which the real-world object is detected, wherein the extracted natural language characters may represent the extracted geolocation indicator.

[0037] Based on the geographic location indicator, a particular embodiment may determine weather data, road condition data, traffic data, and / or event data (e.g., accidents in a certain area) for the geographic location. For example, a particular embodiment may access one or more data structures indexed by a city or other geographic unit (e.g., a lookup data structure). Therefore, after detecting the city in which the vehicle is located, a particular embodiment may use the city identifier as an index in the data structure to read the corresponding value in the same record and / or access additional auxiliary data structures or services for other values. For example, based on the lookup city identifier, a particular embodiment may then conduct a network communication session with a computing node representing a weather, traffic, road condition, and / or other event data service. An embodiment may transmit a query that may include a request to return weather, traffic, and / or other event data for a specific city identifier. Application programming interface (API) logic may responsively obtain weather, traffic, and / or other event data and return it to the requesting computing node. In response, a particular embodiment may then provide weather, road condition, traffic data, and / or event data as additional or alternative input to a machine learning model so that the model presents additional or alternative natural language characters based on this input. For example, the initial input pulled from a weather service might be "City A, High: 90 degrees, Low: 80 degrees." Once ingested, the LLM might produce an output stating "The temperature in this area was high, with a high of 90 degrees."

[0038] In some embodiments, the geographic location indicator is additionally or alternatively used to generate a summary of the information contained in the radio station broadcast associated with a particular geographic area. For example, based on the geographic location indicator, a particular embodiment can access the first audio data associated with the radio station (e.g., a time series of accessing the audio data through a network (e.g., the Internet)). This time series can be determined in any suitable manner, such as extracting the first X seconds or identifying specific keywords (e.g., via voice detection) or detecting pauses (silence) to start and end access to the audio data. For example, in response to the recognition of the keyword "weather", a particular embodiment can responsively initiate the extraction of audio data and then end when there is a threshold time (or threshold pause duration) between sounds. A particular embodiment can then convert the first audio data into a text representation (e.g., a document with natural language text) via speech to text. A particular embodiment can then provide such a document (or a portion thereof) as an additional or alternative input to a machine learning model, wherein a model such as an LLM can then derive a summary of the relevant portion of the radio station broadcast.

[0039] Some embodiments generate a natural language response based on how the operator of the machine verbally responds to such natural language response. For example, using the above description, a particular embodiment receives a second phrase utterance representing a driver response, such as "No, I'll be fine" (spoken by the driver in response to the phrase utterance "Warning: You are falling asleep. Do you want directions to the nearest hotel or rest area?"). Subsequently, some embodiments receive another set of image data representing one or more parts of the operator, such as through the above-mentioned DMS system. Based at least on image pattern features of such image data and / or the second phrase utterance representing the driver's response, a particular embodiment generates a second score indicating a second alertness level of the driver. For example, a particular embodiment may use a Gaussian mixture model (GMM) or a hidden Markov model (HMM) to detect the driver's alertness level based on detecting speech patterns associated with alertness or non-alertness. For example, a GMM can be used to distinguish different speech data for a single driver, such as what sounds the driver makes when alert and non-alert. A GMM is a model that includes generative unsupervised learning or clustering capabilities. For a given data set (e.g., speech utterances), each data point (e.g., a single utterance of multiple phenotypes) is generated by linearly combining multiple "Gaussian" representations (e.g., multiple speech utterance sound distributions of the same user over time). A Gaussian representation is a distribution that is a list of viewing outcomes and the probability associated with each outcome. For example, a Gaussian representation may include a frequency value within a time window of a particular utterance received and a predicted frequency value within the next time window. The output is a predicted class, such as determining or predicting whether two different Gaussian distributions or utterances indicate alertness or drowsiness based on a baseline of driver utterances recorded when the driver is not drowsy and / or a baseline of driver utterances recorded when the driver is drowsy. These indications can then be mapped into a second score indicating an alert level.

[0040] In response, certain embodiments provide a second representation of the second score (e.g., another natural language phrase) as an input to the machine learning model, so that the model generates and presents additional natural language characters representing yet another phrase utterance in the conversation with the driver. For example, the output presented may be "You still sound a little tired. Should I call your wife if I may?" Such an output may be based on receiving input via a GMM that classifies the driver phrase as 1 (indicating very sleepy) and a hand-coded data structure that maps such classification to a natural language phrase (e.g., "The driver is very sleepy").

[0041] Some embodiments additionally or alternatively conduct a dialogue with the operator based on the accessed personalized information. For example, some embodiments access a quiz game from a database or other data storage based on the driver's interests (e.g., football quizzes). In some cases, before driving the vehicle, the driver may have uploaded and registered his interests to the database using natural language (e.g., "I really like football") for reference when driving later. In response, a specific embodiment uses this natural language personalized information as an input to a machine learning model so that the dialogue with the driver is additionally or alternatively based on the personalized information associated with the driver. For example, based on the driver's interest in football and based on detecting the driver's alertness level, this output can be "Do you know the name of the only football player who scored 10 times in a game?".

[0042] Thus, operator assistance may be provided based on the use of one or more language models (e.g., a large language model (LLM)) to generate natural language responses. Some embodiments involve providing context-specific information to an operator as part of a natural language dialog or other natural language output. Additionally or alternatively, some embodiments involve engaging an operator based on operator monitoring (e.g., via a relevant dialog).

[0043] refer to Figure 1 , Figure 1 is an example natural language extraction pipeline 100 according to some embodiments of the present disclosure. It should be understood that this arrangement and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used in addition to or in place of the arrangements and elements shown, and some elements may be omitted entirely. In addition, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in combination with other components, and may be implemented in any suitable combination and location. The various functions performed by the entities described herein may be performed by hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in a memory. In some embodiments, the systems, methods, and processes described herein may be used with Figures 10A-10D Example of autonomous vehicle 1000, Fig.11 The example computing device 1100 and / or Fig.12 The example data center 1200 may be implemented with similar components, features and / or functions.

[0044] exist Figure 1In the illustrated embodiment, the natural language extraction pipeline 100 includes an external perception camera 102, an object detector 106, a natural language extractor 108, a speed detection component 110, a geolocation component 112, an object context generator 120, a destination / travel route information extractor 122, an operator viewing probability component 128, one or more language models 126, a text-to-speech component 134, a display component 136, environmental image data 104, event data 114, and destination / travel route data 124.

[0045] The external perception camera 102 may be responsible for capturing one or more images (e.g., video clips) of an environment, such as the exterior of a machine (e.g., a vehicle, an aircraft, a watercraft (e.g., a ship), a drone, etc.). For example, a camera mounted on top of or inside a car may use optics to focus light onto an image sensor within the camera, which converts the light into electrical signals that are processed and stored as a digital image file. The camera may continuously capture and record specific images (also referred to as frames) as well as other images, and then present a sequence of images to create an indication of motion. For example, a camera may generate a sequence of image data of a driving environment, including streets, street signs, traffic lights, pedestrians, etc. The output of the external perception camera 102 is environmental image data 104. The environmental image data 104 may include one or more digital images or frames, such as a sequence of frames indicating a video feed.

[0046] The object detector 106 may be responsible for determining one or more objects within the environmental image data 104. In some embodiments, the object detector 106 performs its functions via object detection and / or semantic segmentation (e.g., panoptic segmentation). Semantic segmentation refers to the task of assigning and indicating (e.g., via a unique pixel-by-pixel mask color or ID) each pixel to a specific class of real-world objects or backgrounds represented in the input image. For example, the semantic segmentation function may define a first group of pixels of the environmental image data 104 as representing "birds" and a second group of pixels as also representing "birds," wherein both birds are represented by the same mask pixel value. In some embodiments, instance segmentation is performed in addition or alternatively. Instance segmentation may use a unique identifier to assign and define each pixel as an instance of a real-world object corresponding to the pixel. For example, using the above description, a first group of pixels representing a first bird may be assigned an instance ID of 1 and a first color mask pixel value. Similarly, a second group of pixels representing a second detected bird may be assigned an instance ID of 2 and / or a different mask color pixel value.

[0047] Semantic segmentation can be implemented using a deep learning algorithm that associates a label or category with each pixel in an image. Some embodiments label each pixel of an image with a corresponding class of the content represented. This is used to identify sets of pixels that form different categories. For example, a model can be trained to mask objects using the pixel values ​​of vehicles, pedestrians, traffic signs, sidewalks, or other road features. For example, a convolutional neural network (CNN) can perform image-related functions at each layer and then downsample the image using a pooling layer (e.g., green). This process is repeated multiple times for the first half of the network. The output of the first half of the network can be followed by an equal number of anti-pooling layers (e.g., orange).

[0048] In some embodiments, semantic segmentation may be performed via panoptic segmentation. The combination of semantic segmentation and instance segmentation is what is known as panoptic segmentation. Specifically, in panoptic segmentation, some or all pixels of an image may be uniquely assigned to one of the background classes or one of the object instances. For object instances, the panoptic segmentation function may thus classify each pixel in the image as belonging to a particular class and identify which instance of that class the pixel belongs to. For background classes, panoptic segmentation may perform a similar function to semantic segmentation.

[0049] In some embodiments, the object detector 106 additionally or alternatively performs object detection or classifies one or more objects in the input image. In an illustrative example of the object detection function, a particular embodiment uses one or more machine learning models (e.g., CNN) to generate a bounding box that defines the boundaries and contains computer objects representing features in the image (e.g., cars, sky, buildings, people, etc.). These machine learning models can generate classification predictions, that is, computer objects are specific features. In computer vision applications, the output of object detection can be contained by a bounding box or other bounding shape. In terms of the location (e.g., 2D or 3D coordinates) of the bounding box (and the height and width of the bounding box), the bounding box can contain the predicted boundaries of the object. For example, the bounding box can be a rectangular box, which is determined by its X-axis and Y-axis coordinates. This provides an indication of the spatial distinction between objects to the object recognition system to help detect objects in the image. In an illustrative example, on an image, a first bounding box containing a car in the image can be generated and labeled as "car", a second bounding box containing a traffic sign can be generated and labeled as "traffic sign", and a third bounding box containing a mountain object can be generated and labeled as "mountain".

[0050] In some embodiments, one or more machine learning models can be used and trained to generate tighter bounding boxes for each object. In this way, the shape of the bounding box and the confidence level of the classification / prediction may change and may increase with increased training sessions. For example, the output of a CNN or any other machine learning model described herein can be one or more bounding boxes on each object feature (corresponding to a feature in an image), where each bounding box includes a classification prediction (e.g., the object is a building) and a confidence level (e.g., 90% probability).

[0051] In some embodiments, object detector 106 additionally or alternatively performs image classification, object recognition, keypoint detection, edge detection, and / or other functions for identifying different features or objects in an image. For example, with respect to image classification, embodiments may perform pixel-based classification (e.g., minimum-distance-to-mean, maximum likelihood, and minimum-Mahalanobis-distance) or object-based classification to classify the entire image (without determining location information, such as a bounding box). For example, some embodiments perform pre-processing operations, such as converting an image to a vector or matrix, where each value (e.g., an integer or floating point number) represents a corresponding pixel value in the image. In some embodiments, such as in a K-nearest neighbor (KNN) use case, particular embodiments determine the distance between such vectors and other vectors representing training images, where the closest vector indicates that a group of pixels (or the entire image) corresponds to a certain class.

[0052] The speed detection component 110 may be responsible for detecting the speed or velocity of the host machine and returning an indication of the speed to the natural language extractor 108. For example, the speed detection component 110 may be included in wheel speed sensors (ABS sensors), vehicle speed sensors (VSS), global positioning systems (GPS), radar and lidar, wheel tachometers, etc. to detect the speed of the host machine (e.g., in miles per hour).

[0053] The natural language extractor 108 may be responsible for taking the determined object from the object detector 106 and / or the speed indication from the speed detection component 110 as input, thereby triggering the extraction of one or more natural language characters from the environmental image data 104. For example, using the speed data from the speed detection component 110 as input, the natural language extractor 108 may determine whether the vehicle is below and / or exceeds a threshold speed (e.g., above 5 miles / hour and / or below 50 miles / hour). This selective extraction ensures that image data processing resources are not consumed unnecessarily. For example, if the vehicle is completely stopped, it may not be necessary to extract text from the object. Similarly, if the vehicle is traveling too fast, the embodiment does not extract characters because it is more likely that extraction errors will occur and / or because the extraction may not be so useful because the driver may have passed the corresponding object before obtaining the returned extraction data. At least partially in response to the vehicle being below and / or exceeding the speed threshold, the natural language extractor 108 may extract natural language characters in one or more regions of the image data representing the detected object. Additionally or alternatively, using the image data returned by the object detector 106, the natural language extractor 108 may determine (e.g., via edge detection) whether there are natural language characters represented in such an object. At least in part in response to determining that there are natural language characters represented in such an object, the natural language extractor 108 may extract the natural language characters within one or more regions of the image data representing the detected object.

[0054] The natural language extractor 108 may use any suitable function to extract natural language characters from an object. For example, optical character recognition (OCR), handwritten text recognition (HTR), intelligent character recognition (ICR), template matching, rule-based information extraction, and / or other suitable functions may be used. For example, the natural language extractor 108 may represent an OCR function that encodes elements of image data into machine-readable natural language text via a pattern matching or feature extraction method. In some embodiments, OCR includes the following functions: The OCR component may perform an image quality function to change the appearance of the image data 104 by converting one or more color image frames to grayscale, performing desaturation (removing color), changing brightness, changing contrast to obtain contrast correctness, etc. In response, the OCR component may perform a computer process that rotates one or more image frames to a uniform orientation, which is called "deskewing" the image. In some cases, the image frame is slightly rotated or flipped at various angles (e.g., 45°, 90°, etc.) in a vertical or horizontal plane. Therefore, some embodiments de-skew the image to change the orientation of the image to achieve a uniform orientation (e.g., a straight edge profile or a horizontal orientation). In some embodiments, in response to the deskew operation, some embodiments remove background noise (e.g., via Gaussian and / or Fourier transforms). In many cases, one or more image frames contain unnecessary dots or other marks. In order to isolate from the interference of such meaningless noise, some embodiments clean up the image by removing these marks. In response to removing the background noise, some embodiments extract natural characters from the image and place the extracted characters in another format (e.g., JSON). The format (e.g., JSON) can be used as input to other machine learning models (e.g., language model 126 (e.g., LLM)), as described in more detail below.

[0055] The geolocation component 112 may be responsible for generating a geolocation indicator representing the location of the machine in the environment and determining event data associated with the geolocation indicator, such as weather data, road condition data, traffic data, and / or any other event data (e.g., an indication of a traffic accident). The geolocation component 112 includes a geolocation indicator generator 116 and an event data retriever 118. The geolocation indicator generator 116 may be responsible for generating or determining a geolocation indicator and returning it to the event data retriever 118 for further processing. In some embodiments, the geolocation indicator generator 116 generates the geolocation indicator via any suitable method. For example, the machine (and / or a user device of an operator of the machine) may have a built-in GPS module that triangulates or uses trilateration to determine its coordinates based on signals received from one or more satellites, where the coordinates represent the geolocation indicator.

[0056] Alternatively or additionally, a speech-to-text component in the machine may encode audio data representing a radio station and / or user utterance into natural language characters, wherein the natural language characters may indicate the location of the machine (e.g., "Wow, the Grand Canyon is more beautiful than I thought."). In response, natural language processing (NLP), such as NER and semantic analysis, may be performed to determine that the natural language characters indicate the geographic area where the user is located, wherein the natural language characters represent a geographic location indicator. Alternatively or additionally, the geographic location indicator may be generated using beacon technology, magnetic positioning, dead reckoning, etc.

[0057] The event data retriever 118 may be responsible for retrieving any suitable event data 114 based on using as input the generated geographic location indicator from the geographic location indicator generator 116. For example, using the geographic location indicator as an index in a lookup data structure, the event data retriever 118 may locate one or more corresponding records in the event data 114 that specify weather data (e.g., high and low pressure, wind chill, etc.), road condition data (e.g., the state of a road, such as an indication that one or more roads are icy, paved, unpaved, etc.), traffic data (e.g., an indication of wait times for different sections of one or more streets based on traffic conditions), and / or other event data (e.g., indications of automobile accidents, indications of major sporting events) within the geographic area (e.g., zip code, geographic coordinates) represented by the geographic location indicator.

[0058] In some embodiments, event data 114 may include or represent audio data for one or more radio stations, as described herein. The audio data may be used in conjunction with other event data (e.g., weather alerts and geographic-based alert mechanisms). Thus, for example, natural language response generator 132 summarizes alerts and broadcasts messages (e.g., radio station messages) based on certain accidents, road closures, etc.

[0059] In some embodiments, event data 114 may represent one or more remote data sources of one or more network devices that are accessible via an application programming interface (API), such that, for example, event data retriever 118 establishes a communication session with the network device (e.g., via a network handshake) and responsively opens a network communication channel so that the network device can return event data 114 to event data retriever 118 via the API. Event data retriever 118 is also responsible for providing a representation of the derived event data 114 as input to language model 126, as described in more detail below.

[0060] The object context generator 120 may be responsible for collecting or determining metadata, such as context associated with one or more objects detected by the object detector 106. For example, if the detected object includes a stop sign, embodiments may extract parking context, such as the time of day and / or day of the week that the object detector 106 performed its function (e.g., automatically populated in a log file of video images), the machine type of the machine (e.g., based on the machine having broadcasted its ID), and / or an indication of whether a parking space is available for parking (e.g., determined at least in part by performing an object detection or segmentation function on image data representing the parking space). Alternatively, or in addition, other metadata may include travel routes, destinations, health data (e.g., manually registered by an operator), etc. The object context generator 120 is also responsible for returning a representation of its output to the language model 126, which may use this data as input, as described in more detail below.

[0061] The operator view probability component 128 may be responsible for detecting whether the operator of the machine has seen one or more objects detected by the object detector 106. For example, in some embodiments, the operator view probability component 128 includes a DMS working in conjunction with the object detector 106. Thus, for example, the functionality of the operator view probability component 128 may be based on analyzing image data representing a traffic sign (e.g., as indicated by the object detector 106) for image data representing one or more parts of the driver (e.g., as captured by a DMS camera). For example, as described above, an eye tracking module may be used to detect the driver's eye gaze in the image and the timestamp of the eye gaze detection to determine the direction the driver is looking at at a particular timestamp. For example, the gaze may indicate that the driver is looking out of a side mirror at a first time, but the traffic sign may be visible directly forward from the windshield at a first time. A "traffic sign" as described herein may refer to a traffic light, such as a green light, a red light, or a yellow light. Alternatively or in addition, a traffic sign may refer to any visual device or symbol placed beside or above a road, street, or highway to convey specific information, regulations, warnings, or guidance to drivers, pedestrians, and other road users. For example, the traffic sign may be a stop sign, a school zone traffic sign, a parking sign, etc.

[0062] Additionally or alternatively, the operator viewing probability component 128 can programmatically call, for example, a routine that accesses a waterfall data structure indicating all features in the driving environment and their locations (e.g., all connected road names, traffic signs, traffic lights, building structures and their orientations in space) and / or a geolocation indicator (generated by the geolocation indicator generator 116) indicating the location of the vehicle at a first time. In this way, the eye gaze direction and timestamp can be correlated or intersected with the orientation of a particular traffic sign indicated in the waterfall structure and the geolocation indicator and corresponding timestamp. Thus, using the above description, because the driver's eyes are detected looking at the side mirror at a first time, and the vehicle is approaching a traffic sign at a first time, and the waterfall data structure indicates that the traffic sign is approximately 10 yards directly ahead at the first time, a particular embodiment generates a score (e.g., a confidence level) that the driver did not see the traffic sign. In response, as Figure 1 As shown, certain embodiments provide a representation of the scores as input to the language model 126, as described in more detail below.

[0063] The destination / travel route information extractor 122 may be responsible for accessing destination or travel route information associated with the destination or travel route of the present machine from the destination / travel route data 124 and responsively providing a representation of the destination or travel route information as at least a portion of the input to the language model 126. For example, in response to detecting a travel route or destination of the vehicle (e.g., via user input of such data or via the LLM and detecting a driver utterance in a conversation that includes natural language words describing a path to be taken to reach the destination), a particular embodiment may extract video, photos, or original natural language text associated with the information from a user device or via a public computer network (e.g., the Internet). For example, based on detecting that the route includes the city of Tusayan, Arizona, a particular embodiment accesses a dataset of images or other data in 124 via a computer network and detects (e.g., via object detection) all images of the Grand Canyon based on ingesting information indicating that the Grand Canyon is near Tusayan, and also extracts natural language characters describing various landmarks along the travel route, such as accessed via queries on various public web pages. Accordingly, these natural language characters may then be provided as input to the language model 126 , where the language model 126 summarizes each landmark that the driver will encounter during the journey.

[0064] Additionally or alternatively, some embodiments may display (e.g., using a display device in the present machine) an extracted image of the Grand Canyon, other images (e.g., images in image data 104), or a natural language representation. Some embodiments may provide an image (e.g., an image detected by external perception camera 102) as an input to a multimodal machine learning model (e.g., contrastive language-image pre-training (CLIP)), which encodes the image into natural language text describing the image or different features in the image. For example, given an image of the Grand Canyon, CLIP may generate a confidence score and predict a text output representing the most relevant text description describing the Grand Canyon (e.g., "This is a photo of the Grand Canyon"). In this way, the text description may be provided to language model 126, where the language model may generate an output via natural language response generator 132. Example output generated by language model 126 in response to a text description may include, for example, “Did you know that the Grand Canyon is not the deepest canyon in the world, but it is often considered one of the most visually stunning? The Grand Canyon is approximately 6,000 feet (1,800 meters) deep at its deepest point, while the Yarlung Zangbo Grand Canyon in Tibet is even deeper, at over 17,000 feet (5,200 meters) deep.”

[0065] The language model 126 may be responsible for acquiring one or more inputs (or portions of inputs) provided by the destination / route information extractor 122, the natural language extractor 108, the operator viewing probability component 128, the object context generator 120, and / or the geographic location component 112 to generate one or more corresponding natural language outputs. In some embodiments, the language model 126 represents one or more machine learning models or other models that perform NLP. In some embodiments, a "language model" is a set of statistical or probabilistic functions that (e.g., collectively) perform natural language processing (NLP) to understand, learn, and / or generate human natural language content. For example, a language model can be a tool for determining the probability that a given sequence of words appears in a sentence (e.g., via NSP or MLM) or a natural language sequence. In short, it can be a tool that is trained to predict the next word in a sentence or other natural language character set.

[0066] When a language model is trained on a large amount of data, it is called a large language model (LLM). Some examples of LLM include GOOGLE's BERT and OpenAI's generative pre-trained transformer (GPT) network series, which includes GPT-2, GPT-3 and GPT-4. For example, GPT-3 contains 175 billion parameters, which are trained on 570GB of text. The functions of these models range from writing simple articles to generating complex computer codes, all of which are performed under limited supervision or even without supervision. Therefore, LLM is a very large deep neural network (e.g., billions to trillions of parameters) that understands, processes and generates human natural language by training a large amount of text. These models predict future words in sentences based on sentences in the text corpus they are trained on, so that they generate sentences similar to how humans speak and write. In some embodiments, LLM is pre-trained (e.g., learning English via NSP and MLM on a natural language corpus), adjusted by prompts, fine-tuned and / or run via prompt engineering, as described in more detail below.

[0067] In some embodiments, at least one of the language models 126 is stored locally on a network device or node within the machine. This can help keep processing local when real-time decisions need to be made (e.g., when the operator is driving). In these cases, it is desirable to reduce processing delays to meet time constraints associated with near-real-time operator driving and tasks, such as extracting natural language characters from traffic signs to inform the operator of the content of the traffic signs. Alternatively or in addition, in some embodiments, at least one of the language models 126 is hosted on a remote device (e.g., a cloud node or a central server). In these embodiments, for example, such cloud nodes or central services can be contacted via a network (e.g., the Internet) to provide model outputs. This network architecture can be useful in situations where large amounts of data processing or storage of large amounts of data are required.

[0068] The language model 126 includes one or more prompt building blocks 130 and a natural language response generator 132. The prompt building block 130 may be responsible for generating (e.g., automatically) or receiving one or more natural language instructions based on input received from the destination / travel route information extractor 122, the natural language extractor 108, the geographic location component 112, the object context generator 120, and / or the operator view probability component 128. The prompt building block 130 generates natural language characters (or representations thereof, such as soft prompts) as inputs to the language model 126, so that the natural language response generator 132 generates other natural language characters associated with the context (e.g., a natural language response indicating whether the current context allows parking in a detected parking space associated with a detected traffic sign) as output based on at least the inputs received from the destination / travel route information extractor 122, the natural language extractor 108, the geographic location component 112, the object context generator 120, and / or the operator view probability component 128.

[0069] In some embodiments, the natural language response generator 132 performs text translation. Text translation is the process of translating natural language from one language (e.g., Chinese) into another language (e.g., English). Various embodiments provide a sentence or text block in one language as input or prompt via a prompt building block, and the model will generate a corresponding translation in the desired target language. For example, the natural language characters extracted and passed from the natural language extractor 108 can be in a first language, which can then be translated into a second language. In an illustrative example, the natural language response generator 132 can translate the first OCR text (first language) in a traffic sign in a first country into a second language corresponding to a second country.

[0070] In an illustrative example, the prompt generated by the prompt building block 130 may include zero-shot, single-shot, or few-shot examples of representative input-output pairs. As described herein, in some embodiments, an "example" refers to one or more model (e.g., representative or exemplary) inputs and / or outputs associated with a request, where a "model output" indicates, at least in part, how the output should be formatted based on the example input (e.g., via sentence structure or grammar, word selection, length in the output (e.g., number of words), etc.). In some embodiments, an "example" refers to natural language content that a model uses as a guide to construct or design its output, and the model does not typically use the example as a guide to deriving substantial natural language text in the example (e.g., a subject or object in a sentence) to copy to the output. For example, if the instruction is "notify the operator that she missed a traffic sign," the example is an input-output pair, such as a stop sign (or a natural language description of a stop sign) (example input) and "Jane, you just missed the stop sign..." (example output). The output might say "Jack, you just missed the stop sign." Thus, the example used for the format and style of the output "You just missed the stop sign..." (i.e., various syntax and introductory words are copied to the output, but not all words are copied, e.g. "Jane"), because there was a stop sign in the input. The name was changed (from Jane to Jack).

[0071] In some embodiments, the prompt includes entity data, for example, a label that describes a specific entity in the detected object in natural language characters. For example, the label can be generated via named entity recognition (NER). NER is an information extraction technique that identifies tokens / words or "entities" in natural language text and classifies them into predefined categories. Such predefined categories can be indicated in corresponding labels or annotations. An entity can be, for example, a person's name, a specific organization, a specific location, a specific time, a specific number, a specific currency price value, a specific percentage, a specific page, etc. Similarly, the corresponding label or annotation can be a specific person, organization, location, time, price (or other invoice data), etc. In an illustrative example of the NER function, if NER marks an entity (e.g., Thomas Edison) as a "name entity", a phrase in the prompt is triggered, such as "deer [animal] lane [where the deer pass]", where the information in brackets indicates the NER entity to be included in the prompt.

[0072] The language model 126 may ingest the prompt and responsively generate an output of natural language characters and / or words based on the confidence interval via the natural language response generator 132. In an illustrative example of the language model 126 functionality, the operator view probability component 128 may generate a score indicating that the driver did not see the traffic sign - a binary value (e.g., "0") or other value indicating that the driver did not see the traffic sign. The operator view probability component 128 may then responsively map such a score to a natural language input, such as "the driver did not see the stop sign" via a hand-coded data structure (e.g., a hash map). Such a phrase may be responsively returned to the prompt building block 130 to be placed in the prompt, such that the LLM is adjusted by the prompt to output the phrase "You just missed the stop sign. Please be careful. I will let you know before you reach the next stop sign." Examples of various natural language inputs, prompts, and outputs are described in more detail below.

[0073] The text-to-speech component 134 may be responsible for converting the written or visual natural language characters generated by the natural language response generator 132 into corresponding audio data representing the written or visual natural language characters via a speech-to-text function. In these embodiments, such audio data may be presented on a sound device (e.g., a voice assistant speaker or a stereo system), which may be helpful so that the operator of the present machine can focus their eyes on the road without reading text. The display component 136 may be responsible for transmitting the written or visual natural language characters generated by the natural language response generator 132 to a display device (e.g., an LCD screen memory), such as a display screen in the present machine. In this way, the operator can alternatively or additionally view or read the generated output.

[0074] refer to Figure 2 , Figure 2 is an example alert level detector pipeline 200 according to some embodiments of the present disclosure. It should be understood that this arrangement and other arrangements described herein are presented only as examples. In addition to or in place of the arrangements and elements shown, other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used, and some elements may be omitted entirely. In addition, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in combination with other components, and may be implemented in any suitable combination and location. The various functions performed by the entities described herein may be performed by hardware, firmware, and / or software. For example, the various functions may be performed by a processor executing instructions stored in a memory.

[0075] In some embodiments, the alert level detector pipeline 200 includes Figure 1 One or more components of the natural language extraction pipeline 100. For example, in some embodiments, Figure 2 The language model 226 includes an external perception camera 102, environmental image data 104, an object detector 106, a speed detection component 110, a destination / travel route information extractor 122, a geographic location component 112, and / or an object context generator 120. In some embodiments, the language model 226 represents Figure 1 Same language model 126. In some embodiments, the systems, methods, and processes described herein may use Figures 10A-10D Example of autonomous vehicle 1000, Fig.11 The example computing device 1100 and / or Fig.12 The example data center 1200 may be implemented with similar components, features and / or functions.

[0076] exist Figure 2 In the illustrated embodiment, the alertness level detection pipeline 200 includes an internal perception camera 202, operator image data 204, an operator alertness level detector 206, an operator response handler 208, a personalized information extractor 210, personalized information 212, one or more language models 226, an operator viewing probability component 228, a text-to-speech component 234, and a display component 236.

[0077] The interior awareness camera 202 may be responsible for capturing one or more images (e.g., video clips) of one or more portions of the interior of the present machine - which may include the operator of the present machine - and then responsively storing the images as operator image data 204. For example, the interior awareness camera 202 may be included in a DMS that relies on cameras and sensors strategically placed within the vehicle cabin. In some embodiments, the interior awareness camera 202 is located on the dashboard, rearview mirror, or other suitable location. In some embodiments, the interior awareness camera 202 includes an infrared camera for nighttime operation. Thus, for example, sensors within the DMS may capture images of the interior and store them in a data repository as operator image data 204.

[0078] Operator image data 204 is a data repository of one or more images of the internal parts of the machine and / or the operator within the machine. For example, operator image data 204 may be various streaming video sequences of the internal parts of the machine at various time stamps. Operator alertness level detector 206 detects the alertness level of the operator of the machine based on detecting patterns or associations within operator image data 204. For example, DMS may employ computer vision and facial recognition algorithms to monitor the operator's face (or operator image data 204 representing the operator's face) in real time or near real time. It tracks key facial features such as eyes, eyelids, mouth and / or head position. In addition or alternatively, DMS may incorporate eye tracking technology to monitor the driver's eye movements. For example, it may track factors such as blink rate, gaze direction, and duration of eyelid closure. DMS may alternatively or additionally monitor the operator's head position and movement. Sudden twitches or unusual head positions may be signs of distraction or drowsiness.

[0079] The operator alertness level detector 206 is also generally responsible for providing a representation of the alertness level (e.g., a natural language sentence) to the language model 226 for further processing, as described in more detail below. For example, the operator alertness level detector 206 may use the alertness level score as an index into a data structure to look up the corresponding hand-coded natural language characters, such as "This operator has the highest possible alertness level on the KSS scale."

[0080] The operator response handler 208 may be responsible for receiving and handling all operator responses and providing representations of such operator responses as at least partial inputs to the language model 226. For example, after the operator's alertness level has been detected via the operator alertness level detector 206, it may feed a representation of the score as input to the language model 226, which then generates a natural language response, "You are really tired, should I call your spouse?" Subsequently, the operator response handler 208 receives a phrase utterance representing the driver's response, such as "No, I'll be fine". Subsequently, in some embodiments, the operator alertness level detector 206 receives another set of image data representing one or more parts of the driver, such as through the above-mentioned DMS system. Based at least on the image pattern features of such image data and / or the operator responses / utterances handled by the operator response handler 208, a particular embodiment generates a second score indicating a second alertness level of the driver. In an illustrative example, a particular embodiment may use a Gaussian mixture model (GMM) or a hidden Markov model (HMM) to detect the driver's alertness level based on detecting speech patterns associated with alertness or non-alertness, as described above.

[0081] In response, certain embodiments provide another representation of the second score (e.g., another natural language phrase) as input to the language model 226, so that the model generates and presents additional natural language characters representing yet another phrase utterance in the conversation with the operator. For example, the presented output may be "You still sound a little tired. Should I call your wife if I may?" Such output may be based on receiving input via a GMM that classifies the driver phrase as 1 (indicating very sleepy) and a hand-coded data structure that maps such classification to a natural language phrase (e.g., "The driver is very sleepy").

[0082] The personalized information extractor 210 may be responsible for extracting the personalized information 212 and providing a representation of such information as at least a portion of the input to the language model 226. For example, some embodiments access an indication from the personalized information 212 that the operator likes science and plays tennis as a hobby. In response, certain embodiments provide such natural language personalized information as an input to the language model 226 so that the dialogue with the operator is additionally or alternatively based on the personalized information associated with the driver. For example, based on the driver's interest in tennis and based on the driver's alertness level being below a threshold, such an output generated by the natural language response generator 232 may be "Your favorite tennis player is playing today."

[0083] The language model 226 may be responsible for taking one or more inputs (or portions of inputs) provided by the operator alertness level detector 206, the operator response handler 208, the personalized information extractor 210, and / or the operator viewing probability component 228 in order to generate one or more corresponding natural language outputs. The language model 226 includes a prompt building block 230 and a natural language response generator 232. In some embodiments, the prompt building block 230 and the natural language response generator 232 are respectively Figure 1 The prompt building block 130 and the natural language response generator 132 in represent the same functionality, except that the instructions or prompts and outputs are specifically tailored to the outputs produced by the operator response handler 208, the operator alertness level detector 206, the personalized information extractor 210, and / or the operator viewing probability component 228.

[0084] In an illustrative example of the prompt building block 230, it may include natural language personalized interests returned by the personalized information extractor 210, zero-shot, one-shot, or few-shot examples of example conversations conducted by the operator (or another operator) with the language model 226 given various example inputs, a natural language instruction to "start a conversation consistent with the operator's interests", a natural language description of the operator's alertness level returned by the operator alertness level detector 206 (e.g., via a KNN scale), the operator's natural language response provided by the operator response handler 208, and / or a natural language description of whether the operator saw or did not see one or more detected objects (returned via the operator viewing probability component 228). Then, the natural language response generator 232 generates one or more natural language responses in response. For example, the natural language response generator 232 can generate a natural language sentence whose content is "Wake up! Let's play a football quiz game together" (based on the input provided by the operator alertness level detector 206 and the personalized information extractor 210).

[0085] In some embodiments, the operator view probability component 228 includes the same functionality as the operator view probability component 128. Similarly, in some embodiments, the text-to-speech component 234 and the display component 236 represent the same functionality as for the Figure 1 The text-to-speech component 134 and the display component 136 have the same respective functions as described.

[0086] Figure 3 is a block diagram of a large language model 300 (e.g., a BERT model or a GPT-4 model) that uses a specific input to generate a specific natural language response according to some embodiments. In some embodiments, the model 300 represents or includes Figure 1 or Figure 2 The language model 300 may be used to implement the functionality described by the language model 126 and / or the language model 226. In various embodiments, the language model 300 includes one or more encoder and / or decoder blocks 506 (or any transformer or portion thereof).

[0087] At a first time, the input 301 is converted into tokens, then converted into feature vectors and embedded into input embeddings 302 (e.g., to derive the meaning of individual natural language words (e.g., English semantics) during pre-training). In some embodiments, for example, each word or character in the input 301 is mapped to the input embedding 302 in parallel or simultaneously, which is different from existing long short-term memory (LSTM) models. The input embedding 302 maps a word to a feature vector representing the word. But the same word (e.g., "apple") in different sentences may have different meanings (e.g., device vs. fruit). This is why a position encoder 304 may be implemented. The position encoder 304 is a vector that provides context for a word (e.g., "apple") based on its position in the sentence. For example, for the message "I just sent a document", because "I" is at the beginning of the sentence, an embodiment may indicate a position in the embedding that is closer to "just" than "document". Some embodiments use a sign / cosine function to generate a position encoder vector, as shown below:

[0088]

[0089] After passing the input 301 through the input embedding 302 and applying the position encoder 304, the output is a word embedding feature vector that encodes the position information or context based on the position encoder 304. These word embedding feature vectors are then passed to the encoder and / or decoder block 306, where it passes through the multi-headed attention layer 306-1 and the feed-forward layer 306-2. The multi-headed attention layer 306-1 can be responsible for focusing on or processing certain parts of the feature vector representing a specific part of the input 301 by generating an attention vector. For example, in a question-answering system, the multi-headed attention layer 306-1 determines the relevance of the i-th word (or a specific word in a sentence) to answering the question or to other words in the same or other blocks, and its output is an attention vector. For each word, some embodiments generate an attention vector that captures the contextual relationship between other words in the same sentence or other character sequence. For a given word, some embodiments calculate a weighted average of or otherwise aggregate the attention vectors of other words (e.g., other words in the same row or block) that contain the given word to calculate a final attention vector.

[0090] In some embodiments, the single-head attention has abstract vectors Q, K, and V that extract different components of a particular word. These are used to calculate the attention vector for each word, using the following formula:

[0091] Z = softmax(〖QK〗^T / √(the dimension of vector Q, K or V)).V

[0092] For multi-head attention, there may be multiple weight matrices Wq, Wk, and Wv, and therefore multiple attention vectors Z for each word. However, the neural network may only expect one attention vector per word. Therefore, another weight matrix Wz is used to ensure that the output is still one attention vector per word. In some embodiments, after layers 306-1 and 306-2, some form of normalization (e.g., batch normalization and / or layer normalization) is performed to smooth the loss surface, making it easier to optimize while using a larger learning rate.

[0093] Layers 306-3 and 306-4 represent residual connections and / or normalization layers, where normalization re-centers and rescales or normalizes the data on the feature dimension. The feed-forward layer 306-2 is a feed-forward neural network that is applied to each of the attention vectors output by the multi-head attention layer 306-1. The feed-forward layer 306-2 transforms the attention vector into a form that can be processed by the next encoder block or predicted at 308. For example, assuming that the document includes a first natural language sequence "The due date is...", the encoder / decoder block 306 predicts that the next natural language sequence will be a specific date or a specific word based on past documents that include the same or similar language as the first natural language sequence.

[0094] In some embodiments, the encoder / decoder block 306 includes pre-training to learn the language (pre-training) and make corresponding predictions. In some embodiments, the encoder / decoder block 306 learns the language and context of words in pre-training by training two unsupervised tasks (masked language modeling (MLM) and next sentence prediction (NSP)) simultaneously or at the same time. In terms of input and output, during pre-training, the natural language corpus of input 301 can be various historical documents, such as textbooks, journals, network data and / or journals, so as to output the predicted natural language characters in 308 (at this time, the prediction when prompt engineering or prompt adjustment is not performed). The encoder / decoder block 306 receives sentences, paragraphs, or sequences (for example, included in input 301), in which random words are replaced with masks. The goal is to output the value or meaning of the masking flag. For example, if a line of content is "Please [mask] this document immediately", the prediction of the "mask" value is "send". This helps the encoder / decoder block 306 understand the bidirectional context in sentences, paragraphs, or lines in the document. In the case of NSP, the encoder / decoder block 306 takes two or more elements (e.g., sentences, lines, or paragraphs) as input and determines, for example, whether the second sentence in the document actually follows (e.g., is directly located after) the first sentence in the document. This helps the encoder / decoder block 306 understand the context of all elements in the document, not just the context in a single element. Using these two together, the encoder / decoder block 306 can gain a good understanding of natural language during pre-training.

[0095] In pre-training, the output is typically a binary value C (for NSP) and various word vectors (for MLM). Through training, the loss (e.g., cross entropy loss) is minimized. In some embodiments, all feature vectors are of the same size and generated at the same time. Therefore, each word vector can be passed to a fully connected layered output with the same number of neurons as the number of tokens in the vocabulary.

[0096] In some embodiments, once pre-training is performed, the encoder / decoder block 306 performs prompt engineering, prompt adjustment, and / or fine-tuning. For example, for fine-tuning, some embodiments perform the QA task by adding a new question-answer (or prompt-response) header or encoder / decoder block in 306, just as the masked language model head is added (in pre-training) to perform the MLM task, except that the task is part of fine-tuning and is used to add new input data to the input 301 and adjust the weights developed during pre-training. In other words, fine-tuning adds additional input data (i.e., specific prompts in the input 301 that are not part of the pre-training) and performs additional training rounds to further adjust the weights to develop the output 308 that is not part of the pre-training.

[0097] Prompt engineering is the process of guiding and shaping ML model responses (e.g., outputs 308) by relying on users or prompt engineers to design more carefully worded, more specific queries or prompts. With prompt engineering, weights are frozen (i.e., their values ​​remain the same as they were during pre-training) so that they are not adjusted during prompt engineering. "Prompts" as described herein include one or more of the following: natural language requests (e.g., questions, commands, or instructions (e.g., "write a summary of a poem")), one or more datasets (e.g., specific documents or images), code snippets, mathematical equations, one or more examples (e.g., one-shot or two-shot examples), and / or numerical embeddings (e.g., "soft" prompts). In some embodiments, "examples" represent few-shot prompts, which is a technique for guiding large language models (LLMs) (such as GPT-3) to generate desired outputs by providing them with some examples of input-output pairs.

[0098] The prompt engineering process typically involves iteratively asking more and more specific and detailed questions / commands / instructions, or testing different ways of expressing questions / commands / instructions. The goal is to use prompts to elicit better behavior or output from the model. Prompt engineers typically try various types of questions / commands / instructions and formats to find the most ideal and relevant model response. For example, a prompt engineer may initially provide a prompt from the event data retriever 118 (i.e., the "event data prompt" of input 301), prompting "What is the weather like today?", where the "event data output" in output 308 initially displays "The weather is sunny." However, this may not be specific enough, so the prompt engineer may formulate another prompt template, which states "What is the temperature now and in the next 4 hours?" The "event data output" of the response is "The temperature is 68 degrees and will remain at this temperature for the next three hours, and then the temperature will drop to 65 degrees." The prompt engineer may be satisfied with this prompt. After obtaining this satisfactory answer, a specific embodiment saves the corresponding event data prompt as a template (e.g., "What is the temperature now and in the next 4 hours?"). In this way, the prompt template (e.g., a "hard" prompt) can be used at runtime or when deploying a model. In some embodiments, such a template leaves certain words in the prompt template blank, as the blank space may depend on the use case provided by the runtime prompt. For example, using the example template above, the template may read "... for the next __ hours ...".

[0099] Prompt tuning is the process of acquiring or learning the most effective prompts or clues (out of a large number of prompts) and providing them as task-specific context to the encoder / decoder block 306. For example, a common question or phrase - "What is my account balance?" - can be taught to the encoder / decoder block 306 to help optimize the model and guide it towards the most ideal decision or corresponding output in 308. Unlike prompt engineering, prompt tuning is not about the user asking better questions or making more specific requests. Prompt tuning means identifying more frequent or more important prompts (e.g., with higher node activation weight values) and training the encoder / decoder block 306 to respond to these common prompts more effectively. The benefit of prompt tuning is that it can be used to moderately train the model without adding any input 301 or prompts (unlike fine-tuning), thereby saving a lot of time and cost.

[0100] In some embodiments, prompt adjustment may use only soft prompts and may not include the use of hard prompts. Hard prompts are manually handcrafted text prompts (e.g., prompt templates) with discrete input tokens, typically used for prompt engineering. Prompt templating allows prompts to be stored, reused, shared, and programmed. Soft prompts are typically created during the prompt adjustment process. Unlike hard prompts, soft prompts are typically not viewed and edited in text form. Soft prompts typically include an embedding, i.e., a string of numbers, which obtains knowledge from the encoder / decoder block 306 (e.g., via pre-training). Therefore, soft prompts are learnable tensors concatenated with input embeddings that can be optimized for the dataset. In some embodiments, prompt adjustment creates a smaller, lightweight model that is placed in front of a frozen pre-trained model (i.e., a large language model 300 whose weights are set during pre-training). Therefore, prompt adjustment involves using a small trainable model before using the LLM 300. The small model is used to encode the text prompt and generate task-specific virtual tokens. These virtual tokens are pre-appended to the prompt and passed to the LLM 300. When the tuning process is complete, these virtual landmarks are stored in a lookup table (or other data structure) and used during inference, replacing the smaller model.

[0101] like Figure 3 As shown, input 301 and output 308 have specific prompts, as described below. In some embodiments, the prompt of input 301 represents the final prompt adjusted, fine-tuned and / or prompt engineered prompt. The "natural language extractor prompt" is generated by the prompt building block 130 based on Figure 1 The prompt (or portion of the prompt) is generated from the output of the natural language extractor 108 of the present machine. For example, in some embodiments, the natural language extractor prompt may include one or more of the following: a query (e.g., "What is this traffic sign?"), an instruction (e.g., "Tell the driver what the traffic sign says"), one or more zero-shot, one-shot, or few-shot examples of representative inputs (e.g., a natural language description of a stop sign generated via a CLIP model), an output (e.g., "This is a stop sign"), entity data associated with one or more extracted natural language characters (e.g., determined via NER), and / or a hierarchical data structure (e.g., a waterfall data structure) representing multiple features of the environment in which the present machine is located and their spatial relationships (e.g., a particular map of the driving environment and its features). In response, the model generates one or more outputs in 308, such as "natural language extraction output", such as "You just passed a stop sign."

[0102] The "event data prompt" input 301 is generated by the prompt building block 130 based on Figure 1The event data retriever 118 of the present invention generates a prompt (or a portion of a prompt) based on the output of the event data retriever 118. For example, the event data prompt may include weather data for the area where the machine is located, traffic data for the area, radio station feeds for the area, instructions for specific sporting events in the area, instructions (e.g., "tell the driver the specific temperature and whether there are any accidents within the next 3 miles"), zero-sample, single-sample or few-sample examples, etc. In response, the encoder / decoder block 306 provides one or more outputs 308, such as "event data output". For example, the "event data output" may be a text summary of weather data, traffic data, radio station feeds, sporting events, etc.

[0103] The “object context hint” input 301 is generated by the hint building block 130 based on Figure 1 The object context generator 120 may generate a prompt (or a portion of a prompt) based on the output of the object context generator 120. For example, the object context prompt may include parking context, such as the time of day when the element representing the stop sign was detected, the day of the week when the element representing the stop sign was detected, the type of machine, instructions (e.g., "tell the driver if the parking space is available"), and / or an indication of whether the object detector detected a car in the parking space. In response, the encoder / decoder block 306 generates an output at 308, such as parking space availability. For example, such an output may be a natural language sentence that reads, "This parking space is not available for parking because the parking instructions in the sign indicate that the parking space is available after 7pm, but it is not yet 7pm."

[0104] The "operator viewing prompt" input 301 is generated by the prompt building block 130 based on Figure 1 / Figure 2 The prompt (or portion of the prompt) is generated as an output of the operator viewing probability component 128 / 228. For example, the operator viewing prompt may include an indication of a hierarchical data structure as described above (e.g., a "waterfall" structure), a representation of a score representing the probability that the operator did not see the object, an instruction (e.g., "tell the driver if they missed the traffic sign"), and / or zero-shot, one-shot, or few-shot examples of input and / or output. Responsibly, as at least a portion of the output 308, the encoder / decoder block 306 may generate a natural language response representing an "operator viewing response," such as "It looks like you missed the stop sign."

[0105] The "description / travel route tips" of input 301 is generated by the tips building block 130 based on Figure 1The prompt (or portion of the prompt) is generated as an output of the destination / route information extractor 122 of the destination / route information extractor 122. For example, the prompt may include all the destination and / or route information extracted from the destination / route data 124, an instruction to "summarize the destination and / or route information", and / or zero-sample, single-sample, or few-sample examples. In response, the encoder / decoder block 306 generates "summary information about the destination / route" in the output 308. For example, the destination / route data 124 may include all tourist areas along the route and various information about each tourist area. The summary may include a high-level overview of each tourist area, but does not include all information about each tourist area. For example, such a summary might be, "On your route, you will first visit Phoenix, Payson, and then arrive at the Grand Canyon. The Grand Canyon will close around 6 PM today."

[0106] The "operator alert level prompt" input 301 is generated by the prompt building block 130 based on Figure 2 The prompt (or portion of a prompt) is generated based on the output of the operator alertness level detector 206. For example, such a prompt may include personalized information 212, a representation of the calculated operator alertness level (e.g., "The operator is very sleepy"), a natural language instruction (e.g., "Start a conversation with the driver that is in line with the driver's interests (e.g., American football)"), zero-shot, single-shot, or few-shot examples of representative inputs / outputs, and / or "Send a control signal to honk the horn." One or more of these inputs prompt the encoder / decoder block 306 to output a "response to the operator alertness level" at output 308. For example, such a response may be "I just honked the horn because you were falling asleep. Do you know the name of the only quarterback to win the regular season MVP, Super Bowl, and Super Bowl MVP in the same season?"

[0107] The "operator response prompt" input 301 is generated by the prompt building block 130 based on Figure 2The prompt (or portion of a prompt) generated as an output of the operator response handler 208. For example, such a prompt may include a natural language utterance uttered by the operator (e.g., an answer to the Super Bowl question above, such as "Patrick Mahomes"), a new representation of another alertness level detected or a second representation (e.g., "The driver is now more alert and has a KSS value of 4"), an instruction (e.g., "Reply to the driver based on the driver's new KSS value"), a zero-shot, a one-shot, or a few-shot example, etc. In response to the encoder / decoder block 306 ingesting such a prompt, it generates a "response to the operator response" in the output 308. For example, such a response may be "Your answer about Patrick Mahomes is correct. I'm glad you're more alert. Should I call a nearby hotel so you can get some sleep?".

[0108] The "personalized information" input 301 is generated by the prompt building block 130 based on Figure 2 The prompt (or portion of the prompt) is generated at the output of the personalized information extractor 210 of the embodiment of the present invention. For example, such a prompt may include all of the personalized information 212, an instruction (e.g., "Start a conversation with the driver using the personalized information"), a hierarchical data structure as described herein, and / or a zero-shot, few-shot, or single-shot example. In response, the encoder / decoder block 306 generates a response at the output 308. For example, if the personalized information 212 indicates that the driver enjoys bass fishing and the driver is near a particular lake known for bass fishing (e.g., as determined via the hierarchical data structure and event data 114), the model may generate a natural language response such as, "There is a lake about 5 miles ahead - Lake Murray. This lake is one of the best bass fishing lakes in the country."

[0109] Figure 4A 4 and 408 according to some embodiments. In some embodiments, the image data 400 is determined or detected by the external perception camera 102 and / or the object detector 106, such as Figure 1 The image data 400 (e.g., various pixels) includes a bounding box 402 (labeled “Stop Sign”) located on a subset of the image data 400 representing a Stop Sign 404. The image data 400 also includes a bounding box 406 (labeled “Parking Space”) located on a subset of the image data 400 representing a Parking Space 408.

[0110] In some embodiments, object detector 106 detects stop sign 404 (as indicated by bounding box 402) and detects parking space 408 (and no car occupies parking space 408), as indicated by bounding box 406. In response to such detection, in some embodiments, natural language extractor 108 responsively extracts natural language text within stop sign 404, such as information about Figure 1 As described. For example, the natural language extractor 108 may perform an OCR function to encode the image data in 404 (i.e., "15 minute parking for store customers only") into a JSON data structure, wherein the resulting JSON natural language characters are fed as input to the encoder / decoder block 306 as at least a portion of the "natural language extractor prompt" in 301. Additionally or alternatively, the object context generator 120 provides as input an "object context prompt" as indicated in the input 301. For example, the object context prompt may include an indication of "unoccupied parking space" (e.g., detected by the object detector 106). In response, the model generates an output as a natural language response, such as Figure 4B 414 in a natural language response, as described in more detail below.

[0111] Figure 4B The detected Figure 4A Traffic signs 404 and / or Figure 4A 408. That is, for example, natural language response generator 230 generates a response "It looks like you are trying to park in a 'Customers Only' parking lot. As a reminder, you can only park here for 15 minutes." This response can be based on the model being prompt-engineered, prompt-adjusted, and / or fine-tuned to generate this response based on performing natural language processing and otherwise extracting natural language characters from stop sign 404 and / or ingesting prompts indicating that parking space 408 is available. If the parking space is not available (e.g., because the car is parked in parking space 408 or because of special instructions within stop sign 404), the model can additionally or alternatively tell the driver 410 that they cannot park there.

[0112] In response to the natural language response generator 132 generating such output, it returns the output to the text-to-speech component 131, which then outputs an audio message 414 as audio data to the sound device 416 - "It looks like you are attempting to park in the 'customers only' parking lot. As a reminder, you can only park here for 15 minutes."

[0113] Figure 5A5 shows a present machine operator 502 (or image data of the operator 502) in a drowsy state and a responsive natural language response 504 according to some embodiments. In some embodiments, one or more portions of the operator 502 or corresponding image data are analyzed by the operator alertness level detector 206, such as with respect to Figure 2 For example, a DMS camera in a vehicle's steering column may employ an eye tracking module that emits infrared light that should be reflected in the eyes of operator 502. However, operator 502's eyes may be closed or not looking straight ahead. In response, a CNN or other model may classify operator 502 with a confidence level score or prediction of driver drowsiness. In response, particular embodiments map such scores to KNN-level natural language outputs via a lookup data structure, such as "The operator's KSS level score is 8, which is drowsy but requires some effort to be alert."

[0114] In response, the operator alertness level detector 206 sends such natural language phrases to the language model 226, where the prompt building block 226 generates an "operator alertness level prompt" for input 301, which includes a KSS level score natural language phrase. In response, the natural language response generator 132 generates a natural language response, such as "You seem very sleepy! Can I play your favorite fast-paced music?" In response, the natural language response generator 132 automatically transmits such responses to the text-to-speech component 134, which converts such natural language responses to the audio data response 504, so that the audio device 506 (e.g., car speakers) outputs the audio data response 504 in the form of sound waves, which mirrors the response generated by the natural language response generator 232.

[0115] In various examples, the audio data response 504 or other generated response to the calculated alertness level described herein indicates initiating a personalized conversation with the operator based on the KSS level, with the goal of reducing the operator's KSS or other non-attention level to a certain threshold. For example, some embodiments continuously loop to generate natural language sentences indicating a conversation with the operator until the KSS level or threshold is reached (e.g., a KSS level of 4). In some embodiments, this can be used to further refine topics that actually help reduce the KSS level (e.g., personalized topics) rather than those that do not help. For example, reinforcement learning can be used to understand, based on a history of different topics discussed with the operator in the past, that for a particular operator, any discussion about topic A is likely to result in a reduction in the KSS level relative to topic B. As a result, the language model will not use topic B in the conversation with the operator, but will only discuss topic A based on the model reward given in the reinforcement training to generate a natural language response for discussing topic A (and / or a penalty for generating a natural language response for discussing topic B).

[0116] Figure 5B According to some embodiments Figure 5A Operator 502 pairs Figure 5A 504 and the model's response 510 to the operator's 502 response 508. Thus, after the audio device 506 outputs the audio response 504, the operator 502 (whose eyes may now be open and more alert) can be said to utter the audio response 508, namely, "Sure." In response, the operator response handler 208 can take this "Sure" response, perform a speech-to-text function on this audio response 508, and then send the corresponding natural language response / text (no longer audio data) as input to the language model 226. In response, the prompt building block 230 generates the following: Figure 3 The "Operator Response Prompt" shown as input 301 in , includes the text-based word "Of course" (represented by 508).

[0117] Additionally or alternatively, the operator alertness level detector 206 again detects the alertness level of the operator 502 at a second time, such as Figure 5B For example, using DMS, the operator alertness level detector 206 may detect that the KSS level is now 4 “fairly alert” (e.g., somewhat alert), which means that the driver 502 is now relatively Figure 5AThe state of the driver 502 in the input 301 is more alert. This new KSS level 4 can also be provided as an input to the language model 226. In response, the natural language response generator 232 generates a natural language response "Great! I will also point out different travel destinations on your route based on your interests." This natural language response can also be based on obtaining personalized information 212 as input (or part of the input) from the personalized information extractor 210 (e.g., included in the "Personalized Information Hint" of the input 301) and / or obtaining the destination / travel route data 124 from the destination / travel route information extractor 122 (e.g., included in the "Destination / Travel Route Hint" of the input 301). In response, the natural language response generator 132 transmits this natural language response to the text-to-speech component 134, which converts this response into a corresponding audio response 510 via the audio device 506. In response, in some embodiments, the personalized information extractor 210 searches the personalized information 212 and / or contacts the music service to retrieve the operator's 502 favorite fast-paced music and / or the user's interest in travel destinations (e.g., the operator may like the Grand Canyon). When such music and / or other interests are retrieved, the audio device 506 can then automatically play audio data representing the operator's 502 favorite fast-paced music. In addition or alternatively, the destination / travel route information extractor 122 and / or the personalized information extractor 210 can provide input travel routes / destination and / or personalized information so that the operator 502 is informed of different travel destinations along the route according to interest (e.g., the operator 502 is informed that the next tourist area on his route is the Grand Canyon, which he is interested in).

[0118] Reference now Figure 6 , Figure 7 , Figure 8 and Fig. 9 , each process block 600, 700, 800, and 900 described herein includes a computing process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in a memory. The methods can also be embodied as computer-usable instructions stored on a computer storage medium. The methods can be provided by a stand-alone application, a service or a hosted service (standalone or in combination with another hosted service), or a plug-in to another product, to name a few. In addition, the methods 600, 700, 800, and 900 are provided by way of example with respect to Figure 1 The pipeline 100 and / or Figure 2 However, these methods may additionally or alternatively be performed by any one system or any combination of systems, including but not limited to the systems described herein.

[0119] Figure 6is a flow chart of an example process 600 for generating a natural language response based on extracting natural language characters represented in image data according to some embodiments. According to block 602, some embodiments receive image data representing one or more objects in an environment, wherein the image data is generated using one or more sensors (e.g., cameras) of a machine (e.g., a car, a boat, or an airplane). For example, referring back to Figure 1 The object detector 106 may receive image data of a driving environment including street signs, streets, traffic lights, or other physical structures after a camera mounted on the vehicle captures such data and stores such data.

[0120] According to box 604, a particular embodiment extracts the first one or more natural language characters represented in the image data. In some embodiments, "extracting" in this context means identifying and / or capturing natural language characters in the image data and converting the image data natural language characters into machine-readable and editable text that is no longer in the image data format. For example, the conversion can include encoding the image data natural language characters into a data structure including a JSON string matching the natural language characters using pattern recognition and machine learning algorithms. Therefore, the "first one or more natural language characters" represent the output encoded characters (or non-image data) represented in the image data.

[0121] In the illustrative example of block 604, in some embodiments, the one or more objects include a traffic sign. Thus, for example, certain embodiments detect, via object detection (e.g., via object detector 106) and within the image data, one or more regions of image data depicting a traffic sign (e.g., having bounding box 402). Some embodiments extract, based at least on performing optical character recognition (OCR) within the one or more regions of the image data in response to detecting the one or more regions of the image data depicting the traffic sign. For example, this may be performed on the one or more regions of the image data depicting the traffic sign. Figure 1 , where the natural language extractor 108 can extract natural language characters from one or more objects in response to the object detector 106 detecting one or more objects.

[0122] Additionally or alternatively, some embodiments determine that the speed of the host machine is below or above a threshold speed. Thus, the extraction of the one or more first natural language characters represented in the traffic sign occurs at least in part in response to determining that the speed of the host machine is below or above a threshold speed. For example, referring back to Figure 1 , the natural language extractor 108 can extract natural language characters in response to determining whether the speed detected from the speed detection component 110 is below or exceeds a threshold.

[0123] According to box 606, some embodiments provide a representation of the first one or more natural language characters extracted from the image data (e.g., a hard prompt or a soft prompt) as at least a part of the input of one or more machine learning models to generate a second one or more natural language characters responsive to the environment. In some embodiments, "providing" the representation includes transmitting the representation (e.g., over a network and to another computing node) to a device including a machine learning model. Providing may alternatively or additionally include programmatically returning or passing the representation to the machine learning model (e.g., stored to the same computing node). In some embodiments, box 606 generates a second one or more natural language characters based on providing a representation of the first one or more natural language characters extracted from the image data as at least a part of the input of the machine learning model. In some embodiments, "responsive to the environment" means being based on, according to, or otherwise associated with the environment. In other words, the natural language characters output by the model are related to the environment through which the machine passes to some extent.

[0124] In some embodiments, generation of the second one or more natural language characters at box 606 is based on additional or alternative inputs and / or portions of inputs (e.g., additional prompts) as described herein. For example, in some embodiments, the input at box 606 includes a prompt, and the one or more machine learning models include a large language model (LLM). In some embodiments, such a prompt also includes a query (e.g., "Tell the driver what sign they passed and what was on the sign") and one or more of the following: zero-shot, single-shot, or few-shot examples of one or more representative inputs and / or outputs (e.g., "You just passed a stop sign"), entity data associated with the one or more first natural language characters (e.g., NER entities for the characters), or a hierarchical data structure (e.g., a waterfall structure) representing multiple features of an environment (e.g., the orientation of connecting roads, traffic lights, and signs). For example, such data may be included in Figure 3 In the “natural language extractor prompt” indicated in the input 301 of . In this way, the generation of the second natural language character is based on the LLM ingestion prompt. For example, the second natural language character may be “You might want to slow down, you just missed the stop sign.” (e.g., detected by the object detector 106 and / or the speed detection component 110).

[0125] Continuing with block 606, in some embodiments, the one or more objects represent traffic signs (e.g., a yield sign, a stop sign, a traffic light, a deer sign, etc.). In some embodiments, the traffic sign is a stop sign that includes a stop indication. The stop indication indicates how or when a vehicle may stop, e.g., Figure 4A404 of a parking sign. In these embodiments, various embodiments provide a representation of the parking context (e.g., metadata) as at least a second portion of the input to the machine learning model. In some embodiments, the parking context includes at least one of the following: the time of day when the element representing the parking sign is detected (e.g., performed by the object detector 106), the day of the week when the element representing the parking sign is detected (e.g., performed by the object detector 106), the machine type (e.g., brand and model) of the machine, or an indication of whether the parking space is available for parking. For example, all of this information can be provided by the object context generator 120, or represented in the "object context prompt" shown in the input 310. In this way, a particular embodiment outputs a specific natural language phrase (representing a second one or more natural language characters) that indicates whether the parking space is allowed to be occupied by the machine based on at least the parking context and the parking instruction. For example, if the time of day indicated in this prompt exceeds the time window (e.g., "Parking only between 6 and 8"), this can be used to notify the operator that they may not be able to park in the corresponding parking space.

[0126] Continuing with box 606, some embodiments receive a geolocation indicator representing the location of the machine in the environment. Based at least on the geolocation indicator, a particular embodiment (e.g., event data retriever 118) determines at least one of the following: weather data, road condition data, traffic data, or event data associated with the geolocation (e.g., stored to event data 114). In response, a particular embodiment provides at least one of the weather data, road condition data, traffic data, event data, or geolocation indicator as at least a second portion of the input to one or more machine learning models. For example, such event data may include or represent the "event data prompt" in input 301 so as to generate an "event data output" as indicated in output 308.

[0127] Continuing with block 606, some embodiments receive second image data representing one or more portions of an operator of the present machine at a first time. Based at least on the second image data, some embodiments generate a first score (e.g., a KSS score) representing a first alertness level of the operator. An indicator or representation of the first score is provided as at least a second portion of an input to the machine learning model. For example, such a second portion may represent or include an "operator alertness level prompt" as indicated in input 301, so as to produce a "response to operator alertness level" in output 308.

[0128] Continuing with block 606, some embodiments provide a representation of the calculated operator alertness level and / or a representation of the operator's response to the presented one or more second natural language characters (e.g., Figure 5A The audio output 504) of the operator response (e.g., Figure 5B The phrase utterance of the "Sure" response 508) is used as a subsequent input to one or more machine learning models to generate one or more third natural language characters representing a generated response to the operator response (e.g., Figure 5B For example, the subsequent input may include or represent the second “operator alertness level prompt” and / or “operator response prompt” in input 301 so as to generate a “response to alertness level” and / or “response to operator response” as indicated in output 308.

[0129] Continuing with block 606, some embodiments provide a representation of the personalized information related to the operator retrieved from the database (e.g., personalized information 212) as at least a second portion of the input to the machine learning model. For example, such a second portion may represent or include a "personalized information prompt" as indicated in input 301, so as to generate a "response to operator alertness level" and / or a "response to operator response" as indicated in output 308.

[0130] Continuing with box 606, in some embodiments, the object includes a traffic sign. Some embodiments receive a score representing the probability that the operator did not see (or saw) the traffic sign, the score being based at least on the first image data and the second image data representing one or more portions of the operator or occupant. For example, as described herein, some embodiments use a waterfall or other hierarchical data structure to compare near real-time data (e.g., GPS location) of the machine and various objects in the environment (e.g., a stop (STOP) sign) with DMS data (e.g., eye gaze) indicating whether the operator saw the stop sign. In response, a score (e.g., a binary classification Boolean or a binary "yes") is generated and then mapped (e.g., with a hand-coded data structure) to a natural language sequence, such as "the operator did see the traffic sign". This representation of the score is then provided as at least a second portion of the input to the machine learning model. For example, in some embodiments, this representation includes or represents the "operator viewing prompt" in the input 301 to generate the "operator viewing prompt", as shown in the output 308.

[0131] Continuing with block 606, some embodiments access first audio data associated with the radio station (e.g., in event data 114) based at least on the geolocation indicator. The geolocation indicator indicates the location of the present machine in the environment. Various embodiments then provide the representation of the audio data as at least a second portion of an input to one or more machine learning models, wherein one or more second natural language characters include or represent a summary of the representation of the audio data. For example, such a second portion may include an "event data prompt" of input 301 to generate a summary such as Figure 3 The “Event Data Output” shown in output 308 .

[0132] Continuing with block 606, some embodiments access destination or travel route information associated with the destination or travel route of the present machine from one or more data sources (e.g., destination / travel route data 124). Some embodiments then provide a representation of the destination or travel route information as at least a second portion of an input to a machine learning model to generate a summarized representation of the destination or travel route information. For example, such a second portion may include or represent the "destination / travel route hint" of input 301 in order to produce the "summary information about the destination / travel route" of output 308.

[0133] According to block 608, some embodiments cause a representation of the second natural language characters to be presented on a display (e.g., an LCD screen or monitor) or an audio device (e.g., a virtual assistant speaker or a car speaker). For example, such a “representation” may represent a second natural language character represented by Figure 1 The content generated by the text-to-speech component 134 and / or the display component 136. In another example, such a representation may be Figure 4A Audio output 414, Figure 5A 504 and / or Figure 5B 510. According to various embodiments, such representation may reflect any of the embodiments described herein with respect to block 606. For example, presentation of a representation of one or more second natural language characters may include a natural language phrase indicating whether a parking space is predicted to be allowed to be occupied by the present machine based at least on the parking context and the parking instruction. In another example, the representation of the second natural language character includes a natural language phrase representing weather data, road condition data, traffic data, and / or event data. In another example, such representation may include audio data representing a first phrase utterance that initiates a conversation with an operator based at least on a first score (alertness level) (e.g., Figure 5A In another example, the representation may include a first phrase utterance (e.g., audio data 504) that initiates a conversation with the operator based on personalized information associated with the operator. In another example, the representation may include a phrase that represents a response to the operator or occupant not seeing the traffic sign.

[0134] In some embodiments, process 600 is performed by one or more processing units, which include at least one of the following: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation of 3D assets; a system for performing deep learning operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system for generating synthetic data; a system for generating synthetic data using AI; a system including one or more virtual machines (VMs); a system implemented at least in part in a data center; or a system implemented at least in part using cloud computing resources, such as with respect to FIG. 10A to FIG. 12 described.

[0135] Figure 7 is a flow chart of an example process 700 for extracting natural language characters from one or more objects according to some embodiments. In some embodiments, the process 700 is performed by Figure 1 According to block 703, some embodiments detect the speed of the machine and receive image data. For example, returning to reference Figure 1 , the speed detection component 110 can detect that the speed of the machine is traveling at a specific speed, such as 35 miles per hour. In addition, the object detector 106 can receive image data from the environment image data 104.

[0136] According to box 705, some embodiments (e.g., natural language extractor 108) determine whether the speed is below and / or above one or more thresholds. For example, using the above example, it is determined whether 35MPH is above the 10MPH threshold. If it is determined to be "no" (e.g., the speed is 10MPH or lower), process 700 stops. In another example, it can be determined in addition (or alternatively) whether 35MPH is above another threshold, namely the 60MPH threshold. Therefore, if it is detected that the speed of the machine is not below the 60MPH threshold (e.g., it is above 60MPH; determined to be "no"), process 700 stops. It will be understood that in some embodiments, the "no" and "yes" decisions at box 705 are interchanged. For example, when the "yes" decision at box 705 (e.g., the speed is detected to be above the threshold), process 700 stops.

[0137] Continuing with process 700 at block 707, if the speed is below and / or above one or more thresholds, certain embodiments (e.g., natural language extractor 108) perform block 707 by determining whether one or more objects have been detected in the image data. For example, natural language extractor 108 may receive image data with a bounding box in the image data and / or a binary or Boolean value (e.g., "true" or "yes") indicating that a traffic sign or other object has been detected from object detector 106. If the determination is "no," process 700 stops.

[0138] If one or more objects are detected (a "yes" decision), certain embodiments (e.g., the natural language extractor 108) execute box 709 by determining whether one or more natural language characters are detected in the object. For example, in some embodiments, the object detector 106 may include computer vision functionality to check whether any natural language characters are present by performing text detection. "Text detection" is the process of identifying areas in image data where text exists. For example, various techniques can be used for text detection, such as sliding window methods or deep learning-based methods such as convolutional neural networks (CNNs). These techniques analyze the image and generate bounding boxes around the text areas. For example, return to reference Figure 4A , an additional bounding box (not shown) may be generated over the natural language label—“15 minute parking for store customers only.” In response, the object detector 106 may transmit an indication (e.g., a Boolean value or image data with a bounding box) indicating that natural language characters are present in the detected object. If no natural language characters are detected in the detected object, the process 700 stops.

[0139] If one or more natural language characters are detected in the detected object, certain embodiments (eg, the natural language extractor 108 ) perform block 711 by performing optical character recognition (OCR) on the detected natural language characters.

[0140] Figure 8 is a flow chart of an example process 800 for generating a natural language response based on a calculated alertness level of an operator of a present machine according to some embodiments. According to block 802, some embodiments receive image data representing a portion of an operator of the present machine, wherein the image data is generated using (one or more) sensors of the present machine traversing an environment. For example, referring back to Figure 2 , the operator alertness level detector 206 may receive image data from the operator image data 204, such as regarding Figure 2 described.

[0141] Per block 804, certain embodiments generate a representation of the calculated operator alertness level (e.g., a direct KSS sleepiness number or score or a natural language sentence indicating the number or score) based at least on the image data. Figure 2 , an operator alertness level detector 296 may be included in a DMS that detects eye gaze, hand drop, head nodding, etc. in image data, where a CNN has been trained to make a specific classification prediction (e.g., the probability that the operator's KSS sleepiness score is 4 is 0.95). In response, via a lookup table, such a classification score is mapped to a natural language sentence (e.g., representation) that states "the user's alertness level is 4".

[0142] In accordance with box 806, certain embodiments provide a representation of the calculated alertness level as an input to one or more machine learning models to generate one or more natural language characters based at least on the calculated operator alertness level. For example, the representation of the calculated alertness level may be represented or included in the "operator alertness level prompt" indicated in input 301 (e.g., an instruction to "ask the user if they want to listen to their favorite music and tell them their alertness level" and an input that the operator's KSS level is 4) to generate a "response to operator alertness level" output in output 308. In some embodiments, additional or alternative inputs (or portions of inputs) may be provided to the machine learning model. For example, a response to an operator alertness level may be provided. Figure 6 Any input described in block 606 or Figure 3 Any (one or more) input 301. Similarly, any suitable natural language characters can be generated, such as Figure 3 Any output in output(s) 308 of.

[0143] According to block 808, some embodiments cause a representation of the natural language characters to be presented on a display or audio device associated with an operator of the machine. For example, block 808 may include causing a representation of the natural language characters to be presented at the audio device 506. Figure 5A or Figure 5Baudio data 504 and / or 510. In some embodiments, process 800 is performed by one or more processing units, the processing units comprising at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation of 3D assets; a system for performing deep learning operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system for generating synthetic data; a system for generating synthetic data using AI; a system comprising one or more virtual machines (VMs); a system implemented at least in part in a data center; or a system implemented at least in part using cloud computing resources, such as with respect to FIG. 10A to FIG. 12 described.

[0144] Fig. 9 is a flow chart of an example process 900 for generating natural language characters based on providing different image data sets to the LLM according to some embodiments. According to block 903, some embodiments receive first image data representing one or more portions of one or more objects in an environment, wherein the first image data is generated using one or more first sensors (e.g., object detection cameras) of the machine traversing the environment (e.g., driving along a street). For example, returning to reference Figure 1 , the object detector 106 may receive the environment image data 104 .

[0145] According to block 905, some embodiments receive second image data representing one or more portions of an operator (e.g., a driver) of the machine, wherein the second image data is generated using one or more second sensors of the machine (e.g., a DMS infrared sensor that can be used to monitor eye movements and track gaze patterns of the operator). For example, referring back to Figure 2 , the operator alertness level detector 205 can receive the operator image data 204 .

[0146] According to block 907, some embodiments provide a representation of information extracted from the first image data and the second image data as at least part of the input to the LLM to generate one or more natural language characters. For example, the representation of information extracted from the first image data may be an object detected by the object detector 106 and / or a natural language character extracted by the natural language extractor 108. In the example of a representation of the second image data, this may mean, for example, a representation of the calculated alertness level and / or the operator image data 204 itself. For example, the representation of the operator image data 204 may represent the output of a CLIP model, where the CLIP model describes each feature of the second image data in natural language (e.g., "This is a photo of a driver with his eyes closed and his head down while both hands are on the steering wheel"). In some embodiments, the output of the CLIP model may alternatively or additionally be a representation of information extracted from the first image data. For example, based on the object detection function of a stop sign, the CLIP model may be used to generate a natural language output stating "This image data shows a red stop sign." Thus, such a natural language output (or a soft prompt representing such an output) may be directly provided to the LLM to produce an output.

[0147] In some embodiments, the representation at block 907 includes or represents Figure 3 In some embodiments, the LLM generates natural language characters at block 907 based on additional or alternative inputs (or portions of inputs). For example, other portions of the input may include: Figure 6 Any prompts described in block 606 and / or Figure 3 Any prompts described in input 301.

[0148] According to block 909, some embodiments cause a representation of the natural language characters to be presented at a device (e.g., a sound device or a display device) associated with an operator of the machine. For example, such a representation at block 909 may be included in a Figure 6 Block 608 and / or Figure 8 In any representation described in box 808.

[0149] The systems and methods described herein may be used by, but are not limited to, non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more adaptive driver assistance systems (ADAS)), driven and undriven robots or robotic platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, flying boats, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, construction vehicles, trains, underwater boats, remotely operated vehicles such as drones, and / or other vehicle types. Further, the systems and methods described herein may be used for a variety of purposes, by way of example and not limitation, for machine control, machine motion, machine driving, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twins, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twins, data center processing, conversational AI, light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for 3D assets, cloud computing, and / or any other suitable application.

[0150] The disclosed embodiments may be included in a variety of different systems, such as automotive systems (e.g., control systems for autonomous or semi-autonomous machines, perception systems for autonomous or semi-autonomous machines), systems implemented using robots, aerial systems, inboard systems, rowing systems, smart area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using edge devices, systems including one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least in part in a data center, systems for performing conversational AI operations, systems for performing light transport simulations, systems for performing collaborative content creation of 3D assets, systems implemented at least in part using cloud computing resources, and / or other types of systems.

[0151] Example Autonomous Vehicle

[0152] Fig. 10Ais an illustration of an example autonomous vehicle 1000 according to some embodiments of the present disclosure. Autonomous vehicle 1000 (alternatively referred to herein as "vehicle 1000") may include, but is not limited to, a passenger vehicle such as a car, a truck, a bus, a first response vehicle, a shuttle, an electric or motorized bicycle, a motorcycle, a fire truck, a police vehicle, an ambulance, a boat, a construction vehicle, an underwater vessel, a robotic vehicle, a drone, an airplane, a vehicle coupled to a trailer (e.g., a semi-tractor trailer for hauling cargo), and / or another type of vehicle (e.g., a vehicle that is unmanned and / or accommodates one or more passengers). Autonomous vehicles are generally described in terms of automation levels as defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE), “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (Standard No. J3016-201806, issued on June 15, 2018, Standard No. J3016-201609, issued on September 30, 2016, and previous and future versions of the standard). The vehicle 1000 may be capable of implementing one or more functions in Levels 3-5 that meet the autonomous driving level. The vehicle 1000 may be capable of implementing one or more functions in Levels 1-5 of the autonomous driving level. For example, depending on the embodiment, the vehicle 1000 may be capable of driver assistance (Level 1), partial automation (Level 2), conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5). The term “autonomy” as used herein may include any and / or all types of autonomy of a vehicle 1000 or other machine, such as fully autonomous, highly autonomous, conditionally autonomous, partially autonomous, providing assisted autonomy, semi-autonomous, primarily autonomous, or other designations.

[0153] The vehicle 1000 may include components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. The vehicle 1000 may include a propulsion system 1050, such as an internal combustion engine, a hybrid power plant, an all-electric engine, and / or another type of propulsion system. The propulsion system 1050 may be connected to a drive train of the vehicle 1000, which may include a transmission, to achieve propulsion of the vehicle 1000. The propulsion system 1050 may be controlled in response to receiving a flag from a throttle / accelerator 1052.

[0154] A steering system 1054, which may include a steering wheel, may be used to steer the vehicle 1000 (e.g., along a desired path or route) when the propulsion system 1050 is operating (e.g., when the vehicle is in motion). The steering system 1054 may receive a signal from a steering actuator 1056. For fully automated (Level 5) functionality, a steering wheel may be optional.

[0155] Brake sensor system 1046 may be used to operate vehicle brakes in response to receiving indications from brake actuator 1048 and / or brake sensors.

[0156] May include one or more system on chip (SoC) 1004 ( Fig. 10C ) and / or one or more controllers 1036 of one or more GPUs can provide flags (e.g., representing commands) to one or more components and / or systems of the vehicle 1000. For example, the one or more controllers can send flags to operate vehicle brakes via one or more brake actuators 1048, operate steering system 1054 via one or more steering actuators 1056, and operate propulsion system 1050 via one or more throttles / accelerators 1052. The one or more controllers 1036 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor flags and output operating commands (e.g., flags representing commands) to achieve autonomous driving and / or assist a human driver in driving the vehicle 1000. The one or more controllers 1036 may include a first controller 1036 for autonomous driving functions, a second controller 1036 for functional safety functions, a third controller 1036 for artificial intelligence functions (e.g., computer vision), a fourth controller 1036 for infotainment functions, a fifth controller 1036 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 1036 may handle two or more of the above functions, two or more controllers 1036 may handle a single function, and / or any combination thereof.

[0157] The one or more controllers 1036 may provide indicia for controlling one or more components and / or systems of the vehicle 1000 in response to sensor data (eg, sensor input) received from one or more sensors. Sensor data may be received from, for example and without limitation, global navigation satellite system (“GNSS”) sensors 1058 (e.g., global positioning system sensors), RADAR sensors 1060 , ultrasonic sensors 1062 , LIDAR sensors 1064 , inertial measurement unit (IMU) sensors 1066 (e.g., accelerometers, gyroscopes, magnetic compasses, magnetometers, etc.), microphones 1096 , stereo cameras 1068 , wide angle cameras 1070 (e.g., fisheye cameras), infrared cameras 1072 , surround cameras 1074 (e.g., 360 degree cameras), long-range and / or mid-range cameras 1098 , speed sensors 1044 (e.g., for measuring the velocity of the vehicle 1000 ), vibration sensors 1042 , steering sensors 1040 , brake sensors (e.g., as part of a brake sensor system 1046 ), one or more occupant monitoring system (OMS) sensors 1001 (e.g., one or more interior cameras), and / or other sensor types.

[0158] One or more of the controllers 1036 may receive inputs (e.g., represented by input data) from the instrument set 1032 of the vehicle 1000 and provide outputs (e.g., represented by output data, display data, etc.) via a human machine interface (HMI) display 1034, an audible indicator, a speaker, and / or via other components of the vehicle 1000. These outputs may include information such as vehicle speed, velocity, time, map data (e.g., Fig. 10C The HMI display 1034 may include information such as a high definition (“HD”) map 1022 of the vehicle 1000 , location data (e.g., the location of the vehicle 1000 on the map), directions, locations of other vehicles (e.g., an occupancy grid), information about objects and states of objects as sensed by the controller 1036 , and the like. For example, the HMI display 1034 may display information about the presence of one or more objects (e.g., street signs, warning signs, traffic light changes, etc.) and / or information about driving maneuvers that the vehicle has made, is making, or will make (e.g., changing lanes now, leaving 34B in two miles, etc.).

[0159] The vehicle 1000 also includes a network interface 1024 that can communicate via one or more networks using one or more wireless antennas 1026 and / or a modem. For example, the network interface 1024 may be capable of communicating via Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multi-Carrier (“CDMA2000”), etc. The one or more wireless antennas 1026 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using one or more local area networks such as Bluetooth, Bluetooth Low Energy (“LE”), Z-wave, ZigBee, etc., and / or one or more low power wide area networks (LPWANs) such as LoRaWAN, SigFox, etc.

[0160] Fig. 10B For use according to some embodiments of the present disclosure Fig. 10A 1000. The cameras and respective fields of view are an example embodiment and are not intended to be limiting. For example, additional and / or replaceable cameras may be included and / or the cameras may be located at different locations on the vehicle 1000.

[0161] The camera type used for the camera may include, but is not limited to, a digital camera that may be suitable for use with components and / or systems of the vehicle 800. The camera may operate at an automotive safety integrity level (ASIL) B and / or at another ASIL. The camera type may have any image capture rate, such as 60 frames per second (fps), 120fps, 240fps, etc., depending on the embodiment. The camera may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include a red-white-white-white (RCCC) color filter array, a red-white-white-blue (RCCB) color filter array, a red-blue-green-white (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras such as cameras with RCCC, RCCB, and / or RBGC color filter arrays may be used in efforts to improve light sensitivity.

[0162] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-function monocular camera can be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. One or more of the cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).

[0163] One or more of the cameras may be mounted in a mounting assembly such as a custom designed (three-dimensional ("3D") printed) assembly to cut out stray light and reflections from within the car (e.g., reflections from the dashboard reflected in the windshield mirror) that may interfere with the camera's image data capture capabilities. With respect to the wing mirror mounting assembly, the wing mirror assembly may be custom 3D printed so that the camera mounting plate matches the shape of the wing mirror. In some examples, one or more cameras may be integrated into the wing mirror. For side view cameras, one or more cameras may also be integrated into the four pillars at each corner of the cab.

[0164] A camera (e.g., a front-facing camera) having a field of view that includes a portion of the environment in front of the vehicle 1000 can be used for surround vision to help identify the forward path and obstacles, as well as assist in providing information critical to generating an occupancy grid and / or determining a preferred vehicle path with the help of one or more controllers 1036 and / or control SoCs. The front-facing camera can be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. The front-facing camera can also be used for ADAS functions and systems, including lane departure warning ("LDW"), autonomous cruise control ("ACC"), and / or other functions such as traffic sign recognition.

[0165] A variety of cameras may be used in the front-facing configuration, including, for example, a monocular camera platform including a complementary metal oxide semiconductor ("CMOS") color imager. Another example may be a wide-angle camera 1070, which may be used to sense objects (e.g., pedestrians, intersection traffic, or bicycles) entering the field of view from the periphery. Fig. 10B Only one wide-angle camera is shown in the figure, but there may be any number (including zero) of wide-angle cameras 1070 on the vehicle 1000. In addition, any number of long-range cameras 1098 (e.g., a long-view stereo camera pair) may be used for depth-based object detection, especially for objects for which a neural network has not been trained. Long-range cameras 1098 may also be used for object detection and classification and basic object tracking.

[0166] Any number of stereo cameras 1068 may also be included in the front configuration. In at least one embodiment, one or more stereo cameras 1068 may include an integrated control unit including a scalable processing unit that may provide a multi-core microprocessor and programmable logic ("FPGA") with an integrated controller area network ("CAN") or Ethernet interface on a single chip. Such a unit can be used to generate a 3D map of the vehicle environment, including distance estimates for all points in the image. Alternative stereo cameras 1068 may include a compact stereo vision sensor that may include two camera lenses (one on the left and one on the right) and an image processing chip that may measure the distance from the vehicle to the target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 1068 may be used in addition to or alternatively to those described herein.

[0167] A camera having a field of view that includes portions of the environment to the sides of the vehicle 1000 (e.g., a side view camera) can be used for surround viewing, providing information used to create and update an occupancy grid and generate side impact collision warnings. Fig. 10B The four surround cameras 1074 shown in FIG. 1074 may be placed on the vehicle 1000. The surround cameras 1074 may include a wide-angle camera 1070, a fisheye camera, a 360-degree camera, and / or the like. For example, the four fisheye cameras may be placed in front, behind, and on the sides of the vehicle. In an alternative arrangement, the vehicle may use three surround cameras 1074 (e.g., left, right, and rear), and may utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround camera.

[0168] A camera having a field of view that includes a portion of the environment behind the vehicle 1000 (e.g., a rear view camera) can be used to assist with parking, surround view, rear collision warning, and creating and updating an occupancy grid. A variety of cameras can be used, including but not limited to cameras that are also suitable as front-facing cameras as described herein (e.g., long-range and / or mid-range cameras 1098, stereo cameras 1068, infrared cameras 1072, etc.).

[0169] A camera (e.g., one or more OMS sensors 1001) whose field of view includes a portion of the interior environment within the cabin of the vehicle 1000 may be used as part of an occupant monitoring system (OMS), such as, but not limited to, a driver monitoring system (DMS). For example, an OMS sensor (e.g., OMS sensor 1001) may be used (e.g., by controller 1036) to track the gaze direction, head posture, and / or blinking of an occupant and / or driver. The gaze information may be used to determine the attention level of the occupant or driver (e.g., to detect drowsiness, fatigue, and / or distraction), and / or to take responsive actions to prevent harm to the occupant or operator. In some embodiments, data from the OMS sensor may be used to implement gaze control operations triggered by a driver and / or non-driver occupant, such as, but not limited to, adjusting cabin temperature and / or airflow, opening and closing windows, controlling cabin lighting, controlling an entertainment system, adjusting rearview mirrors, adjusting seat positions, and / or other operations. In some embodiments, the OMS may be used for applications such as determining when an object and / or occupant remains in the cabin (e.g., by detecting the presence of an occupant after the driver leaves the vehicle).

[0170] Fig. 10C For use according to some embodiments of the present disclosure Fig. 10A Block diagram of an example system architecture of an example autonomous vehicle 1000. It should be understood that this arrangement and other arrangements described herein are set forth merely as examples. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used in addition to or in place of those shown, and some elements may be omitted entirely. Further, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in combination with other components, and in any appropriate combination and location. The various functions described herein as being performed by an entity may be implemented by hardware, firmware, and / or software. For example, the various functions may be implemented by a processor executing instructions stored in a memory.

[0171] Fig. 10C Each of the components, features, and systems of vehicle 1000 is illustrated as being connected via bus 1002. Bus 1002 may include a controller area network (CAN) data interface (alternatively, referred to herein as a "CAN bus"). CAN may be a network inside vehicle 1000 that assists in controlling various features and functions of vehicle 1000, such as driving of brakes, acceleration, braking, steering, windshield wipers, and the like. The CAN bus may be configured to have tens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus may be read to find steering wheel angle, ground speed, engine speed per minute (RPM), button position, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.

[0172] Although bus 1002 is described as a CAN bus here, this is not intended to be restrictive. For example, in addition to or alternatively to the CAN bus, FlexRay and / or Ethernet can be used. In addition, although bus 1002 is represented by a single line, this is not intended to be restrictive. For example, there can be any number of buses 1002, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more buses of other types using different protocols. In some examples, two or more buses 1002 can be used to perform different functions, and / or can be used for redundancy. For example, a first bus 1002 can be used for a collision avoidance function, and a second bus 1002 can be used for driving control. In any example, each bus 1002 can communicate with any component of the vehicle 1000, and two or more buses 1002 can communicate with the same component. In some examples, each SoC 1004, each controller 1036, and / or each computer within the vehicle may have access to the same input data (e.g., input from sensors of the vehicle 1000) and may be connected to a common bus such as a CAN bus.

[0173] The vehicle 1000 may include one or more controllers 1036, such as those described herein. Fig. 10A Those controllers described. Controller 1036 can be used for a variety of functions. Controller 1036 can be coupled to any other different components and systems of vehicle 1000, and can be used for control of vehicle 1000, artificial intelligence of vehicle 1000, infotainment for vehicle 1000, and / or the like.

[0174] The vehicle 1000 may include one or more system on chip (SoC) 1004. The SoC 1004 may include a CPU 1006, a GPU 1008, a processor 1010, a cache 1012, an accelerator 1014, a data store 1016, and / or other components and features not shown. The SoC 1004 may be used to control the vehicle 1000 in a variety of platforms and systems. For example, one or more SoCs 1004 may be combined with an HD map 1022 in a system (e.g., a system of the vehicle 1000), and the HD map may be downloaded from one or more servers (e.g., a server) via a network interface 1024. Fig. 10D One or more servers 1078) obtain map refreshes and / or updates.

[0175] CPU 1006 may include a CPU cluster or CPU complex (alternatively, referred to herein as "CCPLEX"). CPU 1006 may include multiple cores and / or L2 caches. For example, in some embodiments, CPU 1006 may include eight cores in a coherent multiprocessor configuration. In some embodiments, CPU 1006 may include four dual-core clusters, each of which has a dedicated L2 cache (e.g., 2MB L2 cache). CPU 1006 (e.g., CCPLEX) may be configured to support simultaneous cluster operations so that any combination of clusters of CPU 1006 can be active at any given time.

[0176] CPU 1006 may implement power management capabilities including one or more of the following features: each hardware block may be automatically clock gated when idle to save dynamic power; each core clock may be gated when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core may be independently power gated; each core cluster may be independently clock gated when all cores are clock gated or power gated; and / or each core cluster may be independently power gated when all cores are power gated. CPU 1006 may further implement an enhanced algorithm for managing power states, in which allowed power states and expected wake-up times are specified, and hardware / microcode determines the optimal power state to enter for the core, cluster, and CCPLEX. The processing core may support a simplified power state entry sequence in software, with the work being offloaded to the microcode.

[0177] GPU 1008 may include an integrated GPU (alternatively, referred to herein as an "iGPU"). GPU 1008 may be programmable and efficient for parallel workloads. In some examples, GPU 1008 may use an enhanced tensor instruction set. GPU 1008 may include one or more streaming microprocessors, each of which may include an L1 cache (e.g., an L1 cache with at least 96KB storage capacity), and two or more of these streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB storage capacity). In some embodiments, GPU 1008 may include at least eight streaming microprocessors. GPU 1008 may use a computing application programming interface (API). In addition, GPU 1008 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0178] In the case of automotive and embedded use, GPU 1008 can be power optimized to achieve optimal performance. For example, GPU 1008 can be manufactured on fin field effect transistors (FinFETs). However, this is not intended to be limiting, and GPU 1008 can be manufactured using other semiconductor manufacturing processes. Each streaming microprocessor can merge several mixed precision processing cores divided into multiple blocks. For example and without limitation, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed precision NVIDIA tensor cores for deep learning matrix arithmetic, L0 instruction cache, warp scheduler, dispatch unit and / or 64KB register file. In addition, the streaming microprocessor may include independent parallel integer and floating point data paths to provide efficient execution of workloads using a mix of computation and addressing computations. The streaming microprocessor may include independent thread scheduling capabilities to allow for finer-grained synchronization and cooperation between parallel threads. The streaming microprocessor may include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.

[0179] GPU 1008 can include high bandwidth memory (HBM) and / or a 16GB HBM2 memory subsystem that provides a peak memory bandwidth of approximately 900 GB / s in some examples. In some examples, synchronous graphics random access memory (SGRAM), such as fifth generation graphics double data rate synchronous random access memory (GDDR5), can be used in addition to or in lieu of HBM memory.

[0180] The GPU 1008 may include unified memory technology that includes access counters to allow memory pages to be more accurately migrated to the processor that accesses them most frequently, thereby improving the efficiency of memory ranges shared between processors. In some examples, address translation service (ATS) support may be used to allow the GPU 1008 to directly access the CPU 1006 page table. In such an example, when the GPU 1008 memory management unit (MMU) experiences a miss, an address translation request may be transmitted to the CPU 1006. In response, the CPU 1006 may look up the virtual-physical mapping for the address in its page table and transmit the translation back to the GPU 1008. In this way, unified memory technology may allow a single unified virtual address space to be used for memory of both the CPU 1006 and the GPU 1008, thereby simplifying GPU 1008 programming and porting applications to the GPU 1008.

[0181] In addition, GPU 1008 may include access counters that can track how often GPU 1008 accesses the memory of other processors. Access counters can help ensure that memory pages are moved to the physical memory of the processor that accesses them most frequently.

[0182] SoC 1004 may include any number of caches 1012, including those described herein. For example, cache 1012 may include an L3 cache available to both CPU 1006 and GPU 1008 (e.g., connected to both CPU 1006 and GPU 1008). Cache 1012 may include a write-back cache that may track the state of a line, for example, by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, although smaller cache sizes may also be used.

[0183] The SoC 1004 may include an arithmetic logic unit (ALU) that may be utilized in performing any of a variety of tasks or operations with respect to the vehicle 1000, such as processing a DNN. Additionally, the SoC 1004 may include a floating point unit (FPU) (or other math coprocessor or digital coprocessor type) for performing mathematical operations within the system. For example, the SoC 1004 may include one or more FPUs integrated as execution units within the CPU 1006 and / or GPU 1008.

[0184] SoC 1004 may include one or more accelerators 1014 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, SoC 1004 may include a hardware accelerator cluster, which may include optimized hardware accelerators and / or large on-chip memory. The large on-chip memory (e.g., 4MB SRAM) may enable the hardware accelerator cluster to accelerate neural networks and other calculations. The hardware accelerator cluster may be used to secondary GPU 1008 and offload some tasks of GPU 1008 (e.g., free up more cycles of GPU 1008 for performing other tasks). As an example, accelerator 1014 may be used for targeted workloads (e.g., perception, convolutional neural networks (CNNs), etc.) that are sufficiently stable to easily control acceleration. When used herein, the term "CNN" may include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).

[0185] Accelerator 1014 (e.g., a hardware accelerator cluster) may include a deep learning accelerator (DLA). DLA may include one or more tensor processing units (TPUs) that may be configured to provide an additional 10 trillion operations per second for deep learning applications and reasoning. TPU may be an accelerator configured to perform image processing functions (e.g., for CNN, RCNN, etc.) and optimized for performing image processing functions. DLA may be further optimized for a specific set of neural network types and floating point operations and reasoning. The design of DLA may provide higher performance per millimeter than a general-purpose GPU, and far exceeds the performance of a CPU. TPU may perform several functions, including a single instance convolution function, support, for example, INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.

[0186] The DLA can quickly and efficiently execute neural networks, particularly CNNs, on processed or unprocessed data for any of a wide variety of functions, such as, but not limited to: a CNN for object recognition and detection using data from a camera sensor; a CNN for distance estimation using data from a camera sensor; a CNN for emergency vehicle detection and identification and detection using data from a microphone; a CNN for facial recognition and vehicle owner identification using data from a camera sensor; and / or a CNN for safety and / or security-related events.

[0187] The DLA can perform any function of the GPU 1008, and by using an inference accelerator, for example, the designer can target any function to either the DLA or the GPU 1008. For example, the designer can focus the processing of CNNs and floating point operations on the DLA, and leave other functions to the GPU 1008 and / or other accelerators 1014.

[0188] The accelerator 1014 (e.g., a hardware accelerator cluster) may include a programmable vision accelerator (PVA), which may be alternatively referred to as a computer vision accelerator herein. The PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA may provide a balance between performance and flexibility. For example, each PVA may include, for example and without limitation, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.

[0189] The RISC core can interact with an image sensor (e.g., an image sensor of any camera described herein), an image marker processor, and / or the like. Each of these RISC cores can include any amount of memory. Depending on the embodiment, the RISC core can use any of a number of protocols. In some examples, the RISC core can execute a real-time operating system (RTOS). The RISC core can be implemented using one or more integrated circuit devices, application specific integrated circuits (ASICs), and / or storage devices. For example, the RISC core can include an instruction cache and / or a tightly coupled RAM.

[0190] The DMA may enable components of the PVA to access system memory independently of the CPU 1006. The DMA may support any number of features used to provide optimizations to the PVA, including but not limited to support for multi-dimensional addressing and / or circular addressing. In some examples, the DMA may support addressing in up to six or more dimensions, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.

[0191] The vector processor can be a programmable processor that can be designed to efficiently and flexibly perform programming for computer vision algorithms and provide logo processing capabilities. In some examples, the PVA can include a PVA core and two vector processing subsystem partitions. The PVA core can include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem can operate as the main processing engine of the PVA, and can include a vector processing unit (VPU), an instruction cache, and / or a vector memory (e.g., VMEM). The VPU core can include a digital logo processor, such as, for example, a single instruction multiple data (SIMD), a very long instruction word (VLIW) digital logo processor. The combination of SIMD and VLIW can enhance throughput and rate.

[0192] Each of the vector processors may include an instruction cache and may be coupled to a dedicated memory. As a result, in some examples, each of the vector processors may be configured to execute independently of the other vector processors. In other examples, the vector processors included in a particular PVA may be configured to employ data parallelism. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms on sequence images or portions of an image. Among other things, any number of PVAs may be included in a hardware accelerator cluster, and any number of vector processors may be included in each of these PVAs. In addition, the PVA may include additional error correction code (ECC) memory to enhance overall system security.

[0193] The accelerator 1014 (e.g., a hardware accelerator cluster) may include an on-chip computer vision network and SRAM to provide high bandwidth, low latency SRAM for the accelerator 1014. In some examples, the on-chip memory may include at least 4MB of SRAM consisting of, for example and not limited to, eight field-configurable memory blocks that can be accessed by both the PVA and the DLA. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and DLA may access the memory via a backbone that provides high-speed memory access to the PVA and DLA. The backbone may include an on-chip computer vision network that interconnects the PVA and DLA to the memory (e.g., using APB).

[0194] The on-chip computer vision network may include an interface that determines that both the PVA and DLA provide ready and valid flags before transmitting any control flags / addresses / data. Such an interface may provide separate phases and separate channels for transmitting control flags / addresses / data, as well as burst communication for continuous data transmission. This type of interface may comply with ISO 26262 or IEC 615010 standards, but other standards and protocols may also be used.

[0195] In some examples, SoC 1004 may include a real-time ray tracing hardware accelerator such as described in U.S. Patent Application No. 16 / 101,232 filed on August 10, 2018. The real-time ray tracing hardware accelerator may be used to quickly and efficiently determine the position and range of objects (e.g., within a world model) to generate real-time visualization simulations for RADAR signature interpretation, for sound propagation synthesis and / or analysis, for SONAR system simulations, for general wave propagation simulations, for comparison with LIDAR data for the purpose of positioning and / or other functions, and / or for other purposes. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing related operations.

[0196] The accelerator 1014 (e.g., a hardware accelerator cluster) has a wide range of uses in autonomous driving. The PVA can be a programmable visual accelerator that can be used for key processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are a good match for algorithmic domains that require predictable processing, low power, and low latency. In other words, the PVA performs well on semi-dense or dense rule calculations, and even on small data sets that require predictable runtimes with low latency and low power. Therefore, in the context of a platform for autonomous vehicles, the PVA is designed to execute classic computer vision algorithms because they are effective at object detection and integer math operations.

[0197] For example, according to one embodiment of the technology, PVA is used to perform computer stereo vision. In some examples, algorithms based on semi-global matching can be used, but this is not intended to be limiting. Many applications for level 3-5 autonomous driving require instant motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.). PVA can perform computer stereo vision functions on input from two monocular cameras.

[0198] In some examples, PVA can be used to perform dense optical flow. According to the process raw RADAR data (e.g., using 4D fast Fourier transform) to provide processed RADAR. In other examples, PVA is used for time-of-flight depth processing, which is for example by processing raw time-of-flight data to provide processed time-of-flight data.

[0199] DLA can be used to perform one or more operations of any type of network to enhance control and driving safety, including, for example, a neural network that outputs a confidence measure for each object detection. Such confidence values ​​can be interpreted as probabilities, or as providing a relative "weight" of each detection compared to other detections. The confidence value enables the system to make further decisions about which detections should be considered true positive detections rather than false positive detections. For example, the system can set a threshold for confidence and only consider detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, a false positive detection can cause the vehicle to automatically perform emergency braking, which is obviously undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can execute a neural network for regressing confidence values. The neural network may take as its input at least some subset of parameters, such as bounding box dimensions, a ground plane estimate obtained (e.g., from another subsystem), an inertial measurement unit (IMU) sensor 1066 output related to the orientation and distance of the vehicle 1000, a 3D position estimate of the object obtained from the neural network and / or other sensors (e.g., a LIDAR sensor 1064 or a RADAR sensor 1060), etc.

[0200] SoC 1004 may include one or more data stores 1016 (e.g., memory). Data store 1016 may be on-chip memory of SoC 1004 that may store neural networks to be executed on the GPU and / or DLA. In some examples, data store 1016 may be large enough to store multiple instances of a neural network for redundancy and safety. Data store 1016 may include an L2 or L3 cache 1012. References to data store 1016 may include references to memory associated with a PVA, DLA, and / or other accelerator 1014 as described herein.

[0201] SoC 1004 may include one or more processors 1010 (e.g., embedded processors). Processor 1010 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions and related safety implementations. The boot and power management processor may be part of the SoC 1004 boot sequence and may provide runtime power management services. The boot power and management processor may provide clock and voltage programming, auxiliary system low power state transitions, SoC 1004 thermal and temperature sensor management, and / or SoC 1004 power state management. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to the temperature, and SoC 1004 may use the ring oscillator to detect the temperature of CPU 1006, GPU 1008, and / or accelerator 1014. If it is determined that the temperature exceeds a threshold, the boot and power management processor may enter a temperature fault routine and place SoC 1004 in a lower power state and / or place vehicle 1000 in a driver safety parking mode (e.g., parking vehicle 1000 safely).

[0202] Processor 1010 may also include a set of embedded processors that can be used as an audio processing engine. The audio processing engine can be an audio subsystem that allows full hardware support for multi-channel audio through multiple interfaces and a wide range of flexible audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital logo processor with dedicated RAM.

[0203] Processor 1010 may also include an always-on processor engine that can provide the necessary hardware features to support low-power sensor management and wake-up use cases. The always-on processor engine may include a processor core, tightly coupled RAM, supporting peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0204] Processor 1010 may also include a safety cluster engine that includes a dedicated processor subsystem that handles safety management of automotive applications. The safety cluster engine may include two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In safety mode, the two or more cores may operate in lockstep mode and act as a single core with comparison logic to detect any differences between their operations.

[0205] Processor 1010 may also include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.

[0206] Processor 1010 may also include a high dynamic range logo processor, which may include an image logo processor, which is a hardware engine that is part of the camera processing pipeline.

[0207] The processor 1010 may include a video image compositer, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by the video playback application to produce a final image for the player window. The video image compositer may perform lens distortion correction for the wide-angle camera 1070, the surround camera 1074, and / or for an in-cab monitoring camera sensor. The in-cab monitoring camera sensor is preferably monitored by a neural network running on another instance of the advanced SoC, configured to recognize in-cab events and respond accordingly. The in-cab system may perform lip reading to activate mobile phone service and place a call, dictate an email, change a vehicle destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Certain functions are only available to the driver when the vehicle is operating in autonomous mode and are disabled in other circumstances.

[0208] The video image compositer may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, in the case of motion in the video, the noise reduction appropriately weights the spatial information and reduces the weight of information provided by adjacent frames. In the case where an image or portion of an image does not include motion, the temporal noise reduction performed by the video image compositer may use information from previous images to reduce noise in the current image.

[0209] The video image compositor may also be configured to perform stereoscopic rectification on the input stereoscopic footage frames. The video image compositor may further be used for user interface composition when the operating system desktop is in use and the GPU 1008 does not need to continuously render new surfaces. Even when the GPU 1008 is powered on and active for 3D rendering, the video image compositor may be used to offload the GPU 1008 to improve performance and responsiveness.

[0210] SoC 1004 may also include a mobile industry processor interface (MIPI) camera serial interface for receiving video and input from a camera, a high-speed interface, and / or a video input block that may be used for camera and related pixel input functions. SoC 1004 may also include an input / output controller that may be controlled by software and may be used to receive I / O flags that are not committed to a specific role.

[0211] The SoC 1004 may also include a wide range of peripheral interfaces to enable communication with peripherals, audio codecs, power management and / or other devices. The SoC 1004 may be used to process data from cameras (connected via Gigabit multimedia serial links and Ethernet), sensors (e.g., LIDAR sensor 1064, RADAR sensor 1060, etc., which may be connected via Ethernet), data from the bus 1002 (e.g., speed of the vehicle 1000, steering wheel position, etc.), data from the GNSS sensor 1058 (connected via Ethernet or CAN bus). The SoC 1004 may also include dedicated high-performance mass storage controllers, which may include their own DMA engines, and which may be used to free the CPU 1006 from routine data management tasks.

[0212] SoC 1004 can be an end-to-end platform with a flexible architecture that spans automation levels 3-5, thereby providing a comprehensive functional safety architecture that utilizes and efficiently uses computer vision and ADAS technologies to achieve diversity and redundancy, together with deep learning tools to provide a platform for a flexible and reliable driving software stack. SoC 1004 can be faster, more reliable, and even more energy efficient and space efficient than conventional systems. For example, when combined with CPU 1006, GPU 1008, and data storage 1016, accelerator 1014 can provide a fast and efficient platform for level 3-5 autonomous vehicles.

[0213] The technology thus provides capabilities and functionality that cannot be achieved by conventional systems. For example, computer vision algorithms can be executed on CPUs that can be configured to execute a wide variety of processing algorithms across a wide variety of visual data using high-level programming languages ​​such as the C programming language. However, CPUs often fail to meet the performance requirements of many computer vision applications, such as those related to, for example, execution time and power consumption. In particular, many CPUs are unable to execute complex object detection algorithms in real time, which is a requirement for in-vehicle ADAS applications and for practical Level 3-5 autonomous vehicles.

[0214] In contrast to conventional systems, by providing a CPU complex, a GPU complex, and a cluster of hardware accelerators, the techniques described herein allow multiple neural networks to be executed simultaneously and / or sequentially, and the results to be combined together to achieve Level 3-5 autonomous driving functions. For example, a CNN executed on a DLA or dGPU (e.g., GPU 1020) may include text and word recognition, allowing the supercomputer to read and understand traffic signs, including signs for which a neural network has not been specifically trained. The DLA may also include a neural network that can recognize, interpret, and provide semantic understanding of the signs, and pass that semantic understanding to a path planning module running on the CPU complex.

[0215] As another example, as required for Level 3, 4, or 5 driving, multiple neural networks may be executed simultaneously. For example, a warning sign consisting of "Caution: Flashing Lights Indicate Icing Conditions" along with electric lights may be interpreted by several neural networks independently or collectively. The sign itself may be recognized as a traffic sign by a deployed first neural network (e.g., a trained neural network), and the text "Flashing Lights Indicate Icing Conditions" may be interpreted by a deployed second neural network that informs the vehicle's path planning software (preferably executing on a CPU complex) that icing conditions exist when the flashing lights are detected. The flashing lights may be identified by operating a deployed third neural network over multiple frames that informs the vehicle's path planning software of the presence (or absence) of the flashing lights. All three neural networks may be executed simultaneously, for example, within a DLA and / or on GPU 1008.

[0216] In some examples, a CNN for facial recognition and owner recognition can use data from the camera sensor to identify the presence of an authorized driver and / or owner of the vehicle 1000. The always-on sensor processing engine can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and in security mode, disable the vehicle when the owner leaves the vehicle. In this way, the SoC 1004 provides security against theft and / or carjacking.

[0217] In another example, a CNN for emergency vehicle detection and identification can use data from microphone 1096 to detect and identify emergency vehicle sirens. In contrast to conventional systems that use general classifiers to detect sirens and manually extract features, SoC 1004 uses CNN to classify environmental and urban sounds and to classify visual data. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative closing rate of emergency vehicles (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area in which the vehicle is operating, as identified by the GNSS sensor 1058. Thus, for example, when operating in the European Union, the CNN will seek to detect EU sirens, and when in the United States, the CNN will seek to identify only North American sirens. Once an emergency vehicle is detected, with the assistance of the ultrasonic sensor 1062, the control program can be used to execute emergency vehicle safety routines to slow the vehicle, drive to the side of the road, stop the vehicle, and / or idle the vehicle until the emergency vehicle passes.

[0218] The vehicle may include a CPU 1018 (e.g., a discrete CPU or dCPU) that may be coupled to the SoC 1004 via a high-speed interconnect (e.g., PCIe). The CPU 1018 may include, for example, an X106 processor. The CPU 1018 may be used to perform any of a variety of functions, including, for example, arbitrating potentially inconsistent results between ADAS sensors and the SoC 1004, and / or monitoring the status and health of the controller 1036 and / or the infotainment SoC 1030.

[0219] The vehicle 1000 may include a GPU 1020 (e.g., a discrete GPU or dGPU) that may be coupled to the SoC 1004 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU 1020 may provide additional artificial intelligence functionality, for example, by executing redundant and / or different neural networks, and may be used to train and / or update the neural network based at least in part on input from sensors of the vehicle 1000 (e.g., sensor data).

[0220] The vehicle 1000 may also include a network interface 1024, which may include one or more wireless antennas 1026 (e.g., one or more wireless antennas for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). The network interface 1024 can be used to enable wireless connections to the cloud (e.g., with a server 1078 and / or other network devices), to other vehicles, and / or to computing devices (e.g., a passenger's client device) via the Internet. In order to communicate with other vehicles, a direct link can be established between the two vehicles, and / or an indirect link can be established (e.g., across a network and through the Internet). The direct link can be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide the vehicle 1000 with information about vehicles approaching the vehicle 1000 (e.g., vehicles in front of, to the side of, and / or behind the vehicle 1000). This functionality can be part of the cooperative adaptive cruise control functionality of the vehicle 1000.

[0221] The network interface 1024 may include a SoC that provides modulation and demodulation functions and enables the controller 1036 to communicate over a wireless network. The network interface 1024 may include a radio frequency front end for up-conversion from baseband to radio frequency and down-conversion from radio frequency to baseband. The frequency conversion may be performed by a known process and / or may be performed using a super-heterodyne process. In some examples, the radio frequency front end function may be provided by a separate chip. The network interface may include a wireless function for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0222] The vehicle 1000 may also include data storage 1028, which may include off-chip storage (e.g., outside the SoC 1004). The data storage 1028 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard disk, and / or other components and / or devices that can store at least one bit of data.

[0223] The vehicle 1000 may also include a GNSS sensor 1058. The GNSS sensor 1058 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) is used to assist in mapping, perception, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 1058 may be used, including, for example and without limitation, GPS using a USB connector with an Ethernet to serial (RS-232) bridge.

[0224] The vehicle 1000 may also include a RADAR sensor 1060. The RADAR sensor 1060 may be used by the vehicle 1000 for remote vehicle detection even in darkness and / or inclement weather conditions. The RADAR functional safety level may be ASILB. The RADAR sensor 1060 may use CAN and / or bus 1002 (e.g., to transmit data generated by the RADAR sensor 1060) for control and access to object tracking data, accessing Ethernet in some examples to access raw data. A variety of RADAR sensor types may be used. For example and without limitation, the RADAR sensor 1060 may be suitable for front, rear, and side RADAR use. In some examples, a pulse Doppler RADAR sensor is used.

[0225] The RADAR sensor 1060 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, and the like. In some examples, the long-range RADAR may be used for adaptive cruise control functions. The long-range RADAR system may provide a wide field of view (e.g., within a range of 250m) achieved by two or more independent scans. The RADAR sensor 1060 may help distinguish between static objects and moving objects, and may be used by the ADAS system for emergency braking assistance and forward collision warnings. The long-range RADAR sensor may include a single-station multimode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the central four antennas may create a focused beam pattern designed to record the surroundings of the vehicle 1000 at a higher rate with minimal traffic interference from adjacent lanes. The other two antennas may expand the field of view, making it possible to quickly detect vehicles entering or leaving the lane of the vehicle 1000.

[0226] As an example, a medium-range RADAR system may include a range of up to 1060m (front) or 80m (rear) and a field of view of up to 42 degrees (front) or 1050 degrees (rear). A short-range RADAR system may include, but is not limited to, a RADAR sensor designed to be mounted at both ends of a rear bumper. When mounted at both ends of a rear bumper, such a RADAR sensor system may create two beams that continuously monitor blind spots behind and beside the vehicle.

[0227] Short-range RADAR systems can be used in ADAS systems for blind spot detection and / or lane change assistance.

[0228] The vehicle 1000 may also include ultrasonic sensors 1062. The ultrasonic sensors 1062, which may be placed on the front, rear, and / or sides of the vehicle 1000, may be used for parking assistance and / or creating and updating an occupancy grid. A variety of ultrasonic sensors 1062 may be used, and different ultrasonic sensors 1062 may be used for different detection ranges (e.g., 2.5 m, 4 m). The ultrasonic sensors 1062 may operate at a functional safety level of ASIL B.

[0229] The vehicle 1000 may include a LIDAR sensor 1064. The LIDAR sensor 1064 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 1064 may be ASIL B for functional safety level. In some examples, the vehicle 1000 may include multiple LIDAR sensors 1064 (e.g., two, four, six, etc.) that may use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).

[0230] In some examples, the LIDAR sensor 1064 may be able to provide a list of objects and their distances for a 360-degree field of view. Commercially available LIDAR sensors 1064 may have, for example, an advertised range of approximately 1000m, an accuracy of 2cm-3cm, and support for 1000Mbps Ethernet connections. In some examples, one or more non-protruding LIDAR sensors 1064 may be used. In such examples, the LIDAR sensor 1064 may be implemented as a small device that can be embedded in the front, back, side, and / or corner of the vehicle 1000. In such an example, the LIDAR sensor 1064 may provide a field of view of up to 120 degrees horizontally and 35 degrees vertically, with a range of 200m, even for low reflectivity objects. The front-mounted LIDAR sensor 1064 may be configured for a horizontal field of view between 45 and 135 degrees.

[0231] In some examples, LIDAR technologies such as 3D flash LIDAR can also be used. 3D flash LIDAR uses flashes of laser as an emission source to illuminate the vehicle's surroundings up to about 200m. The flash LIDAR unit includes a receiver that records the laser pulse transmission time and the reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LIDAR can allow highly accurate and distortion-free images of the surrounding environment to be generated with each laser flash. In some examples, four flash LIDAR sensors can be deployed, one on each side of the vehicle 1000. Available 3D flash LIDAR systems include solid-state 3D staring array LIDAR cameras (e.g., non-scanning LIDAR devices) with no moving parts other than fans. The flash LIDAR device can use 5 nanosecond Class I (eye-safe) laser pulses per frame, and can capture the reflected laser light in the form of a 3D range point cloud and co-registered intensity data. By using flash LIDAR, and because flash LIDAR is a solid-state device with no moving parts, the LIDAR sensor 1064 may be less susceptible to motion blur, vibration, and / or shock.

[0232] The vehicle may also include an IMU sensor 1066. In some examples, the IMU sensor 1066 may be located at the center of the rear axle of the vehicle 1000. The IMU sensor 1066 may include, for example and without limitation, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 1066 may include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 1066 may include an accelerometer, a gyroscope, and a magnetometer.

[0233] In some embodiments, the IMU sensor 1066 can be implemented as a miniature high-performance GPS-assisted inertial navigation system (GPS / INS) that combines a micro-electromechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filter algorithm to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 1066 can enable the vehicle 1000 to estimate heading by directly observing and correlating velocity changes from the GPS to the IMU sensor 1066 without the need for input from a magnetic sensor. In some examples, the IMU sensor 1066 and the GNSS sensor 1058 can be combined into a single integrated unit.

[0234] The vehicle may include microphones 1096 positioned in and / or around the vehicle 1000. The microphones 1096 may be used for, among other things, emergency vehicle detection and identification.

[0235] The vehicle may also include any number of camera types, including stereo cameras 1068, wide-angle cameras 1070, infrared cameras 1072, surround cameras 1074, long-range and / or mid-range cameras 1098, and / or other camera types. These cameras can be used to capture image data around the entire periphery of the vehicle 1000. The type of camera used depends on the embodiment and the requirements of the vehicle 1000, and any combination of camera types can be used to provide the necessary coverage around the vehicle 1000. In addition, the number of cameras can vary depending on the embodiment. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. As an example and not limitation, the cameras can support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each of the cameras described herein may be capable of supporting Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Fig. 10A and Fig. 10B Described in more detail.

[0236] The vehicle 1000 may also include a vibration sensor 1042. The vibration sensor 1042 may measure vibrations of a component of the vehicle, such as an axle. For example, changes in vibration may indicate changes in the road surface. In another example, when two or more vibration sensors 1042 are used, the difference between the vibrations may be used to determine friction or slip of the road surface (e.g., when there is a vibration difference between a powered steering axle and a free-wheeling axle).

[0237] The vehicle 1000 may include an ADAS system 1038. In some examples, the ADAS system 1038 may include a SoC. The ADAS system 1038 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.

[0238] The ACC system may use a RADAR sensor 1060, a LIDAR sensor 1064, and / or a camera. The ACC system may include a longitudinal ACC and / or a lateral ACC. The longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the vehicle 1000, and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle in front. The lateral ACC performs distance keeping and suggests that the vehicle 1000 change lanes when necessary. The lateral ACC is related to other ADAS applications such as LCA and CWS.

[0239] CACC uses information from other vehicles, which can be received from other vehicles indirectly via a wireless link or through a network connection (e.g., through the Internet) via a network interface 1024 and / or a wireless antenna 1026. A direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while an indirect link can be an infrastructure-to-vehicle (I2V) communication link. Typically, the V2V communication concept provides information about the vehicle immediately ahead (e.g., the vehicle immediately ahead of the vehicle 1000 and in the same lane as it), while the I2V communication concept provides information about traffic farther ahead. The CACC system can include either or both of the I2V and V2V information sources. Given information about the vehicle ahead of the vehicle 1000, CACC can be more reliable, and it is possible to improve the smoothness of traffic flow and reduce road congestion.

[0240] The FCW system is designed to alert the driver to hazards so that the driver can take corrective action. The FCW system uses a front camera and / or RADAR sensor 1060 coupled to a dedicated processor, DSP, FPGA and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker and / or vibrating component. The FCW system can provide warnings in the form of, for example, sound, visual warnings, vibrations and / or rapid brake pulses.

[0241] The AEB system detects an impending forward collision with another vehicle or other object and can automatically apply the brakes if the driver does not take corrective action within specified time or distance parameters. The AEB system can use a front camera and / or RADAR sensor 1060 coupled to a dedicated processor, DSP, FPGA and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, the AEB system can automatically apply the brakes in an effort to prevent or at least mitigate the effects of the predicted collision. The AEB system may include technologies such as dynamic brake support and / or collision approach braking.

[0242] The LDW system provides visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle 1000 crosses a lane marking. When the driver indicates an intention to leave the lane, by activating a turn sign, the LDW system is not activated. The LDW system may use a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.

[0243] The LKA system is a variation of the LDW system. If the vehicle 1000 begins to leave the lane, the LKA system provides steering input or braking to correct the vehicle 1000.

[0244] The BSW system detects and warns the driver of vehicles in the car's blind spot. The BSW system can provide visual, auditory and / or tactile alerts to indicate that it is unsafe to merge or change lanes. The system can provide additional warnings when the driver uses a turn sign. The BSW system can use a rear-facing camera and / or RADAR sensor 1060 coupled to a dedicated processor, DSP, FPGA and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker and / or vibration component.

[0245] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the range of the rear camera while the vehicle 1000 is in reverse. Some RCTW systems include AEB to ensure that the vehicle brakes are applied to avoid a crash. The RCTW system can use one or more rear RADAR sensors 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component.

[0246] Conventional ADAS systems may be prone to false positive results, which may annoy and distract the driver, but are typically not catastrophic because the ADAS system alerts the driver and allows the driver to decide whether the safety condition actually exists and take action accordingly. However, in the autonomous vehicle 1000, in the case of conflicting results, the vehicle 1000 itself must decide whether to pay attention to the results from the main computer or the auxiliary computer (e.g., the first controller 1036 or the second controller 1036). For example, in some embodiments, the ADAS system 1038 can be a backup and / or auxiliary computer for providing perception information to the backup computer rationality module. The backup computer rationality monitor can execute redundant and diverse software on hardware components to detect faults in perception and dynamic driving tasks. The output from the ADAS system 1038 can be provided to the supervisory MCU. If the outputs from the main computer and the auxiliary computer conflict, the supervisory MCU must determine how to coordinate the conflict to ensure safe operation.

[0247] In some examples, the master computer can be configured to provide a confidence score to the supervisory MCU, indicating the master computer's confidence in the selected result. If the confidence score exceeds a threshold, the supervisory MCU can follow the direction of the master computer, regardless of whether the auxiliary computer provides conflicting or inconsistent results. In the case where the confidence score does not meet the threshold and the master computer and the auxiliary computer indicate different results (e.g., conflicts), the supervisory MCU can arbitrate between these computers to determine the appropriate result.

[0248] The supervisory MCU may be configured to execute a neural network that is trained and configured to determine conditions under which the auxiliary computer provides a false alarm based, at least in part, on outputs from the primary computer and the auxiliary computer. Thus, the neural network in the supervisory MCU may learn when the output of the auxiliary computer may be trusted and when it may not. For example, when the auxiliary computer is a RADAR-based FCW system, the neural network in the supervisory MCU may learn when the FCW system is identifying a metal object that is not actually dangerous, such as a drainage grid or manhole cover that triggers an alarm. Similarly, when the auxiliary computer is a camera-based LDW system, the neural network in the supervisory MCU may learn to ignore the LDW when a cyclist or pedestrian is present and lane departure is actually the safest strategy. In an embodiment that includes a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or GPU suitable for running the neural network with associated memory. In a preferred embodiment, the supervisory MCU may include and / or be included as a component of SoC 1004.

[0249] In other examples, the ADAS system 1038 can include an auxiliary computer that uses traditional computer vision rules to perform ADAS functions. In this way, the auxiliary computer can use classic computer vision rules (if-then), and the presence of a neural network in the supervisory MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity make the entire system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functions. For example, if there is a software vulnerability or error in the software running on the main computer and the non-identical software code running on the auxiliary computer provides the same overall result, the supervisory MCU can be more confident that the overall result is correct and that the vulnerability in the software or hardware on the main computer does not cause a substantial error.

[0250] In some examples, the output of the ADAS system 1038 can be provided to the perception block of the main computer and / or the dynamic driving task block of the main computer. For example, if the ADAS system 1038 indicates a forward collision warning due to an object immediately ahead, the perception block can use this information when identifying the object. In other examples, the auxiliary computer can have its own neural network that is trained and thus reduces the risk of false positives as described herein.

[0251] The vehicle 1000 may also include an infotainment SoC 1030 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more discrete components. The infotainment SoC 1030 may include a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), phone (e.g., hands-free calling), network connectivity (e.g., LTE, WiFi, etc.), and / or information services (e.g., navigation system, rear parking assistance, radio data system, vehicle-related information such as fuel level, total distance covered, brake fuel level, oil level, door open / closed, air filter information, etc.) to the vehicle 1000. For example, the infotainment SoC 1030 may include a radio, a disc player, a navigation system, a video player, USB and Bluetooth connectivity, an onboard computer, onboard entertainment, WiFi, steering wheel audio controls, hands-free voice controls, a head-up display (HUD), an HMI display 1034, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. The infotainment SoC 1030 may further be used to provide information (e.g., visual and / or auditory) to a user of the vehicle, such as information from an ADAS system 1038, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0252] The infotainment SoC 1030 may include GPU functionality. The infotainment SoC 1030 may communicate with other devices, systems, and / or components of the vehicle 1000 via a bus 1002 (e.g., a CAN bus, Ethernet, etc.). In some examples, the infotainment SoC 1030 may be coupled to a supervisory MCU so that in the event of a failure of a master controller 1036 (e.g., a primary and / or backup computer of the vehicle 1000), the GPU of the infotainment system may perform some self-driving functions. In such an example, the infotainment SoC 1030 may place the vehicle 1000 in a driver-safe parking mode as described herein.

[0253] The vehicle 1000 may also include a meter group 1032 (e.g., a digital meter panel, an electronic meter group, a digital meter panel, etc.). The meter group 1032 may include a controller and / or a supercomputer (e.g., a discrete controller or a supercomputer). The meter group 1032 may include a set of instruments, such as a speedometer, a fuel level, an oil pressure, a tachometer, an odometer, a turn indicator, a shift position indicator, a seat belt warning light, a parking brake warning light, an engine fault light, an airbag (SRS) system information, a lighting control, a safety system control, navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 1030 and the meter group 1032. In other words, the meter group 1032 may be included as part of the infotainment SoC 1030, or vice versa.

[0254] Fig. 10D For a cloud-based server and Fig. 10A 1076. System 1076 may include server 1078, network 1090, and vehicles including vehicle 1000. Server 1078 may include multiple GPUs 1084(A)-1084(H) (collectively referred to herein as GPU 1084), PCIe switches 1082(A)-1082(H) (collectively referred to herein as PCIe switches 1082), and / or CPUs 1080(A)-1080(B) (collectively referred to herein as CPU 1080). GPU 1084, CPU 1080, and PCIe switch may be interconnected with a high-speed interconnect and / or PCIe connection 1086 such as, for example and without limitation, the NVLink interface 1088 developed by NVIDIA. In some examples, GPU 1084 is connected via NVLink and / or NVSwitch SoC, and GPU 1084 and PCIe switch 1082 are connected via PCIe interconnect. Although eight GPUs 1084, two CPUs 1080, and two PCIe switches are illustrated, this is not intended to be limiting. Depending on the embodiment, each of the servers 1078 may include any number of GPUs 1084, CPUs 1080, and / or PCIe switches. For example, each of the servers 1078 may include eight, sixteen, thirty-two, and / or more GPUs 1084.

[0255] Server 1078 may receive image data over network 1090 and from a vehicle, the image data representing images showing unexpected or changed road conditions, such as a recently begun road project. Server 1078 may transmit neural network 1092, updated neural network 1092, and / or map information 1094, including information about traffic and road conditions, over network 1090 and to the vehicle. Updates to map information 1094 may include updates to HD map 1022, such as information about construction sites, potholes, curves, flooding, or other obstacles. In some examples, neural network 1092, updated neural network 1092, and / or map information 1094 may have been generated from new training and / or data received from any number of vehicles in the environment and / or based on experience from training performed at a data center (e.g., using server 1078 and / or other servers).

[0256] Server 1078 can be used to train a machine learning model (e.g., a neural network) based on training data. The training data can be generated by the vehicle, and / or can be generated in a simulation (e.g., using a game engine). In some examples, the training data is labeled (e.g., when the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not labeled and / or preprocessed (e.g., when the neural network does not require supervised learning). Training can be performed according to any one or more categories of machine learning techniques, including but not limited to the following categories: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, joint learning, transfer learning, feature learning (including main component and cluster analysis), multilinear subspace learning, manifold learning, representation learning (including alternate dictionary learning), rule-based machine learning, anomaly detection, and any variants or combinations thereof. Once the machine learning model is trained, the machine learning model can be used by the vehicle (e.g., transmitted to the vehicle via network 1090), and / or the machine learning model can be used by server 1078 to remotely monitor the vehicle.

[0257] In some examples, server 1078 can receive data from the vehicle and apply the data to the latest real-time neural network for real-time intelligent reasoning. Server 1078 may include a deep learning supercomputer and / or a dedicated AI computer powered by GPU 1084, such as DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 1078 may include a deep learning infrastructure of a data center powered only by CPUs.

[0258] The deep learning infrastructure of server 1078 may be capable of rapid real-time inference, and may use this capability to assess and verify the health of the processors, software, and / or associated hardware in vehicle 1000. For example, the deep learning infrastructure may receive periodic updates from vehicle 1000, such as a sequence of images and / or objects located in the sequence of images that vehicle 1000 has located (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure may execute or operate its own neural network to identify objects and compare them to the objects identified by vehicle 1000, and if the results do not match and the infrastructure concludes that the AI ​​in vehicle 1000 has failed, then server 1078 may transmit a flag to vehicle 1000 instructing the fail-safe computer of vehicle 1000 to take control, notify passengers, and complete a safe parking maneuver.

[0259] For reasoning, the server 1078 may include a GPU 1084 and one or more programmable reasoning accelerators (e.g., NVIDIA's TensorRT). The combination of GPU-powered servers and reasoning acceleration can make real-time responses possible. In other examples, such as when performance is not so important, CPU, FPGA, and other processor-powered servers can be used for reasoning.

[0260] Example computing device

[0261] Fig.11 1 is a block diagram of an example computing device 1100 suitable for implementing some embodiments of the present disclosure. The computing device 1100 may include an interconnect system 1102 that directly or indirectly couples the following devices: a memory 1104, one or more central processing units (CPUs) 1106, one or more graphics processing units (GPUs) 1108, a communication interface 1110, an input / output (I / O) port 1112, an input / output component 1114, a power supply 1116, one or more presentation components 1118 (e.g., one or more displays), and one or more logic units 1120. In at least one embodiment, one or more computing devices 1100 may include one or more virtual machines (VMs), and / or any of their components may include virtual components (e.g., virtual hardware components). For non-limiting examples, one or more of the GPUs 1108 may include one or more vGPUs, one or more of the CPUs 1106 may include one or more vCPUs, and / or one or more of the logic units 1120 may include one or more virtual logic units. As such, one or more computing devices 1100 may include discrete components (eg, a full GPU dedicated to computing device 1100 ), virtual components (eg, a portion of a GPU dedicated to computing device 1100 ), or a combination thereof.

[0262] although Fig.11 The various blocks of are shown as being connected via interconnect system 1102 using wires, but this is not intended to be limiting and is for clarity only. For example, in some embodiments, presentation component 1118 (such as a display device) may be considered to be I / O component 1114 (e.g., if the display is a touch screen). As another example, CPU 1106 and / or GPU 1108 may include memory (e.g., memory 1104 may represent a storage device in addition to the memory of GPU 1108, CPU 1106, and / or other components). In other words, Fig.11 The computing devices described herein are illustrative only. No distinction is made between categories within this category such as "workstation," "server," "laptop," "desktop," "tablet," "client device," "mobile device," "handheld device," "game console," "electronic control unit (ECU)," "virtual reality system," and / or other device or system types, as all are considered within this category. Fig.11 within the range of computing devices.

[0263] The interconnection system 1102 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnection system 1102 may include one or more bus or link types, such as an industry standard architecture (ISA) bus, an extended industry standard architecture (EISA) bus, a video electronics standard association (VESA) bus, a peripheral component interconnect (PCI) bus, a peripheral component interconnect express (PCIe) bus, and / or another type of bus or link. In some embodiments, there is a direct connection between components. As an example, the CPU 1106 may be directly connected to the memory 1104. Further, the CPU 1106 may be directly connected to the GPU 1108. In the case where there is a direct or point-to-point connection between components, the interconnection system 1102 may include a PCIe link to perform the connection. In these examples, the PCI bus does not need to be included in the computing device 1100.

[0264] The memory 1104 may include any of a variety of computer-readable media. Computer-readable media may be any available media that can be accessed by the computing device 1100. Computer-readable media may include volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, computer-readable media may include computer storage media and communication media.

[0265] Computer storage media may include volatile and nonvolatile media and / or removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules, and / or other data types. For example, memory 1104 may store computer readable instructions (e.g., representing one or more programs and / or one or more program elements, such as an operating system). Computer storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by computing device 1100. As used herein, computer storage media does not include the logo itself.

[0266] Computer storage media may embody computer readable instructions, data structures, program modules, and / or other data types in modulated data symbols such as carrier waves or other transmission mechanisms, and include any information delivery media. The term "modulated data symbol" may refer to a symbol that has one or more characteristics set or changed in a manner that encodes information in the symbol. By way of example and not limitation, computer storage media may include wired media (such as a wired network or a direct wired connection) and wireless media (such as acoustic, RF, infrared, and other wireless media). Combinations of any of the above should also be included within the scope of computer readable media.

[0267] The CPU 1106 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1100 to perform one or more of the methods and / or processes described herein. The CPUs 1106 may each include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of handling numerous software threads simultaneously. The CPU 1106 may include any type of processor, and may include different types of processors (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server) depending on the type of computing device 1100 implemented. For example, depending on the type of computing device 1100, the processor may be an Advanced RISC Machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). The computing device 1100 may also include one or more CPUs 1106 in addition to one or more microprocessors or secondary coprocessors (such as a math coprocessor).

[0268] In addition to or in place of one or more CPUs 1106, one or more GPUs 1108 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1100 to perform one or more of the methods and / or processes described herein. One or more of the GPUs 1108 may be an integrated GPU (e.g., with one or more of the CPUs 1106) and / or one or more of the GPUs 1108 may be a discrete GPU. In an embodiment, one or more of the GPUs 1108 may be a coprocessor of one or more of the CPUs 1106. The GPU 1108 may be used by the computing device 1100 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, the GPU 1108 may be used for general-purpose computing on a GPU (GPGPU). The GPU 1108 may include hundreds or thousands of cores capable of handling hundreds or thousands of software threads simultaneously. The GPU 1108 may generate pixel data for an output image in response to a rendering command (e.g., a rendering command received from the CPU 1106 via a host interface). GPU 1108 may include graphics memory (e.g., display memory) for storing pixel data or any other suitable data (e.g., GPGPU data). Display memory may be included as part of memory 1104. GPU 1108 may include two or more GPUs operating in parallel (e.g., via a link). The link may connect the GPUs directly (e.g., using NVLINK) or may connect the GPUs through a switch (e.g., using NVSwitch). When combined together, each GPU 1108 may generate pixel data or GPGPU data for different portions of output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU may include its own memory, or may share memory with other GPUs.

[0269] In addition to or in lieu of the CPU 1106 and / or GPU 1108, the logic unit 1120 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1100 to perform one or more of the methods and / or processes described herein. In embodiments, the one or more CPUs 1106, the one or more GPUs 1108, and / or the one or more logic units 1120 may discretely or jointly perform any combination of methods, processes, and / or portions thereof. One or more of the logic units 1120 may be a part of and / or integrated into one or more of the CPU 1106 and / or GPU 1108 and / or one or more of the logic units 1120 may be discrete components or otherwise external to the CPU 1106 and / or GPU 1108. In embodiments, one or more of logic units 1120 may be co-processors for one or more of CPUs 1106 and / or one or more of GPUs 1108 .

[0270] Examples of logic unit 1120 include one or more processing cores and / or components thereof, such as a data processing unit (DPU), a tensor core (TC), a tensor processing unit (TPU), a pixel vision core (PVC), a vision processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multiprocessor (SM), a tree transverse unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application-specific integrated circuit (ASIC), a floating point unit (FPU), an input / output (I / O) element, a peripheral component interconnect (PCI) or a peripheral component interconnect express (PCIe) element, etc.

[0271] The communication interface 1110 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 1100 to communicate with other computing devices via an electronic communication network (including wired and / or wireless communications). The communication interface 1110 may include components and functions that implement communication through any of a plurality of different networks, such as a wireless network (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), a wired network (e.g., via Ethernet or InfiniBand communication), a low power wide area network (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, the one or more logic units 1120 and / or the communication interface 1110 may include one or more data processing units (DPUs) to send data received through the network and / or through the interconnect system 1102 directly to one or more GPUs 1108 (e.g., memory).

[0272] I / O ports 1112 may enable computing device 1100 to be logically coupled to other devices including I / O components 1114, one or more presentation components 1118, and / or other components, some of which may be built into (e.g., integrated into) computing device 1100. Illustrative I / O components 1114 include microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dishes, scanners, printers, wireless devices, and the like. I / O components 1114 may provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by a user. In some cases, the input may be transmitted to an appropriate network element for further processing. NUI may implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, gesture recognition on and near the screen, air gestures, head and eye tracking, and touch recognition (as described in more detail below) associated with the display of computing device 1100. Computing device 1100 may include a depth camera for gesture detection and recognition, such as a stereo camera system, an infrared camera system, an RGB camera system, touch screen technology, and combinations of these. Additionally, computing device 1100 may include an accelerometer or gyroscope that enables detection of motion (e.g., as part of an inertial measurement unit (IMU)). In some examples, computing device 1100 may use the output of the accelerometer or gyroscope to render immersive augmented reality or virtual reality.

[0273] The power supply 1116 may include a hardwired power supply, a battery power supply, or a combination thereof. The power supply 1116 may provide power to the computing device 1100 to enable the components of the computing device 1100 to operate.

[0274] The presentation component 1118 may include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), a speaker, and / or other presentation components. The presentation component 1118 may receive data from other components (e.g., GPU 1108, CPU 1106, etc.) and output the data (e.g., as an image, video, sound, etc.).

[0275] Sample Data Center

[0276] Fig.12 An example data center 1200 is shown that can be used in at least one embodiment of the present disclosure. The data center 1200 can include a data center infrastructure layer 1210, a framework layer 1220, a software layer 1230, and / or an application layer 1240.

[0277] like Fig.12 As shown, data center infrastructure layer 1210 may include resource coordinator 1210, grouped computing resources 1214, and node computing resources ("node CRs") 1216(1)-1216(N), where "N" represents any complete positive integer. In at least one embodiment, node CRs 1216(1)-1216(N) may include, but is not limited to, any number of central processing units ("CPUs") or other processors (including DPUs, accelerators, field programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules and / or cooling modules, etc. In some embodiments, one or more node CRs from node CRs 1216(1)-1216(N) may correspond to a server having one or more of the above-mentioned computing resources. Furthermore, in some embodiments, node CRs 1216(1)-12161(N) may include one or more virtual components, such as vGPU, vCPU, etc., and / or one or more of node CRs 1216(1)-1216(N) may correspond to a virtual machine (VM).

[0278] In at least one embodiment, the grouped computing resources 1214 may include individual groups of node CRs 1216 housed in one or more racks (not shown), or many racks housed in data centers at different geographic locations (also not shown). Individual groups of node CRs 1216 within the grouped computing resources 1214 may include grouped computing, network, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, several node CRs 1216 including CPUs, GPUs, and / or other processors may be grouped in one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches in any combination.

[0279] Resource coordinator 1222 may configure or otherwise control one or more node CRs 1216(1)-1216(N) and / or grouped computing resources 1214. In at least one embodiment, resource coordinator 1222 may include a software design infrastructure ("SDI") management entity for data center 1200. Resource coordinator 1222 may include hardware, software, or some combination thereof.

[0280] In at least one embodiment, Fig.12 As shown, the framework layer 1220 may include a job scheduler 1232, a configuration manager 1234, a resource manager 1236, and / or a distributed file system 1238. The framework layer 1220 may include a framework that supports the software 1232 of the software layer 1230 and / or one or more applications 1242 of the application layer 1240. The software 1232 or the application 1242 may include network-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 1220 may be, but is not limited to, a free and open source software network application framework (such as Apache Spark) that can utilize the distributed file system 1238 for large-scale data processing (e.g., "big data"). TM(hereinafter referred to as "Spark")). In at least one embodiment, the job scheduler 1232 may include a Spark driver to facilitate scheduling workloads supported by different layers of the data center 1200. The configuration manager 1234 may be able to configure different layers, such as the software layer 1230 and the framework layer 1220 (which includes Spark and a distributed file system 1238 for supporting large-scale data processing). The resource manager 1236 may be able to manage clustered or grouped computing resources mapped to the distributed file system 1238 and the job scheduler 1232 or allocated to support the distributed file system 1238 and the job scheduler 1232. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 1214 at the data center infrastructure layer 1210. The resource manager 1236 may coordinate with the resource coordinator 1210 to manage these mapped or allocated computing resources.

[0281] In at least one embodiment, the software 1232 included in the software layer 1230 may include software used by at least a portion of the node CRs 1216(1)-1216(N), the grouped computing resources 1214, and / or the distributed file system 1238 of the framework layer 1220. The one or more types of software may include, but are not limited to, Internet web search software, email virus scanning software, database software, and streaming video content software.

[0282] In at least one embodiment, the applications 1242 included in the application layer 1240 may include one or more types of applications used by at least a portion of the node CRs 1216(1)-1216(N), the grouped computing resources 1214, and / or the distributed file system 1238 of the framework layer 1220. The one or more types of applications may include, but are not limited to, any number of genomic applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.

[0283] In at least one embodiment, any of configuration manager 1234, resource manager 1236, and resource coordinator 1210 can implement any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. The self-modification actions can save a data center operator of data center 1200 from making potentially poor configuration decisions and potentially avoiding underutilized and / or poorly performing portions of the data center.

[0284] According to one or more embodiments described herein, data center 1200 may include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information. For example, one or more machine learning models may be trained by calculating weight parameters according to a neural network architecture using the software and / or computing resources described above with respect to data center 1200. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 1200 using weight parameters calculated by one or more training techniques, such as, but not limited to, those described herein.

[0285] In at least one embodiment, the data center 1200 may use a CPU, an application-specific integrated circuit (ASIC), a GPU, an FPGA, and / or other hardware (or virtual computing resources corresponding thereto) to perform training and / or inference using the above resources. In addition, one or more software and / or hardware resources described above may be configured to allow a user to train or perform services that infer information, such as image recognition, speech recognition, or other artificial intelligence services.

[0286] Example network environment

[0287] A network environment suitable for implementing embodiments of the present disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be configured to communicate with the client device, server, or other device types. Fig.11 The backend device 1200 may be implemented on one or more instances of one or more computing devices 1100 of the present invention - for example, each device may include similar components, features and / or functions of one or more computing devices 1100. In addition, in the case of implementing a backend device (e.g., a server, NAS, etc.), the backend device may be included as part of the data center 1200, and an example of the data center 1200 is described herein with reference to FIG. Fig.12 Describe in more detail.

[0288] The components of the network environment can communicate with each other via a network, which can be wired, wireless, or both. The network can include a plurality of networks or one of the networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or a public switched telephone network (PSTN)), and / or one or more private networks. In the case where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connections.

[0289] Compatible network environments may include one or more peer-to-peer network environments (in which case the server may not be included in the network environment) and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the functionality described herein for the server may be implemented on any number of client devices.

[0290] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, and the like. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, and the server may include one or more core network servers and / or edge servers. The framework layer may include a framework for software supporting the software layer and / or one or more applications of the application layer. The software or application may include network-based service software or applications, respectively. In an embodiment, one or more client devices may use network-based service software or applications (e.g., by accessing the service software and / or application via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a free and open source software network application framework that can use a distributed file system for large-scale data processing (e.g., "big data").

[0291] A cloud-based network environment can provide cloud computing and / or cloud storage that performs any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these different functions can be distributed across multiple locations from a central or core server (e.g., one or more data centers that can be distributed across a state, region, country, global, etc.). If the connection to the user (e.g., client device) is relatively close to an edge server, the core server can assign at least a portion of the functionality to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0292] One or more client devices may include the Fig.11At least some of the components, features, and functionality of one or more example computing devices 1100 are described. By way of example and not limitation, a client device may be implemented as a personal computer (PC), a laptop computer, a mobile device, a smart phone, a tablet computer, a smart watch, a wearable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or device, a video player, a camera, a surveillance device or system, a vehicle, a boat, a spacecraft, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or system, an entertainment system, a vehicle computer system, an embedded system controller, a remote control, an appliance, a consumer electronic device, a workstation, an edge device, any combination of the depicted devices, or any other suitable device.

[0293] The present disclosure may be described in the general context of machine-usable instructions or computer code executed by a computer or other machine such as a personal digital assistant or other handheld device, including computer-executable instructions such as program modules. Typically, program modules including routines, programs, objects, components, data structures, etc. refer to code that performs a specific task or implements a specific abstract data type. The present disclosure may be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The present disclosure may also be practiced in a distributed computing environment in which tasks are performed by remote processing devices linked through a communication network.

[0294] As used herein, the statement of "and / or" with respect to two or more elements should be interpreted as referring to only one element or a combination of elements. For example, "element A, element B, and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. In addition, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0295] The subject matter of the present disclosure is described in detail herein to meet statutory requirements. However, the description itself is not intended to limit the scope of the present disclosure. On the contrary, the present inventors have contemplated that the claimed subject matter may also be embodied in other ways to include steps different from the steps described herein in conjunction with other current or future technologies or combinations of similar steps. Moreover, although the terms "step" and / or "block" may be used herein to imply different elements of the method employed, these terms should not be interpreted as implying any particular order among or between the various steps disclosed herein, unless the order of the steps is explicitly described.

Claims

1. One or more processors, including: One or more processing units for: receiving image data representing one or more portions of an operator of the present machine, the image data generated using one or more sensors of the present machine traversing an environment; generating a representation of a calculated alertness level of the operator based at least on the image data; providing a representation of the calculated alertness level as an input to one or more machine learning models to generate one or more natural language characters based at least on the calculated alertness level of the operator; as well as A representation of the one or more natural language characters is caused to be presented using a display device or a sound device associated with an operator of the present machine.

2. One or more processors according to claim 1, wherein: The one or more processing units are further configured to: Providing a representation of a phrase utterance representing an operator response of the operator to the presented representation of the one or more natural language characters as a subsequent input to the one or more machine learning models to generate one or more second natural language characters representing a generated response to the operator response.

3. One or more processors according to claim 1, wherein: The one or more processing units are further configured to: Providing a representation of personalized information associated with the operator retrieved from a database as at least a portion of the input to the one or more machine learning models, wherein the representation of the one or more natural language characters generated by the one or more machine learning models is presented using the display device or sound device associated with the operator of the machine and includes a first phrase utterance that initiates a conversation with the operator based at least on the personalized information associated with the operator.

4. One or more processors according to claim 1, wherein: The input includes a prompt, and the one or more machine learning models include a large language model (LLM), and wherein the prompt also includes at least one of the following: personalized information of the operator, a representation of the operator's calculated alertness level, natural language instructions to initiate a conversation with the operator that is consistent with the operator's interests, zero-sample, single-sample, or few-sample examples of representative input or output, or instructions to send a control signal to cause a specific action of the machine.

5. One or more processors according to claim 1, wherein: The one or more processing units are further configured to: receiving a geographic location indicator representing a location of the machine in the environment; determining, based at least on the geographic location indicator, at least one of weather data, road condition data, traffic data, or event data associated with the geographic location; and At least one of the weather data, the road condition data, the traffic data, the event data, or the geographic location indicator is provided as at least a portion of the input to the one or more machine learning models, wherein a representation of the one or more natural language characters presented using the display device or sound device associated with an operator or occupant of the machine includes a natural language phrase representing a summary of the at least one of the weather data, the road condition data, the traffic data, or the event data.

6. One or more processors according to claim 1, wherein: The one or more processing units are further configured to: accessing first audio data associated with a radio station based at least on a geographic location indicator, wherein the geographic location indicator represents a location of the local machine in the environment; and A representation of the audio data is provided as at least a portion of the input to the one or more machine learning models, wherein the one or more natural language characters generated by the one or more machine learning models comprise a summary of the audio data representation.

7. One or more processors according to claim 1, wherein: The one or more processing units are further configured to: accessing, from one or more data sources, destination or travel route information associated with a destination or travel route of the host machine; as well as Providing a representation of the destination or travel route information as at least a portion of the input to the one or more machine learning models to generate a summarized representation of the destination or travel route information.

8. One or more processors according to claim 1, wherein: The one or more processors are included in at least one of the following: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; A system for performing simulation operations; Systems for performing digital twin operations; A system for performing light transport simulations; A system for performing collaborative content creation of 3D assets; Systems for performing deep learning operations; A system for performing real-time streaming; Systems for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; Systems for performing conversational AI operations; Systems for generating synthetic data; Systems that use AI to generate synthetic data; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.

9. A system comprising one or more processing units for causing a device associated with an operator of a machine to present a representation of one or more natural language characters generated using one or more machine learning models based at least on a representation of a calculated alertness level of the operator of the machine.

10. The system according to claim 9, wherein: The one or more processing units are further configured to: Providing a representation of a phrase utterance representing an operator response of the operator to the presented representation of the one or more natural language characters as at least a portion of data to the one or more machine learning models to generate one or more second natural language characters representing a generated response to the operator response.

11. The system according to claim 9, wherein: The one or more processing units are further configured to: Providing a representation of personalized information associated with the operator retrieved from a database as at least a portion of an input to the one or more machine learning models, wherein the representation of the one or more natural language characters generated by the one or more machine learning models is presented using a device associated with the operator of the machine and includes a first phrase utterance that initiates a conversation with the operator based at least on the personalized information associated with the operator.

12. The system according to claim 9, wherein: The one or more machine learning models include a large language model (LLM), and wherein the prompts provided to the LLM include at least one of: personalized information of the operator, a representation of the operator's calculated alertness level, natural language instructions to initiate a conversation with the operator that is consistent with the operator's interests, zero-sample, single-sample, or few-sample examples of representative inputs or outputs, or instructions to send a control signal to cause a specific action of the machine.

13. The system according to claim 9, wherein: The one or more processing units are further configured to: receiving a geographic location indicator representing a location of the machine; determining, based at least on the geographic location indicator, at least one of weather data, road condition data, traffic data, or event data associated with the geographic location; and providing at least one of the weather data, the road condition data, the traffic data, the event data, or the geographic location indicator as at least a portion of an input to the one or more machine learning models, wherein a representation of the one or more natural language characters presented using a device associated with an operator of the machine includes a natural language phrase representing at least one of the weather data, the road condition data, the traffic data, or the event data.

14. The system according to claim 9, wherein: The one or more processing units are further configured to: accessing first audio data associated with a radio station based at least on a geographic location indicator, wherein the geographic location indicator represents a location of the local machine; and A representation of the audio data is provided as at least a portion of the input to the one or more machine learning models, wherein the one or more natural language characters generated by the one or more machine learning models comprise a summary of the audio data representation.

15. The system according to claim 9, wherein: The one or more processors are included in at least one of the following: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; A system for performing simulation operations; Systems for performing digital twin operations; A system for performing light transport simulations; A system for performing collaborative content creation of 3D assets; Systems for performing deep learning operations; A system for performing real-time streaming; Systems for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; Systems for performing conversational AI operations; Systems for generating synthetic data; Systems that use AI to generate synthetic data; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.

16. A method comprising: receiving an indication of the calculated alertness level of an operator of the machine; receiving one or more indications of interests of the operator; providing the representation of the calculated alertness level and the representation of the one or more interests of the operator as at least a portion of an input to one or more machine learning models to generate one or more natural language characters based at least on the calculated alertness level of the operator and the one or more interests of the operator; as well as A representation of the one or more natural language characters is caused to be presented using a device associated with the operator of the present machine.

17. The method according to claim 16, further comprising: Providing a representation of a phrase utterance representing an operator response of the operator to the presented representation of the one or more natural language characters as a subsequent input to the one or more machine learning models to generate one or more second natural language characters representing a generated response to the operator response.

18. The method according to claim 16, wherein: Presenting a representation of the one or more natural language characters generated by the one or more machine learning models at a device associated with the operator of the machine includes: presenting a first phrase utterance using a voice device, the first phrase utterance initiating a conversation with the operator based at least on the one or more interests.

19. The method according to claim 16, wherein: The input includes a prompt, and the one or more machine learning models include a large language model (LLM), and wherein the prompt also includes at least one of the following: the one or more interests of the operator, a representation of the operator's calculated alertness level, a natural language instruction to initiate a conversation with the operator consistent with the one or more interests, a zero-sample, a single-sample, or a few-sample example of a representative input or output, or an instruction to send a control signal to cause a specific action of the machine.

20. The method according to claim 16, wherein: The method is performed by at least one of the following: control systems for autonomous or semi-autonomous machines; Perception systems for autonomous or semi-autonomous machines; A system for performing simulation operations; Systems for performing digital twin operations; A system for performing light transport simulations; A system for performing collaborative content creation of 3D assets; Systems for performing deep learning operations; A system for performing real-time streaming; Systems for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; Systems implemented using edge devices; Systems implemented using robots; Systems for performing conversational AI operations; Systems for generating synthetic data; Systems that use AI to generate synthetic data; A system comprising one or more virtual machines VM; A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.

Citation Information

Patent Citations

  • Method for programmable timeouts of tree traversal mechanisms in hardware

    US10885698B2

  • Machine operation assistance using language model-augmented perception

    US20250136130A1