Systems and methods for automated driving event classification using targeted image compression and intermediary state predictions
A codeword-based compression system and encoder-decoder architecture improve computational efficiency and accuracy in machine learning-based driving systems by selectively processing image portions relevant to vehicle operation, enabling faster and more accurate collision detection.
Patent Information
- Application Number
- PCT/US2025/019923
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-18
- Filing Date
- 2025-03-14
- Publication Date
- 2025-09-25
AI Technical Summary
Existing machine learning-based driving systems face challenges in efficiently processing large and diverse driving data sets due to complex environmental conditions and variability in road infrastructure and human behavior, leading to inefficient use of computational resources and inaccurate alert generation.
Implementing a codeword-based compression system that differentiates between image portions relevant to vehicle operation and those that are not, using an encoder-decoder architecture to predict future states and a classification model for efficient event detection, reducing computational resources and improving accuracy.
Enhances computational efficiency and accuracy in detecting potential collisions by focusing on relevant image portions, allowing for faster and more accurate alert generation in complex driving environments.
Smart Images

Figure US2025019923_25092025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR AUTOMATED DRIVING EVENT CLASSIFICATION USING TARGETED IMAGE COMPRESSION AND INTERMEDIARY STATE PREDICTIONSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority to US Provisional Application No. 63 / 566,762, filed March 18, 2024, the entirety of which is incorporated by reference herein.TECHNICAL FIELD
[0002] Certain aspects of the present disclosure generally relate to generative artificial intelligence (Al) models, and more particularly to systems and methods of improving driving safety using a generative Al model that is adapted to driving data.BACKGROUND
[0003] Driving systems have garnered significant interest and investment in recent years due to their potential to revolutionize transportation by enhancing safety, efficiency, and convenience. These systems utilize various sensors, such as cameras, LiDAR, and radar, in conjunction with sophisticated algorithms, to perceive tire surrounding environment.
[0004] One important aspect of driving systems is their ability to provide timely alerts to drivers to prevent accidents or mitigate risks. These alerts can range from simple notifications about upcoming obstacles to more complex alerts that indicate a driver may be driving unsafely.
[0005] While machine learning techniques have shown promise in improving the capabilities of automated alert generation in driving systems, several challenges persist, hindering their widespread adoption and effectiveness. For example, one challenge with using machine learning for alert generation is that machine learning models often rely heavily on large and diverse datasets for training. However, obtaining high-quality and diverse data representative of all possible driving scenarios is challenging. Variability in environmental conditions, road infrastructure, and human behavior adds complexity to tire data collection process. Additionally, even when there is a substantial amount of data available for automatic alert decision-making, it can be difficult to identify the portions that are the most relevant, or relevant at all, to generating an alert. Another challenge with using machine learning for alert generation is that driving environments are inherently complex, with numerous factorsinfluencing driving decisions or whether it is proper to generate an alert. Machine learning algorithms navigate through a vast array of scenarios, including adverse weather conditions, unpredictable pedestrian or cyclist behavior, and complex traffic patterns. Designing machine learning algorithms or any other type of algorithms that are capable of handling this complexity while ensuring safety and reliability is a significant challenge.SUMMARY
[0006] Systems may attempt to use camera-based monitoring for predicted collision detection and alert generation in vehicles. For example, a computing system located on a vehicle may continuously receive image or video data of the environment surrounding or inside the vehicle captured by one or more cameras. The computing system may apply image processing techniques, such as by using neural networks, to analyze the images and determine whether to generate an alert or control the vehicle. However, processing all portions of an image with equal importance can lead to inefficient use of computational resources given the large amount of data that may be included in the image and may not adequately prioritize features that are relevant to detecting driving events that may trigger in-vehicle alerts or vehicle control. Thus, the computing system may not accurately predict potential collision events before they occur, particularly in complex environmental conditions that include many objects or features.
[0007] A computing device implementing the systems and methods described herein may overcome the aforementioned technical problems by implementing a codeword-based compression system that differentiates between image portions containing objects that are relevant to vehicle operation and image portions that do not. To do so, the computing device can execute an object classifier to detect objects within captured images and identity which portions of the images contain objects related to vehicle operation (e.g., a stop sign, another vehicle on the road, a pedestrian crossing the road, a traffic light, etc ). The computing device can convert an image into learned codewords, including vehicle operation-related codewords, from a first codeword space and non-vehicle operation-related codewords from a second codeword space. The first codeword space may include more codewords than the second codeword space. In doing so, the computing device can generate codewords that represent the image such that portions of the image that include objects that are relevant to operation of the vehicle have more codewords than other portions of the image (e.g., because the computing device may have more codewords available to describe or represent the portions of the imagethat include objects that are relevant to operation of tire vehicle). The computing device can execute a machine learning model on these learned codewords to efficiently detect predicted events and generate alerts within the vehicle, improving computational efficiency and accuracy by increasing the focus on the portions of the image that are relevant to operation of the vehicle.
[0008] In some implementations, the computing device uses an encoder-decoder architecture (e.g., of a language processing machine learning model, such as a large language model) that processes images representing the current state of tire environment around the vehicle and predicts potential future states of the environment. The encoder can generate an intermediary state indicating multiple candidate future states (e.g., multiple probabilities or confidence scores for candidate future states) of the environment surrounding the vehicle at time steps subsequent to when the images were captured. The decoder of the architecture can process the intermediary state to select or generate an image or other representation of the future state of the environment. The computing device can process the future state to determine whether the future state indicates a collision or another driving event will occur. In doing so, the computing system can anticipate collision events before they develop, rather than only reacting to current conditions. The computing system can train the encoder using a loss function based on such future states generated by the decoder, increasing the accuracy of intermediary states that the encoder generates to accurately represent, correspond to, or include confidence scores for different future states of the environment surrounding the vehicle.
[0009] Responsive to training the encoder in the above-described manner (e.g., responsive to training the encoder above an accuracy threshold), the computing device can use a classification machine learning model (e.g., a classification decoder) to detect predicted collisions based on intermediary states generated by the trained encoder. For instance, the encoder can generate an intermediary state based on a sequence of images depicting the surroundings of the vehicle, and the computing device can execute the classification machine learning model using tire intermediary state to generate an output indicating a likelihood that a driving event will occur. In doing so, the computing device can avoid or stop using the decoder to generate a predicted future state and directly use the classification model for event detection instead, which can substantially reduce the computational resources and time that are required to predict events without the extra decoder processing step. In doing so, the computing device can detect risks of collisions faster than other systems, which can aid a self-driving system avoid such collisions or facilitate earlier alert generation to alert the driver of the upcoming collision to give the driver more time to react.
[0010] In one embodiment, a method can include receiving, by one or more processors, one or more images of an environment surrounding a vehicle or within the vehicle captured by a camera mounted to or in the vehicle; executing, by the one or more processors, an object classifier using the one or more images to detect one or more objects depicted within at least one of tire one or more images; identifying, by the one or more processors, a first set of portions of the at least one image that each depict at least one of the detected one or more objects that relate to operation of the vehicle in the environment and a second set of portions of the at least one image that do not depict any of the detected one or more objects that relate to operation of the vehicle in the environment; converting, by the one or more processors, the at least one image into a plurality of learned codewords in which the first set of portions of the at least one image are each represented by a first set of learned codewords selected from a first codeword space corresponding to objects that relate to operation of the vehicle in tire environment, and the second set of portions of the at least one image are each represented by a second set of learned codewords selected from a second codeword space corresponding to portions of images that do not contain objects that relate to operation of the vehicle in the environment; and responsive to executing, by the one or more processors, a machine learning model using the plurality of learned codewords for each of the at least one image to detect a predicted event, generating, by the one or more processors, an alert within the vehicle.
[0011] Executing the object classifier to detect the one or more objects can include executing, by the one or more processors, the object classifier using the one or more images to detect one or more objects that correspond to the vehicle getting into a collision.
[0012] The method can further include generating, by the one or more processors, a training data set comprising a first set of images each captured within a defined time period of a vehicle getting into a collision and second set of images that are not captured within the defined time period of a vehicle getting into a collision; labeling, by tire one or more processors, each of tire first set of images to indicate each of the first set of images corresponds to a collision and each of the second set of images to indicate each of tire second set of images does not correspond to a collision; and training, by the one or more processors, the object classifier using labels of the first set of images and the second set of images.
[0013] Converting the at least one image into tire plurality of learned codewords can include converting, by the one or more processors, the at least one image into the plurality of learned codewords using a second machine learning model. The method can further includetraining, by the one or more processors, tire second machine learning model using a plurality of training images and identifications of objects detected by the object classifier within the plurality of training images labeled based on whether the images correspond to a collision.
[0014] Training the second machine learning model can include training, by the one or more processors, the second machine learning model using a loss function that corresponds to higher weights for portions of images containing at least one object detected by the object classifier within the plurality’ of training images than for portions of images for which the object classifier does not detect objects that relate to operation of the vehicle in the environment.
[0015] The first codeword space corresponding to objects that relate to operation of the vehicle in the environment has a larger size than the second codeword space corresponding to portions of images that do not contain objects that relate to operation of the vehicle in the environment.
[0016] Converting the at least one image into a plurality of learned codewords can include generating, by the one or more processors, codewords that represent different grid locations of a defined grid segmenting each of the at least one image.
[0017] Converting the at least one image into a plurality of learned codewords can include identifying, by the one or more processors, a first set of grid locations of an image that each contain at least one object that relates to operation of the vehicle in the environment and a second set of grid locations of the image that do not contain any objects that relate to operation of the vehicle in the environment; and using, by the one or more processors, the first codeword space to generate codewords for each of the first set of grid locations and the second codeword to generate codewords for each of the second set of grid locations.
[0018] Executing the machine learning model using the plurality of learned codewords for each of the at least one image to detect a predicted event can include executing, by the one or more processors, a large language model using the first set of learned codewords and the second set of learned codewords as input.
[0019] The method can further include converting, by the one or more processors, the first set of learned codewords and the second set of learned codewords into a compressed at least one image containing portions that correspond to the first set of learned codewords at a higher resolution then portions that correspond to the second set of learned codewords, wherein executing the machine learning model using the plurality’ of learned codewords for each of theat least one image to detect the predicted event comprises executing, by the one or more processors, the machine learning model using the compressed at least one image as input.
[0020] The method can include executing, by the one or more processors, the machine learning model using a sequence of the plurality of learned codewords for each of the at least one image ordered and grouped based on a time of capture of each of the at least one image from which the plurality of learned codewords were generated to detect the predicted event.
[0021] In one embodiment, a system can include one or more processors configured by machine-readable instructions to receive one or more images of an environment surrounding a vehicle or within the vehicle captured by a camera mounted to or in the vehicle; execute an object classifier using the one or more images to detect one or more objects depicted within at least one of the one or more images; identify a first set of portions of the at least one image that each depict at least one of the detected one or more obj ects that relate to operation of the vehicle in the environment and a second set of portions of the at least one image that do not depict any of the detected one or more objects that relate to operation of the vehicle in the environment; convert the at least one image into a plurality' of learned codewords in which the first set of portions of the at least one image are each represented by a first set of learned codewords selected from a first codeword space corresponding to objects that relate to operation of the vehicle in tire environment, and the second set of portions of the at least one image are each represented by a second set of learned codewords selected from a second codeword space corresponding to portions of images that do not contain objects that relate to operation of the vehicle in the environment; and responsive to executing a machine learning model using the plurality of learned codewords for each of the at least one image to detect a predicted event, generate an alert within the vehicle.
[0022] The one or more processors can be configured to execute the object classifier to detect tire one or more objects by executing the object classifier using tire one or more images to detect one or more objects that correspond to the vehicle getting into a collision.
[0023] The one or more processors can be further configured to generate a training data set comprising a first set of images each captured within a defined time period of a vehicle getting into a collision and second set of images that are not captured within the defined time period of a vehicle getting into a collision; label each of the first set of images to indicate each of the first set of images corresponds to a collision and each of the second set of images toindicate each of the second set of images does not correspond to a collision; and train the object classifier using labels of tire first set of images and the second set of images.
[0024] The one or more processors can be configured to convert the at least one image into the plurality of learned codewords by converting the at least one image into the plurality of learned codewords using a second machine learning model, the one or more processors further configured to train the second machine learning model using a plurality of training images and identifications of objects detected by the object classifier within the plurality of training images labeled based on whether the images correspond to a collision.
[0025] The one or more processors can be configured to train the second machine learning model by training the second machine learning model using a loss function that corresponds to higher weights for portions of images containing at least one object detected by tire object classifier within the plurality of training images than for portions of images for which the object classifier does not detect objects that relate to operation of the vehicle in the environment.
[0026] The first codeword space corresponding to objects that relate to operation of the vehicle in the environment can have a larger size than the second codeword space corresponding to portions of images that do not contain objects that relate to operation of the vehicle in tire environment.
[0027] The one or more processors can be configured to convert the at least one image into a plurality of learned codewords by generating codewords that represent different grid locations of a defined grid segmenting each of the at least one image.
[0028] The one or more processors can be configured to convert the at least one image into a plurality of learned codewords by identifying a first set of grid locations of an image that each contain at least one object that relates to operation of the vehicle in the environment and a second set of grid locations of the image that do not contain any objects that relate to operation of the vehicle in the environment; and using the first codeword space to generate codewords for each of the first set of grid locations and the second codeword to generate codewords for each of the second set of grid locations.
[0029] The one or more processors can be configured to execute the machine learning model using the plurality of learned codew ords for each of the at least one image to detect apredicted event by executing a large language model using the first set of learned codewords and the second set of learned codewords as input.
[0030] In one embodiment, a method can include receiving, by one or more processors, a first sequence of images of a first environment surrounding a vehicle or within the vehicle captured by a camera mounted to or in the vehicle; executing, by the one or more processors, an encoder of a machine learning model using the first sequence of images to generate a first intermediary state indicating a first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at one or more time steps subsequent to a first time period in which the first sequence of images were captured; executing, by the one or more processors, a future-state decoder of the machine learning model based on the first intermediary state to generate a future state from the first plurality of candidate future states and of the first environment surrounding the vehicle or within the vehicle at a time step subsequent to the first time period in which the first sequence of images were captured; training, by the one or more processors, the encoder using a loss function based on the generated future state of the first environment surrounding the vehicle; subsequent to the training the encoder, receiving, by the one or more processors, a second sequence of images of a second environment surrounding the vehicle or within the vehicle captured by the camera mounted to or in the vehicle; executing, by the one or more processors, tire trained encoder of the machine learning model using the second sequence of images to generate a second intermediary state indicating a second plurality of candidate future states of the second environment surrounding the vehicle or within the vehicle at one or more second time steps subsequent to a second time period in which the second sequence of images were captured; and responsive to executing, by the one or more processors, a classification decoder of the machine learning model based on the second intermediary state to detect a predicted event, generating, by the one or more processors, an alert within the vehicle.
[0031] Executing the encoder of the machine learning model using the first sequence of images can include executing, by the one or more processors, the encoder of the machine learning model to generate the first intermediary state indicating a first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle each at a first time step subsequent to the first time period in which the first sequence of images were captured, and wherein executing the future-state decoder of the machine learning model based on the first intermediary state can include executing, by the one or more processors, the futurestate decoder of the machine learning model to generate a future state from the first pluralityof candidate future states and of the first environment surrounding the vehicle or within the vehicle at the first time step subsequent to the first time period in which the first sequence of images were captured.
[0032] Executing the encoder of the machine learning model using the first sequence of images to generate the first intermediary state can include executing, by the one or more processors, the encoder of the machine learning model to generate tire first intermediary state comprising the first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at the one or more time steps subsequent to the first time period in which the first sequence of images were captured.
[0033] Executing the encoder of the machine learning model using the first sequence of images to generate the first intermediary state can include executing, by the one or more processors, the encoder of the machine learning model to generate the first intermediary state comprising the first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at each of a plurality of time steps subsequent to the first time period in which the first sequence of images were captured.
[0034] The method can further include labeling, by the one or more processors, a third sequence of images to indicate the third sequence of images corresponds to a collision with the vehicle; and training, by the one or more processors, the classification decoder based on whether the classification decoder generates a prediction that a collision will occur using the label of the third sequence of images.
[0035] Labeling the first sequence of images can include labeling, by the one or more processors, the third sequence of images to indicate the collision with the vehicle occurred at a first time step subsequent to the first time period in which the first sequence of images were captured; and wherein training tire encoder can include training, by the one or more processors, tire encoder based on whether the classification decoder generates a prediction that a collision will occur using the label of the third sequence of images.
[0036] Executing the encoder of the machine learning model using the first sequence of images to generate the first intermediary state can include executing, by the one or more processors, the encoder of the machine learning model to generate a numerical vector representation of the first plurality of candidate future states of the first environmentsurrounding the vehicle or within the vehicle at the one or more time steps subsequent to the first time period in which the first sequence of images were captured.
[0037] The method can include executing, by the one or more processors, the classification decoder of the machine learning model based on the second intermediary state to detect a confidence score for a predicted event; and detecting, by the one or more processors, the predicted event responsive to determining the confidence score exceeds a threshold.
[0038] Executing the classification decoder of the machine learning model can include executing, by the one or more processors, the classification decoder of the machine learning model based on the second intermediary state to detect a confidence score for a predicted collision with the vehicle; and wherein detecting the predicted event can include comparing, by the one or more processors, the confidence score for the predicted collision with the vehicle with the threshold; and determining, by the one or more processors, the confidence score for the predicted collision with the vehicle exceeds the threshold.
[0039] Executing the encoder of the machine learning model to generate tire second intermediary state can include executing, by the one or more processors, the encoder of the machine learning model based on one or more sensor values collected from sensors of the vehicle and the second sequence of images to generate the second intermediary state.
[0040] In one embodiment, a system can include one or more processors configured by machine-readable instructions to receive a first sequence of images of a first environment surrounding a vehicle or within the vehicle captured by a camera mounted to or in the vehicle; execute an encoder of a machine learning model using the first sequence of images to generate a first intermediary state indicating a first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at one or more time steps subsequent to a first time period in which the first sequence of images were captured; execute a futurestate decoder of tire machine learning model based on the first intermediary state to generate a future state from the first plurality of candidate future states and of the first environment surrounding the vehicle or within the vehicle at a time step subsequent to the first time period in which the first sequence of images were captured; train the encoder using a loss function based on the generated future state of the first environment surrounding the vehicle; subsequent to the training the encoder, receive a second sequence of images of a second environment surrounding the vehicle or within the vehicle captured by the camera mounted to or in the vehicle; execute the trained encoder of the machine learning model using tire second sequenceof images to generate a second intermediary state indicating a second plurality of candidate future states of the second environment surrounding the vehicle or within the vehicle at one or more second time steps subsequent to a second time period in which the second sequence of images were captured; and responsive to executing a classification decoder of the machine learning model based on the second intermediary state to detect a predicted event, generating, by the one or more processors, an alert within the vehicle.
[0041] The one or more processors can be configured to execute tire encoder of the machine learning model using the first sequence of images by executing the encoder of the machine learning model to generate the first intermediary state indicating a first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle each at a first time step subsequent to the first time period in w hich the first sequence of images were captured. The one or more processors can be configured to execute tire future-state decoder of the machine learning model based on the first intermediary state by executing the future-state decoder of the machine learning model to generate a future state from the first plurality of candidate future states and of the first environment surrounding the vehicle or within the vehicle at the first time step subsequent to the first time period in which the first sequence of images were captured.
[0042] The one or more processors can be configured to execute the encoder of the machine learning model using the first sequence of images to generate the first intermediary state by executing the encoder of tire machine learning model to generate the first intermediary state comprising the first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at the one or more time steps subsequent to the first time period in which the first sequence of images were captured.
[0043] The one or more processors can be configured to execute the encoder of the machine learning model using the first sequence of images to generate the first intermediary state by executing the encoder of the machine learning model to generate the first intermediary state comprising the first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at each of a plurality of time steps subsequent to the first time period in which the first sequence of images were captured.
[0044] The one or more processors can be further configured to labeling, by the one or more processors, a third sequence of images to indicate the third sequence of images corresponds to a collision with the vehicle; and training, by the one or more processors, theclassification decoder based on whether the classification decoder generates a prediction that a collision will occur using the label of the third sequence of images.
[0045] The one or more processors can be configured to label the third sequence of images by labeling the third sequence of images to indicate the collision with the vehicle occurred at a first time step subsequent to the first time period in which the first sequence of images were captured; and wherein the one or more processors are configured to train, using the label of tire third sequence of images, the encoder based on whether the classification decoder generates a prediction that a collision will occur.
[0046] The one or more processors can be configured to execute the encoder of the machine learning model using the first sequence of images to generate the first intermediary state by executing the encoder of the machine learning model to generate a numerical vector representation of the first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at the one or more time steps subsequent to the first time period in which the first sequence of images were captured.
[0047] In one embodiment, a method can include receiving, by one or more processors, a sequence of images of a first environment surrounding a vehicle or within the vehicle captured by a camera mounted to or in the vehicle; executing, by the one or more processors, an encoder of a machine learning model using the sequence of images to generate an intermediary state indicating a plurality of candidate future states of the first environment surrounding tire vehicle or within the vehicle at one or more time steps subsequent to a time period in which the sequence of images were captured, wherein the encoder of the machine learning model is trained based on predictions generated by a future-state decoder configured to generate a future state of a second environment surrounding or within the vehicle or a second vehicle based on a set of candidate future states of the second environment surrounding or within the vehicle or the second vehicle; and responsive to executing, by the one or more processors, a classification decoder of tire machine learning model based on the intermediary state to detect a predicted event, generating, by the one or more processors, an alert within the vehicle.
[0048] Executing the encoder of the machine learning model using the sequence of images can include executing, by the one or more processors, the encoder of the machine learning model to generate the intermediary state indicating a first plurality’ of candidate futurestates of the first environment surrounding the vehicle or within the vehicle each at a first time step subsequent to the first time period in which the sequence of images were captured.
[0049] Executing the encoder of the machine learning model using the sequence of images to generate the intermediary state can include executing, by tire one or more processors, the encoder of the machine learning model to generate the intermediary state comprising the plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at tire one or more time steps subsequent to the first time period in which the sequence of images were captured.
[0050] In one embodiment, a method for vehicle risk assessment can include receiving an image; dividing the image into a plurality of patches; identifying objects of interest within the image using an Al perception engine; processing the identified objects of interest separately with a Vector Quantization (VQ) model; assigning codewords from a predefined vocabulary to each patch; and integrating the output of the VQ model with a risk prediction model. The objects of interest can include other vehicles within the image. The VQ model can process the patches containing the identified objects of interest with a higher level of detail compared to the rest of the image. The codeword assignment process can use different parts of the codeword vocabulary for the regions of interest and the rest of the image.
[0051] The risk prediction model can predict the risk of a collision based on the codeword representation of the image. The risk prediction model can use the spatial relationships between the codewords representing the ego vehicle and other vehicles to make its prediction. The method can further include processing sequences of images and including codewords that represent changes over time. The method can further include implementing a feedback loop where the risk prediction model influences tire VQ model. The Al perception engine, tire VQ model, and the risk prediction model can be implemented as part of an autonomous driving system.
[0052] In one embodiment, a non-transitory computer-readable medium can have stored thereon instructions that, when executed by a processor, cause the processor to perform the one or more of the methods described herein.
[0053] In one embodiment, a method can include receiving image data from a vehiclemounted camera; processing tire image data with a generative driving model, where thegenerative driving model was trained to predict a future image based on one or more images captured from a vehicle-mounted camera: and training a generative model with the video data.
[0054] In one embodiment, a method for predicting driving scenarios in an autonomous vehicle can include using a generative driving model trained on event-specific instances to predict future driving scenarios; employing a classifier model to process the hidden state of the generative driving model to classify potential outcomes; and providing a notification to the driver or other user based on the classified outcomes.
[0055] In one embodiment, a system for enhancing the safety of autonomous driving can include a perception engine that detects various elements in the driving environment and distinguishes between safe and unsafe driving maneuvers; a generative driving model trained on data derived from the perception engine to predict future driving scenarios; a classifier model that processes the hidden state of the generative driving model to classify potential outcomes; a response mechanism that reacts based on the classified outcomes.
[0056] In one embodiment, a method for predicting and preventing accidents in an autonomous vehicle can include detecting elements in the driving environment and distinguishing between safe and unsafe driving maneuvers using a perception engine; training a generative driving model on event-specific instances derived from the perception engine; predicting future driving scenarios using the generative driving model; processing the hidden state of the generative driving model using a classifier model to classify potential outcomes; generating a notification, assisting the driver, or taking control of the vehicle when a likelihood of collision exceeds a certain threshold.
[0057] These and other aspects and implementations are discussed in detail below. The foregoing information and the following detailed description include illustrative examples of various aspects and implementations and provide an overview or framework for understanding tire nature and character of tire claimed aspects and implementations. The drawings provide illustration and a further understanding of the various aspects and implementations and are incorporated in and constitute a part of this specification. Aspects can be combined, and it will be readily appreciated that features described in the context of one aspect of the invention can be combined with other aspects. Aspects can be implemented in any convenient form. For example, by appropriate computer programs, which may be carried on appropriate carrier media (computer-readable media), which may be tangible carrier media (c.g., disks) or intangible carrier media (e.g., communications signals). Aspects may also be implementedusing suitable apparatus, which may take tire form of programmable computers running computer programs arranged to implement the aspect. As used in the specification and in the claims, the singular form of ‘a’, ‘an’, and ‘the’ include plural referents unless the context clearly dictates otherwise.BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Non-limiting embodiments of the present disclosure are described by way of example with reference to the accompanying figures, which are schematic and are not intended to be drawn to scale. Unless indicated as representing the background art, the figures represent aspects of the disclosure. For purposes of clarity, not every component may be labeled in every drawing. In the drawings:
[0059] FIG. 1 illustrates an example environment showing a computing system for a generative driving model, according to an embodiment.
[0060] FIG. 2 illustrates a block diagram of machine learning models for event detection, according to an embodiment.
[0061] FIG. 3 illustrates a sequence diagram of a sequence for using an image compression model, according to an embodiment.
[0062] FIG. 4 illustrates a sequence diagram of a sequence for tokenizing sensor data, according to an embodiment.
[0063] FIG. 5 illustrates a sequence diagram of using a generative driving model, according to an embodiment.
[0064] FIG. 6A illustrates a flowchart of a method for compressing an image, according to an embodiment.
[0065] FIG. 6B illustrates a segmentation of an image for image compression, according to an embodiment.
[0066] FIG. 7 illustrates a flowchart of a method for implementing a generative driving model, according to an embodiment.
[0067] FIG. 8 illustrates an intermediary state of a generative driving model, according to an embodiment.
[0068] FIGS. 9A-9G illustrate example driving images that can be used for event detection, according to an embodiment.
[0069] FIGS. 10A-10D illustrate example driving images that can be used for event detection, according to an embodiment.
[0070] FIG. 11 illustrates a sequence diagram of a sequence for using a generative driving model, according to an embodiment.
[0071] FIGS. 12A-12F illustrate example driving images that can be used for event detection, according to an embodiment.
[0072] FIGS. 13A-13F illustrate example driving images that can be used for event detection, according to an embodiment.
[0073] FIG. 14 illustrates example driving images that are used for event detection, according to an embodiment.DETAILED DESCRIPTION
[0074] Reference will now be made to the illustrative embodiments depicted in the drawings, and specific language will be used here to describe the same. It will nevertheless be understood that no limitation of the scope of the claims or this disclosure is thereby intended. Alterations and further modifications of the inventive features illustrated herein, and additional applications of the principles of the subject matter illustrated herein, which would occur to one skilled in the relevant art and having possession of this disclosure, are to be considered within the scope of the subject matter disclosed herein. Other embodiments may be used and / or other changes may be made without departing from the spirit or scope of the present disclosure. The illustrative embodiments described in the detailed description are not meant to be limiting of the subject matter presented.
[0075] Autonomous vehicle or assisted vehicle guidance can use image analysis to generate alerts or make self-driving decisions. However, while images can play a role in the functioning of self-driving technology, real-time image analysis can be improved upon for more accurate and / or safe self-driving and / or alert generation. For example, images can be blunrfor various reasons, such as because of adverse weather conditions (e.g., heavy rain, snow, or fog), which can limit the system's ability to accurately interpret the surroundings. Additionally, variations in lighting conditions, such as glare from the sun or darkness at night,can also affect image quality and interpretation. Finally, identifying and accurately categorizing objects in real-time from images poses other challenges, particularly with complex scenarios in which there may be a large number of objects and only a subset of the objects may be pertinent to an event (e.g., a driving event, such as collision with another vehicle) that will occur.
[0076] For example, a computing system located on a vehicle (e.g., operating within or otherwise controlling tire vehicle) may continuously receive image or video data of the environment surrounding or inside the vehicle. The computing system may apply a rules engine or a machine learning algorithm to the images to determine whether to generate an alert (e.g., an audible or visual alert, such as to alert a vehicle ahead is within a proximity' threshold of the vehicle) or an operational control to use to control the vehicle (e.g., turn right onto a new7road or apply sudden or emergency brake). However, given the amount of data that is typically in any singular image, let alone multiple images of a video or a sequence of images, as well as the variation in quality in images that may be present depending on the environment, tire rules engine or machine learning algorithm may not be able to account for every situation, scenario, or object depicted in the image or video. In some cases, the rules engine or machine learning algorithm may attempt to take each object into account for the processing regardless of whether the objects are relevant to whether to generate an alert or generate a self-driving decision, which can require a large amount of processing resources and time. Thus, the computing system may generate false alerts or make improper driving decisions and / or incur a large amount of latency when generating such driving decisions, which may cause vehicles relying on such systems to be more dangerous. A system that is configured to generate alerts and / or decisions to selectively process the data depicted within the images can improve functioning of an automatic alert generation system and / or self-driving system by reducing the processing resources required and latency incurred for alert generation and / or self-driving decisionmaking.
[0077] To address these technical challenges, a computer or computing device implementing the systems and methods described herein may compress images used to detect events to focus on aspects of the images that are relevant for alert generation and / or self-driving decision-making (e.g., relevant to detecting events). For example, a computing device of a vehicle may receive one or more images of an environment surrounding a vehicle or w ithin the vehicle captured by a camera mounted to or in die vehicle. The computing device may execute an object classifier (e g., a neural network or another type of machine learning model) usingeach of the one or more images to detect one or more objects depicted within at least one of the one or more images.
[0078] The computing device can use the detected objects to selectively compress the at least one image. To do so, the computing device can identify portions of an image that each depict at least one of the detected one or more objects that relate to operation of the vehicle. For example, the object classifier can identify objects within the environment depicted in the image. The object classifier can classify’ whether the objects relate to operation of the vehicle. Such objects can be learned or defined and include objects that may have an impact on a travel direction, speed, or otherwise control of the vehicle. For example, objects that relate to operation of the vehicle can include stop signs, traffic lights, a tire (e.g., a tire that fell off of another vehicle), other vehicles on the road, pedestrians, animals, etc., because the vehicle may stop, change speed, or turn or swerve based on the objects. The portions of the image may be defined grid locations or defined shaped areas within tire image. The portions can depict such objects if at least one or a portion of one (e.g., a defined percentage of one) object classified as relating to operation of the vehicle. The computing device can identify or determine portions of the image that depict at least one object that relates to operation of the vehicle and portions of the image that do not depict any of such objects. The computing device can convert the image into learned codewords by using codewords selected from a first codeword space for the portions of the image that depict at least one object relating to the operation of the vehicle and using codewords selected from a second codeword space for the other portions of the image.
[0079] The computing device can execute a machine learning model (e.g., a classification model or a large language model) using the learned codewords representing the image to detect a predicted event. Responsive to detecting the predicted event, the computing device can generate an alert within the vehicle.
[0080] By compressing the image into codew ords in this manner, the computing device may cause the machine learning model to more accurately detect predicted events from the image. The computing device may do so because the first codeword space may be larger (e.g., include more codewords) than the second codeword space. For example, because the first codeword space can include more codewords than the second codeword space, the computing device can use more codewords to describe or that correspond to tire portions of the image that contain the objects that relate to operation of the vehicle than the other portions of the image. Because more codewords may be used, the computing device may have more details of theportions of the image to process than the other portions. This can reduce the processing power and time that the computing device uses to process the portions that do not contain data that is relevant to generating an alert or vehicle control, which can increase the accuracy of the event detection and reduce the amount of time tire processing takes.
[0081] A computing device may use a machine learning model (e.g., a large language model or a classification model or classification decoder) to detect events from compressed images or non-compressed images. For example, the computing device may receive a first sequence of images of a first environment surrounding a vehicle or within the vehicle captured by a camera mounted to or in the vehicle. The computing device can execute an encoder of a large language model including the encoder and a decoder using the first sequence of images or compressed versions of the images (e.g., codewords generated from the images or images generated from codewords of the images) to generate a first intermediary state. The first intermediary state can be an embedding, a numerical vector, a string, or another representation indicating a first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at one or more time steps subsequent to a first time period in which the first sequence of images were captured. The first intermediary state can include or indicate a confidence score for each of the first plurality of candidate future states indicating the likelihood that the candidate future state is correct. The computing device can execute a future-state decoder of the machine learning model based on the first intermediary state to generate a future state from the first plurality of candidate future states and of tire first environment surrounding the vehicle or within the vehicle at a time step subsequent to the first time period in which the first sequence of images was captured. The computing device can train the encoder using a loss function based on the accuracy of the generated future state of the first environment surrounding the vehicle.
[0082] Subsequent to training the encoder (e.g., subsequent to training the encoder until the encoder is accurate to a threshold), the computing device can receive a second sequence of images of a second environment surrounding the vehicle or within the vehicle captured by the camera mounted to or in the vehicle. The computing device can execute the trained encoder of the machine learning model using the second sequence of images to generate a second intermediary state indicating a second plurality' of candidate future states of the second environment surrounding the vehicle or within the vehicle at one or more second time steps subsequent to a second time period in which the second sequence of images were captured. The computing device can then execute a classification decoder of the machine learning model(e.g., instead of the future-state decoder) based on the second intermediary state to generate an alert within the vehicle or otherwise control the vehicle.
[0083] In this way, the computing device can avoid or stop using the future-state decoder to generate a predicted future state for event detection, alert generation, and / or automatic driving decision-making, and directly use the classification decoder instead, which can substantially reduce the computational resources and time that is required to predict events (e.g., driving events), such as collisions, without the extra decoder processing step. In doing so, the computing device can detect risks of collisions faster than other systems, which can aid a self-driving system avoid such collisions or facilitate earlier alert generation to alert the driver of tire upcoming collision to give the driver more time to react.
[0084] FIG. 1 depicts an example environment that includes example components of a system in which such a computer can selectively generate instructions to generate alerts for driving events. Various other system architectures may include more or fewer features and / or may utilize the techniques described herein to achieve the results and outputs described herein. Therefore, the system depicted in FIG. 1 is a non-limiting example.
[0085] FIG. 1 illustrates a system 100, which includes components of a generative driving model system 105 for using a generative artificial intelligence model on images captured by a camera attached to or integrated on or within a vehicle 110 for autonomous driving or alert generation. The system 100 can include the generative driving model system 105, the vehicle 110, and / or a cloud computing system 115. The generative driving model system 105 can include a computing device 120. a camera 125. and / or a communication interface 130. The generative driving model system 105 may include an alert device, such as an audio alarm, a warning light, another type of visual indicator, and / or a haptic feedback device (e.g., one or more vibrating seats or a vibrating steering wheel). The generative driving model system 105 can be mounted on a dashboard, windshield, or other area inside the vehicle 110. The system 100 is not confined to the components described herein and may include additional or other components, not shown for brevity, which are to be considered within the scope of the embodiments described herein.
[0086] The vehicle 110 can be any type of vehicle, such as a car, truck, van, sportutility-vehicle (SUV), motorcycle, semi-tractor trailer, or other vehicle that can be driven on a road or another environment. The vehicle 110 can be operated by a user, or, in some implementations, can include an autonomous vehicle control system (not pictured) thatnavigates the vehicle 110 or provides navigation assistance to an operator of the vehicle 110. The vehicle 110 can be a vehicle of a fleet of vehicles that are owned and / or operated by an entity (e.g., a business or organization) to transport goods, materials, and / or individuals.
[0087] The vehicle 110 can include the generative driving model system 105, which can be used to detect objects within images captured by the camera 125 when the vehicle is parked and / or as tire vehicle is driving down a road and generate alerts or control the vehicle 110 based on the detected objects. As outlined above, the generative driving model system 105 can include the computing device 120. The computing device 120 can be mounted on or in the vehicle 110. In some cases, the computing device 120 is a computing device that automatically controls the vehicle for self-driving or generates audible or visual alerts via devices within the vehicle 110.
[0088] The computing device 120 can include tire storage 135, which can store images and / or video captured by the camera 125. machine learning model(s) 140, and an application 150. The storage 135 can be a computer-readable memory that can store or maintain any of the information described herein that is generated, accessed, received, transmitted, or otherwise processed by the computing device 120. The storage 135 can maintain one or more data structures, which may contain, index, or otherwise store each of the values, pluralities, sets, variables, vectors, numbers, or thresholds described herein. The storage 135 can be accessed using one or more memory addresses, index values, or identifiers of any item, structure, or region of memory maintained by the storage 135.
[0089] The storage 135 may be internal to the computing device 120 or may exist externally to the computing device 120 and accessed via a suitable bus or interface. In some implementations, the storage 135 can be distributed across many different storage elements. The computing device 120 (or any components thereof) can store, in one or more regions of tire memory of the storage 135, the results of any or all computations, determinations, selections, identifications, generations, constructions, or calculations in one or more data structures indexed or identified with appropriate values.
[0090] The computing device 120 can include or be in communication with a communication interface 130 that can communicate wirelessly with other devices. The communication interface 130 of the computing device 120 can include, for example, a Bluetooth communications device, a Wi-Fi communication device, or a 5G / LTE / 3G cellular data communications device. The communication interface 130 can be used, for example, totransmit any information described herein to the cloud computing system 115, including images and / or videos that the computing device 120 receives from the camera 125. The communication interface 130 can also be used, for example, to receive indications of driving events that the processor detects from tire images or videos, in some cases with the images or videos from which the processor detects the driving events.
[0091] The camera 125 can include any type of camera or multiple cameras that are capable of capturing images or videos of the environment surrounding the vehicle 110 and / or within the vehicle 110, including a configuration having an external camera view of the surrounding environment and a cabin-facing camera view of the driver. The camera 125 may periodically capture images or video while the vehicle 110 is turned on and parked and / or as the vehicle 110 is moving. For example, the camera 125 can capture an image or video of a stop sign 180 as the vehicle 110 approaches or passes the stop sign 180. In another example, the camera 125 can capture an image or video of the driver before, during, and / or after a driving event. In some cases, the camera 125 can capture images or video when the vehicle 110 is turned off (e.g., when the vehicle is off but in a surveillance mode). The camera 125 may capture images or videos and transmit the images or videos to the computing device 120. The computing device 120 may receive the images or videos and store images or videos in the storage 135 and / or transmit tire images or videos to other vehicles or to the cloud computing system 115.
[0092] The camera 125 can communicate with the computing device 120 via a vehicle interface, which may include a CAN bus or an on-board diagnostics interface. The camera 125 can capture images or video and transmit the captured images or video to the computing device 120. The computing device 120 can receive the captured images or video and transmit the images or video to the cloud computing system 115 to use to detect driving events for a driver of tire vehicle 110.
[0093] The machine learning model(s) 140 can include one or more machine learning models or computer models that the computing device 120 can use for autonomous driving and / or alert generation. For example, referring now to FIG. 2, the machine learning model(s) 140 can include an object classifier 141, an image compressor 142, a data tokenizer 143, a large language model (LLM) 144, and / or an event detector 145. Each of the machine learning model(s) 141-145 can be or include one of a neural network, a support vector machine, a random forest, a clustering algorithm, etc.
[0094] The object classifier 141 can be or include a machine learning model (e.g.. a neural network) that is configured to use image processing techniques (e.g., as a convolutional neural network) to detect objects in individual images when using the individual images as input. The object classifier 141 can detect the objects in respective bounding boxes. The object classifier 141 can additionally or instead determine a type of the detected objects, in some cases. The computing device 120 can execute the object classifier 141 using a sequence of images (e g., set of sequentially received images) received from the camera 125 as input. The computing device 120 can execute the object classifier 141 using each image of the sequence of images as input at once or separately execute the object classifier 141 for each image of the sequence of images. The object classifier 141 can output identifications of objects the object classifier 141 detects from the sequence of images, in some cases with identifications of the respective images, locations of the objects within the images, sizes of the objects, orientations of the objects, types of the detected objects, and / or any other data regarding the objects. In some cases, the computing device 120 can use another type of machine learning model for such image processing and / or object detection.
[0095] In some cases, the object classifier 141 can additionally or instead classify the objects that the object classifier 141 detects as being related to operation of the vehicle 110. The object classifier 141 can do so, for example, based on the weights or parameters (e.g., trained weights or parameters) of the object classifier 141. In some cases, the object classifier 141 can use a table for such classification. For example, tire object classifier 141 can identify the types of the detected objects and compare the identified types to a table (e.g., a look-up table) that contains mappings of types of objects to indications of whether the objects are related to operation of the vehicle 110. The object classifier 141 can compare identifications of the detected objects to the table to determine whether the detected objects relate to operation of tire vehicle.
[0096] The image compressor 142 can be a machine learning model or another type of computer model that can compress received images. The image compressor 142 can compress images by converting the images into codewords (e.g.. learned codewords). In some cases, the image compressor 142 can convert the codewords of an image back into an image (e.g., a compressed version of the image with less detail or a lower resolution than the initial image). By compressing images in this way, tire image compressor 142 can reduce the amount of data further processing of the images may incur.
[0097] An example of a sequence 300 of using the image compressor 142 to compress an image is shown in FIG. 3. In the sequence 300, the image compressor 142 can include an encoder 302 (e.g., a VQ-GAN encoder) and a decoder 304 (e.g., a VQ-GAN decoder). The encoder 302 can be trained to convert images into codewords (e.g., numerical values, alphabetical values, alphanumerical values, tokens, etc.). The codewords can be stored in a data structure or dictionary (e g., a codeword space) in memory of tire computing device 120. The encoder 302 can be configured to convert portions or regions of images into the codewords, such as by retrieving the codewords from storage and assigning the codewords to the corresponding locations of the images. The decoder 304 can be configured or trained to convert a set of generated codewords into an image (e.g., a compressed image). Because the decoder 304 converts codewords back into an image and the codewords may not capture all of the details of the image, or otherwise may genericize different details, the decoder 304 may generate tire image to have less detail or a lower resolution than the original image.
[0098] The data structure or dictionary containing tire codewords can be divided into separate codeword spaces. For example, die data structure or dictionary can include a first codeword space and a second codeword space. The first codeword space may correspond to codewords that describe or correspond to portions or regions of images that depict an object that relates to operation of a vehicle. The second codeword space may correspond to codewords that describe or correspond to portions or regions of images diat do not depict any object that relates to operation of a vehicle. The first codeword space may be larger than the second codeword space. Because the first codeword space may be larger, when the encoder 302 uses the first and second codeword spaces to compress an image, the encoder 302 may include more codewords for portions or regions of images that depict objects that relate to operation of a vehicle from (e.g., retrieved from) the first codeword space than other portions of the images. Accordingly, more detail about the image may be captured with the codewords for decoding or further processing. Because more details may be included for portions that include relevant portions of images than other portions, any further processing may cause the computing device 120 to more accurately detect alerts and / or use less processing power on portions of images that are not relevant to operation or generating an alert.
[0099] For example, in the sequence 300, the image compressor 142 can receive an image 306 of the environment surrounding a vehicle. The image compressor 142 can perform pre-processing on the image 306 to format the image to match the inputs (e.g., input nodes) of the encoder 302 of tire image compressor. The image compressor 142 can do so by performingone or more cropping or resizing operations 308, for example, of the image 306 into a reformatted image 310. The computing device 120 can execute the encoder 302 using the reformatted image 310 to convert the reformatted image 310 into codewords 312 from a codeword vocabulary or dictionary (e.g., containing 4096 discrete codewords). The computing device 120 can do so after using the object classifier 141 to identify objects depicted in the image 306 or the reformatted image 310 that relate to operation of a vehicle and including identifications of the classified objects in the input into the encoder 302 with the reformatted image 310. The computing device 120 can input the codewords 312 into the decoder 304 and execute the decoder 304 to generate a compressed image 314.
[0100] In some cases, when encoding the reformatted image 310 into the codewords 312, the encoder 302 can do so by assigning or determining codewords for separate defined portions or regions of the reformatted image 310. For example, the encoder 302 can divide the reformatted image 310 into a grid with separate portions separated by lines. The portions can be equally sized squares or rectangles, for example. The encoder 302 can determine whether the portions contain a portion (e.g., at least a defined percentage) of at least one object identified as related to operation of a vehicle. For example, the encoder 302 can identify portions of the reformatted image 310 that contain a pedestrian crossing the road and a bird flying overhead. For portions of the reformatted image 310 that the encoder 302 identifies as containing a portion of at least one object identified as related to operation of a vehicle, such as a pedestrian crossing tire road, the encoder 302 can identify or determine codewords from the first codeword space. The encoder 302 can use codewords from the second codeword space for the other portions of the reformatted image 310, such as the bird flying overhead.
[0101] Thus, the image compressor 142 can convert visual data into a form that the LLM 144 or the event detector 145 can interpret and learn from. In doing so, the image compressor 142 can process visual data into a format that tire LLM 144 or the event detector 145 can interpret. The format can be or include tokens from a dictionary or data structure including 4096 discrete codewords, each representing a different visual feature or aspect. The visual data can be divided into 8x16 patches, and each patch can then be associated with one of the 4096 codewords, creating a 4096-token vocabulary. The vision tokenizer may be trained using detection-guided vector quantization (VQ) for driving images in some cases.
[0102] The data tokenizer 143 can be or include a machine learning model or a computer model that is configured to generate tokens from data collected from different sensorsof the vehicle 110. The data tokenizer 143 can convert data values collected from different sensors of the vehicle 110 into tokens or values that represent the data values. In doing so, the data tokenizer 143 can convert the data values into tokens that include less data (e.g., fewer bytes or bits) than the data values themselves. In this way, the image compressor 142 can reduce the amount of data further processing of the data values may incur.
[0103] An example of a sequence 400 of using the data tokenizer 143 to tokenize the data values collected from sensors of tire vehicle 110 is shown in FIG. 4. In the sequence 400, the data tokenizer 143 can receive sensor data from different sensors of the vehicle 110. Examples of such sensor data can include speed or velocity values from a velocity' or speed sensor of the vehicle 110, gyroscopic values from a gyroscope of the vehicle 110, and / or accelerometer values from an accelerometer of the vehicle 110. The data tokenizer 143 can perform a resampling operation 404 on the received values to tie values to individual captured images or frames (e.g., images or frames captured by the camera 125). To do so, the data tokenizer 143 can identify timestamps of the received values and received images and filter out any values with timestamps that are not within a defined threshold of timestamps of the images or frames. The data tokenizer 143 can use a minimum RMS quantizer 406 on the remaining values to generate a quantized representation of the values that minimize the root mean square error between tire original remaining values and the quantized representation. The values of tire quantized representation can be tokens that can be used for processing for event detection. In some cases, the data tokenizer 143 can convert the remaining values into other tokens according to a data structure with a defined vocabulary 408 for different types of data values. For example, the data tokenizer 143 can convert speed data values into tokens using a token dictionary' with a 40 token vocabulary for speed data values and gyroscopic data values into tokens using a token dictionary with a 30 token vocabulary per axis.
[0104] Accordingly, the data tokenizer 143 can process data from various sensors on the vehicle 110. In doing so, the data tokenizer 143 can translate the sensor data into a tokenized format that the LLM 144 or the event detector 145 can understand, allowing the LLM 144 or the event detector 145 to integrate sensor information into its understanding of the driving environment.
[0105] Referring again to FIG. 2, The LLM 144 can be or include a large language model or small language model that is configured to receive a prompt containing content such as text, videos, and / or images as input. The LLM 144 can be a transformer, a neural network,or any other type of machine learning configured to receive content as input and apply learned weights and / or parameters to the content to generate a response to the received content. The LLM 144 can include an encoder 146 and a future-state decoder 147. The encoder 146 can be or include a neural network or nodes of a neural network that can receive content and generate an intermediary state (c.g., an embedding, a vector, a string, or a value) based on the content. The intermediary state may be, include, or correspond to one or more potential outputs of the future-state decoder 147. The future-state decoder 147 can process the intermediary state to generate one or more (e.g., one or a plurality of) potential responses or answers based on the intermediary state or the one or more potential outputs representing the intermediary state.
[0106] For example, the encoder 146 can receive as input a sequence of compressed images depicting the environment surrounding tire vehicle 110 from the image compressor 142 and / or codewords generated from a sequence of images depicting the environment surrounding the vehicle 110 by the image compressor 142. The encoder 146 can additionally or instead receive tokens of sensor data generated by the data tokenizer 143 that correspond to the sequence of images. The computing device 120 can execute the encoder 146 based on the input to generate an intermediary state. The intermediary state can be, include, or correspond to one or more candidate future states of the environment surrounding the vehicle that can be output or generated by the future-state decoder 147. The intermediary state can include or correspond to confidence scores representing probabilities that each of tire one or more candidate future states is the correct future state of the environment surrounding the vehicle.
[0107] The computing device 120 can input the intermediary state generated by the encoder 146 into the future-state decoder 147 and execute the future-state decoder 147. The execution can cause the future-state decoder 147 to generate a representation of the future state of tire environment surrounding or within the vehicle 110. The representation can be or include one or more tokens (c.g., values of a vector) or a vector representing or one or more images depicting the environment at each of one or more (e.g., a plurality of) time steps into the future after a time period in which the sequence of images was captured (e.g.. a representation or image of the environment .5 seconds after the time period, 1 second after the time period, f .5 time periods after the time period, etc.) The time steps can be at any interval and / or include any number of time steps.
[0108] In one example, referring now to FIG. 5, a sequence diagram of a sequence 500 for using a generative driving model is shown, according to some embodiments. A dataprocessing system can perform the operation of the sequence 500. The data processing system can be or include a computing system of a vehicle (e.g., the generative driving model system 105) and / or a remote computing system (e.g., a cloud server or the cloud computing system 115) for using a generative artificial intelligence model for alert generation, in accordance with an embodiment.
[0109] In the sequence 500, the data processing system can generate and / or input tokens 502 representing the surrounding environment of a vehicle into a machine learning model 504. The tokens 502 may be codewords or textual tokens or values depicting or representing a sequence of images captured by a camera of a vehicle driving down the road generated by the image compressor 142 and / or the data tokenizer 143. The tokens 502 may be inserted into a feature vector or a prompt in sequence such that tokens representing the initial images of the sequence are first then tokens representing the other images of the sequence are after the initial tokens in tire feature vector or prompt, thus capturing tire sequence of the images. The tokens 502 can additionally or instead include tokens generated from data values collected from sensors of the vehicle, such as tokens generated from speed data or gyroscopic data of the vehicle. The tokens generated from the sensor-generated data values may correspond to the same time period as the time period in which the sequence of images was captured.
[0110] The data processing system can execute the machine learning model 504 based on the tokens 502. The execution can cause the machine learning model 504 to generate tokens 506 that represent or depict a future state of the environment surrounding the vehicle at one or more time steps subsequent the time period of the sequence of images. The tokens 506 can include predicted sensor tokens representing predicted values of sensor data at the one or more time steps and / or visual tokens representing the state of the environment at the one or more time steps. The tokens 506 can additionally or instead include environment tokens, which arc designed or trained to represent aspects of the driving environment. Other types of tokens that can be included in the tokens 506 can include ego vehicle class tokens, which categorize a vehicle (e.g., a self-driving vehicle into one of eight classes, micro weather tokens describing the current weather conditions (clear, rain, fog, snow), and time of day tokens indicating whether it's day, dawn, dusk, or night. The environment token vocabulary may include 16 tokens. Separator tokens can be used to distinguish between different time frames in the data. There may only be one token in this vocabulary.
[0111] The data processing system can process the tokens 506 to detect an event, such as a collision. The data processing system can generate an alert based on the detected event.
[0112] In an example, the machine learning model 504 can be multi-modal, meaning the machine learning model 504 can process and integrate different types of data, such as visual inputs, sensor readings, and temporal information. These additional sensor streams may enable greater coverage for alert generation or autonomous driving, as different data types provide complementary information about the driving environment.
[0113] To manage the different data modalities, the data processing system can synchronize the data in the different modalities to ensure that all information corresponds to the same time frame. This synchronization allows the machine learning model 504 to generate a coherent understanding of the current driving scenario. A separator token can be used between time frames to clearly delineate different moments in time. This aids in maintaining the temporal order and separating the prediction for one time frame from the next.
[0114] The output of the machine learning model 504 can provide a prediction for the next frame, based on tire current and past information. This prediction can include, for example, the positions and movements of other vehicles, pedestrians, and road signs, as may be detected by the machine learning models or applications running on an edge safety device.
[0115] Moreover, the machine learning model 504's hidden states (e.g.. intermediary states, as described herein) can encapsulate probabilities over future scenarios. This means the machine learning model 504 not only predicts a single outcome but considers a range of possible futures, which can be useful for handling the inherent uncertainty and variability in real-world driving scenarios. As described herein, the hidden states may be accessed by other networks or processing modules, such as an accident prediction module. The architecture of the machine learning model 504 can also include a tokenization vocabulary, which may translate different types of input data into a representation that the machine learning model 504 can understand.
[0116] Referring again to FIG. 2, the computing device 120 can execute the event detector 145 to detect driving events based on outputs of the LLM 144, such as based on an output indicating or depicting the future state of the objects and / or the environment (e.g., in the images at the time steps). For example, the event detector 145 can analyze the output of the future-state decoder 147 (e.g., tire future state of the objects depicted in tire sequence of images)using one or more rules or a machine learning model (e.g., a neural network, a support vector machine, a random forest, etc.). When using a rule-based approach, the event detector 145 can compare the future state (e.g., detected objects within the future state), generated codewords, and / or the sequence of images to one or more rules to determine whether to generate an alert or make an autonomous driving decision. For example, the event detector 145 can determine a vehicle in front of tire vehicle 110 will be within a threshold distance of the vehicle 110 in the image two seconds into the future based on the images or codewords satisfying a rule configured to be satisfied if codewords generated from a sequence of images indicate a vehicle is within a threshold distance of the vehicle 110. As applied to generating an autonomous driving decision, the event detector 145 can determine to automatically slow the vehicle 110 down or change lanes responsive to determining the vehicle in front of the vehicle 110 is within a threshold distance of the vehicle 110 in the image of the future state of tire objects in the environment surrounding the vehicle 110.
[0117] The computing device 120 can use outputs of the future-state decoder 147 to train the encoder 146. For example, the computing device 120 can execute the encoder 146 using a sequence of images (e g., a sequence of images each compressed by the image compressor 142) to generate an intermediary state indicating or including a plurality of candidate future states (e.g., confidence scores or probabilities for a plurality of candidate future states) of the environment surrounding tire vehicle or within the vehicle at one or more time steps subsequent to a time period in which the sequence of images were captured. The computing device can execute the future-state decoder 147 using the intermediary state to generate a future state of the environment surrounding or within the vehicle at one or more time steps subsequent the time period in which the sequence of images was captured. The future state can include an image or vector at each of the one or more time steps. The computing device 120 can identify images captured by the camera 125 at the time steps. The computing device 120 can adjust tire weights and / or parameters of tire encoder 146 and / or the future-state decoder 147 based on any differences (e.g., using a loss function) between the future states and the images captured by the camera 125 at the time steps. In some cases, the computing device 120 may freeze the future-state decoder 147 and only train the encoder 146 to prioritize training the encoder 146. The computing device 120 can train the encoder 146 and / or the future-state decoder 147 in this manner over time until determining the encoder 146 and / or future-state decoder 147 arc generating outputs at an accuracy above a threshold.
[0118] Responsive to determining the outputs generated by the future-state decoder 147 are at an accuracy above a threshold, the computing device 120 can stop using the futurestate decoder 147 for event detection and instead use the intermediary states generated by the encoder 146 directly for event detection. For instance, tire event detector 145 can include a machine loaming model or rules engine that is configured or trained to detect events (e.g., collisions or future collisions) based on intermediary states generated by the encoder 146. The machine learning model can be a classification machine learning model (e.g.. a classification decoder) configured to output indications of whether a collision will occur based on intermediary states generated by the encoder 146. The computing device can execute the machine learning model of the event detector 145 using intermediary states generated by the encoder 146 to generate classifications of whether the intermediary states correspond to events or collisions.
[0119] In some cases, responsive to determining to stop using the future-state decoder 147, tire computing device 120 can initiate training of the machine learning model of the event detector 145 configured to process intermediary states generated by the encoder 146. The computing device 120 can do so using labeled training data of sequences of images. For instance, the sequences of images may be labeled based on whether an event or a collision occurred subsequent to (e.g., within a defined time period of, such as within one second of) the time periods in which the sequences of images were respectively captured. The computing device 120 can feed the sequences of images into the encoder 146 and execute the encoder 146 to generate intermediary states for each of the sequences of images. The computing device 120 can execute the event detector 145 using the intermediary states to generate outputs indicating whether a collision or another type of event will occur subsequent the time periods of the sequences of images. The computing device 120 can use backpropagation techniques with a loss function based on differences between the outputs of the machine learning model of the event detector 145 and tire labels of the sequences of images based on which the outputs were generated to adjust the weights or parameters of the machine learning model of the event detector 145. In some cases, the computing device 120 can train the encoder 146 using the same or a different loss function based on the differences. In doing so, the computing device 120 may train the machine learning model of the event detector 145 and / or the encoder 146 itself based on the trained state of the encoder 146, which may be necessary given tire outputs of the encoder 146 may change based on the training of the encoder 146 with the future-state decoder 147, and the machine learning model of the event detector 145 may not be able toprocess or accurately process the outputs without being trained based on outputs of the trained encoder 146.
[0120] Referring again to FIG. 1, The cloud computing system 115 can be or include one or more computing devices that are configured to train and distribute machine learning models for object detection from images and / or decisions for alert generation or autonomous driving decisions. The cloud computing system can include a processor 170 and / or a memory 175. The cloud computing system 115 may receive training sequences of images from computing devices of multiple vehicles and train an individual machine learning language processing model to generate tokens or images indicating or depicting a future state of objects depicted in a sequence of images (e.g., the future state of the objects in a new image). The tokens can be a group of words, subwords (e.g., portions of individual words or across different sequential words), or characters. The cloud computing system 115 can train the machine learning language processing model to do so based further on sensor data (e.g., acceleration, speed, gyroscope data, weather data, etc.) that correspond to the individual sequences of images, identifications of types of drivers that were driving when the sequences of images were captured, and / or identifications of the types of vehicles that were being driven. The cloud computing system 115 may train one or more machine learning language processing models using such sequences of images and / or other data collected from different vehicles (e.g., vehicles of a fleet of vehicles) and transmitted to the cloud computing system 115 over a network. The cloud computing system 115 can train the machine learning language processing model to do so and transmit the machine learning language processing models (e.g.. copies of the machine learning language processing models) to computing devices of vehicles (e.g., the computing device 120) once the machine learning language processing model or models are sufficiently trained. After transmitting the machine learning language processing models to the computing devices, the cloud computing system 115 can continue to train a local version of tire machine learning language processing models with training images to improve the machine learning language processing models and / or account for changes in new objects that appear in the environment and / or changes in the cameras that are capturing the images. The cloud computing system 115 can provision the updated models to the computing devices of vehicles after each training iteration and / or at defined time intervals such that the computing devices can continue to use a more accurate machine learning language processing model over time.
[0121] The cloud computing system 115 may use parallel processing techniques or have different computers perform different tasks to facilitate the operations described herein. For example, one computer of the cloud computing system 115 can establish connections with the computers of vehicles to receive sequences of images from generative driving model systems of the vehicles (e.g., the generative driving model system 105), another computer of the cloud computing system 115 can process the received sequences of images using a machine learning language processing model to generate tokens or depictions of objects or features depicted in the sequences of images in a future state and train the machine learning language processing model based on the prediction (e.g., by using backpropagation techniques), and another computer of the cloud computing system 115 can establish connections with different computing devices of vehicles and transmit the trained machine learning language processing model to the different computing devices to use for autonomous driving and / or alert generation. Any combination of one or more of the computers of the cloud computing system 115 and / or the generative driving model system 105 may perform such processes.
[0122] A computing device 155 can be a computer that is configured to receive messages and / or present messages or other data on a user interface. The computing device 155 can include a processor 160 and a memory' 165. The computing device 155 can be any type of computing device, such as a mobile phone, laptop computer, desktop computer, smart watch, gaming console, personal data assistant, dashboard computer, or other computing device. A driver associated with the vehicle 110 can own or otherwise access the computing device 155 to access a platform provided by the cloud computing system 115 that manages accounts for different drivers (e.g., drivers of a driver fleet managed by the cloud computing system 115). The driver can access the platform through the account for the driver that the cloud computing system 115 stores in the memory 165.
[0123] The computing device 120 can execute the application 150 to determine whether to generate an alert and / or to determine or generate an autonomous driving decision based on the future state of the objects and / or the environment (e.g., in the image). For example, the application 150 can be configured to execute the machine learning models 140 over time using sequences of images and / or sensor data collected by sensors of the vehicle 110 as input to generate outputs indicating whether an event was detected. Responsive to detecting an event, the application 150 can generate an audible or visual alert within the vehicle 110 and / or adjust operation or control the vehicle 110.
[0124] For example, responsive to detecting an event in which a vehicle is too close (e.g., below a threshold) to a vehicle in front of the vehicle 110 or will be within a threshold distance of the vehicle 110 two seconds into the future, the application 150 can generate an alert. As applied to generating an autonomous driving decision, the application 150 can determine to automatically slow the vehicle 110 down or change lanes responsive to determining the vehicle in front of the vehicle 110 is within a threshold distance of the vehicle 110. In another example, the application can identify an output indicating a collision with the vehicle 110 is imminent (e.g., was detected based on a future state of the environment (e.g.. the objects in the environment) surrounding the vehicle 110). Responsive to identifying the output indicating the collision, the application 150 can generate an alert or execute a driving operation to reduce the chances of a collision, such as by controlling the vehicle 110 to automatically slow down, stop, or avoid an object in the path of the vehicle 110.
[0125] The computing device 120 can operate to autonomously drive the vehicle 110 and / or generate audible and / or visual alerts within the vehicle. The computing device 120 can do so using a machine learning architecture on images or video captured by the camera 125. For example, the camera 125 can generate or capture a plurality of images (e.g., a sequence of images or images of a video) depicting the environment surrounding the vehicle 110 and / or within the vehicle 110. The camera 125 can capture the individual images of a sequence of images at a defined interval or pseudo-randomly. The camera 125 can transmit or send the images to the computing device 120 responsive to capturing the images and / or in a batch at defined time intervals. The camera 125 can attach timestamps to the individual images as metadata indicating the times in which the camera 125 captured the images and / or the times in which the camera 125 transmitted or sent tire images to the computing device 120. Any number of cameras of the vehicle 110 similar to the camera 125 can transmit such images to the computing device 120 over time as the vehicle 110 is driving or is at rest.
[0126] For example, the computing device 120 can receive a sequence of images depicting the environment surrounding tire vehicle 110 from the camera 125 and / or other cameras of the vehicle 110. The sequence of images may have been captured within a defined time period. The computing device 120 can use the object classifier 141 to detect objects that relate to operation of the vehicle (e.g., objects that may relate to causing the vehicle 110 to crash or collide with another object). The computing device 120 may use detected objects as input into tire image compressor 142 to generate codewords for regions or portions within eachof the images and / or decode such codewords into compressed versions of each of tire sequence of images.
[0127] The computing device 120 can input the codewords or the sequence of compressed images into the LLM 144. The computing device 120 can execute the encoder 146 of the LLM 144 based on the input to generate an intermediary state that corresponds to or represents multiple candidate future states of the environment around tire vehicle 110 at each of one or more time steps subsequent the time period in which tire sequence of images was captured. The computing device 120 can execute the event detector 145 using the intermediary state to generate a classification that an event or collision will occur in the future or at a defined time step into the future. Responsive to the output, the computing device 120 can generate an audible or visual alert within the vehicle 110 or otherwise control the vehicle 110 to avoid the predicted collision or event.
[0128] FIG. 6A illustrates a flow of a method 600 for compressing an image, according to an embodiment. A data processing system can perform the operation of the method 600. The data processing system can be or include a computing system of a vehicle (e.g., the generative driving model system 105) and / or a remote computing system (e.g., a cloud server or the cloud computing system 115) for using a generative artificial intelligence model for alert generation and / or self-driving decision-making, in accordance with an embodiment. The data processing system can perform the method 600 to compress images to be used by the generative artificial intelligence model in the system. The method 600 is shown to include steps 602-612. However, other embodiments may include additional or alternative steps, or may omit one or more steps altogether. Different steps can be performed by different computing systems (e.g., the computing system of the vehicle can perform one or more of the steps 602-612 and / or the remote computing system can perform one or more of the steps 602-612) and / or the different computing systems and can operate together to perform individual steps of the steps 602-612.
[0129] In step 602, the data processing system can receive one or more images. The one or more images can be or include a sequence of images or images of a video (e.g., video data) of an environment surrounding a vehicle or the environment within the vehicle. The sequence of images can be consecutively captured images by the same camera. The images of the sequence can be captured at a set time interval (e.g., every five seconds) or pseudo- randomly. The images can be images of a road on which the vehicle is driving or of areas adjacent to the road. The data processing system can receive tire images from a cameramounted to or within the vehicle (e.g., mounted on the dashboard of the vehicle). The images can be or include a JPEG, RAW, or another image file type. The visual data can be or include one or more standalone images or one or more frames of a video. The data processing system can receive the images as a driver is driving the vehicle, when the driver has parked the vehicle, and / or when the vehicle is in a standby mode. The data processing system can receive such images at defined intervals in the case of static standalone images or as the camera streams a video to the data processing system in individual frames.
[0130] In step 604, the data processing system can execute an object classifier using the one or more images (e g., the sequence of images) as input. The data processing system can execute tire object classifier (e.g., an object detection machine learning model, such as a neural network) to detect one or more objects depicted in or w ithin at least one of the sequence of images. The data processing system can execute the object classifier using the sequence of images as input to output or detect one or more objects or features that are depicted in the one or more images (e.g.. in each of the one or more images). In doing so, the data processing system can execute the object classifier using individual images of the sequence of images for each execution or the entire sequence or a defined number of images of the sequence at once. For each execution, the object classifier can output identifications of objects (e.g., pedestrians, animals, cyclists, other vehicles, the sidew alk, trees, etc.) that are depicted in the images with data regarding the objects. The output data can include the location of the objects (e.g., the location or position of the objects within the respective images and / or relative to the vehicle), the orientation of the objects, tire size of the objects, etc., within the respective images, for example. The object classifier can determine the locations or orientations of objects relative to the vehicle based on an identification of tire camera that captured the images and / or the data processing system can determine such data using a set of rules based on the camera that captured the images. Such data can provide context regarding how' objects depicted in images correspond with the location or position of the vehicle. The data processing system can label the identifications of the objects with the corresponding data for the objects with an identification of the images from which the data processing system detected the objects. The data processing system can use any type of machine learning model to detect objects in images or sequences of images.
[0131] In some cases, the object classifier can identify objects in the images that relate to operation of a vehicle. For example, objects may relate to operation of a vehicle if the objects impact a direction or speed of the vehicle. Examples of such objects can include a stopsign, a pedestrian, another vehicle, a stoplight, lane lines, a rolling tire, etc. The object classifier may be configured to identify such objects in one of a few manners. For example, the object classifier may be configured to use a table storing mappings of types of objects to indications of whether they relate to operation of a vehicle. The object classifier can classify detected objects and compare identifications of the types of the detected objects with the tabic to identify indications of whether tire objects relate to operation of the vehicle. In another example, the object classifier may be trained or configured to only identify’ objects that relate to operation of the vehicle.
[0132] In some cases, the object classifier may identify (e.g., may only identify) objects that pertain to the vehicle getting into a collision as the objects that relate to operation of the vehicle. For example, the object classifier can be trained to specifically identify’ objects that relate to vehicles getting into a collision. To do so, the data processing system can generate a training data set that includes a first set of images that were each captured within a defined time period of a vehicle getting into a collision and a second set of images that are not captured within the defined time period of a vehicle getting into a collision. The data processing system can label each of the first set of images to indicate each of the first set of images corresponds to a collision and each of the second set of images to indicate each of the second set of images does not correspond to a collision. The data processing system can train the object classifier using labels of the first set of images and the second set of images.
[0133] For tire training, the data processing system can input the individual images into the object classifier and execute the object classifier to identify one or more objects in each of the images. The data processing system can use the labels with a loss function to adjust the weights or parameters of the object classifier based on the labels indicating whether a collision occurred within a time period of capture of the images. For instance, the data processing system can adjust the weights to make it more likely to detect objects within the images labeled to indicate tire collision and less likely to detect objects within the images labeled to indicate that a collision did not occur within a time period of the images (e.g.. within a time period of the last image of the sequence of images). By training the object classifier in this way over time, the data processing may train the object classifier to only detect objects that relate to collisions. The data processing system can similarly train the object classifier to detect any type of event, all types of events, and / or a set of one or more events.
[0134] By training tire object classifier in this manner, the data processing system may train the object classifier to detect objects that are relevant to operation of the vehicle in a more nuanced manner. For example, the object classifier may be trained to detect pedestrians crossing the road in front of the vehicle while not detecting pedestrians that are walking on the sidewalk because the pedestrian may cause the vehicle to stop while the pedestrian walking on the sidewalk may not affect how tire vehicle operates. Training a machine learning model to detect such nuances may enable the object classifier to more accurately detect objects that are relevant to operation of the vehicle or to detecting a potential collision than a system that uses an object mapping approach, for example.
[0135] At operation 606, the data processing system can identify a first set of portions of the at least one image that each depict at least one of the detected one or more objects that relate to operation of the vehicle in the environment and a second set of portions of the at least one image that do not depict any of the detected one or more objects that relate to operation of the vehicle in the environment. To do so, the data processing system can divide each of the at least one image into defined portions. In doing so, the data processing system can divide or segment each of the at least one image into a grid with defined sections. The defined sections can be same shape and / or size or can have varying sizes or shapes. In some cases, the defined sections can be of a stored template. The data processing system can identify the regions or locations of the portions of each of the at least one image. The data processing system can compare tire locations of the objects depicted within images that relate to operation of a vehicle. Based on the comparison, the data processing system can identify portions that depict at least one object (or at least a threshold number of objects) that relates to operation of a vehicle. The data processing system can identify the other portions of the at least one image as not depicting any objects that relate to operation of a vehicle.
[0136] The data processing system can identify’ portions of the at least one image that depict at least one object that relates to operation of a vehicle using one or more rules. For example, for an image, the data processing system can identify tire objects that are depicted within the image that relate to operation of a vehicle. The data processing system can identify portions of the image that depict any portion or part of at least one of the identified objects as depicting at least one object that relates to operation of a vehicle. In another example, the data processing system may only determine a portion of an image depicts at least one object that relates to operation of a vehicle if at least a defined portion of such an object is depicted in the portion of the image, at least a defined portion of the portion of the image depicts one or moreobjects that relate to operation of a vehicle, the portions of the image depicting at least a defined number of objects that relate to operation of a vehicle, etc. For each image, the data processing system can identify each portion of the image that depicts at least one object that relates to operation of a vehicle and / or each portion of the image that does not depict at least one object that relates to operation of a vehicle.
[0137] For example, referring now to FIG. 6B, the data processing system may operate to compress an image 601. In doing so. the data processing system may detect objects represented by bounding boxes 628. 629, and 630 as objects that relate to operation of a vehicle. The data processing system can divide the image 601 into portions or sections 613-627. The data processing system can identify the locations of the bounding boxes 628-630 within the image 601. The data processing system may compare the locations of the bounding boxes 628- 630 with the portions 613-627 of the image 601. Based on tire comparison, tire data processing system can determine one or more, or all, of the portions 616, 617. 619, 620, 621, 622, 623, and 624 depict at least one object that relates to operation of a vehicle, such as by using the rules or techniques described herein. The data processing system can determine the remaining portions of the image 601 do not depict any objects that relate to operation of a vehicle.
[0138] Referring again to FIG. 6A, at step 608, the data processing system can convert the at least one image into a plurality of learned codewords. The data processing system can convert the at least one image into the plurality’ of learned codewords using a machine learning model (e.g.. a compression machine learning model), such as a neural network. The codewords can be or include numerical values, text, words, vectors, etc. The codewords may be learned because the machine learning model may learn the codewords during a training phase in which an encoder of the machine learning model is trained to generate the codewords for different images for compression. The data processing system can generate the codewords by inputting the images into an encoder of the machine learning model and executing the encoder to generate learned codewords representing tire details or objects depicted within the images.
[0139] In some cases, the encoder can use a defined set of codewords (e.g., learned codewords) to convert the at least one image into codewords. For example, the data processing system may store a data structure or dictionary that includes one or more codewords. The data processing system may generate the data structure or dictionary by training the encoder to generate the codewords over time and storing the generated codewords in the data structure or dictionary when the encoder learns the codewords. In some cases, the data structure ordictionary can have a defined size, so the data processing system can remove codewords (e.g.. based on age or based on the codewords being the oldest) from the data structure or dictionary as the encoder learns to generate new codewords. In some cases, the codewords can be generated separately from any training of the encoder and stored and / or updated in the data structure. The codewords can represent or correspond to different features, characteristics, objects, scenes, colors, tones, or any other features tlrat may be depicted in an image. For example, different codewords may represent different types of objects such as dogs, cats, pedestrians, stop lights, stop signs, lane lines, etc. In some cases, the codewords can represent entire separated portions (e.g., generically describe or correspond to the respective portions) of an image.
[0140] The data structure can store the codewords as separate sets of codewords. For example, the data structure can store a first set of codewords (e.g., a first codeword space) that the encoder can use to encode portions of images tlrat depict at least one (or at least a threshold number) objects that relate to operation of a vehicle. Examples of such codewords can include a pedestrian crossing the road, a stop sign, a stoplight, a lane line, etc. The data structure can store a second set of codewords (e.g., a second codeword space) that the encoder can use to encode portions of images that do not depict or contain objects that relate operation a vehicle. Examples of such codewords can include a pedestrian on a sidewalk, a bird flying overhead, a building on the side of the road, etc.
[0141] The first set of codewords can be larger or have a higher number of codewords (e.g., a higher size) than the second set of codewords. In some cases, the first set of codewords can be larger because the encoder may have been trained with a weighted loss function that causes the encoder to generate more codewords for portions of images that depict objects that relate to operation of a vehicle than for other portions of images. In some cases, the first set of codewords can be larger because the data processing system inserted more codewords into the first set of codewords than the second set of codewords. In some cases, the encoder may use a defined set of codewords that are not stored, but that the encoder is configured to generate for input images based on learned weights or parameters of the encoder.
[0142] In some cases, the data processing system can train the encoder to generate more codewords for portions of images that depict objects that relate to operation of a vehicle than other portions. For example, the data processing system can train the encoder using a training data set that includes training images and identifications of objects detected by the objectclassifier from the training images. The encoder can generate codewords based on such identification of objects and the training images. The encoder can train the encoder based on the outputs using a loss function that causes the encoder to allocate or generate more codewords for portions of the training images containing the detected object that relate to operation of a vehicle. In some cases, the encoder may not receive the identifications from the object classifier and be trained to generate more codewords for regions of images that depict objects that relate to operation of a vehicle only based on tire images without the object identifier. The encoder may be trained using a weighted loss function that causes the encoder to allocate or generate more codewords for portions of the training images containing at least one detected object that relates to operation of a vehicle. Thus, the data processing system can train the encoder to generate codewords that represent more detailed representations of images for portions of the images that are relevant for alert generation or vehicle control than other portions of the images, reducing the processing resources and / or time that are required for such processing.
[0143] The data processing system can use the encoder to convert the at least one image into codewords of the defined set of codewords. For example, the data processing system can input the at least one image into the encoder with identifications of the objects related to operation of the vehicle and / or other metadata (e.g., location data of the bounding boxes representing the objects) detected by the object classifier in each of the at least one images. The data processing can execute the encoder based on tire input. Based on the execution, the encoder can identify a first set of portions that depict at least one object that relates to operation of the vehicle and a second set of portions that do not depict at least one object that relates to operation of the vehicle. The encoder can use codewords from the first codeword space to represent each of the first set of portions of the at least one image and codewords from the second codeword space to represent each of the second set of portions of the at least one image. In doing so. because the first codeword space may be larger than the second codeword space, the encoder may use more codewords to represent the first set of portions than the second set of portions of the at least one image. Accordingly, the portions of the at least one image depict at least one object that relates to operation of the vehicle may be represented in more detail than the other portions. The data processing system can similarly convert each image of the sequence of images into codewords or sets of codewords.
[0144] In one example, the encoder can encode an image depicting a stop sign, a vehicle, and a pedestrian walking on the sidewalk. The object classifier can detect the stopsign and the vehicle as objects that relate to operation of the vehicle. The encoder can divide the image into sections or portions. The encoder can retrieve or generate one or more codewords from tire first codeword space to represent or describe the portions of the image that contain the stop sign and the vehicle detected by the object classifier. The encoder can retrieve or generate one or more codewords from the second codeword space to represent or describe the remaining portions of the image. In doing so, the encoder may retrieve or generate more codewords for tire portions of the image that depict the stop sign and the vehicle than the other portions of the image. Accordingly, when the codewords are used for further processing, there may be more data available describing or representing the stop sign and the vehicle than for other portions of the image, such as any portions that depict the pedestrian on the sidewalk. Thus, any processing for event detection or prediction can focus on the stop sign and vehicle, which can reduce the amount of data that may processed overall and / or reduce the amount of unrelated data that may cause inaccurate results.
[0145] For example, referring again to FIG. 6B. the encoder can use more codewords to represent or define each of the portions 616, 617, 619, 620, 621, 622, 623, and 624 than for portions 613, 614, 615, 618, 626, 626, and 627. Accordingly, the portions 616, 617, 619, 620, 621, 622, 623, and 624 may be represented in more detail than the portions 613, 614, 615, 618, 625, 626, and 627.
[0146] Referring again to FIG. 6A, in step 610, the data processing system can determine whether the sequence of images corresponds to an event or a collision based on the codewords generated from the sequence of images. The data processing system can do so in one of a few manners. For example, the data processing system can feed the codewords generated from each of the sequence of images into an event detection classification machine learning model (e.g., a classification decoder, such as a neural network). The event detection classification machine learning model can be a neural network, support vector machine, or random forest. The event detection classification machine learning model can be configured or trained to generate confidence scores or classifications of whether an event or collision is imminent based on codewords generated by the encoder. As described herein, learned codewords and codewords can be used interchangeably. The data processing system can execute the event detection classification machine learning model using the codewords generated by the encoder based on tire sequence of images to generate an output indicating whether an event is imminent or a confidence score indicating a likelihood that an event is imminent. In some cases, the data processing system can do so by inputting the codewordsinto the machine learning model in sequence (e.g.. in tire order of the sequence of images, such as in groups of keywords) to cause the machine learning model to process the codewords in the same sequence as the sequence of images or in parallel to cause the machine learning model to process the codewords in parallel. When input in sequence, the codewords can be ordered and grouped based on a time of capture of each of the images from which the codewords were generated. The data processing system can input the codewords into the machine learning model in any sequence. The data processing system can generate a confidence score for a detected event or otherwise detect the event based on a pattern of the sequence by applying weights to the sequence of codewords. The data processing system can determine an event is detected responsive to determining the confidence score for an event exceeds the threshold or the output indicates the event is imminent.
[0147] In some cases, tire data processing system can use images generated from the codewords to determine whether the sequence of images corresponds to an event or a collision. For example, the data processing system can input the codewords generated from each of the sequence of images into a decoder (e.g., a neural network or another type of machine learning model) configured to decode such codewords into compressed images. The data processing system can execute the decoder based on input codewords to recreate the sequence of images into compressed versions of each of the sequence of images (e.g., a sequence of compressed images). Because tire portions of the image depict objects that are relevant to operation of the vehicle correspond to more codewords, the decoded or compressed images may include more detail or have a higher resolution in such portions than other portions of the compressed images.
[0148] The event detection classification machine learning model can be configured or trained to detect or predict events or collisions based on images or compressed images (e.g., instead of or in addition to codewords). The data processing system can execute the event detection classification machine learning model using the sequence of images or the sequence of compressed images to generate an output indicating whether an event is imminent or a confidence score indicating a likelihood that an event is imminent. The data processing system can determine an event or predicted event is detected responsive to determining the confidence score for an event exceeds the threshold or the output indicates the event is imminent.
[0149] When no event is detected, the data processing system can return to step 602. The data processing system can repeat steps 602-610 any number of times over time until determining an event or predicted event is detected.
[0150] When an event (e.g.. a predicted event) is detected, at operation 612, the data processing system can generate an alert. The alert can be an audible alert that the data processing system generates through a speaker within the vehicle and / or a visual alert that the data processing system generates by activating a light or sequence of user interfaces on a display within the vehicle, for example. For example, responsive to a predicted event being detected, the data processing system can activate a light within the vehicle or generate an audible alert. The alert can indicate to the driver to slow down or move out of the way of the vehicle. In some cases, the data processing system can update a user interface displayed on a display device of the vehicle to indicate the alert, the type of the alert, and / or any data that caused the alert to be generated. In doing so, the data processing system can make the driver more aware of safety hazards or other reasons to adjust operation of the vehicle.
[0151] In some cases, the data processing system can generate instructions to control the vehicle based on the driver alert. For example, responsive to detecting the predicted event, the data processing system can control the vehicle to slow down, stop, or otherwise change direction. In doing so, the data processing system may control the vehicle to avoid or reduce the changes of the predicted event occurring (e.g., the vehicle entering into a collision). Different criteria can correspond to different control operations for the vehicle, and the data processing system can generate instructions to operate tire vehicle according to tire criteria that is satisfied.
[0152] In a non-limiting example, the data processing system can receive a sequence of images captured by a camera mounted on or within a vehicle as the vehicle approaches an intersection. The images can depict the surrounding environment of the vehicle including traffic flow, pedestrians on sidewalks, changing traffic signals, and various urban structures along the street. The data processing system can execute the object classifier on the sequence of images to detect different objects depicted in each of tire sequence of images. In doing so, the object classifier can identify a pedestrian moving toward a crosswalk against the signal, a bus changing lanes in adjacent traffic, construction barriers ahead, as well as storefronts, decorative planters, and parked vehicles in peripheral areas.
[0153] The data processing system or the object classifier can divide the detected objects into two categories, a first category containing objects that relate to operation of the vehicle and a second category containing objects that do not relate to operation of the vehicle. The first category can contain objects such as the pedestrian approaching the crosswalk, thetraffic signal, the adjacent bus, and the construction barriers. The second category can contain objects such as storefronts, decorative planters, and distant parked vehicles. The data processing system or the object classifier can assign labels to the objects indicating the categories or otherwise discard the objects that do not relate to operation of the vehicle.
[0154] The data processing system can convert each of the sequence of images into a set of codewords (e.g., learned codewords). In doing so, the data processing system can divide each image of the sequence of images into a grid. The data processing system can identify objects depicted in the images that the object identifier detected as relating to operation of the vehicle and the spaces in the grid depicting such objects. The data processing system can identify such spaces as a first set of spaces that relate to operation of the vehicle. The data processing system can identify the other spaces as a second set of spaces that do not relate to operation of the vehicle. The data processing system can execute an encoder to convert the first set of spaces into a first set of codewords from a first codeword space that corresponds to spaces that depict objects that relate to operation of the vehicle and the second set of spaces of the image into a second set of codewords from a second codeword space that does not depict objects that relate to operation of the vehicle. In doing so, the data processing system may use more codewords to represent spaces of the grid that depict the pedestrian, traffic signal, adjacent bus, and construction barriers from the images than the spaces of tire grid that depict the storefronts, decorative elements, and distant parked vehicles.
[0155] The data processing system can execute an event detection or collision detection computer model or machine learning model to detect whether a collision is imminent (e.g.. within a defined time period of two seconds). The data processing system can do so using the generated codewords representing the sequence of images or by decoding the set of codewords generated for each of the sequence of images into compressed images and using the compressed images as input. Responsive to detecting a predicted event or collision (e.g., a collision with the pedestrian) based on the execution, the data processing system can generate an alert. The alert can appear on the dashboard display accompanied and / or include an audible warning tone, and can enable the driver to initiate braking before the pedestrian actually enters the roadway.
[0156] Referring now to FIG. 7 illustrates a flow of a method 700 for implementing a generative driving model, according to an embodiment. A data processing system can perform the operation of the method 700. The data processing system can be or include a computing system of a vehicle (e g., the generative driving model system 105) and / or a remote computingsystem (e.g., a cloud server or the cloud computing system 115) for using a generative artificial intelligence model for alert generation, in accordance with an embodiment. In some cases, the data processing system can perform the method 700 to generate alerts using a sequence of images each compressed according to tire method 600, such as by using codewords generated for each of the sequence of images or compressed images decoded from such codew ords. The method 700 is shown to include steps 702-716. However, other embodiments may include additional or alternative steps, or may omit one or more steps altogether. Different steps can be performed by different computing systems (e.g.. the computing system of the vehicle can perform one or more of the steps 702-716 and / or the remote computing system can perform one or more of the steps 702-716) and / or the different computing systems and can operate together to perform individual steps of the steps 702-716.
[0157] In step 702, the data processing system can receive one or more images. The one or more images can be or include a sequence (e.g., a first sequence) of images or images of a video (e.g., video data) of an environment surrounding a vehicle or the environment within the vehicle. The sequence of images can be consecutively captured images by the same camera. The images of the sequence can be captured at a set time interval (e g., every five seconds) or pseudo-randomly. The images of the sequence can be captured within a time period (e.g., within a two second time period, a five second time period, a 10 second time period, etc.). The time period can be a first time period. The time period can begin with the time the first image of the first sequence was captured and end with the time the last image of the sequence was captured, in some cases. The images can be images of a road on which the vehicle is driving or of areas adjacent to the road. The data processing system can receive the images from a camera mounted to or within the vehicle (e g., mounted on the dashboard of the vehicle). The images can be or include a JPEG, RAW, or another image file type. The visual data can be or include one or more standalone images or one or more frames of a video. The data processing system can receive the images as a driver is driving the vehicle, when the driver has parked the vehicle, and / or when the vehicle is in a standby mode. The data processing system can receive such images at defined intervals in the case of static standalone images or as the camera streams a video to the data processing system in individual frames.
[0158] In step 704, the data processing system can execute an encoder of a machine learning model. The machine learning model can be an encoder of a language model, such as a large language model, that is configured or trained to generate intermediary states based on inputs. The intermediary states can be hidden states of the machine learning model. Theintermediary state can be or include one or more values, a numerical vector, a text string, an embedding, etc. The intermediary states can be, include, or represent one or more (e.g., a plurality of candidate responses to an input. In some cases, the intermediary states can each include or correspond to probabilities or confidence scores that the candidate responses represented by the intermediary states arc correct. For example, the encoder can receive the sequence of images received in step 702 and convert the sequence of images into an intermediary state (e.g., a first intermediary state) that can be, include, or represent one or more future states of the environment surrounding the vehicle or within the vehicle at one or more time steps subsequent to the time period (e.g., first time period) in which the sequence of images were captured.
[0159] In some cases, a future state can indicate what the environment would look like in a new image. For example, the machine learning model can output a prediction of characteristics of the one or more objects from the sequence of images at a time step (e.g., time) subsequent to the last image of tire sequence of images. In doing so, the machine learning model can output the future state of the objects and / or the environment surrounding and / or within the vehicle a defined time (e.g., a defined time step) after the last image of the sequence of images. For instance, the future state may depict or indicate positions of objects in an image one second subsequent to the last image of tire sequence of images. In some cases, the machine learning model can output an image depicting tire objects and / or environment in the future state.
[0160] For example, the data processing system can receive the sequence of images depicting a vehicle in front of the vehicle housing the data processing system, a vehicle in the lane to the right of the vehicle housing the data processing system, and an open lane to the left of the vehicle housing the data processing system. The data processing system can perform the method 600 to generate compressed versions of each image of the sequence. The data processing system can input the sequence of compressed images into the encoder and execute the encoder. The execution can cause the encoder to output an intermediary state indicating a plurality of candidate future states with various updated or predicted locations of the two vehicles relative to the vehicle housing the data processing system. For instance, in one candidate future state the vehicle in front of the vehicle housing the data processing system can be closer to the vehicle housing the data processing system than in the last image of the sequence of images, and the same vehicle can be behind the vehicle housing the data processing system in another candidate future state. The intermediary state can represent any number ofcandidate future states with the vehicle in any position. Each candidate future state can indicate the change in state of any number of objects as the change in state of the objects would be depicted in an image of the future state of the objects.
[0161] The data processing system can input codewords generated from the sequence of images generated as described with respect to the method 600 into tire machine learning model. The codewords can represent different features of each of the sequence of images or portions (e.g.. defined portions, such as portions of a grid) of each of the sequence of images. The data processing system can input the codewords into the encoder in a feature vector. The feature vector can be a prompt into the machine learning model (e.g., large language model or machine learning language processing model). The data processing system can generate the feature vector by combining or concatenating the codewords in sequence, such as in the same sequence as the sequence of images. In doing so, the data processing system can group the codewords by tire image from which the objects were detected, in some cases. The data processing system can include codewords of the most recently captured image first and add the codewords generated for other images in reverse order accordingly, or vice versa (e.g., include codewords generated for the most recently captured image last). The data processing system can similarly combine any number of codewords into the feature vector to use as input into the encoder.
[0162] In some cases, the data processing system can include other data in the feature vector that is input into the encoder. For example, the data processing system can include sensor data regarding the environment surrounding the vehicle in tire feature vector. Examples of such sensor data can include an outside air temperature, an amount of moisture in the air, a wind reading and / or direction, etc. The data processing system can additionally or instead include data about the vehicle collected by one or more sensors of the vehicle, such as the tire pressure of the vehicle, the current speed of the vehicle, the acceleration of the vehicle, gyroscopic data of tire vehicle, etc., in the feature vector. In some cases, the data processing system can tokenize such data into tokens using a quantizer (e.g., a minimum RMS quantizer) and input the tokens into the feature vector. In some cases, the data processing system can additionally or instead include data regarding the occupants of the vehicle, such as indications of whether the driver is paying attention to the road (which may be determined using machine learning techniques, for example). In doing so, the data processing system can include data that was collected at the same time or within a defined time period of the times in which the respective images were captured (e.g., based on timestamps of the data or images). Such datacan provide the encoder with further context regarding changes in the objects depicted in the sequence of images to improve the accuracy of the candidate future states of the intermediary state that the encoder generates. The data processing system can generate the feature vector with the sequence of images and / or any data regarding the state of the vehicle or the environment surrounding the vehicle to generate tire second plurality of tokens representing the future state of the environment surrounding tire vehicle and / or within the vehicle.
[0163] In some cases, the data processing system can include an identification of a vehicle type in the feature vector. The vehicle type can be an indication of a type of the vehicle for which the data processing system is generating the feature vector (e.g., the type of vehicle that is housing the data processing system), such as a sedan, a pick-up truck, a truck, etc. The encoder may be trained to take into account tire type of vehicle when generating an intermediary state indicating or representing a plurality of candidate future states of objects in the environment surrounding tire vehicle or inside the vehicle because different types of vehicles may have different turn radiuses or acceleration or breaking capacities. Such differences may impact locations of objects external to the vehicle relative to the vehicle over a defined time period. The data processing system may be configured to input the identification of the type of vehicle into the feature vector to use as input into the machine learning model for execution.
[0164] In some cases, the data processing system can include an identification of a driver type in the feature vector. The driver type can be an indication of a type of the driver driving the vehicle for which the data processing system is generating the feature vector. The type of driver can be, for example, an aggressive driver, a slow driver, a fast driver, a careful driver, etc. Each type can correspond to a different identification. The machine learning model may be trained to take into account the type of driver driving the vehicle w hen generating an intermediary state indicating or representing a plurality’ of candidate future states of objects in the environment surrounding the vehicle or inside the vehicle because different types of drivers may navigate obstacles or situations differently that may impact the locations of external objects relative to the vehicle over time. The data processing system may be configured to input the identification of the type of driver into the feature vector to use as input into the machine learning model for execution.
[0165] In step 706, the data processing system can execute a future-state decoder of the machine learning model (e g., the large language model). The future-state decoder can betrained or configured to generate future states (e.g., images, vectors, tokens, or other representations) of environments based on intermediary states generated from sequences of images and / or other types of data. The data processing system can execute the future-state decoder based on or using as input the intermediary state generated by the encoder. For example, the data processing system can input the intermediary state including or representing one or more candidate future states of the environment surrounding the vehicle or within the vehicle at one or more time steps subsequent to the time period in which the sequence of images was captured. The data processing system can execute the future-state decoder based on the input. The execution can cause the decoder to generate a future state from the intermediary state indicating or representing the first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at a time step subsequent to the first time period in which the first sequence of images was captured. The future state can be or include one or more vector representations of the future state generated based on the first plurality- of candidate future states, one or more images of such a future state, one or more tokens or values representing the future state, etc.
[0166] In some cases, the intermediary state can include confidence scores for each of the plurality of candidate future states. The future-state decoder can identity’ the confidence scores and select the candidate future state that corresponds to the highest confidence score of the identified confidence scores and select or generate a representation of the selected candidate future state as an output. In some cases, the future-state decoder can use the confidence scores and / or representations of the candidate future state to generate a future state. For example, the future-state decoder can weight the candidate future states according to the confidence scores and use the weights to generate a future state that is more similar to the candidate future states that are more highly weighted than other future states. The generated future state may not be an exact match with any of the candidate future states.
[0167] In some cases, the future-state decoder may be configured to generate a future state of the environment surrounding the vehicle at a time step subsequent the time period of the sequence of images (e.g., a defined time after a timestamp of the last image of tire sequence of images). The future-state decoder may be trained to do so, for example, using a loss function and backpropagation techniques with labeled or ground truth images each captured at a particular time step after the last images of sequences of training images. The data processing system can feed the sequences of training images into the encoder to generate intermediary states for the sequences of training images. The data processing system may feed theintermediary states into the future-state decoder to cause the future-state decoder to generate a prediction of a future state of the environments surrounding the vehicles of the future-state decoder for each intermediary state. The data processing system can compare the predictions with the ground truth images or representations of tire environment surrounding the vehicle at the particular time step and adjust the weights and / or parameters of the encoder and / or the future-state decoder based on the differences, such as by using a loss function. The data processing system can train the future-state decoder and / or encoder in this manner over time such that tire future-state decoder can be trained to generate predicted future states of environments from sequences of images at the particular time step. The data processing system can perform this training for any time step (e.g., .5 seconds, one second, two seconds, three seconds, etc., after the last image of sequences of images). In some cases, tire data processing system can use a similar training process to train the future-state decoder to generate future states for multiple time steps subsequent the last images of sequences of images (e.g., by using multiple images at defined time steps as labels).
[0168] In step 708, the data processing system can train the encoder. The data processing system can train the encoder using a loss function based on the generated future state of tire environment surrounding or within the vehicle generated by the decoder. For example, the data processing system can train the encoder to generate intermediary states representing a plurality of candidate future states of objects and / or environments in real time.
[0169] For example, tire data processing system can feed codewords generated from a sequence of images, the sequence of images as compressed (e.g., decoded from the codewords), or the sequence of images into the encoder to generate an intermediary state indicating a plurality of candidate future states of the objects. The data processing system can capture another image (e.g., a label image) at the time for which the machine learning model is configured to generate the future state of objects relative to the final image in the sequence of images. The data processing system can generate codewords from the captured image or otherwise and compare the generated codewords or the captured image to the future state of the objects generated by the machine learning model. The data processing system can use a loss function to adjust the weights and / or parameters of the encoder and / or the decoder based on any differences between the captured image and the future state of the objects. The data processing system can repeat this process over time (e.g., at set intervals or pseudo-randomly) to train the encoder and decoder of the machine learning model to operate to predict future states of objects in the environment surrounding or within the vehicle. The data processingsystem can train the encoder and / or decoder in this manner using real-time data (e.g.. images captured while a driver is driving in the normal course) or in a static training period, such as with curated training data.
[0170] The data processing system can train the encoder and / or decoder over time. For each training iteration, the data processing system can determine an accuracy of the output, such as based on a difference between the predicted future state and the actual future state or ground truth future state. The data processing system can compare each detennined accuracy to a threshold (e.g., an accuracy threshold or a defined accuracy threshold). The data processing system can determine when the accuracy of a prediction or at least a defined percentage or number of accuracies exceeds or otherwise satisfies the threshold or satisfy any other criteria.
[0171] Subsequent to training the encoder (e.g., subsequent to determining the accuracy of the machine learning model or the encoder exceeds or satisfies the threshold or satisfies another criterion), in step 710, the data processing system can receive a second sequence of images of the environment surrounding the vehicle or within the vehicle captured by the camera mounted to or in the vehicle. The data processing system can receive the second sequence of images in tire same manner as tire data processing system received the first sequence of images at step 702.
[0172] In step 712, tire data processing system can execute the trained encoder of the machine learning model using the second sequence of images to generate a second intermediary state. The data processing system can execute the trained encoder using the second sequence of images by using the second sequence of images directly as input, compressing each of the second sequence of images according to the method 600 and inputting the compressed images into the decoder, and / or inputting codewords generated from each of die second sequence of images according to the method 600. In some cases, the data processing system can additionally or instead include values generated from sensors of the vehicle, such as speed values or gyroscopic values, in the input to tire encoder with the second sequence of images. The data processing system can execute the encoder based on such inputs to generate the second intermediary state. The second intermediary state can indicate a second plurality of candidate future states of the second environment surrounding the vehicle or within the vehicle at one or more second time steps subsequent to a time period (e.g., a second time period) in which tire second sequence of images were captured. The time period can begin at the time inwhich the first image of the second sequence of images was captured and end at the time in which the last image of the second sequence of images was captured.
[0173] In step 714, the data processing system can determine whether the second sequence of images corresponds to an event or a collision based on the second intermediary state generated by the encoder from tire second sequence of images. The data processing system can do so in one of a few manners. For example, the data processing system can feed the second intermediary state into an event detection classification machine learning model (e.g., a classification decoder of the machine learning model of the encoder and the future-state decoder). The event detection classification machine learning model can be a neural network, support vector machine, a random forest, or any other type of machine learning model or computer model. The event detection classification machine learning model can be configured or trained to generate confidence scores or classifications of whether an event or collision is imminent based on intermediary states generated by the encoder. The data processing system can execute the event detection classification machine learning model using the second intermediary state generated by the encoder based on the second sequence of images to generate an output indicating whether an event is imminent or a confidence score indicating a likelihood that an event is imminent. The data processing system can generate a confidence score for a detected event or otherwise detect the event by applying weights to the sequence of codewords. The data processing system can determine a predicted event is detected responsive to determining the confidence score for a predicted event (e.g., a collision) exceeds the threshold or the output indicates the event is imminent.
[0174] In some cases, the data processing system may train the event detection machine learning model and / or the encoder based on outputs of the event detection machine learning model. The data processing system can do so after training the encoder based on outputs of the future-state decoder, for example. The data processing system can do so using labeled training data with back-propagation techniques. The labeled training data may include sequences of images and labels indicating whether the sequences of images correspond to an event or collision. The labels may be generated based on real-world data indicating whether the sequences of images were captured within a time period of an event or collision. In some cases, tire labeled training data may be labeled based on whether a collision occurred at the same time step for which the future-state decoder is configured to predict the future state or w ithin the time period betw een the time of the last image of a sequence of images and the time step. The data processing system may execute the encoder and event detection machinelearning model based on the sequences of images, such as by generating intermediary embeddings from the sequences of images (e.g., from the sequences of images themselves, from compressed versions of the sequences of images, or based on codewords of the sequences of images) using the encoder. The data processing system can execute the event detection machine learning model using the intermediary states as input to generate outputs each indicating whether an event or a collision is imminent or will occur.
[0175] In some cases, subsequent to training the encoder (e.g., to an accuracy above the threshold), tire data processing system can activate a setting of the machine learning model to stop using the future-state decoder to predict the future state. For instance, the data processing system or the machine learning model may have a setting that causes intermediary states generated by the encoder to be automatically input into the decoder to generate an output. The data processing system can change the setting such that the output stops being input into the decoder and is instead input into the event detection machine learning model for event detection. In doing so, the data processing system can avoid the processing by the decoder to reduce the computational resources and / or time incurred in predicting the future states by the decoder and instead directly perform predicted event detection based on the intermediary state generated by the encoder.
[0176] The data processing system can train the event detection machine learning model and / or the encoder using a loss function based on differences between the outputs generated from the sequences of images and the labels. The data processing system can repeat this process over time to train the event detection machine learning mode to detect events or collisions based on outputs from the encoder. This can be important because the weights and / or parameters may change during the training of the encoder based on outputs of the future-state decoder, so the event detection machine learning model may need to be trained to process outputs generated by the encoder according to such trained weights, which may not be possible until the encoder is trained to the accuracy threshold or otherwise done training based on outputs of the future-state decoder. The data processing system can additionally train the encoder to fine-tune or refine the weights and / or parameters to the configuration or weights and / or parameters of the event detection machine learning model.
[0177] Responsive to determining no event was detected from the second sequence of images or the second intermediary state, the data processing system can return to step 710. Thedata processing system can repeat steps 710-714 any number of times over time until determining a predicted event is detected.
[0178] Responsive to determining a predicted event is detected, at operation 716, the data processing system can generate an alert. The data processing system can do so in the same manner as described with reference to operation 612 of the method 600. The alert can be an audible alert that the data processing system generates through a speaker within the vehicle and / or a visual alert that the data processing system generates by activating a light or sequence of user interfaces on a display within the vehicle, for example. For example, responsive to a predicted event being detected, the data processing system can activate a light within the vehicle or generate an audible alert. The alert can indicate to the driver to slow down or move out of the way of tire vehicle. In some cases, the data processing system can update a user interface displayed on a display device of the vehicle to indicate the alert, the type of the alert, and / or any data that caused the alert to be generated. In doing so, the data processing system can make the driver more aware of safety hazards or other reasons to adjust operation of the vehicle.
[0179] In some cases, the data processing system can generate instructions to control the vehicle based on tire driver alert. For example, responsive to detecting the predicted event, the data processing system can control tire vehicle to slow down, stop, or otherwise change direction. In doing so, tire data processing system may control the vehicle to avoid or reduce the changes of the predicted event occurring (e.g., the vehicle entering into a collision). Different criteria can correspond to different control operations for the vehicle, and the data processing system can generate instructions to operate the vehicle according to the criteria that is satisfied.
[0180] In a non-limiting example, tire data processing system can receive a first sequence of images captured by a dashboard-mounted camera in a vehicle. The first sequence of images can depict the roadway ahead of the vehicle with various objects including other vehicles, pedestrians, traffic signals, and changing road conditions as the vehicle navigates through a busy city street during daytime hours. Upon receiving the images, the data processing system can compress the images of the sequence into a sequence of compressed images.
[0181] The data processing system can execute an encoder (e.g., an encoder of a large language model) using tire sequence of compressed images and / or other sensor data to generatea first intermediary state containing a vector representation of multiple candidate future states of the environment surrounding the vehicle or objects depicted in the environment surrounding the vehicle, such as pedestrians crossing streets, vehicles ahead braking, traffic lights changing from green to yellow or from yellow to red, or road conditions potentially changing due to approaching weather events. The data processing system can then execute a future-state decoder based on this intermediary state to generate a future state or a prediction of the future state of the environment or the objects within the environment 1 second into the future. In doing so, the future-state decoder can generate a future state predicting that a pedestrian waiting at a crosswalk will begin crossing in 1 second with 87% confidence.
[0182] The data processing system can train the encoder based on the prediction by the decoder. The data processing system can do so using a loss function that compares the predicted future state against an actual observed outcome. For instance, the future-state decoder may predict a pedestrian would cross but the pedestrian may remain on the sidewalk, the data processing system can adjust the encoder weights to reduce similar false predictions in future iterations. This training process can utilize supervised learning techniques where ground truth data can be collected from numerous recorded driving scenarios across various environments and conditions.
[0183] After completing the training phase, the data processing system can receive a second sequence of images from the same dashboard camera as the vehicle navigates through a different environment, such as a school zone with individuals playing near the street and a ball rolling toward the road. The data processing system can execute the trained encoder using the second sequence of images to generate a second intermediary state that can indicate potential future scenarios, such as predicting that an individual can chase after the ball and enter the roadway one second after the last image of tire second sequence of images.
[0184] The data processing system can then execute a classification machine learning model or decoder using the second intermediary state to detect an imminent collision or predicted event with the child. Responsive to this prediction, the data processing system can generate both audible and visual alerts within the vehicle, warning the driver to slow down immediately and / or initiate a braking or stopping to avoid the collision.
[0185] Accordingly, by implementing the systems and methods described herein, the data processing system can enable the autonomous driving system or alert generation system to anticipate potentially dangerous situations before they fully develop, which can provideextra seconds for preventative action or alert generation that can significantly enhance vehicle safety. In doing so, the data processing system can additionally reduce the processing resources required for such predictions compared with systems that generate predicted states of the environment and detect events based on the predicted states of the environment.Example Embodiments
[0186] Certain aspects of tire present disclosure generally relate to generative Al models, and more particularly to systems and methods of improving driving safety using a generative Al model that is adapted to driving data. Certain aspects of the present disclosure generally relate to providing, implementing, and using a generative Al model in the context of driving. The operations or steps of the discussion below can be performed by a data processing system. The data processing system can be or include a computing system of a vehicle (e.g., tire generative driving model system 105) and / or a remote computing system (e.g., a cloud server or the cloud computing system 115) for using a generative artificial intelligence model for alert generation, in accordance with an embodiment.
[0187] Generative artificial intelligence (Al) models are a class of Al models that includes large language models (LLMs). At present, various challenges, including challenges relating to the training or execution of generative Al models, make it difficult to apply in the context of driving. Certain aspects arc directed to overcoming challenges and limitations of generative Al models so that they may be used effectively in driving applications.
[0188] For drivers today, a foundational / generative driving model may enable a new generation of advanced driver assistance systems (ADAS) features. These include generalized collision warning, which may include functions that can detect and predict pedestrian and driver intent. It may also enable generalized unsafe driving warnings. These may include warnings of running lights or stop signs, being cut off in traffic, and the like.
[0189] For autonomous vehicles, there may be a long tail data challenge. A generative driving model, however, may be a way to effectively utilize billions of miles of observable traffic scenarios, both to identify rare but critical events, and to plan for them. Accordingly, a foundational model may enable a form of path planning that incorporates a generalized understanding of driving vocabulary and semantics.
[0190] The development of a generative driving model may be similar to development of large language models (LLMs). LLMs systems are typically trained on a large corpus of textdata using self-supervised training on next token prediction. This training approach yields models that exhibit emergent capabilities and generalization, even for tasks they were not explicitly trained for.
[0191] In an embodiment of certain aspects of the present disclosure, a similar approach can be applied to train a generative driving model on billions of miles of real-world driving data. As explained below, tire model itself may only be exposed to a curated sample that is based on processing of billions of miles of real-world driving data. This vast amount of data, along with a similar self-supervised training strategy, allows the model to learn and generalize from diverse driving scenarios, conditions, and nuances.
[0192] The training data for such a model can be multi-modal, incorporating a variety of data types such as video data, Inertial Measurement Unit (IMU) data, Global Positioning System (GPS) data, vehicle-specific data, and Al event detections. This rich, varied input data provides a comprehensive foundation for the model, enabling it to understand and predict driving scenarios and context-awareness.
[0193] For certain embodiments, the generative driving model can be used to control the Ego vehicle. Integrating the model with the vehicle's control system allows it to respond in real-time to the model's predictions and decisions. This feature enables the model to transition from being predictive to actively participating in the driving process.
[0194] The training of such a model can be facilitated by using, for example, NVIDIA A100 GPUs, leveraging tire NCCL (NVIDIA Collective Communications Library) for distributed training, or other suitable computational substrates and software. Hardware and software infrastructure that has been designed to support the training and execution of large language models (LLMs) can support efficient and effective training of the generative driving model and may make practical training of such models on extensive and diverse datasets, thereby enabling robust learning and high performance of the generative driving model. The model's training is facilitated by using advanced GPUs and distributed training software. This infrastructure supports the efficient training of the model on massive and diverse datasets, ensuring its robust learning and high performance. The model is not just capable of understanding and predicting driving scenarios but also efficiently learns from the data it is trained on.
[0195] Acquiring billions of miles of real-world driving data is a significant challenge. However, this can be achieved by leveraging an existing driver safety platform that specialize in capturing and analyzing driving data to improve safer driving.
[0196] In such systems, rich, high-definition driving data is analyzed and categorized, providing a visual context-based understanding of various driving scenarios. This data includes driving scenarios in all weather conditions, road types, localities, and vehicle categories.
[0197] The dataset also encompasses a wide range of situations and events such as accidents, near-miss incidents, construction zones, encounters with pedestrians and bicyclists, traffic light and stop sign interactions, lane changes, and the like.
[0198] This extensive dataset may be considered to surpass what is typically available in tire autonomous vehicle (AV) industry, which often has limited mileage data, less than 50 million miles in total across companies, and is usually restricted to a handful of cities. Thus, by leveraging comprehensive datasets, a generative driving model can be trained on a scale and diversity of real-world driving scenarios that addresses the long-tail risks that continue to impede progress on Autonomous vehicles.Evidence of Emergent Understanding
[0199] FIG. 8 illustrates the operation of a generative driving model, which may be referred to as a driving prediction model, in accordance with certain aspects of the present disclosure. The model accepts as input a driving context, illustrated in FIG. 8 as a series of three images captured by a road-facing camera that is attached to the windshield of a vehicle being driven through an urban environment. Based on the context, the driving prediction model predicts three different possible future states, each future state having a corresponding probability. The most likely future state in this illustration is assigned a probability of 0.5. The image of the future state shows a dark car in front of the vehicle having the camera. This image is one of the outputs of the driving prediction model. Alternatively, the image may be a refinement of an output of tire driving prediction model, where the refinement involves a second neural network that produces a higher resolution image based on a lower 2D image of the model. The refinement may also be a neural network that transforms the output of the Driving Prediction model into a 3D simulator representation of the driving scene. The 3D simulator may then produce a 2-D representation from a point-of-view corresponding to thevehicle-mounted camera. The next most likely future states depicted in this image are shown as having a probability of 0.3 and 0.2, respectively.
[0200] While the three possible future outcomes are illustrated in FIG. 8 are illustrated as each corresponding to a single image, the various future outcomes may also correspond to tire likelihood of various scenarios. For example, FIG. 8 might instead be understood to mean that future driving scenarios corresponding to scenario A account for 50% of the predicted scenarios. Scenario A may be that there is a dark colored car in the same lane as the vehicle having the camera and there is a second light colored car that is nearer to the vehicle and in the adjacent lane to tire left. Similarly, scenario B, which may account for 30% of the predicted scenarios, may correspond to a variety of predicted scenarios for which the light-colored car has entered the vehicle’s lane from the left so that it puts itself between the vehicle and the dark colored vehicle.
[0201] FIGS. 9A-G depict a scenario that demonstrates the generative driving model's abil ity to imagine multiple potential outcomes. Specifically, these figures present a situation where a detected vehicle either crosses into the driver's lane or does not.
[0202] Each of FIGS. 9A-G represents a different "imagined future" that the model generates based on its analysis of the current scenario. The model's ability to predict multiple potential outcomes is useful for safe and effective autonomous driving, as it allows the system to anticipate and plan for various situations.
[0203] The visualizations of these outputs provide insight into the model's emergent understanding of the road environment. The model identifies cars as distinct objects and predicts possible lane changes, showing that it understands these fundamental aspects of driving.
[0204] Moreover, by predicting whether or not the detected vehicle will cross into the driver's lane, the model demonstrates an understanding of driving rules and behaviors. This ability to predict and understand driving rules is a critical step towards developing a safe and efficient autonomous driving system.
[0205] In greater detail, FIG. 9A illustrates four panels. The top left panel depicts an image taken from a vehicle-mounted camera that is processed by a vision tokenizer, described below, which includes an encoder and decoder. The result of the encoder-decoder may be lower resolution or in some ways different from the original image, even though the purpose of theencoder-decoder may be to preserve the contents of the original image. In each of the subsequent FIGS. 9B-9G, the top left panel depicts an encoded-decoded image that was captured by a vehicle-mounted camera.
[0206] In FIG. 9A, the other three panels are also the result of the encoder-decoder tokenizer of the image that was captured by the vehicle-mounted camera, but the subsequent FIGS. 9B-9G, the top right and bottom panels depict the output of the driving prediction model. The images are encoded-decoded representations of the original in FIG. 9A because the driving prediction model may use a series of images as context before it can make useful predictions of the next driving frame. In this example illustrated, the driving behavior model operates on 9 images of visual context. FIG. 9B illustrates 5 images frames after FIG. 9A. In FIG. 9B, all of the panels are similar to each other, but the images all show that the white car in the adjacent right lane has moved farther ahead of tire vehicle that has the camera that captured the original images.
[0207] FIG. 9C depicts 5 frames further still. At this point, the driving prediction model has enough context to start making useful predictions. In this example, the driving prediction model was instantiated three different times, each model taking as inputs 10 images corresponding to the top left panel, and then making a prediction the image that will be captured in tire next frame. At the moment depicted in FIG. 9C all of the predictions are similar to each other, although not exactly the same.
[0208] FIG. 9D depicts 10 frames further ahead, at frame 20. The first 10 frames were used as context to predict the images depicted in FIG. 9C. The predicted frame was then appended to the context frames and 9 frames plus 1 predicted frame were used to predict the next image frame after that. This process is repeated until the entire 10 frame context is comprised of predicted frames. As can be seen in FIG. 9D, the predicted futures start to diverge. The white car in the top right panel is noticeably closer to the lane line than it is in the other panels.
[0209] FIG. 9E depicts 10 frames further ahead, at frame 30. In the predicted future of the driving prediction model corresponding to the top right, tire white car is crossing into the vehicle lane. The white car in the true image frame, top right, is still in the adjacent lane. The predicted futures depicted in the bottom two panels also show the white car still in the adjacent lane. In this example, the predicted future in the top right is a driving scenario that could plausibly happen, even though it did not actually happen.
[0210] FIG. 9F depicts 10 frames further ahead, at frame 40. At this point, the white car in the imagined future in the top right has completed the lane change and is drifting further to the left. A second car appears in the lane that was formerly occupied by the white car, suggesting that the driving prediction model predicted a lane change by the white car to get around a slower moving vehicle in its lane. The two bottom panels have finally started to diverge from each other. The white car in the bottom left panel is farther ahead of the vehicle having the camera than it is in the bottom right panel.
[0211] FIG. 9G depicts 19 frames further ahead, at frame 59. The bottom panels have continued to diverge. The imagined future in the bottom right panel corresponds to the vehicle having the camera passing the white car. In the bottom left panel, the vehicle having the camera is maintaining a similar distance. This imagined future also corresponds to the true image scene that is depicted in tire top left, although the details of the contents on the side of the road have diverged from the actual scene.
[0212] Taken together, FIGS. 9A-G illustrate how a driving prediction model, in accordance with certain aspects of the present disclosure, can predict or imagine a variety of ways that a future driving scene could unfold. This capability may be useful in a number of ways, including accident prediction, which could enable generalized collision warnings or avoidance, etc.Evidence of Generalization
[0213] FIGS. 10A-D aim to demonstrate the generative driving model's capability to generalize and learn in-context, even when exposed to unfamiliar environments. These figures show the model's performance when applied to driving scenarios in India, despite the model being trained on US data.
[0214] The top left panel in each of FIGS. 10A-D displays the original video, which captures the real-world driving scenario in India. This sen es as the input data for the model.
[0215] The top right panel in each of FIGS. 10A-D presents the ground truth video through a tokenizer. This is essentially a processed version of the original video, serving as a reference for what the model should ideally predict or generate.
[0216] The bottom left panel in each of FIGS. 10A-D and the bottom right panel in each of FIGS. 10A-D both show the foundational model's output, represented in red. These figures illustrate how the model interprets and responds to the original video scenario.
[0217] Despite not being trained on India-specific data, tire foundational driving model demonstrates impressive generalization capabilities. It successfully generates outputs that are consistent with the ground truth, indicating its emergent understanding of the road environment and the objects within it.
[0218] For instance, the model recognizes an auto-rickshaw as a vehicle object, even though it has never encountered an auto-rickshaw in the US training data. This indicates that the model has effectively learned tire concept of a "vehicle" and can apply it to unfamiliar objects in different environments.
[0219] However, it's also observed that the video tokenizer, which processes the original video into a format the model can understand, has room for improved fidelity. Improving the tokenizer could enhance the accuracy and detail of the model's predictions and responses.Accident Prediction
[0220] A generative driving model may be used to predict if an accident may occur with high probability in the near future. For example, the generative driving model may be used to determine if the likelihood of a collision within the next two seconds exceeds a predetermined threshold.
[0221] As explained below, the generative driving model is trained to predict future outcomes given some context. The context may include multiple image frames captured prior to the present moment by a road-facing camera. The context may also include representation of other sensor data, the vehicle class, and the like. The generative driving model may process this context data and produce a prediction one time step in the future. The generative model may also predict a distribution of outcomes, which may each have an associated probability. The most likely prediction that is one step in the future may then be fed back into the generative driving model again and a prediction that is two steps into the future may be produced.
[0222] To fully explore future possibilities, a common approach might be to unroll all different future possibilities and check if any of them lead to an accident. This would involve generating a distribution of next-time step predictions, then feeding each prediction back into the generative model and proceeding. However, this method is computationally expensive, and it might not be feasible to explore all possibilities in this way.
[0223] An alternative, potentially more efficient approach could involve using a beam search. A beam search is a heuristic search algorithm that explores a graph (unrolled future possibilities) by expanding the most promising nodes. The beam search starts by generating all possible next steps from the current position (the present time). In the context of a generative driving model, all possible next steps could refer to a predetermined number of alternative predictions. Then, instead of exploring all the alternative predictions, the algorithm only keeps the top 'b' possibilities, where 'b' is a parameter known as the 'beam width1. These are selected because they are the 'b' most promising or likely possibilities, according to the heuristic measure. The heuristic measure might be a likelihood generated by the driving model itself, based on clustering of tire generated possible next steps, or based on an analysis of the generated driving scenes, among other possibilities. This process is repeated for each of the 'b' possibilities, generating a tree of possibilities. The search continues until it has found a solution or until it has explored a predetermined number of levels in the tree.
[0224] In the example of using the generative driving model to predict if an accident may occur, the search may continue until it has explored the most likely scenarios stretching two seconds into the future. If none of the likely scenarios lead to a collision, then the process may end without any notification to the driver. If, however, one or more explored scenarios leads to a predicted accident, the system may generate a notification to the driver, may assist tire driver by, for example, priming the brakes, may take control of tire vehicle, and the like. When the collision is predicted, the process may also stop earlier (before a two second future horizon has been explored) so that the system can make a timelier notification to the driver, assist the driver, and the like.
[0225] For the generative driving model, the hidden state (e.g., an intermediary state, as described herein) of the model may implicitly contain the probabilities of various predicted futures, including predicted futures that arc more than only one time step in the future. This may be expected from a consideration of LLMs that predict one alphanumeric character at a time. Suppose a Language Learning Model (LLM) has generated the string "accid." Given the context, it's highly probable that the next three characters will be 'e', 'n'. 'f to form "accident" or continue to words like "accidental" or "accidentally." The LLM's internal hidden state likely includes a representation that corresponds to these potential word completions. In a typical operation of an LLM as a generative language model, the LLM would generate the letter "e" after processing tire context "accid," update tire context to "accide," and then predict the next letters ('n', 't') one at a time. However, instead of running the LLM multiple times, a classifiertrained to make inferences from tire hidden state of the LLM could potentially determine the next several letters ('e', 'n', 't') in one step based on the hidden state of the LLM after it processed the context "accid." This approach could avoid the computationally expensive operation of repeatedly running the LLM, providing a more efficient prediction mechanism, especially in scenarios where rapid accident prediction is critical.
[0226] In a well-trained generative driving model, the hidden state is anticipated to implicitly represent the most probable future states. If these expected future states are highly certain, alternative future states might not be represented at all. However, in ambiguous situations, such as when a vehicle slows down at a four-way intersection, the model's hidden state may simultaneously represent multiple distinct future states.
[0227] Since the hidden state representation is expected to change in a way that predicts the likelihood of various driving futures, a separate classifier can be trained to process the hidden state of the model and classify these outcomes. Hence, according to certain aspects of the present invention, a classifier can be trained on the hidden state of the generative driving model to forecast whether an accident might occur in the near future.
[0228] For instance, the model could be trained to classify hidden state representations of a variety of driving scenarios, some of which result in a collision shortly after the given context. This method enables the system to make quick and accurate accident predictions, enhancing the safety and efficiency of autonomous driving systems.
[0229] This classifier model processing may be more efficient and / or faster to compute than an approach of expanding future driving scenarios and performing a beam search. In some cases, it may even be faster than expanding only the most likely future driving scenario. By using the classifier model, therefore, the processing associated with unrolling different possible futures may be avoided.
[0230] As previously mentioned, when the classifier predicts that the likelihood of a collision exceeds a certain threshold, the system may react in a variety of ways. This could include generating a notification for the driver, assisting the driver by priming the brakes, or even taking control of the vehicle, among other actions.Detection-guided vector quantization (VQ) for driving images
[0231] Certain aspects of the present disclosure relate to a method of employing vector quantization (VQ) guided by object detection and classification for images collected by a camera in a vehicle, which may be applied to the assessment of risk in vehicular environments.
[0232] Image compression and representation plays a role in many artificial intelligence (Al) applications, especially those involving computer vision. Vector Quantization (VQ) is a technique that has been used for image compression, where an image is divided into non-overlapping blocks or "patches." In VQ, each patch may be represented by a codeword from a pre-defined codebook or vocabulary.
[0233] In vehicular applications, such as autonomous driving, Al perception engines may be used to detect and classify objects in the vehicle's environment and may further process images and these detections to driver safety metrics, collision risk, and the like. According to certain aspects, the output of these Al perception engines can be incorporated into the design, training, and / or functioning of the VQ module. Doing so can lead to more precise representations of information about the objects of interest in the image, or at least the detected objects that may have an impact on safety and risk in the context. Accordingly, certain aspects disclosed herein address a need for a more effective method of image representation that can preserve the details of objects of interest in the image that are relevant to a driving scene, thus enhancing the performance of Al applications in vehicular environments.
[0234] The present disclosure describes a vector quantization (VQ) model that is guided by object detection and classification. It aims to address the problem in traditional deep learning VQ models, which is that they treat all parts of an image equally, wasting computational power on uninteresting or irrelevant parts of the image. In certain embodiments, object detectors and classifiers, or labeled data, are used to guide the VQ model to focus more on objects that affect driving behavior, like vehicles, traffic signs, and traffic lights. This way, the model creates higher fidelity reconstructions around these important objects while ignoring less relevant parts of the image.
[0235] Certain aspects of the present invention provide a method of object detection / classification guided vector quantization (VQ) for vehicle risk assessment. The method involves using an Al perception engine to detect and classify objects in an image, and then using this information to guide the VQ process.
[0236] In one example, an image is divided into 8x16 patches. The Al perception engine is used to identify objects of interest within the image, such as other vehicles. In some embodiments, these specific regions are then processed separately with the VQ model, with a higher level of detail compared to the rest of the image. The codeword assignment process is modified to use different parts of the codeword vocabulary for the regions of interest and the rest of the image. The separate processing of these regions of interest may result from training separate VQ layers for the regions of interest and the rest of the image, or using a conditional VQ model that takes a region-of-interest mask or (masks) as an additional input.
[0237] This disclosure provides a method for processing visual data for use in a vehicle-mounted generative Al model. The process begins with the capture of a high-resolution image, specifically at 1080p, by a camera mounted on a vehicle. The captured image is then subjected to a cropping and resizing process, wherein the image's dimensions are adjusted to a standardized size of 128 pixels by 256 pixels. This standardized size may be chosen to ensure consistency across all processed images, and to reduce the computational load of the following steps, enabling the system to operate efficiently on an edge device installed in the vehicle. The resized image is then processed by a vector quantization generative adversarial network (VQ- GAN) (e.g., the encoder 302). This is a type of generative Al model that is trained to generate new images that resemble the input images. The VQ-GAN operates by dividing the image into patches and replacing each patch with a codeword from a predefined vocabulary .
[0238] In this example, the vocabulary consists of 4096 discrete codewords, effectively creating a 4096-token vocabulary. The image is divided into patches of 8x16. and each patch is replaced by the corresponding codeword in the vocabulary. This results in a tokenized representation of the image, where the complexity of the image data is significantly reduced while still retaining the essential information. This reduction in data complexity is critical for enabling the system to operate on an edge device, where computational resources may be limited.
[0239] The tokenized image is then (optionally) processed by the VQ-GAN decoder (e.g., the decoder 304), which is trained to reconstruct the original image from the tokenized representation. The decoder generates a second image with dimensions of 128x256, which is a high-fidelity reconstruction of the original image. In real-time operation, the image reconstruction itself may be skipped, because in real-time operation, the encodedrepresentation of each interest may be fed directly into a generative driving model. The generative driving model may operate on these VQ codewords
[0240] This process of capturing, resizing, tokenizing, and reconstructing images allow s the generative Al model to efficiently process visual data on an edge device installed in a vehicle. By reducing the complexity of the image data and focusing on the most relevant features, the system can operate in real-time, providing valuable insights to assist with vehicle operation.
[0241] The output of the VQ model may be integrated with a risk prediction model. This risk prediction model is trained to predict the risk of a collision based on the codeword representation of the image. The risk prediction model may use the spatial relationships between the codew ords representing the ego vehicle and other vehicles to make its prediction. The risk prediction model may additionally consider the temporal aspect of tire codew ords in determining the risk assessment. For VQ. this consideration of temporal components may be facilitated by processing sequences of images and including codewords that represent changes over time. In some embodiments, a feedback loop may be implemented w'here the risk prediction model influences the VQ model during training to ensure that risk-predictive elements are well represented, but not dining run-time, since the VQ may be paired with a different risk prediction module compared to the one used for training and feedback from such a module may have unintended consequences. Still, to accommodate and benefit from run-time feedback, the modules may be configured to communicate with each other along defined dimensions, such as the accident prediction module may have access to a defined interface where it can communicate to the VQ an indication of salient objects for accident prediction.
[0242] This object detection / classification guided VQ method may allow' for a more accurate and detailed representation of the objects of interest in the image, leading to more effective risk assessment in vehicular environments.
[0243] The vision tokenizer may be considered an application of tokenization, which is a common technique in Natural Language Processing (NLP), to vision-based Al models. In NLP, tokenization is the process of breaking down text into smaller pieces called tokens. In the context of vision, tokenization would involve breaking dow n an image into smaller segments or 'tokens', each representing a certain feature or part of the image. By breaking down visual input into tokens, the Al model might be able to better understand and predict complex driving scenarios.
[0244] The Al perception engine, capable of detecting and classifying various objects including vehicles, traffic signs, and traffic lights from the images, can be used to guide the vector quantization (VQ) model to create higher fidelity’ reconstructions around these objects. In an “attention mechanism” approach, the VQ model can incorporate an attention mechanism that focuses on the areas of the image where relevant objects have been detected by the Al perception engine. For such embodiments, tire attention mechanism may weigh the regions of the image containing the detected objects more heavily during the VQ process. This would result in those regions being reconstructed with higher fidelity, as more codewords from the vocabulary would be allocated to represent them.
[0245] In a “region-specific quantization” approach, after the Al perception engine identifies the objects of interest within the image, these specific regions can be isolated and processed separately with the VQ model. This allows for a larger portion of the codeword vocabulary to be used for these important areas, thus achieving a higher fidelity reconstruction for the objects of interest. In one embodiment, the identified regions containing the objects of interest could first be segmented or "masked out" from the rest of the image. After the Al perception engine identifies the objects of interest within the image, these specific regions can be isolated. The isolating process can create what is effectively a 'mask' on the original image, highlighting the areas of interest. The VQ model would then process these highlighted areas separately with a higher proportion of the codeword vocabulary, ensuring a high-fidelity reconstruction for these regions. For the rest of the image, a separate, less detailed quantization could be performed. This could involve using fewer codewords from the vocabulary or a lower resolution patch grid for the VQ process. This would allow the system to save computational resources while still providing a basic reconstruction of the less important parts of the image. By combining the high-fidelity reconstructions of the regions of interest with the basic reconstruction of the rest of the image, a full image reconstruction can be produced that maintains high detail for tire important driving objects while being computationally efficient.
[0246] In another embodiment, region-specific quantization may be done without explicit masking. The process can be facilitated through a form of “differential quantization” where the system uses a more detailed or extensive part of the codeword vocabulary' for the regions of the image identified as containing objects of interest, and a less detailed or smaller part of the codeword vocabulary for the other regions. In this approach, after the Al perception engine identifies the objects of interest, the VQ model can be guided to focus more on these regions during tire quantization process, allocating more codewords or using a finer patch gridfor these areas. For the rest of the image, fewer codewords or a coarser patch grid can be used. This method would also enable the VQ model to create higher fidelity reconstructions through the objects of interest, while still providing a basic reconstruction of the less important parts of the image, although this way, the need for explicit masking is eliminated. Skipping an explicit making step may potentially improve the run-time efficiency of the system. This approach still maintains high-quality reconstructions of the objects of interest and provides a computationally efficient method suitable for implementation on an edge device.
[0247] Configuring the V Q processing with region-specific quantization with masking vs without masking may be based on the available computational resources and the required level of detail in the reconstructed image. For the “with masking” approach, the Al perception engine identifies regions containing objects of interest. These areas are then masked and isolated from the rest of the image. Each region (masked and unmasked) is then processed separately by the VQ model. This method allows for potentially more precise and therefore more focused processing of tire regions of interest, potentially resulting in a higher level of detail in the reconstructed objects of interest. The masking process, however, may add an additional computational step. Moreover, there could be challenges in seamlessly integrating the high-fidelity7reconstructions of the objects of interest with the lower fidelity reconstruction of tire rest of the image. There may be, for example, real-time challenges that arise when there are a large number of detected objects in the visual scene.
[0248] For tire “without masking” approach, the Al perception engine identifies regions with objects of interest. The VQ model then applies differential quantization, using a more detailed part of the codeword vocabulary for these regions and a less detailed part for the remaining areas. This method eliminates the need for explicit masking, simplifying the overall process and potentially improving computational efficiency. It also naturally allows for a more seamless integration of the high and low fidelity’ regions in the final reconstructed image. However, the differentiation between areas of interest and tire rest of the image might not be as distinct as with the masking method. The lower granularity may also result in less detail in the reconstructed objects of interest compared to the masking method. This lower granularity might be considered a benefit in some scenarios, such as where there is an interest in obscuring identifiable details from third parties in the driving scene, as may be a general privacy consideration.
[0249] In both methods of guiding the VQ. the idea is to focus more computational resources on the objects of interest, ensuring that these are reconstructed with high fidelity while the less important parts of the image are reconstructed with lower fidelity, thus optimizing the use of computational resources. The choice between the two methods would depend on the specific requirements and constraints of the system, as described above.
[0250] In both methods, by focusing the VQ process on the regions of the image that contain relevant driving objects, the system improves the ability for the overall system to provide more accurate and detailed information to assist with vehicle operation, even when operating on an edge device with limited computational resources.Additional Codeword representation and training considerations
[0251] A goal of vector quantization (VQ) is to learn a set of codewords (or tokens) that can represent different parts of tire image. These codewords are learned during the training phase of the VQ model, where it is trained on a large dataset of images. The model learns to represent each patch of an image with the most similar codeword from its learned vocabulary.
[0252] In operation, a single detected object, such as a white car, may appear in the image in a location and at a size that spans multiple quantization patches. If a white car spans four patches, each of those patches would be represented by one codeword, resulting in four codewords total to represent tire car.
[0253] Each patch may be represented by a single codeword. The combination of codewords come together to represent the whole image. In the case of a vehicle spanning multiple patches, each patch of the vehicle may be represented by a single codeword, and the overall representation of the vehicle would be a combination of these codewords. In one example, the VQ model could use different codewords to represent different parts of the vehicle. The codewords for the different patches containing a portion of the vehicle may represent different types of visual features that are found in the image patches, such as certain shapes, colors, textures, or patterns. Depending on how the VQ was trained, it may be possible that the same or similar combination of codewords could be used to represent other objects with similar features in other images, with the specific codewords selected based on the closest match in the model's vocabulary'.
[0254] In a region-specific quantization approach, different levels of detail may be applied to different regions of the image, but not to add additional codewords for the objectsof interest. If the Al perception engine identifies certain objects of interest within the image, these areas could be processed separately with a higher level of detail, meaning, using a more detailed part of the codeword vocabulary (or a finer patch grid). The rest of the image would be processed with a lower level of detail, using a less detailed part of tire codeword vocabulary (or a coarser patch grid). The result would still be a single set of codewords representing the entire image, but the distribution of codewords would be biased towards the objects of interest. For example, in an image with 8x16 patches, more codewords may represent the objects of interest and fewer codewords may represent the rest of the image.
[0255] This approach does not require adding additional codewords for the objects of interest. Instead, it involves selectively allocating the learned codewords in a way that emphasizes the objects of interest. This allows the VQ model to maintain high fidelity in the areas of the image that matter most, while reducing the amount of data needed to represent the less important parts of the image.
[0256] Modify ing a vector quantizer to apply region-specific quantization may involve deviating from a standard model architecture and training process to incorporate additional information about the regions of interest. These changes may include training the modified VQ model on a large dataset of images, using a loss function that encourages the model to allocate more codewords to the regions of interest. This could involve using a weighted loss function that applies a higher weight to the regions of interest, or a multi-task loss function that includes a separate term for the object detection task.
[0257] The codeword vocabulary could be designed or trained to include more codewords that represent features of vehicles. This may be achieved by training the VQ model on a dataset that includes a large number of images of vehicles, which itself may be achieved by preferentially selecting image data from vehicles running a safety system that detects such objects.
[0258] In some embodiments, the output of the VQ model (the set of codewords representing the image) could be integrated with a risk prediction model. This model could be a machine learning model that is trained to predict the risk of a collision based on the codeword representation of the image. The risk prediction model could use the spatial relationships between the codewords representing the ego vehicle and other vehicles to make its prediction. It may also include a risk prediction model that may influence tire VQ model, as mentioned above. For instance, if tire risk prediction model determines that a particular configuration ofvehicles is high-risk, it could signal the VQ model to allocate more codewords to similar configurations in future images. In addition, since the risk of collision is not only determined by the current state but also the motion and trajectory of tire vehicles, temporal information may be incorporated into the VQ model. This could be achieved by processing sequences of images and including codewords that represent changes over time.
[0259] Designing a VQ model for collision risk warnings would involve focusing on accurately representing the features and configurations of objects in the scene that may be involved in collisions or near-miss collision avoidance events. By integrating the VQ and risk prediction models, the system may more quickly and reliably predict the risk of a collision and provide appropriate warnings or implement risk mitigation steps.
[0260] Accordingly, one aspect of this disclosure lies in the integration of object detection / classification with Vector Quantization (VQ) for the purpose of vehicle risk assessment. The VQ process is guided by object detection and classification, which allows the model to give more emphasis to the objects of interest (such as other vehicles) in the image. This is a novel approach to image representation and compression in the context of vehicular risk assessment. The technical effects include improved image representation, efficient resource utilization, enhanced risk prediction, and adaptive learning. By using object detection / classification to guide the VQ process, this method can achieve a more accurate and detailed representation of tire objects of interest in the image. This can lead to improved performance in downstream tasks such as risk prediction. By allocating more codewords to the objects of interest and fewer codewords to the less important parts of the image, this method can achieve a more efficient use of computational and storage resources. This may be particularly enabling for real-time applications such as autonomous driving. By integrating the output of the VQ model with a risk prediction model, this method can provide a more effective way of predicting the risk of a collision. The risk prediction model can use the spatial relationships between the codewords representing the ego vehicle and other vehicles, as well as temporal information, to make more accurate predictions. The inclusion of a feedback loop where the risk prediction model influences the VQ model allows the system to adapt and learn over time, improving its performance in changing environments.Controlling the Driving World Model with an external policy
[0261] FIG. 11 illustrates an example of how a policy may be used in conjunction with the driving world model, according to certain aspects of the present disclosure.
[0262] A part of autonomous driving is controlling the "Ego vehicle" or the self-driving vehicle. A driving world model may also be designed to receive steering and other movement directions from an external policy. The model will then be directed to imaging a future scenario in which the ego-vehicle is driven in accordance with the external policy.
[0263] The policy may be a decision-making component that determines the actions of the Ego vehicle. It interprets the tokenized sensor data, which represents the current state of the vehicle and its environment and determines the most appropriate action or motion for the vehicle to take.
[0264] Alternatively, the policy may be an external module that operates according to a desired trajectory, and which may not be responsive to changes in the driving scene.
[0265] The next frame visual tokens are generated based on this policy, which is expressed in sensor tokens. This means tire policy not only determines the vehicle's immediate actions but also influences the model's predictions of the visual environment in the next time frame.
[0266] In summary, FIG. 11 demonstrates how the driving world model and an external policy may work together to simulate control of tire Ego vehicle and predict future states of the environment. This can be a form of predicting “What If’ scenarios.
[0267] FIGS. 12A-F illustrate an example of driving images that are processed through a quantizer and influenced by an external policy, along with three instantiations of a driving prediction model, in line with certain aspects of the present disclosure.
[0268] In this example, the focus is on controlling the Ego vehicle, specifically at an intersection having a stop sign. The images represent different actions forced by tire control policy at this intersection.
[0269] The top left panels in FIGS. 12A-F shows the Ground Truth, where the vehicle naturally turns right. The top right panels in FIGS. 12A-F demonstrates a situation where the control policy forces the vehicle to make a left turn, contrary to the Ground Truth.
[0270] The bottom left panels in FIGS. 12A-F depicts a scenario where the control policy forces the vehicle to go straight at the intersection. Lastly, the bottom right panels in FIGS. 12A-F shows a situation where the control policy forces the vehicle to make a right turn after a stop.
[0271] These figures underscore the flexibility and adaptability of the driving prediction model. It can respond to different external policies and adjust the vehicle's actions accordingly. It can then generate an expected visual scene that results from the different actions that are considered.
[0272] FIGS. 13A-13F illustrate an example of driving images that are processed through a quantizer and influenced by an external policy, along with three instantiations of a driving prediction model, in line with certain aspects of the present disclosure.
[0273] In this example, the focus is on controlling the Ego vehicle, specifically at an intersection having a stop sign. The images represent different actions forced by the control policy at this intersection.
[0274] The top left panels in FIGS. 13A-13F shows the ground truth, where the vehicle naturally turns right. The top right panels in FIGS. 13A-13F demonstrates a situation where the control policy forces the vehicle to make a left turn, contrary to the ground truth.
[0275] The bottom left panels in FIGS. 13A-13F depicts a scenario where tire control policy forces the vehicle to go straight at the intersection. Lastly, the bottom right panels in FIGS. 13A-13F shows a situation where the control policy forces the vehicle to make a right turn after a stop.
[0276] These figures underscore the flexibility and adaptability of the driving prediction model. It can respond to different external policies and adjust the vehicle's actions accordingly. It can then generate an expected visual scene that results from the different actions that are considered.Using the Driving World Model as an Autonomous Vehicle Policy
[0277] The driving world model can be used as a policy for controlling an autonomous vehicle, in line with certain aspects of the present disclosure.
[0278] Here, the driving world model is itself used as a policy for controlling the Ego vehicle, or autonomous vehicle. This implies that instead of having a separate external policy, the driving world model directly controls the motion of the vehicle based on the sensor tokens.
[0279] The sensor tokens, which are generated by the sensor tokenizer, represent the current state of the vehicle and its environment. These tokens are used to control tire motion ofthe Ego vehicle in real-time or "on-line." This allows for instant response to changes in the driving environment and immediate action based on the current state of the vehicle.
[0280] Furthermore, in accordance with certain aspects, the model may be aligned to safe driving. This means it is designed and trained to make decisions that prioritize safety. This could involve obeying traffic rules, avoiding obstacles, maintaining a safe distance from other vehicles, and so on. Various methods of aligning a generative driving model to safe driving are disclosed in US Provisional Application No. 63 / 498.493 filed April 26, 2023, and PCT Application No. PCT / US2024 / 023397. filed April 5, 2024, each of which is incorporated herewith in its entirety for all purposes.
[0281] The driving world model can directly control an autonomous vehicle, emphasizing its potential as a comprehensive tool for autonomous driving that integrates perception, prediction, and control.Curation of Training Data
[0282] Training the generative driving model (and accident prediction the classifier) may be strongly influenced by the selection of the training data used. This may be referred to as curation of training data. The data is derived from a higher-level perception engine that detects various elements in the driving environment such as vehicles, traffic signs, and driving maneuvers. This perception engine is capable of distinguishing between safe and unsafe driving maneuvers. For instance, a safe driving maneuver could be a lane change executed with proper signaling or coming to a complete stop at a stop sign. An unsafe maneuver, on the other hand, might involve driving through a stop sign without stopping. The system also detects and records instances of accidents.
[0283] The training data may be curated to comprise instances where specific events or 'happenings' are detected, be it safe or unsafe maneuvers or accidents. Filtering for video samples associated with these specific events may reduce the dataset size, which is both more manageable and potentially more likely to improve the resulting trained models. In comparison to using all observed data (all data that may be observed by a vehicle-mount edge device), this event-specific data filtering can reduce the dataset by 80 - 99%. This method ensures that the model is trained on relevant data, enhancing its ability to predict future driving scenarios and potential accidents effectively.
[0284] Data curation may enhance a machine learning model’s ability to generalize from the training data to unseen scenarios. However, a common challenge in training these models is the risk of overfitting to the training data. Overfitting happens when a model learns the training data too well, including its noise and outliers, resulting in poor performance on unseen data.
[0285] In the context of a generative driving model, overfitting could occur if the model excessively learns from 'boring' or average scenarios that represent routine, uneventful driving situations. These scenarios, while common, do not provide much valuable information for predicting accidents or other significant driving events.
[0286] By focusing the training data on specific events of interest, such as accidents or unsafe driving maneuvers, the model may be more likely to avoid overfitting to these average scenarios. Instead of learning to represent the mundane aspects of driving, the model can allocate more of its capacity to understanding and predicting these events of interest.
[0287] This selective focus on critical events allows the model to build a more robust and nuanced understanding of the complex patterns and correlations that lead up to accidents. It can learn to pick up subtle cues and anomalies that may suggest a higher risk of an accident, such as erratic lane changes, sudden braking, or failing to stop at a stop sign.
[0288] Furthermore, by reducing the dataset size through filtering for event-specific instances, we not only make the training process more efficient but also improve the quality of the data the model learns from. This curated dataset is more likely to represent the diversity and complexity of real-world driving scenarios that the model will encounter in operation.
[0289] Ultimately, this approach of focusing on events of interest in the training data contributes to a generative driving model that is more capable of accurately predicting future driving scenarios and potential accidents. It ensures that the model is well-equipped to handle complex and high-risk situations, thereby enhancing tire safety and reliability of autonomous driving systems.
[0290] An approach to training data curation for the driving world model is described below, according to certain aspects of the present disclosure. The goal of this curation process is to create a smaller foundational model that can run in real-time for edge computing use cases. Edge computing refers to processing data directly on tire devices that produce or collect the data (in this case, tire autonomous vehicle), rather than sending the data to a remote server forprocessing. In this case, the edge computing devices are leveraged to help curate the data. The edge device may help curate the data by processing visual and other data and only transmitting a small percentage of the video data to a remote server. The driving model may then be trained on a further sub-selection of tire video and other sensor data that was transmitted to the remote server from a fleet of edge devices.
[0291] One of the challenges in training autonomous driving models is that most driving data is uninteresting or mundane, while interesting or critical events are sparse. If too much of the model's capacity is used for capturing the fidelity of frequent, mundane driving, there may be insufficient capacity left for interesting or important events. This could lead to sub-optimal performance when the model encounters unusual or critical scenarios.
[0292] Therefore, the motivation behind this curated training dataset is to ensure that tire model is trained on a balanced mix of mundane and interesting driving events. This would enable the model to handle both common and uncommon driving scenarios effectively.
[0293] Through carefully curated and cleaned data, the performance of LLMs can be enhanced.
[0294] FIGS. 13A-13F show an example of driving images that are processed through a quantizer and two instances of a driving prediction model, where the models were trained with differently curated training data, in accordance with certain aspects of the present disclosure.
[0295] The example may illustrate a "Lane Change Hallucination." It compares the performance of a 90 million parameter model trained on curated data in handling rare events, like forced lane changes, vs. one that is trained on training data that is less curated.
[0296] In the left case, the model w as trained a higher proportion of mundane driving data. As a result, tire trained model starts "hallucinating" (predicting incorrectly) when it is forced to make a lane change, which is a less common event. However, this model would be expected to have better pixel fidelity in common, mundane driving scenarios because it was trained on a large amount of such data.
[0297] In contrast, the right case shows a model trained on more curated data, meaning the training data was carefully selected to include a balance of mundane and interesting events.This model successfully completes the lane change in its imagined future, demonstrating its effectiveness in handling rare or unusual events.
[0298] One strategy' to curate the training data is to use Al triggers to sub-select or curate the data, focusing on interesting segments out of the ten billion miles of driving data. This way, tire model can be trained to handle a wide range of driving scenarios, including both common and rare events, enhancing its overall performance and reliability.
[0299] FIG. 14 illustrates a variety of uncommon driving scenarios that could benefit from a generalized collision warning system, in accordance with certain aspects of the present disclosure.
[0300] Given the vast amount of driving data, there are numerous rare events that can occur, including accidents. FIG. 14 demonstrates the importance of including these rare events in the training data of autonomous driving models.
[0301] Two trained models - one trained without accident data (No Accident Model) and the other with accident data (Accident Model) may also be compared. The model trained with accident data may be trained using accident recordings for which all personal identifiers have been removed and the like.
[0302] It may be observed that the Accident Model accurately predicted collisions and their follow-on effects, whereas tire No Accident Model may not.
[0303] This comparison underscores the importance of covering "long tail" events - rare events that can have significant consequences - in the training data. By including many examples of each rare event in the training data, the model can learn to handle these situations effectively.Autonomous Driving 2.0
[0304] Autonomous Driving 2.0 may refer to a future generation of autonomous vehicle (AV) technology, which may be based on a foundational driving model as disclosed herein that has been trained on billions of miles of real-world driving data.
[0305] The creation of this foundational driving model involves curating data from billions of miles driven in tire real world. This data forms the basis for training the model, and tire architecture has been designed to create and control this driving model.
[0306] The model shows emergent abilities and generalization. Emergent abilities refer to the model's ability to handle new, unseen scenarios based on its training, while generalization refers to its ability to apply what it has learned to a wide range of scenarios.
[0307] Real-world data and filtering by edge devices may ultimately enable reliable handling the "long tail" of comer cases - rare events that can have significant consequences. By training the model on real-world data, it can learn to handle these rare but important scenarios effectively.
[0308] The driving model's policy may align with safe driving. This means it is designed to make decisions that prioritize safety, such as obeying traffic rules, avoiding obstacles, and maintaining a safe distance from other vehicles.
[0309] To further enhance the model's performance and safety, additional sensors can be added to tire system. These can include RADAR, multiple cameras, and other sensing technologies that can provide the model with more detailed and comprehensive information about its environment.
[0310] Accordingly, systems and methods of training and using a generative driving model are disclosed herein. The generative driving model may predict future images that may be captured from a road-facing vehicle-mounted camera, based on processing of one or more images from the same, and optionally further based on processing of other sensor data, such as accelerometer, speed, radar, LiDAR. and the like. The generative model may incorporate detection-guided vector quantization. The generative driving model may be used for advanced driver assistance functions, such as generalized collision warning. Certain aspects may be used as an autonomous driving controller or as one or more components of an autonomous driving controller system.
[0311] A computer program product may include a non-transitory computer-readable medium having instructions stored thereon, the instructions being executable by one or more processors configured to carry out any of the systems and / or methods described herein.
[0312] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whethersuch functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of this disclosure or the claims.
[0313] Embodiments implemented in computer software may be implemented in software, firmware, middleware, microcode, hardware description languages, or any combination thereof. A code segment or machine-executable instructions may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardw are circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.
[0314] The actual software code or specialized control hardware used to implement these systems and methods is not limiting of the claimed features or this disclosure. Thus, the operation and behavior of the systems and methods were described without reference to the specific software code being understood that software and control hardware can be designed to implement the systems and methods based on the description herein.
[0315] When implemented in software, the functions may be stored as one or more instructions or code on a non-transitory computer-readable or processor-readable storage medium. The steps of a method or algorithm disclosed herein may be embodied in a processorexecutable software module, which may reside on a computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable media includes both computer storage media and tangible storage media that facilitate transfer of a computer program from one place to another. A non-transitory processor-readable storage media may be any available media that may be accessed by a computer. By way of example, and not limitation, such non-transitory processor-readable media may comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other tangible storage medium that may be used to store desired program code in the form of instructions or data structures and that may be accessed by a computer or processor. Disk and disc, as used herein, include compact disc (CD), laser disc,optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media. Additionally, the operations of a method or algorithm may reside as one or any combination or set of codes and / or instructions on a non-transitory processor-readable medium and / or computer-readable medium, which may be incorporated into a computer program product.
[0316] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0317] The terms “data processing apparatus”, “data processing system”, “client device”, “client computing device”, “computing platform”, “computing device”, “computing system”, “user device”, or “device” can encompass all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, a system on a chip, or multiple ones, or combinations, of the foregoing. The apparatus can include special purpose logic circuitry, e.g., an FPGA or an ASIC. The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them.
[0318] A computer program (also known as a program, softw are, softw are application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub-programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0319] Processors suitable for the execution of a computer program include, by w ay of example, both general and special purpose microprocessors, and any one or more processorsof any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The elements of a computer include a processor for performing actions in accordance with instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a GPS receiver, a digital camera device, a video camera device, or a portable storage device (e.g., a universal serial bus (USB) flash drive), for example. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0320] To provide for interaction with a user, implementations of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube), plasma, or LCD monitor, for displaying information to the user: a keyboard; and a pointing device, e.g., a mouse, a trackball, or a touchscreen, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can include any form of sensory- feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user.
[0321] In certain circumstances, multitasking and parallel processing system may be advantageous. Moreover, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. For example, the computing devices described herein can each be a single module, a logic device having one or more processing modules, one or more servers, or an embedded computing device.
[0322] Having now described some illustrative implementations and implementations, it is apparent that the foregoing is illustrative and not limiting, having been presented by wayof example. In particular, although many of the examples presented herein involve specific combinations of method acts or system elements, those acts and those elements may be combined in other ways to accomplish the same objectives. Acts, elements, and features discussed only in connection with one implementation are not intended to be excluded from a similar role in other implementations or implementations.
[0323] The phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of "‘including,” “comprising.” “having,” “containing,” “involving.” “characterized by.” “characterized in that.” and variations thereof herein, is meant to encompass the items listed thereafter, equivalents thereof, and additional items, as well as alternate implementations consisting of the items listed thereafter exclusively. In one implementation, the systems and methods described herein consist of one, each combination of more than one, or all of the described elements, acts, or components.
[0324] Any references to implementations or elements or acts of tire systems and methods herein referred to in the singular may also embrace implementations including a plurality of these elements, and any references in plural to any implementation or element or act herein may also embrace implementations including only a single element. References in the singular or plural form are not intended to limit the presently disclosed systems or methods, their components, acts, or elements to single or plural configurations. References to any act or element being based on any information, act or element may include implementations where the act or element is based at least in part on any information, act, or element.
[0325] Any implementation disclosed herein may be combined with any other implementation, and references to “an implementation,” “some implementations,” “an alternate implementation,” “various implementation,” “one implementation,” or the like are not necessarily mutually exclusive and are intended to indicate that a particular feature, structure, or characteristic described in connection with the implementation may be included in at least one implementation. Such terms as used herein are not necessarily all referring to the same implementation. Any implementation may be combined with any other implementation, inclusively or exclusively, in any manner consistent with the aspects and implementations disclosed herein.
[0326] References to “or” may be construed as inclusive so that any terms described using “or” may indicate any of a single, more than one, and all of the described terms.
[0327] Where technical features in the drawings, detailed description or any claim are followed by reference signs, the reference signs have been included for the sole purpose of increasing the intelligibility of the drawings, detailed description, and claims. Accordingly, neither the reference signs nor their absence have any limiting effect on the scope of any claim elements.
[0328] The preceding description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the embodiments described herein and variations thereof. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other embodiments without departing from the spirit or scope of the subject matter disclosed herein. Thus, the present disclosure is not intended to be limited to the embodiments shown herein but is to be accorded tire widest scope consistent with the following claims and the principles and novel features disclosed herein.
[0329] While various aspects and embodiments have been disclosed, other aspects and embodiments are contemplated. The various aspects and embodiments disclosed are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the following claims.
Claims
CLAIMSWhat is claimed is:
1. A method comprising: receiving, by one or more processors, one or more images of an environment surrounding a vehicle or within the vehicle captured by a camera mounted to or in the vehicle; executing, by the one or more processors, an object classifier using the one or more images to detect one or more objects depicted within at least one of the one or more images; identifying, by the one or more processors, a first set of portions of the at least one image that each depict at least one of the detected one or more objects that relate to operation of tire vehicle in the environment and a second set of portions of the at least one image that do not depict any of the detected one or more objects that relate to operation of the vehicle in the environment; converting, by tire one or more processors, the at least one image into a plurality of learned codewords in which tire first set of portions of the at least one image are each represented by a first set of learned codewords selected from a first codeword space corresponding to objects that relate to operation of the vehicle in the environment, and the second set of portions of the at least one image are each represented by a second set of learned codewords selected from a second codeword space corresponding to portions of images that do not contain objects that relate to operation of the vehicle in tire environment; and responsive to executing, by the one or more processors, a machine learning model using the plurality of learned codewords for each of the at least one image to detect a predicted event, generating, by the one or more processors, an alert within the vehicle.
2. The method of claim 1 , wherein executing the object classifier to detect the one or more objects comprises: executing, by the one or more processors, the object classifier using the one or more images to detect one or more objects that correspond to the vehicle getting into a collision.
3. The method of claim 2, further comprising: generating, by the one or more processors, a training data set comprising a first set of images each captured within a defined time period of a vehicle getting into a collision and second set of images that arc not captured within the defined time period of a vehicle getting into a collision;labeling, by the one or more processors, each of the first set of images to indicate each of the first set of images corresponds to a collision and each of the second set of images to indicate each of the second set of images does not correspond to a collision; and training, by the one or more processors, the object classifier using labels of the first set of images and the second set of images.
4. The method of claim 3, wherein converting the at least one image into the plurality of learned codewords comprises converting, by the one or more processors, the at least one image into the plurality of learned codewords using a second machine learning model, the method further comprising: training, by the one or more processors, the second machine learning model using a plurality of training images and identifications of objects detected by the object classifier within the plurality of training images labeled based on whether the images correspond to a collision.
5. The method of claim 4, wherein training the second machine learning model comprises: training, by the one or more processors, the second machine learning model using a loss function that corresponds to higher weights for portions of images containing at least one object detected by the object classifier within the plurality of training images than for portions of images for which the object classifier does not detect objects that relate to operation of the vehicle in the environment.
6. The method of claim 1 , wherein the first codeword space corresponding to objects that relate to operation of the vehicle in the environment has a larger size than the second codeword space corresponding to portions of images that do not contain objects that relate to operation of the vehicle in the environment.
7. The method of claim 1. wherein converting the at least one image into a plurality of learned codewords comprises: generating, by the one or more processors, codewords that represent different grid locations of a defined grid segmenting each of the at least one image.
8. The method of claim 7, wherein converting the at least one image into a plurality of learned codewords comprises:identifying, by the one or more processors, a first set of grid locations of an image that each contain at least one object that relates to operation of the vehicle in the environment and a second set of grid locations of the image that do not contain any objects that relate to operation of tire vehicle in the environment; and using, by the one or more processors, the first codeword space to generate codewords for each of the first set of grid locations and tire second codeword to generate codewords for each of the second set of grid locations.
9. The method of claim 1, wherein executing the machine learning model using the plurality of learned codewords for each of the at least one image to detect a predicted event comprises: executing, by the one or more processors, a large language model using the first set of learned code ords and the second set of learned codewords as input.
10. The method of claim 1, further comprising: converting, by the one or more processors, the first set of learned codewords and the second set of learned codewords into a compressed at least one image containing portions that correspond to the first set of learned codewords at a higher resolution then portions that correspond to the second set of learned codewords, wherein executing the machine learning model using the plurality of learned codewords for each of the at least one image to detect the predicted event comprises executing, by the one or more processors, the machine learning model using the compressed at least one image as input.
11. The method of claim 1, comprising: executing, by the one or more processors, the machine learning model using a sequence of the plurality of learned codewords for each of the at least one image ordered and grouped based on a time of capture of each of the at least one image from which the plurality of learned codewords were generated to detect the predicted event.
12. A system, comprising: one or more processors configured by machine -readable instructions to: receive one or more images of an environment surrounding a vehicle or within the vehicle captured by a camera mounted to or in the vehicle;execute an object classifier using the one or more images to detect one or more objects depicted within at least one of the one or more images: identify a first set of portions of the at least one image that each depict at least one of the detected one or more objects that relate to operation of the vehicle in the environment and a second set of portions of the at least one image that do not depict any of tire detected one or more objects that relate to operation of the vehicle in the environment; convert the at least one image into a plurality of learned codewords in which the first set of portions of the at least one image are each represented by a first set of learned codewords selected from a first codeword space corresponding to objects that relate to operation of the vehicle in the environment, and the second set of portions of the at least one image are each represented by a second set of learned codewords selected from a second codeword space corresponding to portions of images that do not contain objects that relate to operation of the vehicle in tire environment: and responsive to executing a machine learning model using the plurality of learned codewords for each of the at least one image to detect a predicted event, generate an alert within the vehicle.
13. The system of claim 12, wherein the one or more processors are configured to execute the object classifier to detect tire one or more objects by: executing the object classifier using the one or more images to detect one or more objects that correspond to the vehicle getting into a collision.
14. The system of claim 13, wherein tire one or more processors are further configured to generate a training data set comprising a first set of images each captured within a defined time period of a vehicle getting into a collision and second set of images that are not captured within the defined time period of a vehicle getting into a collision; label each of the first set of images to indicate each of the first set of images corresponds to a collision and each of the second set of images to indicate each of the second set of images does not correspond to a collision: and train the object classifier using labels of the first set of images and the second set of images.
15. The system of claim 14. wherein the one or more processors are configured to convert the at least one image into the plurality of learned codewords by converting the at least one image into the plurality of learned codewords using a second machine learning model, the one or more processors further configured to: train the second machine learning model using a plurality’ of training images and identifications of objects detected by the object classifier within the plurality of training images labeled based on whether the images correspond to a collision.
16. The system of claim 15 wherein the one or more processors are configured to train the second machine learning model by: training the second machine learning model using a loss function that corresponds to higher weights for portions of images containing at least one object detected by the object classifier within the plurality of training images than for portions of images for which the object classifier does not detect objects that relate to operation of the vehicle in the environment.
17. The system of claim 12, wherein the first codeword space corresponding to objects that relate to operation of the vehicle in the environment has a larger size than the second codeword space corresponding to portions of images that do not contain objects that relate to operation of the vehicle in the environment.
18. The system of claim 12, wherein the one or more processors are configured to convert the at least one image into a plurality of learned codewords by: generating codewords that represent different grid locations of a defined grid segmenting each of the at least one image.
19. The system of claim 18, wherein the one or more processors are configured to convert the at least one image into a plurality of learned codewords by: identifying a first set of grid locations of an image that each contain at least one object that relates to operation of the vehicle in the environment and a second set of grid locations of the image that do not contain any objects that relate to operation of the vehicle in the environment; and using the first codeword space to generate codewords for each of the first set of grid locations and the second codeword to generate codewords for each of the second set of grid locations.
20. The system of claim 12, wherein the one or more processors are configured to execute the machine learning model using the plurality of learned codewords for each of the at least one image to detect a predicted event by: executing a large language model using the first set of learned codewords and the second set of learned codewords as input.
21. A method comprising: receiving, by one or more processors, a first sequence of images of a first environment surrounding a vehicle or within the vehicle captured by a camera mounted to or in the vehicle; executing, by the one or more processors, an encoder of a machine learning model using the first sequence of images to generate a first intermediary state indicating a first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at one or more time steps subsequent to a first time period in which the first sequence of images were captured; executing, by the one or more processors, a future-state decoder of the machine learning model based on the first intermediary state to generate a future state from the first plurality of candidate future states and of the first environment surrounding the vehicle or within the vehicle at a time step subsequent to the first time period in which the first sequence of images were captured; training, by the one or more processors, the encoder using a loss function based on the generated future state of die first environment surrounding the vehicle; subsequent to the training the encoder, receiving, by the one or more processors, a second sequence of images of a second environment surrounding the vehicle or within the vehicle captured by the camera mounted to or in die vehicle; executing, by the one or more processors, the trained encoder of the machine learning model using the second sequence of images to generate a second intermediary state indicating a second plurality of candidate future states of the second environment surrounding the vehicle or within the vehicle at one or more second time steps subsequent to a second time period in which die second sequence of images were captured; and responsive to executing, by the one or more processors, a classification decoder of the machine learning model based on the second intermediary state to detect a predicted event, generating, by the one or more processors, an alert within the vehicle.
22. The method of claim 21. wherein executing the encoder of the machine learning model using the first sequence of images comprises: executing, by the one or more processors, the encoder of the machine learning model to generate the first intermediary state indicating a first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle each at a first time step subsequent to the first time period in which the first sequence of images were captured, and wherein executing the future-state decoder of the machine learning model based on the first intermediary state comprises: executing, by the one or more processors, the future-state decoder of the machine learning model to generate a future state from the first plurality of candidate future states and of the first environment surrounding the vehicle or w ithin the vehicle at the first time step subsequent to the first time period in which the first sequence of images were captured.
23. The method of claim 21. wherein executing the encoder of the machine learning model using the first sequence of images to generate the first intermediary state comprises: executing, by the one or more processors, the encoder of the machine learning model to generate the first intermediary state comprising the first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at the one or more time steps subsequent to the first time period in w hich the first sequence of images were captured.
24. The method of claim 21, wherein executing the encoder of the machine learning model using the first sequence of images to generate the first intermediary state comprises: executing, by the one or more processors, the encoder of the machine learning model to generate the first intermediary state comprising the first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at each of a plurality of time steps subsequent to the first time period in which the first sequence of images w ere captured.
25. The method of claim 21 , further comprising: labeling, by the one or more processors, a third sequence of images to indicate the third sequence of images corresponds to a collision with the vehicle; andtraining, by the one or more processors, the classification decoder based on whether the classification decoder generates a prediction that a collision will occur using the label of the third sequence of images.
26. The method of claim 25, wherein labeling the first sequence of images comprises: labeling, by the one or more processors, the third sequence of images to indicate the collision with the vehicle occurred at a first time step subsequent to the first time period in which the first sequence of images were captured: and wherein training the encoder comprises: training, by the one or more processors, the encoder based on whether the classification decoder generates a prediction that a collision will occur using the label of the third sequence of images.
27. The method of claim 21. wherein executing the encoder of the machine learning model using the first sequence of images to generate the first intermediary state comprises: executing, by the one or more processors, the encoder of the machine learning model to generate a numerical vector representation of the first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at the one or more time steps subsequent to the first time period in which the first sequence of images were captured.
28. The method of claim 21. comprising: executing, by the one or more processors, the classification decoder of the machine learning model based on the second intermediary state to detect a confidence score for a predicted event; and detecting, by the one or more processors, tire predicted event responsive to determining the confidence score exceeds a threshold.
29. The method of claim 28, wherein executing the classification decoder of the machine learning model comprises: executing, by the one or more processors, the classification decoder of the machine learning model based on the second intermediary state to detect a confidence score for a predicted collision with the vehicle; and wherein detecting the predicted event comprises:comparing, by the one or more processors, the confidence score for the predicted collision with the vehicle with the threshold: and determining, by the one or more processors, the confidence score for the predicted collision with the vehicle exceeds the threshold.
30. The method of claim 21, wherein executing tire encoder of the machine learning model to generate the second intermediary state comprises: executing, by the one or more processors, the encoder of the machine learning model based on one or more sensor values collected from sensors of the vehicle and the second sequence of images to generate the second intermediary state.
31. A system, comprising : one or more processors configured by machine -readable instructions to: receive a first sequence of images of a first environment surrounding a vehicle or within the vehicle captured by a camera mounted to or in the vehicle; execute an encoder of a machine learning model using the first sequence of images to generate a first intermediary state indicating a first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at one or more time steps subsequent to a first time period in which the first sequence of images were captured: execute a future-state decoder of the machine learning model based on the first intermediary state to generate a future state from the first plurality of candidate future states and of the first environment surrounding the vehicle or within the vehicle at a time step subsequent to the first time period in which tire first sequence of images were captured; train the encoder using a loss function based on the generated future state of the first environment surrounding the vehicle; subsequent to the training the encoder, receive a second sequence of images of a second environment surrounding the vehicle or within the vehicle captured by the camera mounted to or in the vehicle; execute the trained encoder of the machine learning model using tire second sequence of images to generate a second intermediary state indicating a second plurality of candidate future states of the second environment surrounding the vehicle or withinthe vehicle at one or more second time steps subsequent to a second time period in which the second sequence of images were captured; and responsive to executing a classification decoder of the machine learning model based on the second intermediary state to detect a predicted event, generating, by the one or more processors, an alert within the vehicle.
32. The system of claim 31, wherein the one or more processors are configured to execute the encoder of the machine learning model using the first sequence of images by: executing the encoder of the machine learning model to generate the first intermediary state indicating a first plurality of candidate future states of the first enviromnent surrounding the vehicle or within the vehicle each at a first time step subsequent to the first time period in which the first sequence of images were captured, and wherein the one or more processors are configured to execute the future-state decoder of the machine learning model based on the first intermediary state by: executing the future-state decoder of the machine learning model to generate a future state from the first plurality of candidate future states and of the first environment surrounding the vehicle or within the vehicle at the first time step subsequent to the first time period in which the first sequence of images were captured.
33. The system of claim 31, wherein the one or more processors are configured to execute the encoder of the machine learning model using the first sequence of images to generate the first intermediary state by: executing the encoder of the machine learning model to generate the first intermediary state comprising the first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at the one or more time steps subsequent to the first time period in which the first sequence of images were captured.
34. The system of claim 31, wherein the one or more processors are configured to execute the encoder of the machine learning model using the first sequence of images to generate the first intermediary state by:executing the encoder of the machine learning model to generate the first intermediary state comprising the first plurality of candidate future states of the first environment surrounding tire vehicle or within the vehicle at each of a plurality of time steps subsequent to the first time period in which the first sequence of images were captured.
35. The system of claim 31, wherein the one or more processors are further configured to: labeling, by the one or more processors, a third sequence of images to indicate the third sequence of images corresponds to a collision with the vehicle; and training, by the one or more processors, the classification decoder based on whether the classification decoder generates a prediction that a collision will occur using the label of the third sequence of images.
36. The system of claim 35, wherein the one or more processors are configured to label the third sequence of images by: labeling the third sequence of images to indicate the collision with the vehicle occurred at a first time step subsequent to the first time period in which the first sequence of images were captured; and wherein the one or more processors are configured to: train, using the label of the third sequence of images, the encoder based on whether the classification decoder generates a prediction that a collision will occur.
37. The system of claim 31, wherein the one or more processors are configured to execute the encoder of the machine learning model using the first sequence of images to generate the first intermediary state by: executing the encoder of the machine learning model to generate a numerical vector representation of the first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at the one or more time steps subsequent to the first time period in which the first sequence of images were captured.
38. A method comprising: receiving, by one or more processors, a sequence of images of a first environment surrounding a vehicle or within the vehicle captured by a camera mounted to or in the vehicle; executing, by the one or more processors, an encoder of a machine learning model using the sequence of images to generate an intermediary state indicating a plurality of candidatefuture states of the first environment surrounding the vehicle or within the vehicle at one or more time steps subsequent to a time period in which the sequence of images were captured, wherein the encoder of the machine learning model is trained based on predictions generated by a future-state decoder configured to generate a future state of a second environment surrounding or within the vehicle or a second vehicle based on a set of candidate future states of the second environment surrounding or within the vehicle or the second vehicle; and responsive to executing, by the one or more processors, a classification decoder of the machine learning model based on the intermediary state to detect a predicted event, generating, by the one or more processors, an alert within the vehicle.
39. The method of claim 38, wherein executing the encoder of the machine learning model using the sequence of images comprises: executing, by the one or more processors, the encoder of the machine learning model to generate the intermediary state indicating a first plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle each at a first time step subsequent to the first time period in which the sequence of images were captured.
40. The method of claim 38, wherein executing the encoder of the machine learning model using the sequence of images to generate the intermediary state comprises: executing, by the one or more processors, the encoder of the machine learning model to generate the intermediary state comprising the plurality of candidate future states of the first environment surrounding the vehicle or within the vehicle at the one or more time steps subsequent to the first time period in which the sequence of images were captured.
Citation Information
Patent Citations
Display panel and display device
US20240023397A1
Vehicle collision detection method, electronic equipment and computer readable storage medium
CN117593725A
Sensor device and signal processing method
EP3985958A1
Electronic device and method of detecting driving event of vehicle
US10803323B2
US202363498493P
Cited By
Systems and methods for mitigating false alarms in a building management system
US12626584B2
Systems and methods for mitigating false alarms in a building management system
US20260057763A1