Systems and methods for providing driver assistance alerts using end-to-end artificial intelligence collision avoidance systems and advanced driver assistance systems
By using an end-to-end conditional imitation learning model, combined with imitation learning and memory enhancement techniques, the challenges of autonomous driving technology in terms of data requirements and security are addressed, enabling efficient autonomous driving in complex and dynamic environments, handling rare and extreme scenarios, and meeting safety standards.
Patent Information
- Application Number
- CN202480049155.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2024-06-29
- Publication Date
- 2026-03-03
AI Technical Summary
Existing autonomous driving technologies face challenges in terms of data requirements, scalability, and safety, especially when dealing with rare and extreme driving scenarios. Traditional methods rely on large amounts of expensive data labeling and complex map data, and lack effective training and verification methods.
An end-to-end conditional imitation learning model is adopted, which combines imitation learning and memory enhancement techniques. By imitating human driving behavior, the model is trained to process sensor inputs and generate vehicle control actions, reducing reliance on complex map data and improving training efficiency and safety.
It enables more efficient and safer autonomous driving in complex and dynamic driving environments, can handle rare and extreme scenarios, reduces data and computing costs, and meets safety standards.
Smart Images

Figure CN121605441A_ABST
Abstract
Description
[0001] Priority application
[0002] This application claims priority to and benefit of U.S. Patent Application No. 18 / 731,115 (Attorney's File No. HYPR 1002-1), filed May 31, 2024, entitled "System and Methods for Providing Driver Assistance Alerts Using an End-To-End ArtificiallyIntelligent Collision Avoidance System and Advanced Driver Assistance Systems," which claims priority to U.S. Provisional Application No. 63 / 524,213 (Attorney's File No. HYPR 1001-1), filed June 29, 2023, entitled "Scalable Training and Validation for an End-To-End Autonomous Driving Model."
[0003] Related Cases
[0004] This application relates to U.S. CIP Application No. ______ (Attorney's File No. HYPR1002-3) entitled "System and Methods For Providing Driver Assistance Alerts Using an End-To-End ArtificiallyIntelligent Collision Avoidance System and Advanced Driver Assistance Systems," filed concurrently with this application, which is incorporated herein by reference for all purposes.
[0005] This application also relates to the following jointly owned applications, all of which are incorporated herein by reference for all purposes:
[0006] U.S. Patent Application No. 18 / 431,827, entitled "Multi-Functional Inventory Storage and Delivery System," filed February 2, 2024 (Attorney's File No. HYPR 1000-2); and
[0007] U.S. Provisional Application No. 63 / 443,342, entitled “Multi-Functional Inventory Storage and Delivery System”, filed on February 3, 2023 (Agent's File No. HYPR 1000-1). Technical Field
[0008] The disclosed technology relates to end-to-end neural networks configured for autonomous and semi-autonomous driving. Specifically, the disclosed technology relates to scalable methods and apparatus for training and validating end-to-end networks configured for autonomous and semi-autonomous driving. Background Technology
[0009] The objects discussed in this section should not be assumed to be prior art simply because they are mentioned in this section. Similarly, problems mentioned in this section or associated with objects provided as background art should not be assumed to have been previously recognized in the prior art. The objects in this section represent only different methods, which themselves may correspond to embodiments of the disclosed technology.
[0010] Autonomous driving technology, attractive for its benefits in driver satisfaction and safety, is already evident in semi-autonomous advanced driver assistance systems (ADAS) used for tasks such as lane changing, speed control, and parking. These advancements not only enhance driver convenience and comfort but also offer hope for public safety, infrastructure, and vehicle durability by reducing accidents. Furthermore, autonomous driving technology is expanding into various robotic applications, such as space probes, industrial robots, military drones, and delivery robots, addressing concerns about efficiency, cost, quality, and environmental impact. For example, the e-commerce industry can benefit from the use of autonomous delivery robots, which improve the efficiency, cost, quality, and environmental impact of traditional delivery methods.
[0011] Despite decades of research into the development of autonomous vehicles, there are still no fully autonomous vehicles available for personal use on the market. Waymo has made progress with its autonomous fleet, but so far, it has only been used for taxi services. Despite significant progress, safety and reliability remain lacking. Traditional autonomous driving systems, characterized by aggregates of independent sub-modules, are challenging due to the vast amounts of data required to train these models. Furthermore, the manual labeling of this data required for configuring AI systems for traditional autonomous driving is expensive. Many data formats required by traditional autonomous driving systems (such as pre-built maps) are not only costly to construct and label, but also pose risks to safety and general applicability due to their limited responsiveness in situations where the real-world environment is not correlated with the expected map.
[0012] The drawbacks associated with traditional methods create opportunities for the development of end-to-end (E2E) learning approaches for autonomous driving. E2E autonomous driving typically consists of a single, self-contained deep learning model that maps sensed inputs, such as image frames from a camera or maps generated via light detection and ranging (LiDAR), to steering wheel and accelerometer actuations for vehicle control. E2E autonomous driving systems and methods can be configured to learn via reinforcement learning methods, such as imitation learning, rather than relying on manually designed task aggregations. Successfully training E2E autonomous driving methods using imitation learning must overcome certain challenges, such as the varying quality of human agent driving actions, the difficulty of validating and testing E2E models, scalability, and the implementation of regulatory safety standards.
[0013] Despite regulatory hurdles, supply chain feasibility, and consumer skepticism regarding fully autonomous driving, the development of improved semi-autonomous ADAS technologies continues. The advantages of E2E autonomous driving methods are easily translated into semi-autonomous driving methods. Therefore, there is a desire for ADAS technologies that are compatible with the increasing automation of driving tasks without requiring significant changes to hardware or software components. Opportunities have emerged for collision avoidance systems (CAS) and other ADAS features that utilize E2E neural networks configured for autonomous driving tasks. Summary of the Invention
[0014] The disclosed technology relates to systems and methods for providing driver assistance alerts to drivers. The disclosed technology may include receiving a series of environmental data related to driving conditions, including at least video from a camera, return data from an optical sensor, and position data from a GNSS receiver. The camera, the optical sensor, and the GNSS receiver are coupled to a processor carried by the vehicle. The technology includes processing the environmental data as input to an end-to-end neural network, training the end-to-end neural network to generate specified steering and speed control actions in response to the current driving condition. It can be extended to analyzing hidden layer data and output data from the end-to-end neural network to estimate collision avoidance data. This real-time collision avoidance data includes at least one or more detected objects in the video from the camera, directional cues, and a risk measure based at least in part on the dissimilarity between the generated specified steering and speed control actions and received driver steering and speed control actions. The directional cues may be projected onto a head-up display. Other driver assistance alerts may also be generated in real time based on the collision avoidance data.
[0015] Specific aspects of the disclosed technology are described in the claims, description, and drawings. Attached Figure Description
[0016] The accompanying drawings are for illustrative purposes and are intended only to provide examples of possible structures and processes of one or more embodiments of this disclosure. These drawings are in no way intended to limit any changes in form and detail that may be made by those skilled in the art without departing from the spirit and scope of this disclosure. A more complete understanding of the subject matter can be derived by referring to the detailed description and claims when considered in conjunction with the following drawings, wherein similar reference numerals throughout the drawings refer to similar elements.
[0017] Figure 1 This is an architectural diagram of an end-to-end conditional imitation learning model for autonomous driving.
[0018] Figure 2 illustrates examples of multiple possible driving states within a trajectory according to certain embodiments of this disclosure.
[0019] Figure 3 is an architecture-level schematic diagram of an end-to-end conditional learning model for autonomous driving including a memory-enhanced converter, according to certain embodiments of the present disclosure.
[0020] Figure 4 It is a flowchart describing the process of determining when to initiate autonomous vehicle control in response to high-risk driving scenarios.
[0021] Figure 5 This is a schematic diagram illustrating the generation of collision avoidance data from a specified output of an end-to-end autonomous driving model according to certain embodiments of this disclosure.
[0022] Figure 6A , 6B The first example of a collision avoidance system graphical user interface is shown in 6C and 6D.
[0023] Figure 7A , 7B The second instance of the collision avoidance system is shown in 7C and 7D.
[0024] Figure 8 This describes computer systems that can be used to implement the disclosed technologies according to certain embodiments of this disclosure. Detailed Implementation
[0025] The following detailed description is based on the accompanying drawings. Sample embodiments are described to illustrate the disclosed technology, but not to limit its scope, which is defined by the claims. Those skilled in the art will recognize various equivalent variations described below.
[0026] Researchers across academia, government, and industry have long focused on developing semi-autonomous and autonomous driving technologies. The pursuit of machine automation dates back to pre-autonomous times, exemplified by watermills, windmills, and Leonardo da Vinci's self-propelled cart. In response to the development of automobiles, imaginative ideas about autonomous vehicles have emerged. Advances in robotics, cameras and sensors, network infrastructure, and artificial intelligence (AI) have driven the evolution of autonomous and semi-autonomous driving technologies, including so-called Advanced Driver Assistance Systems (ADAS).
[0027] Modern vehicles frequently incorporate ADAS features such as Collision Avoidance Systems (CAS), lane keeping assist, and dynamic cruise control. ADAS features increase usability, especially for the elderly, people with mobility impairments, and those with sensory disabilities. Improving and increasing the adoption of ADAS can alleviate several sources of traffic congestion, such as suboptimal driving behavior and collisions that block roads. Furthermore, ADAS technology offers significant safety advantages. The World Health Organization estimates that approximately 1.35 million people die in car accidents, and the U.S. National Highway Traffic Safety Administration reports that up to 94% of serious accidents are attributed to user error.
[0028] Despite the rapid advancements in ADAS, challenges persist, including the scalability, cost-effectiveness, and feasibility of implementing AI-based ADAS training and validation. Traditional ADAS features highly specialized components that work collaboratively within complex systems to fully or partially automate driving tasks. AI-enhanced ADAS can surpass previous practices that relied on simpler algorithms, such as relay-type (bang-bang) controllers. However, AI-based approaches require vast amounts of training data, which is expensive and time-consuming to obtain. Many training techniques that promise to improve the usability and accuracy of ADAS are still in early stages of development, such as semantic segmentation in computer vision.
[0029] Emerging E2E learning approaches offer a scalable and efficient alternative to AI-based ADAS systems. E2E autonomous driving, in this context, is understood to involve a single system, such as a deep learning classifier, configured to automatically process sensor inputs, such as camera images, and generate actions to control vehicle actuators, such as steering and acceleration / braking. For further information regarding the training and validation of E2E deep learning models configured for autonomous driving, see the co-owned U.S. patent application cited above in the Relevant Applications section.
[0030] The proposed E2E approach to ADAS is advantageous in terms of its potential for seamless integration into increasingly autonomous functionality without requiring inconvenient and costly updates to existing sensor hardware.
[0031] The disclosed technology provides a system and method for providing both active and passive ADAS to a driver. The E2E driving model can be used for fully autonomous driving, or alternatively, as an active, standby ADAS feature. While the driver maintains control of the vehicle, the E2E driving model operates in shadow mode until an impending risk is detected. The E2E driving model automatically takes over to mitigate the risk (e.g., collision avoidance).
[0032] E2E driving models can also be used in passive ADAS. Active ADAS features can override or correct driver actions, while passive ADAS features assist drivers in making safer decisions. The disclosed technology includes presenting driver alerts and warnings notified by designated outputs from an E2E driving model operating in shadow mode. In various embodiments, one or more of the following can be conveyed to the driver to aid in informing driving decisions: designated outputs from the E2E driving model, approaching objects detected by the E2E driving model, and / or directional cues indicating recommended steering wheel orientation to avoid road hazards or follow a planned route. Driver alerts can include visual, auditory, and tactile alerts. Many embodiments of the disclosed technology include a head-up display that presents driver alerts without requiring the driver to take their eyes off the road. One form of display can be a heatmap of objects identified as important by the E2E driving model projected onto the display. Another is a steering guidance directional cue, which can be presented on the display as dynamic “whisker” arrows indicating the corresponding current and / or recommended steering wheel orientation. The estimated risk level of the current driving action can also be visually displayed to the driver. Passive ADAS alerts can also be used for driver feedback or driver education based on specified outputs and driver behavior.
[0033] End-to-end imitation learning model for autonomous driving
[0034] Despite the focus on developing autonomous vehicle fleets by resources, funding, and public interest, fully autonomous driving technology has yet to emerge in a form that meets the safety and feasibility requirements for deployment in homes. Technological advancements have led to a range of semi-autonomous driving technologies. These achievements possess varying degrees of accuracy and reliability, such as the safety features of lane-keeping assist and collision prevention systems, features that have become commonplace in modern vehicles and in controversial semi-autonomous driving functions that allow drivers to partially relinquish steering and acceleration decisions. However, autonomous driving technology still faces significant hurdles in terms of scalability and safety. As defined by the SAE, the degree of driving automation applicable to a vehicle is:
[0035] Level 0 – No automation and entirely controlled manually by humans.
[0036] Level 1 – The vehicle possesses a single automated system feature, such as cruise control.
[0037] Level 2 – For tasks such as steering and acceleration, partial automation is achieved through an advanced driver assistance system, where the partially automated tasks are fully monitored by the driver, who can intervene at any time.
[0038] Level 3 – Conditional automation in response to the detection of appropriate environmental factors, involving some human override control.
[0039] Level 4 – High level of automation, characterized by fully automated driving where the driver can still exercise overriding control when necessary.
[0040] Level 5 – Full autonomous control of the vehicle under all conditions, without human intervention.
[0041] While rare instances of Level 3 and Level 4 autonomous vehicles exist, at the time of this disclosure, mainstream production for consumer use has not exceeded Level 2. Further advancements in safety and reliability must be demonstrated to support further progress. Safety standards applicable to the functionality and performance of autonomous vehicles are primarily indicated by standards from the American National Standards Institute (ANSI) and ISO standards from the International Organization for Standardization.
[0042] The evaluation of autonomous driving products (including vehicles) utilizes the ANSI / UL 4600 safety standard. ANSI / UL 4600, the first widely adopted safety standard applied to the operation of autonomous vehicles, evaluates fully autonomous driving products independent of human supervision. ANSI / UL 4600 establishes broad, technology-neutral safety guidelines for risk analysis, data integrity, autonomy verification, lifecycle resilience, and compliance assessment. In contrast, ISO standards such as ISO 26262, ISO 21488, and ISO / SAE 21434 define safety requirements specific to autonomous vehicle technologies. ISO 26262 assesses the functional safety of electrical / electronic systems in vehicles, particularly safety management in the event of system failure or malfunction. ISO 21488 covers Safety of Intended Functionality (SOTIF), which addresses unintended system behavior in the absence of ISO 26262 system failures. ISO / SAE 21434 covers cybersecurity risk management from the conceptual design, development, and manufacturing processes of road vehicles, through operation, maintenance, and decommissioning phases.
[0043] The primary obstacle hindering satisfactory compliance of autonomous vehicles with the aforementioned safety standards is scale. A vast amount of data is necessary for autonomous vehicles to be safe, reliable, and scalable to complex and dynamic driving scenarios. While data availability limitations are not solely responsible for all remaining technology gaps, data requirements are closely linked to all aspects of autonomous driving system development. Improving technology areas related to autonomous vehicles, such as computer vision and sensor development, are limited in their growth potential without sufficient data, both in quantity and variety, available for learning. Furthermore, the additional technological scalability challenges related to time, cost, and resources cannot be addressed without the information available to do so.
[0044] Traditional autonomous driving technologies and end-to-end learning methods heavily rely on artificial intelligence and deep learning systems to assess the environment, predict future changes in the environment, and make decisions in response to the environment. The development of robust and generalizable deep learning models capable of learning complex feature spaces and patterns depends heavily on abundant data for training, validation, fine-tuning, and further evaluation / testing processes.
[0045] The autonomous driving system and method described in this disclosure can solve this problem using an E2E approach. The E2E architecture of the disclosed system improves scalability by reducing reliance on up-to-date, highly complex map data and is configured to handle driving conditions not previously seen during training. Using an E2E approach is significantly more efficient in terms of data usage and computational cost, partly due to configuring deep learning models to extract useful features directly from the input data and directly translating input data processing into driving actuation.
[0046] However, scalability concerns are not fully addressed by the improvements provided through implementing an E2E approach. Sufficient data is still needed for both the learning and validation processes, providing adequate training for rare and challenging scenarios. For example, extreme weather, very close-range collisions, and corner cases where pedestrians or stray objects spontaneously block the road are rare events where sufficient training data is difficult to obtain, but learning these scenarios is crucial for autonomous driving models due to the significant safety risks and potential consequences if mishandled. Beyond the obvious ethical importance, safety standards such as the SOTIF guidelines within ISO 21448 essentially focus on assessing risk levels in response to hazardous events.
[0047] In contrast to extreme driving scenarios that typically refer to rare and potentially dangerous situations, it is also crucial to ensure that the model can adequately handle edge cases. While edge cases often overlap with extreme cases, they present unique challenges to the computational system compared to human drivers. For example, autonomous vehicles may respond poorly to edge cases due to limitations in computer vision technology or highly personalized scenarios. Although situations like heavy rain or congested school drop-off / pick-up lanes are typically stressful or challenging for human drivers, the complexity is amplified for autonomous driving models that may not generalize from simple driving to routinely encountered stressful situations.
[0048] The difficulty of training autonomous driving models not only to become familiar with a sufficiently diverse range of driving scenarios but also to extend to unfamiliar scenarios can be addressed using reinforcement learning and imitation learning methods. By utilizing driving demonstrations performed by human drivers, autonomous driving models, such as the E2E system disclosed in this paper, can be trained to learn feature distributions, feature patterns, and overall behavioral strategies. This enables the model to handle driving scenarios and determine a plan of optimal action in response to input data from the environment that is independent of prior exposure to scenarios, locations, or routes.
[0049] Further details regarding how the disclosed system and method can provide solutions to scalability and performance challenges by combining the advantages of E2E learning and imitation learning strategies with additional deep learning methods that implement context-aware learning and scalable approaches for data collection, training, and validation can be found in the co-owned U.S. patent application listed above in the relevant application filings. The architecture of the disclosed E2E model will now be discussed in further detail.
[0050] System Architecture
[0051] Figure 1 This is an architecture-level schematic diagram 100 of an end-to-end conditional imitation learning model 101 for autonomous driving. The conditional imitation learning model 101 is illustrated within schematic diagram 100 according to an exemplary embodiment of the disclosed technology, including a converter architecture. At a high level, the conditional imitation learning model 101 processes states corresponding to the driving environment. 102. Environmental data is used to specify appropriate response actions 124. Specified response actions may include actuating the steering wheel and accelerator / brake, which changes the vehicle's speed 124a, orientation 124b, and thus position 124c.
[0052] In fully autonomous driving mode, the vehicle's operation is controlled via one or more actuators to execute a specified response action 124. In semi-autonomous mode, the conditional imitation learning model 101 operates in a so-called "shadow mode." The human driver manually operates the vehicle while the conditional imitation learning model 101 operates in the background, generating outputs as instructed by the specified response action 124 without actuating the brakes, steering wheel, etc. The specified response action 124 is used to notify one or more ADAS functions of the vehicle. For example, a passive ADAS tool can utilize the conditional imitation learning model 101 for CAS purposes by warning the driver to brake when the specified response action 124 requests brake actuation, thereby mitigating the risk of an impending collision with another object. An active ADAS tool can utilize the conditional imitation learning model 101 for CAS purposes by performing autonomous braking when the specified response action 124 is triggered. This braking actuation mitigates the risk of an impending collision with another object. Further reference Figure 4 The use of the conditional imitation learning model 101 for collision avoidance is extended. The input data and architecture of the conditional imitation learning model 101 will now be described in further detail.
[0053] Input status 102 is represented by observations including image 102a and multiple non-camera environmental data 202b (e.g., LiDAR and GNSS data). In addition to describing the state... In addition to the observation 102, some embodiments also provide an indication condition 102c (e.g., GPS indication guiding the vehicle along a intended route). Therefore, the conditional imitation learning model 101 predicts an action (e.g., braking) in response to a state (e.g., rapidly approaching the rear of another vehicle). In another embodiment, the conditional imitation learning model 101 predicts an action (e.g., turning the vehicle right) in response to a state including an indication condition (e.g., a route indication to turn right onto a perpendicular street to stay on the navigation route) (e.g., approaching an intersection with another perpendicular street). A conditional route may refer to a route that directs the vehicle to a specific destination end position, or a shorter conditional route, such as a driving route for the next three, five, or ten seconds. In some embodiments, the route is based on a static destination end position. In other embodiments, the destination end position may be dynamic and shift in response to previous route progress.
[0054] Besides corresponding to the current state In addition to the data in 102, compressed memory data is extracted from the storage device in the frame buffer, which contains information corresponding to several previous states in a given trajectory. For simplicity and clarity, schematic diagram 100 illustrates a total of five previous memory frames, i.e., compressed memory states. 122. State via compressed storage 142. Status after compression storage 162. State via compressed storage 182, and the state after compression storage. 192. However, in many embodiments of this technology, more than five previous memory frames are stored in the frame buffer as input to the current state, such as ten, fifteen, twenty, or more previous memory frames. These memory frames may cover two, three, five, or more seconds of history at a frame rate lower than that of standard video capture. A memory frame refers to a “snapshot” or latent representation of a previous state processed by the conditional imitation learning model 101. The generation and storage of compressed memory states into the frame buffer will be further elaborated later in the discussion of the converter architecture of Figure 3. As previously described, the segmentation of driving state data into states within a trajectory is variable. In some embodiments of the disclosed technology, the number of states corresponds to the current state. At least three seconds of history before 102.
[0055] Before the second processing phase performed by the conditional imitation learning model 101, the state The observation data of image 102 undergoes preprocessing by preprocessor module 103 in the first-stage processor phase. Preprocessor 103 embeds corresponding input data from image 102a, non-camera environment data 102b, and indication condition 102c. In some embodiments, image 102a undergoes image processing specific to deep learning analysis of image data, as indicated by the dashed shading of units within preprocessor 103 adjacent to image 102a. In some embodiments, this image processing is performed by a convolutional neural network. In one embodiment, the processing model responsible for preprocessing the data contained in image 102a is a pre-trained module that has been transferred or fine-tuned for use in conjunction with conditional imitation learning model 101. In some embodiments, the preprocessing of image 102a or other spatially mapped data (e.g., LiDAR data) involves generating location embeddings that maintain the integrity of the location information corresponding to the data. In one embodiment, the location embedding data can later be used to construct heatmaps, such as attention maps using attention weights from conditional imitation learning model 101, for object detection purposes to inform CAS operations. For example, constructing an attention map using positional embeddings from image data 102a and attention weights from transformer 104 enables the detection of regions within the camera view that are considered important by the conditional imitation learning model 101 when generating a specified output action 124 including brake actuation. When the attention map is projected back onto image data 102a, important regions may overlap with pedestrians walking in a crosswalk or another vehicle ahead. Terms related to the “importance” of an input feature or token, or the “attention” to said input feature or token, are terms that will be readily recognized by those skilled in the art.
[0056] Following preprocessing, the processing stack, including the conditional imitation learning model 101, uses a transformer 104 and a compression layer 106 to process the embedded output from the preprocessor 103 and the compressed memory states 122, 142, 162, 182, and 192. The compression layer 106 of the described processing stack generates input states. The memory frame of layer 102. In other words, the output of compression layer 106 is the input state. 102's compressed memory state 108. Compressed memory state 108 will use a FIFO (First-In-First-Out) stored procedure to store the data in the frame buffer, so that during processing... At that time, the frame buffer will contain frames corresponding to the states respectively. , , , and The compressed memory represents 108, 122, 142, 162, and 182.
[0057] In response to input status 102 Generates predicted response action 124, which is then processed via compressed memory. 108 is processed by classification head 110 in the third-level processor to generate a specified response action 124. Specifically, this refers to the compressed memory state. 108 is processed to generate actuation of the steering wheel and accelerator / brake, which can change the vehicle's speed 124a, orientation 124b, and thus position 124c.
[0058] The conditional imitation learning model 101 is an end-to-end autonomous driving model that can be used for full automation of vehicle operation by executing specified response actions 124. Furthermore, driving automation via the conditional imitation learning model 101 can be translated into partial automation of vehicle operation, such as Level 2 or Level 3 automation tasks. As described above, the conditional imitation learning model 101 can operate in shadow mode. When the conditional imitation learning model 101 operates in shadow mode, the specified response actions 124 can be used for ADAS purposes. In passive ADAS, driver assistance alerts and warnings can be generated based on the specified response actions 124. In active ADAS, task automation may involve executing the specified response actions 124. In one implementation, the specified response actions 124 are processed to calculate a risk level associated with the current driving state to inform a decision on whether manual or autonomous driving is more appropriate. Further References Figure 4 and 5 The expansion covers various implementation schemes involving passive ADAS, active ADAS, and risk calculation, with example use cases in... Figure 6A Presented in D and 7A to D.
[0059] To establish a foundation for certain implementations of the systems and methods disclosed herein, a deep learning framework will now be described in further detail. The disclosed conditional imitation learning model 101 is an end-to-end autonomous driving model that uses imitation learning and memory enhancement techniques to mimic the driving behavior of a skilled and safe human driver. Imitation learning and memory enhancement are briefly described below. For a more detailed description of systems and methods for scalable training and validation of end-to-end autonomous driving models (e.g., conditional imitation learning models), refer to the co-owned U.S. patent application cited above in the relevant application section.
[0060] First, a high-level introduction to the concepts of imitation learning related to state-action pairs within a driving trajectory and driving behavior strategies learned from the driving trajectory is provided. Next, the discussion turns to the disclosed memory-enhanced transducer. In contrast to processing with instantaneous state data lacking recent past information to provide useful context for driving decisions, the memory-enhanced transducer generates a specified action in response to a processed state containing previous states within the trajectory.
[0061] Imitation learning for autonomous driving
[0062] A conditional imitation learning model 101 is trained using a large database of driving demonstrations to estimate driving behavior strategies. The driving demonstrations may include one or more driving tasks, such as lane merging, handling four-way parking, and sudden braking in response to an impending collision risk. The driving tasks within the demonstrations can be performed under various environmental conditions, including different times of day, weather conditions, road and highway types, and traffic levels. Driving demonstrations can be collected from convoys of vehicles manually operated by human drivers, autonomous vehicles, or driving simulations. The driving demonstrations are explained in further detail below with reference to Figure 3, followed by a discussion of model training methods for the imitation learning model used in autonomous driving.
[0063] Conditional imitation learning model 101 learns driving behavior policies from training driving demonstrations. A driving behavior policy is a probability distribution of actions given a state, also known as a state-to-action mapping. Driving behavior policies mimic the subjective rules and logic humans apply to driving. The trained imitation learning model responds to a state (e.g., a red light in video input) using the estimated behavior policy embodied in the model coefficients to compute a specified action (e.g., braking). Imitation estimations of such policies from large training datasets (e.g., hundreds of thousands or millions of instances) inevitably outperform lists of driving rules. Furthermore, a sufficiently comprehensive list of driving rules would be infeasibly long and complex. Instead, using imitation learning models to estimate behavior policies is better equipped to extract this complexity from driving demonstration instances.
[0064] A well-trained imitation learning model trained on a large number of driving demonstrations can be generalized to a wide range of driving states and environmental stimuli, both experienced and not experienced by the model during training, because the model has learned the basic principles of driving decisions rather than memorizing specific actions to be performed in response to a particular state. As the title suggests, the disclosed conditional imitation learning model 101 is trained to predict actions in response to a state based on conditions that restrict the vehicle's driving trajectory. For example, an autonomous vehicle may need to make a left turn at an intersection to stay on the intended route (conditions restricting the vehicle's driving trajectory). Based on the learned behavioral policy estimated by the trained conditional imitation learning model 101, the autonomous vehicle knows to yield to oncoming traffic before making the left turn. Furthermore, if the road conditions are icy, the autonomous vehicle will adjust its deceleration and acceleration rates during the turn accordingly, as instructed by the learned behavioral policy. Terms such as "action," "state," "trajectory," and "environment" will be used herein according to their meaning as understood in the field of conditional imitation learning rather than other common meanings. The terminology of conditional imitation learning is further elaborated below with reference to Figure 2.
[0065] Figure 2 illustrates an example 200 of multiple possible driving states within a trajectory according to certain embodiments of this disclosure. Example 200 occurs within the context of a driving environment 202. The driving environment 202 includes a specific layout of roads with various orientations and relevant legal guidelines for the use of roads, surrounding structures and objects, weather and atmospheric conditions, other vehicles, pedestrians, and vehicle 212. Those skilled in the art will recognize that a driving trajectory, or simply a trajectory, refers to the driving route of vehicle 212. The environment 202 contains a large, nearly infinite number of possible trajectories equivalent to the total combination and arrangement of possible routes that can be taken within the environment, each of which may be very long. The actual trajectory taken by vehicle 212 can be described by a series of states represented by the illustrated numerical lines.
[0066] The digital lines described are in their current state. Centered on 227. Current state. The sequence preceding 227 is a series of previous states, including the state... 226. Status 225. Status 224. Status 223. Status 222. Status 221 and status 220. Current Status Following 227 will be a series of future states, including states 228. Status 229. Status 230. Status 231. Status 232. Status 233 and status 234. The trajectory of vehicle 212 is also indicated by a gray dashed arrow in the schematic diagram of environment 202, indicating that vehicle 212 is about to turn left at the approaching intersection. As shown in the schematic diagram of environment 202, vehicle 212 is in a state 220 is approaching the stop sign and is currently in the following state. There are 227 arrival stop signs. If the future status quo is implemented as expected, then the vehicle will be in the status quo. Turn left at point 234 and position yourself in the middle of the described intersection.
[0067] When vehicle 212 is in the current state At point 227, a vehicle traveling along the intersecting street and crossing the indicated intersection did not encounter a stop sign and could proceed directly. Therefore, vehicle 212 is expected to comply with the stop sign and is in the current state. At point 227, come to a complete stop, give way to crossing vehicles, and remain completely stopped at the stop sign until there is sufficient clearance to safely begin a left turn to avoid a collision with crossing vehicles.
[0068] Within instance 200, two trajectories are described: trajectory 200.1 and trajectory 200.2. Trajectory 200.1 is described with reference to three representative driving states within the trajectories, including the previous state. 220.1 Current Status 227.1 and future status 234.1. Within each driving trajectory, the time occurring... Each state It can be determined by time This describes multiple features associated with the vehicle and / or its surrounding driving environment. This data is collected by hardware coupled to the vehicle, such as one or more cameras and / or LiDAR sensors. For example, in the state... At 220.1, sensors coupled to vehicle 212 record environmental data within environment 202, such as a stop sign ahead and another vehicle approaching the intersection via intersecting streets.
[0069] responsive actions Responds to each state And execute. Respond to the state. 220.1, the operator does not perform any action to change the steering wheel angle from the 0° neutral position, but instead applies the brakes to properly comply with the stop sign. For convenience and clarity, reference angle measurements are used to describe steering wheel orientation, where 0° indicates the vehicle is traveling straight forward, a positive angle of +45° indicates a right turn, and a negative angle of -45° indicates a left turn. In some embodiments, the training data will further include information about the operator's eye movements using retinal tracking data.
[0070] In the current state of trajectory 200.1 At point 227.1, vehicle 212 has approached the boundary of the area designated as the appropriate location for stopping according to the stop sign. (In condition) At 227.1, the sensor coupled to vehicle 212 indicates that the vehicle has reached the stop sign, and as the crossing vehicle crosses the intersection along the intersecting street, the crossing vehicle is now directly in front of vehicle 212. Therefore, the operator of vehicle 212 does not perform any actions to change the vehicle's speed or steering orientation. When vehicle 212 reaches the future state... At 234.1, the sensor coupled to vehicle 212 indicates that the vehicle is now within the intersection while performing its turn. In response to the state... 234.1, Operator performs actions This includes turning the steering wheel to the left at an appropriate angle, such as -45°, and accelerating away from the turn to return to the normal speed limit once the vehicle 212 has completed the turn.
[0071] Driving trajectory 200.1 represents a series of safe state-action pairs applicable to a given driving task, but many other trajectories are possible within environment 202. Furthermore, the trajectory is not predetermined. With the first state... Transition to a new state There exists a response to a state. A possible set of total actions. The choice of action determines the state transition and the resulting state. The variables. Many embodiments of the disclosed technique involve using an estimated behavioral policy learned by a conditional imitation learning model 101 to generate a specified action in response to a state, and the execution of the specified action will naturally affect the properties of subsequent states. Therefore, it is not guaranteed that two different trajectories will continue to overlap simply because they begin with overlapping states and actions. As an extension, trajectories may converge or diverge from each other at any state or at any point in time. To help illustrate the transition from one driving state to the next driving state, and the relationship between the executed action and the resulting state transition, a second trajectory 200.2 is also provided in example 200.
[0072] Trajectory 200.2 represents a series of unsafe, catastrophic state-action pairs. Its reference trajectory includes three representative driving states, encompassing the previous state. 220.2 Current Status 227.2 and future state 234.1 Description. In state At point 220.2, the state and actions are similar to the state in trajectory 200.1. The state and actions within trajectory 220.1 are the same, so details will not be elaborated here. However, trajectory 200.2 is in a different state. The trajectory at point 227.2 diverges from the trajectory at point 200.1. The operator is not adhering to the status. The stop sign at 227.2 does not allow vehicles crossing the intersection to safely pass before the operator executes its turn. Instead, the operator responds to the status... 327.2 The action performed includes initiating a left turn without stopping.
[0073] Therefore, vehicle 212 is in state At some point after 227.2, a collision occurs with a vehicle crossing the intersection, causing a catastrophic failure and prematurely ending the trajectory. Because the trajectory ends with the collision event, trajectory 200.2 will not reach its future state. 234.2.
[0074] Actions that lead to catastrophic failures are not the only cause of divergence between different trajectories. Within Instance 200, two trajectories share the same intended route, driving task, and objective, differing only in the quality and success of their execution. In a driving situation where one driver turns left and the other turns right, two previously overlapping driving trajectories may diverge.
[0075] In one embodiment of the disclosed technology, the conditional imitation learning model 101 can be trained to clone the behavior of an operator within trajectory 200.1, with the aim of successfully simulating the safe driving techniques demonstrated by the operator of vehicle 212. In another embodiment, the conditional imitation learning model 101 is trained using reinforcement learning methods with trajectory 200.1 as a positive example, where similarity to actions within trajectory 300.1 is rewarded, and trajectory 200.2 is used as a negative example, where similarity to actions within trajectory 200.1 is penalized. Other alternative training methods can also lead to the conditional imitation learning model 101 successfully learning the safe turning behavior demonstrated by instance 200.
[0076] The conditional imitation learning model 101 is trained on driving demonstrations (such as the trajectory shown in example 200) and more complex scenarios to learn what can be applied to the current driving state. The environmental data 102a, 102b and condition 102c are used to formulate behavioral strategies and generate specified response actions. 124, at this point the trajectory changes to the subsequent driving state. .
[0077] Given the influence and persistent effects of previous states and actions on future trajectory states, modeling the local and global dependencies between trajectory states is highly beneficial when predicting driving behavior. Therefore, autonomous driving models are at a disadvantage if they cannot store any prior information in a memory cache or retrieve such information for context-aware decision-making. Deep learning architectures configured to enable memory-notified predictions, such as recurrent neural networks or multi-head attention mechanisms in transformer models, are typically computationally expensive. Given the complexity of the autonomous driving problem, recurrent neural networks or multi-head transformers may operate slower than expected. In some implementations of the disclosed techniques, the complexity of the specific learning problem and / or available computational power guarantee the use of these models. However, in most cases, the associated time, monetary, and computational costs of these models can be very high and may affect the model's ability to achieve safety standards, such as the SOTIF guidelines described in ISO standard 21448 previously.
[0078] The discussion now turns to introduce memory-enhanced converters, which utilize input enhancements with memory-cached data to enable the use of local and global dependency patterns within a driving trajectory, achieving improved efficiency compared to traditional recurrent neural networks or converter models.
[0079] Memory Enhancement Converter
[0080] Figure 3 is an architecture-level schematic 300 of an end-to-end conditional learning model 101 for autonomous driving including a memory-enhanced converter according to certain embodiments of the present disclosure. Schematic 300 is equivalent to schematic 100, wherein processing for four separate time steps is illustrated in a so-called unfolded state. In contrast to multi-head converter models configured to repeatedly process large amounts of input data corresponding to multiple consecutive states within a trajectory, the memory-enhanced converter illustrated in schematic 300 utilizes a first-in-first-out frame buffer, which stores cached memory states of previously processed states in the trajectory. Each frame or memory state within the frame buffer contains a compressed latent space representation of the corresponding previous state generated by the compression layer 106 of schematic 100. Given the compression layer 106 for specific states... Process to receive status The corresponding compressed representation includes information about the previous state. Processing n frames within a frame buffer, assuming a constant frame buffer size, in response to Predicted actions It is in response to indicating the previous state It is generated from compressed data.
[0081] Schematic diagram 300 at time points Data processing begins, and is based on time points. The processing is complete. Assume the state... A set of observations 302 is the earliest state processed in the trajectory, therefore at time point... No frame is currently stored in the frame buffer. Preprocessor 103 embeds data from... The data from 302 is then processed by converter 104 and the state is generated by compressor 106. The compressed memory state is represented as 304. State The compressed storage state representation 304 is processed by the classification head 110 to generate predicted actions. 306.
[0082] For time points Preprocessor 103 is embedded from Data from 322. Besides data from the state. In addition to the embedded data, the frame buffer now stores the state. The frames of the compressed memory state 304 are processed by converter 104, and then by compressor 106 to generate the state. The compressed memory state is represented by 324. State The compressed storage state representation 324 is processed by the classification head 110 to generate predicted actions. 326.
[0083] For time points Preprocessor 103 is embedded from Data from 342. Besides data from the state. In addition to the embedded data, the frame buffer now stores the state. The compressed memory state 304 and state The compressed memory states represent the corresponding frames of both 324. These combined inputs are processed by converter 104, and then by compressor 106 to generate states. The compressed memory state is represented as 344. State The compressed storage state representation 344 is processed by the classification head 110 to generate predicted actions. 346.
[0084] For time points Preprocessor 103 is embedded from 362's data. Besides data from the status. In addition to the embedded data, the frame buffer now stores the state. The compressed storage state 304, state The compressed memory state representation 324 and state The corresponding frames of the compressed memory state representation 344 are then processed by converter 104, and subsequently by compressor 106 to generate the state. The compressed memory state is represented by 364. State The compressed storage state representation 364 is processed by classification head 110 to generate predicted actions. 366.
[0085] If the illustration is intended to demonstrate the future time steps of the memory-enhanced converter model, then the frame buffer will eventually reach its capacity and begin to discard the oldest frames, one at a time, every time step, to make room for the storage of the newest frames.
[0086] The system executes specified actions generated by the conditional imitation learning model 101 to control the autonomous vehicle. For partially autonomous vehicles, the conditional imitation learning model 101 can operate in shadow mode. When operating in shadow mode, it references... Figure 1 The process described in section 3 is still performed, but the specified actions generated by the model are not automatically executed by the vehicle's actuators. Instead, the human driver maintains manual control of the vehicle, while ADAS uses the model-generated output to provide driver alerts.
[0087] Advanced driver assistance system with end-to-end AI
[0088] The disclosed conditional imitation learning model 101 can be used to implement additional advanced driver assistance systems (ADAS) within autonomous or semi-autonomous vehicles. In some embodiments, the ADAS configured to act as a collision avoidance system can be designed to emit warning signals (auditory and / or visual notifications) to the human agent operating the vehicle in response to the detection of a predicted dangerous driving state. In one embodiment, the detection of potential hazards by the conditional imitation learning model 101 (or a separately trained model associated with the conditional imitation learning model 101) is performed in response to interaction with an accelerator / brake actuator deviating from a specified speed. In another embodiment, the detection of potential hazards by the conditional imitation learning model 101 (or a separately trained model associated with the conditional imitation learning model 101) is performed in response to processing characteristics of one or more driving states, such as the detection of an object very close to or rapidly approaching the vehicle, changing traffic signals, or lane departure.
[0089] In some implementations, advanced driver assistance systems (ADAS) configured via data collection, learning, and statistical analysis performed in association with the methods and systems disclosed herein can be implemented in semi-autonomous vehicles to initiate a transition from manual to autonomous control, or vice versa. In one instance, an ADAS (e.g., CAS response) can be configured to respond to a predicted collision by overriding manual control of the vehicle and initiating automatic braking (i.e., in response to an object very close to the vehicle or a lack of response to traffic signals by the operator). In another instance, a so-called "adaptive cruise control" system can be configured to respond to vehicles exceeding predefined permissible thresholds for object proximity (e.g., a predefined distance between the operator's vehicle and a single vehicle directly in front of the operator's vehicle, such as a minimum distance of thirty feet, fifteen meters, or two vehicle lengths between vehicles) or speed (e.g., a predefined speed limit for the vehicle, such as 80 mph, 7 mph higher than the currently detected speed limit, or a speed 10% higher than the currently detected speed limit).
[0090] In the third example, the advanced driver assistance system (ADAS) can provide the operator with a series of statistical analyses of the vehicle's current performance and behavior, independent of the degree of autonomous vehicle operation. These analyses can be useful to the operator in terms of driving behavior, safety warnings and feedback, or potentially necessary vehicle maintenance. Due to the high-dimensional, high-volume data collected by the conditional imitation learning model 101, this data (e.g., risk metrics or suggested adjustments to driving behavior) can be more informative than typical driving data presented to the operator. The analysis may include data related to the frequency, speed, and maneuvering trends of high-risk driver behaviors (e.g., sharp turns, unsafe lane changes, etc.), the frequency of close-range collisions, and / or events requiring autonomous control to avoid collisions.
[0091] In the above-described example implementations, and in several other scenarios where those skilled in the art will recognize that an advanced driver assistance system (ADAS) can be implemented within the technology disclosed herein, the ADAS can utilize driving demonstration data from both human and autonomous driving agents, as well as any correlation analysis used in training, validation, fine-tuning, or transfer learning. The ADAS can also utilize pattern recognition and risk analysis data extracted from the trained autonomous driving model, as well as external data input through further expert feedback and / or computational analysis of the data extracted from the trained autonomous driving model. Furthermore, the ADAS can utilize data and data analysis obtained from driving trajectories executed after the initial model deployment. In operation, the vehicle continues to collect and monitor data from the operator after the deployment of the trained autonomous driving model. The collected data can be used both to correct the operator's actions and to further fine-tune the model.
[0092] Figure 4 This is a flowchart describing a semi-autonomous process 400 for determining when to initiate autonomous vehicle control in response to high-risk driving scenarios. Process 400 includes a comparison of collected data (including the model's designated response action 112) from a conditional imitation learning model 101 operating in shadow mode with manual driver response actions 402. In this semi-autonomous driving mode, the driver manually controls the vehicle by executing driver response actions 402, including speed control 402a, steering control 402b, and optionally following a direction based on navigation route 402c. The conditional imitation learning model 101 operates in shadow mode such that designated response actions 112 (e.g., speed control 112a, steering control 112b, and optionally, direction based on navigation route 112c) are still generated, but not always executed automatically. Designated response actions 112 are used to notify the advanced driver assistance system. Figure 4 The process of demonstrating active ADAS includes autonomous vehicle control in response to risky driving behavior from human operators. Figure 5 The description includes additional processes for passive ADAS that present driver alerts and warnings to human operators.
[0093] Return to Figure 4As described above, process 400 includes operation 422 for identifying the deviation between driver response action 402 and a specified response action 112 generated by conditional imitation learning model 101. In many embodiments, operation 422 involves calculating a cross-entropy value between driver response action 402 and specified response action 112. Cross-entropy calculation is a loss function that can be recognized by those skilled in the art as a method for calculating the risk associated with an output (e.g., driver response action 402) given an expected value (e.g., specified response action 112). In the context of operation 422, risk can be similarly considered as a measure of how far a driver's behavior deviates from the behavior specified by conditional imitation learning model 101. For example, in an E2E AI CAS system, a driver who continues to accelerate as they approach a parked vehicle would be assigned a greater risk value compared to a specified response action that brakes immediately in response to an impending collision with a parked vehicle. In operation 442, the cross-entropy may be normalized (e.g., on a scale of [0, 1], which increases proportionally with risk) to obtain an interpretable risk value. A predetermined risk metric threshold (e.g., 0.5, 0.6, or 0.7) may be assigned to the risk metric. In decision 462, the risk metric output associated with the cross-entropy value between the driver's response action 402 and the assigned response action 112 may be compared with the predetermined threshold. If the risk value is higher than the threshold, the risk is classified as unacceptable. If the risk value is lower than the threshold, the risk is classified as acceptable. When an unacceptable risk is detected, the autonomous driving model 100, which includes the conditional imitation learning model 101, acquires autonomous control of the vehicle in operation 481 to eliminate or mitigate the risk. In one embodiment, the autonomous driving model 100 maintains autonomous control for a predetermined time length (e.g., 3 seconds or 30 seconds). When the risk is determined to be acceptable, the human driver maintains manual control of the vehicle in operation 483.
[0094] In some implementations, a first predetermined threshold is assigned to a risk metric defining medium risk, and a second, higher predetermined threshold is assigned to a risk metric defining high risk. For example, the medium risk threshold could be 0.4, and the high risk threshold could be 0.6. If the output risk metric of a pair of driver response actions 402 and designated response action 112 is below 0.4, the driver maintains manual control. If the output risk metric of a pair of driver response actions 402 and designated response action 112 is above 0.4 but less than 0.6, the driver maintains manual control, and a driver alert (e.g., a visual warning or audible alarm) is presented to the driver to warn of increased risk and prompt the driver to correct their behavior. If the output risk metric of a pair of driver response actions 402 and designated response action 112 is above 0.6, the autonomous driving model 100 acquires autonomous control of the vehicle.
[0095] In many implementations, the semi-autonomous driving process 400 is used as a CAS tool. Further references are provided below. Figure 5 , 6A The discussion of D and 7A to D includes other features of the disclosed E2E AI CAS, such as driver alerts and passive ADAS features with displayed driving guidance.
[0096] Collision avoidance system and driver alarm
[0097] Figure 5 This is a schematic diagram 500 illustrating the generation of collision avoidance data 528 from a specified output 112 of an end-to-end autonomous driving model 100 according to certain embodiments of the present disclosure. Similar to the previous description of model 100, the autonomous driving model 100 processes environmental driving data with reference to the current driving state, including image data 102a, non-camera data 102b (such as LiDAR data), and indication conditions 102c related to the expected path. The input data is preprocessed in layer 103 to generate a set of input embeddings 502, including both token embeddings (e.g., image data tokenized by convolutional layers) and location embeddings to keep the location data within the input data. The preprocessing step further includes tokenizing the environmental data to generate environmental data tokens, mapping the environmental data tokens to a reduced-dimensional vector space to produce environmental data embeddings, and adding the environmental data embeddings to the location embeddings to generate the input embeddings 502 of the disclosed E2E neural network, wherein the location embeddings preserve the spatial information of the environmental data tokens. The input embedding is processed by a conditional imitation learning model 101, further comprising an input embedding 502 generated by processing with an encoding layer 104, along with compressed embeddings from nine or more previous driving states within at least three seconds, and a compressed embedding of the current driving state 108 generated as output using a compression layer 106. A classification head 110 generates a specified steering and speed control action 112 in response to the current driving state.
[0098] Many embodiments of the disclosed CAS system and method described by process 500 further include extracting a set of attention weights 524 from model 101 and generating an attention map 526 containing the projection of the extracted attention weights using location embedding. Those skilled in the art will recognize the implications of attention weights within the transformer model, utilizing these weights to infer the importance of various input features and tokens, and generating attention maps to visualize attention as an interpretability technique for transformer processing. The magnitude of a particular attention weight 524 increases proportionally to its importance in generating a specified steering and velocity action. By using location embedding, attention weights associated with a particular token can be used to infer specific regions within data (e.g., LiDAR maps or camera images) that are important to the generated output 112.
[0099] In various implementations, attention weights can be extracted and attention maps generated to detect important objects or areas around the vehicle (e.g., detected objects very close to the vehicle, red lights, or ending lanes) and present the detection information to the driver to assist in decision-making. Objects can be implicitly detected within areas of the real space around the vehicle based on: (i) a comparison of the average attention weight value within said area with another average attention weight in one or more neighboring areas; and (ii) location embedding. (See reference) Figure 6A The explanations to D and 7A to D further illustrate instances of using attention maps generated by projecting attention weights onto an image or display for object detection.
[0100] Additional collision avoidance data 528 may be generated from a combination of the specified response action 112, the current response action 512 (e.g., the driver's response action during manual control, such as response action 412 of process 400), and / or the attention map 526. For example, a risk metric calculated as described with reference to process 400 may be presented to the driver as a scaling value, scale, or classification (e.g., projecting the word "RISK!" or emitting a beep). The specified response action 112 may be presented to the driver, such as a suggested speed, acceleration, or braking instruction, or a suggested correction to the steering wheel orientation for lane keeping or collision avoidance.
[0101] In some implementations, a video feed from one or more cameras coupled to the vehicle is presented to the driver, similar to a "rearview camera" or "360-degree camera" view commonly included in a vehicle's dashboard, or a map of the vehicle and its surrounding area indicating sensor data from LiDAR (e.g., a proximity indicator showing the distance to another object in front of or behind the vehicle). The camera feed or regenerated map presented to the driver may include an attention map as a heatmap overlaid on the camera feed or map, coloring areas around the vehicle that are very close to detected objects or objects detected very close to the vehicle. In other implementations, the disclosed technology includes a head-up display (HUD) projected from the vehicle's dashboard, such that the generated attention map can be transformed and projected directly onto the HUD, allowing the driver to observe the detected important objects without taking their eyes off the road. This can be considered augmented reality. Regardless of whether the graphical user interface display is a screen display on the driver's dashboard or a HUD, the display may further include useful information, such as specified actions or risk metrics.
[0102] Figure 6A D shows the first example of a collision avoidance system with a graphical user interface. (Reference) Figure 6AThe collision avoidance user interface 600A in the video demonstrates the first instance. In the video capturing this frame, there is a vehicle 608 ahead, located in the same left fast lane as the driver, and a vehicle on the right that is leaving the road. The collision avoidance user interface 600A is a HUD projected onto the driver above the dashboard. In the lower left corner, we see a set of three gauges corresponding to the specified acceleration of AI 612 (i.e., autonomous driving model 100), the current acceleration performed by the human driver 614, and the risk 616. Figure 6A As shown, AI 612 currently recommends a moderate acceleration value, while driver 614, despite being in the fast lane, is currently slightly decelerating. The risk metric calculated using process 400 is presented as a very high risk level in risk meter 616. Therefore, a large "ACCELERATE" warning 604 is displayed to the driver in all capital letters within a bright red box at the top of the screen. Intuitively, this is reasonable, as sudden deceleration or stopping on a busy road is dangerous. At the bottom center, we see navigation direction prompts 618 presented to the driver in the style of directional "mustache." Black, solid-line arrows indicate the direction indicated by the navigation direction. Green arrows indicate the specified orientation from model 100 to both avoid collisions and stay on the line. Dashed, white arrows indicate the vehicle's current direction in response to the driver's current steering wheel orientation. Although in the example Figure 6A In this model, the current driver's actions and the designated AI actions are similar for steering, but the driver can use navigation direction cues 618 as an information tool to guide vehicle handling. A heatmap is projected from the attention weights extracted from model 101 onto display 600A. Therefore, vehicles ahead in the same lane are brightly illuminated as red zones 608. Zone 608 is considered important to model 101, and this information is conveyed to the driver via the heatmap.
[0103] refer to Figure 6BThe collision avoidance user interface 600B presents a second instance, containing many of the same features as shown in interface 600A, but in a different driving scenario. Again, in the lower left corner, we see the acceleration gauges for AI 632 and human driver 634 operating in shadow mode. In the instance illustrated in interface 600B, the specified acceleration of AI 632 and the actual acceleration of driver 634 are similar, therefore no acceleration or braking instructions are presented to the driver. However, navigation direction prompt 638 indicates that as the road curves to the left, the driver's steering angle is to the right. The driver's behavior deviates significantly from the specified output action from model 101. The navigation guides the vehicle to remain straight, tilting slightly to the right, as indicated by the solid black arrow. The solid green arrow illustrating the specified output steering direction from model 101 is similarly aligned with the solid black arrow. As indicated by the dashed white arrow curving far to the right, the human driver is currently making a sharp right turn. If the driver continues along this route, they may veer sharply out of their lane or delay exiting from the exit seen on the right side of display 600B. The potential lane departure introduced by the driver's steering action introduces an impending collision risk, as indicated by a large risk value shown by risk meter 636. Therefore, a large "GO LEFT" warning 624 is presented to the driver in all capital letters within a red box at the top of the screen. The warning signifies turning the steering wheel to the left relative to the current direction. Again, the driver can also see a heatmap of the critical areas displayed in bright red zone 628.
[0104] refer to Figure 6CThe collision avoidance user interface 600C presents a third instance, containing many of the same features as those shown in interfaces 600A and 600B, but in a different driving scenario. In the video from which this frame is taken, the driver is still moving towards a pedestrian crossing despite a red light. In 600C, the vehicle is approaching an intersection with traffic lights. Again, in the lower left corner, we see the acceleration meters for AI 652 and human driver 654 operating in shadow mode. AI 652's specified acceleration will decelerate according to the red light 647 indicated by the heatmap. Driver meter 654 shows the driver is currently still slightly accelerating. The AI's specified action in 652 is braking. There is a low risk (656) associated with the driver ignoring braking, which will increase as they approach the pedestrian crossing. A large "BRAKE" warning 644 in all capital letters anticipates the need to stop accelerating and begin braking. However, navigation direction indicator 638, which shows the navigation command (solid black arrow), AI-specified direction (solid green arrow), and the current driver's steering wheel direction (dashed white arrow), all point to a straight line. Furthermore, the driver also sees a heatmap of important areas, including a bright red area 648 overlaid on a "No Right Turn" sign, another bright red area 647 overlaid on a red traffic light, and a lighter-colored area 646 indicating importance in the pedestrian crossing ahead of the vehicle. Areas 647 and 648 appear in a darker shade than area 646, corresponding to a higher level of importance from the AI.
[0105] exist Figure 6D The fourth example, 600D, has a similar environment to 600C, but there is a pedestrian (668) at the crosswalk, and the user (674) accelerates (665) through the red light. The vehicle appears to be stopped or approaching the same intersection as in 600C and 600D, with a crosswalk directly in front of it, and according to navigation direction 678, the vehicle is still oriented forward. AI meter 672 shows the specified response action from the AI is still deceleration, but driver meter 674 shows the driver is accelerating the vehicle. In example 600D, the pedestrian is now in the crosswalk. According to the heatmap of example 600D, traffic light zone 665 is still bright red, corresponding to high importance. The crosswalk is still highlighted in the heatmap, marked in zone 667, but for the AI's decision-making process, most of the crosswalk zone is still a cooler hue than traffic light zone 665 (i.e., less important). However, the bright red zone 668 on the heatmap overlaps with the feet of pedestrians in the crosswalk, indicating that pedestrians are also identified as a zone of high importance for the AI in deciding to slow down. Therefore, when pedestrians are in the crosswalk (e.g., 600D), the risk meter 676 shows a higher risk level at the same intersection than at the same intersection where there are no pedestrians in the crosswalk (e.g., 600C).
[0106] Figure 7AD presents a second instance of a collision avoidance system with a graphical user interface. While further described in the co-owned U.S. patent application identified above, Figure 7A The graphical user interface shown in D corresponds to a local delivery robot transporter, but it should be understood that a similar user interface can also be applied to road vehicles. Furthermore, those skilled in the art will recognize that... Figure 6A The example interfaces in D and 7A to D are provided for illustrative purposes and should not be considered limiting. Many additional various displays in other implementations can be inferred from the examples provided.
[0107] exist Figure 7A In the example interface 700A, similar specified and actual acceleration meters for AI 720 and human operator 721, navigation direction prompts 722, and risk meters 724 are displayed. The specified and actual directions are perfectly aligned, so only one direction prompt is visible. In the display, we see a transport robot approaching pedestrian 702, with no space to pass on either side due to luggage and parked cars. The human operator continues to accelerate according to human operator acceleration meter 721 and the green directional arrow 722. Without immediate braking intervention, the robot will soon collide with pedestrian 702. AI acceleration meter 720 shows that the specified output action generated by the AI includes braking to avoid a collision with pedestrian 702. As shown in risk meter 724, the risk associated with the current operator's behavior is very high due to the deviation between the human operator's behavior and the specified output generated by the AI.
[0108] Next, in Example 700B, we see the same scenario as in 700A, but in Example 700B, the human operator has performed braking based on recommended AI actions, as shown by meters 740 and 741. Therefore, the robotic transporter is less likely to collide with pedestrian 702, and the risk shown in meter 744 is reduced compared to that previously in Example 700A.
[0109] Figure 7CExample D similarly corresponds to a robotic transporter traveling on a sidewalk, but examples 700C and 700D each show a scenario where the robotic transporter is turning left and will collide with parked vehicle 752 if the path is not corrected. Unlike example 700A, the problematic behavior can be corrected by changing the robot's orientation without changing its speed. In example 700C, the AI-recommended acceleration and human operator acceleration control are similar according to metrics 760 and 761, but as shown by navigation tip 762, there is a substantial difference between the specified steering orientation and the actual current steering orientation. The white arrow shows the actual current steering orientation from manual human operator control, which is tilted to the left and directly facing parked vehicle 752. The green arrow shows the AI's suggestion to turn the steering orientation far to the right to correct the route and avoid a collision with parked vehicle 752. As shown in metric 764, the risk is currently very high due to the difference between the actual current steering orientation and the AI-generated specified steering orientation. Human operators can use the driver alert information provided on the screen (e.g., navigation direction prompts 762 and risk meter 764) to correct their current driving behavior, thereby mitigating the risk of an impending collision.
[0110] Therefore, in Example 700D, navigation direction cues 762 now show that the green arrow (indicating the AI-recommended steering orientation) and the white arrow (indicating the human operator's steering orientation) turn right in unison, and the robot is no longer facing the parked vehicle 752. As shown in meter 764, the risk is now much lower. In some implementations, the disclosed techniques can perform active ADAS measurements on either of the driving tasks provided above. For example, in the application of passive collision avoidance ADAS in the form of driver warnings. Figure 7C In the example driving task shown in D, the high level of risk associated with driver behavior in example 700C can trigger a shift to autonomous control, enabling the execution of AI-specified braking actions.
[0111] Computer System
[0112] Figure 8This description describes a computer system 800, according to certain embodiments of the present disclosure, that can be used to implement the disclosed technology. The computer system 800 includes at least one central processing unit (CPU) 852, which communicates with a plurality of peripheral devices via a bus subsystem 842. These peripheral devices may include a storage subsystem 802, including, for example, memory devices and a file storage subsystem 836, a user interface input device 838, a user interface output device 856, and a network interface subsystem 854. The input and output devices allow user interaction with the computer system 800. The network interface subsystem 854 provides an interface to an external network, including interfaces to corresponding interface devices in other computer systems.
[0113] In one embodiment, model 100 may be communicatively linked to storage subsystem 802 and user interface input device 838. In another embodiment, warehouse control unit 826 and conveyor control unit 832 may also be communicatively linked to storage subsystem 802 and user interface input device 838. User interface input device 838 may include a keyboard; pointing devices such as a mouse, trackball, touchpad, or graphics tablet; a scanner; a touchscreen integrated into a display; audio input devices such as a voice recognition system and microphone; and other types of input devices. Generally, the term "input device" is used to encompass all possible types of devices and methods for inputting information into computer system 800.
[0114] User interface output device 856 may include a display subsystem, a printer, a fax machine, or a non-visual display, such as an audio output device. The display subsystem may include an LED display, a cathode ray tube (CRT), a flat panel device (e.g., a liquid crystal display (LCD)), a projection device, or other mechanisms for creating visible images. The display subsystem may also provide non-visual displays, such as audio output devices. Generally, the term "output device" is used to encompass all possible types of means and methods for outputting information from computer system 800 to a user or another machine or computer system.
[0115] Storage subsystem 802 provides the functionality for programming and data construction of some or all of the modules and methods described herein. These software modules are typically executed by processor 858. Processor 858 may be a graphics processing unit (GPU), field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), and / or coarse-grained reconfigurable architecture (CGRA). Processor 858 may be provided by deep learning cloud platforms such as Google Cloud Platform™, Xilinx™, and Cirrascale™. The 878 processor instances include Google's Tensor Processing Units (TPUs)™, rack-mount solutions such as the GX4 rack series™ and GX16 rack series™, NVIDIA DGX-1™, Microsoft Stratix V FPGA™, Graphcore's Intelligent Processor Units (IPUs)™, Qualcomm's Zeroth Platform™ with Snapdragon processors™, NVIDIA's Volta™, NVIDIA's DRIVE PX™, NVIDIA's JETSON TX1 / TX2 MODULE™, Intel's Nirvana™, Movidius VPU™, Fujitsu DPI™, ARM's DynamicIQ™, IBM TrueNorth™, and Lambda GPU servers with Testa V100s™, among others.
[0116] The memory subsystem 812 used in the storage subsystem 802 may include several memories, including a main random access memory (RAM) 832 for storing instructions and data during program execution and a read-only memory (ROM) 834 for storing fixed instructions therein. The file storage subsystem 836 may provide permanent storage for program and data files and may include hard disk drives, floppy disk drives, and associated removable media, CD-ROM drives, optical disk drives, or removable media cartridges. Functional modules implementing some embodiments may be stored by the file storage subsystem 836 in the storage subsystem 802 or in other machines accessible by the processor. The bus subsystem 842 provides a mechanism for allowing various components and subsystems of the computer system 800 to communicate with each other as intended. Although the bus subsystem 842 is schematically shown as a single bus, alternative embodiments of the bus subsystem may use multiple buses.
[0117] The computer system 800 itself can be of various types, including personal computers, portable computers, workstations, computer terminals, network computers, televisions, mainframes, server farms, widely distributed loosely networked computer groups, or any other data processing system or user device. Due to the constantly evolving nature of computers and networks, Figure 8The description of the computer system 800 depicted herein is intended only as a specific example for illustrating preferred embodiments of the invention. Many other configurations of the computer system 800 may have... Figure 8 The computer system depicted has more or fewer components.
[0118] Each of the processors or modules discussed herein may contain an algorithm (e.g., instructions stored on a tangible and / or non-transitory computer-readable storage medium) or sub-algorithm for performing a particular process. Model 100 is conceptually described as a collection of modules, but may be implemented using any combination of dedicated hardware boards, DSPs, processors, etc. Alternatively, system 100 may be implemented using an off-the-shelf PC with a single processor or multiple processors, wherein functional operations are distributed among the processors. As another option, the modules described below may be implemented using a hybrid configuration, where some modular functions are executed using dedicated hardware, while the remaining modular functions are executed using an off-the-shelf PC, etc. Modules may also be implemented as software modules within a processing unit.
[0119] The various processes and steps of the described methods can be implemented using a computer. The computer may include a processor, which is part of the detection device and may be networked with or separate from the detection device for obtaining data processed by the computer. In some embodiments, information (e.g., image data) may be transmitted directly or via a computer network between components of the system disclosed herein. A local area network (LAN) or wide area network (WAN) may be a corporate computing network that includes access to the Internet to which the computer and computing devices including the system are connected. In one embodiment, the LAN conforms to the Transmission Control Protocol / Internet Protocol (TCP / IP) industry standard. In some examples, information (e.g., image data) is input to the system disclosed herein via an input device (e.g., a disk drive, optical disc player, USB port, etc.). In some examples, information is received, for example, by loading information from a storage device (e.g., a disk or flash drive).
[0120] The processor used to run the algorithms or other processes described herein may include a microprocessor. The microprocessor may be any conventional, general-purpose single-chip or multi-chip microprocessor, such as a Pentium™ processor manufactured by Intel Corporation. A particularly useful computer may utilize an Intel Ivy Bridge dual 16-core processor, an LSI RAID controller, 168 GB of RAM, and a 2 TB solid-state drive. Alternatively, the processor may include any conventional special-purpose processor, such as a digital signal processor or a graphics processor. Processors typically have conventional address lines, conventional data lines, and one or more conventional control lines.
[0121] The foregoing description is intended to enable the manufacture and use of the disclosed technology. Various modifications to the disclosed embodiments will be readily apparent, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of the disclosed technology. Therefore, the disclosed technology is not intended to be limited to the illustrated embodiments, but should be given the widest scope consistent with the principles and features disclosed herein. The scope of the disclosed technology is defined by the appended claims.
[0122] Some specific implementation schemes
[0123] We describe specific implementation schemes and features that can be used to provide driver assistance alerts to drivers using E2E models trained for autonomous and semi-autonomous driving. Many implementation schemes include a method for providing driver assistance alerts to drivers. The method includes receiving a series of environmental data related to the driving state, including at least video from a camera, return data from optical sensors, and position data from a GNSS receiver, wherein the camera, optical sensors, and GNSS receiver are coupled to a processor onboard the vehicle and process the environmental data as input to an end-to-end neural network. The end-to-end neural network is trained to generate specified steering and speed control actions in response to the current driving state. The method includes analyzing hidden layer data and output data from the end-to-end neural network to estimate collision avoidance data, wherein the collision avoidance data includes at least one or more detected objects and directional cues from the video from the camera. The directional cues are projected onto a head-up display as a superimposed image of the specified steering control actions, and a risk measure quantifying the dissimilarity between the generated specified steering and speed control actions and the received driver steering and speed control actions. The method also includes presenting a user interface to the driver containing driver assistance alerts based on the collision avoidance data.
[0124] We describe some specific implementations related to the methods used to provide driver alerts. Directional cues can be projected onto a head-up display. Within the user interface, a dynamic whisker arrow can be positioned relative to the current vehicle orientation based on a generated specified steering control action indication. The specified whisker arrow can be juxtaposed with the current human-action whisker arrow.
[0125] In some implementations, the risk acquisition method further includes calculating the cross-entropy between the generated specified steering and speed control actions and the current driver's steering and speed control actions, and standardizing the cross-entropy calculation to generate a risk metric output. The risk metric output is an indicator of the impending collision risk and increases proportionally as the current driver's steering and speed control actions deviate further from the generated specified steering and speed control actions.
[0126] One embodiment includes maintaining a manual driving mode in response to a response metric output when the risk metric output is less than a predetermined threshold, wherein the manual driving mode includes allowing the vehicle to apply the current driver's steering and speed control inputs; and entering an autonomous driving mode when the risk metric output is equal to or greater than the predetermined threshold, wherein the autonomous driving mode includes enabling the vehicle to apply specified steering and speed control inputs. One embodiment further includes classifying a risk metric as a risk level by defining a specific risk level as a risk metric output that falls within a predetermined range between a lower limit and an upper limit.
[0127] The method and other embodiments of the disclosed technology may include one or more of the features described below and / or in combination with the additional features described in the disclosed method. For the sake of brevity, the combinations of features disclosed in this application are not individually listed and do not repeat with each basic feature set. The reader will understand how easily the features identified in this section can be combined with the basic feature sets identified as embodiments.
[0128] In some implementations, the user interface presents the driver with quantitative risk including risk metric outputs or categorized risk including risk levels based on risk metric values. In other implementations, the user interface further presents the driver with the generated specified steering and speed control actions. Driver alerts (e.g., risk or specified response actions) may include visual displays, audio signals, or haptic signals. Visual displays may be displays or HUDs.
[0129] The method and other embodiments of the disclosed technology may include one or more of the features described below and / or in combination with the additional features described in the disclosed method. For the sake of brevity, the combinations of features disclosed in this application are not individually listed and do not repeat with each basic feature set. The reader will understand how easily the features identified in this section can be combined with the basic feature sets identified as embodiments.
[0130] One implementation includes preprocessing environmental data before it is provided to the end-to-end neural network. This includes tokenizing environmental data of the current driving state to generate environmental data tokens, mapping the environmental data tokens to a reduced-dimensional vector space to generate environmental data embeddings, and adding the environmental data embeddings to the location embeddings to generate the input embeddings of the end-to-end neural network, wherein the location embeddings preserve the spatial information of the environmental data tokens.
[0131] The method and other embodiments of the disclosed technology may include one or more of the features described below and / or in combination with the additional features described in the disclosed method. For the sake of brevity, the combinations of features disclosed in this application are not individually listed and do not repeat with each basic feature set. The reader will understand how easily the features identified in this section can be combined with the basic feature sets identified as embodiments.
[0132] Some implementations include a disclosed end-to-end neural network configured for autonomous driving, wherein the end-to-end neural network is a transformer model trained for end-to-end autonomous driving, and the transformer model processes environmental data by further including processing the generated input embeddings together with compressed embeddings from nine or more previous driving states within at least three seconds, and generating compressed embeddings of the current driving state and specified steering and speed control actions as outputs in response to the current driving state.
[0133] Many implementations further include extracting a set of attention weights from the transformer model and generating an attention map containing the extracted attention weights using position embedding, wherein the magnitude of a particular attention weight increases proportionally to the importance of that particular attention weight in generating a specified steering and speed maneuver, and is based on the following: implicit object detection within a region of real space surrounding the vehicle; a comparison of the average attention weight value within the region with another average attention weight value in one or more neighboring regions; and position embedding. In one implementation, presenting one or more detected objects on a head-up display further includes color-coding the attention weights within the attention map to achieve visual recognition of the implicitly detected objects, and projecting the superimposed color-coded attention map onto the head-up display.
[0134] The method may also include storing the history of video from the camera and driver assistance alerts presented to the driver in a driving database, which can be used for additional data analysis and data auditing after the driving activity is completed.
[0135] In some implementations, the disclosed methods are used in driver education tools that provide feedback to drivers as they learn new driving tasks. For example, driver education tools can be used to teach new drivers or specific forms of driving, such as racing. The disclosed methods can also be implemented in ADAS tools for safe driving monitoring by recording in-vehicle risk history and recorded driving behavior.
[0136] Another method of the disclosed technology involves training a neural network to generate driver assistance alert data. The method includes receiving environmental data from a series of driving states generated by human driving, including at least video from a camera, return data from optical sensors, and position data from a GNSS receiver, wherein the camera, optical sensors, and GNSS receiver are coupled to a processor carried by the vehicle. It includes processing the environmental data as input to the imitation training of an end-to-end neural network, including training the end-to-end neural network to generate specified steering and speed control actions in response to the current driving state. Training includes analyzing hidden layer data and output data from the end-to-end neural network to estimate collision avoidance data, and imitation training of the hidden layer data. The collision avoidance data includes at least directional and speed control cues. The directional and speed cues are suitable for generating specified steering and speed control actions projected onto the vehicle's head-up display or other dashboard display.
[0137] Training generates attention weights in the hidden layer data of an end-to-end neural network, the attention weights indicating the regions from the video from the camera that contribute the most to the generated specified steering and speed cues.
[0138] This method may further include configuring a system comprising an end-to-end neural network and further comprising a risk metric generator. Training is extended to the parameters of the risk metric generator for generating a normalized risk metric that quantifies the dissimilarity between a generated specified steering and speed control cue and a received driver steering and speed control action that differs from the generated specified steering and speed control action, thereby normalizing the risk metric onto the vehicle's head-up display or other dashboard display.
[0139] This method can be combined with features disclosed in the above embodiments and throughout this application, which are not individually enumerated or duplicated with the training feature set. The reader will understand how easily the features identified in this section can be combined with the basic feature set identified as an embodiment.
[0140] The disclosed technology can be practiced as a system, method, or article of manufacture. For example, the disclosed technology can be practiced as a system having: a hardware processor; a memory coupled to the processor; and instructions executable on the processor, which, when executed, cause the system to perform any of the described methods. Such a system may include a connected sensor from which environmental data is received. Similarly, the disclosed technology can be practiced as a computer-readable medium having instructions executable on a hardware processor, which, when executed, cause a system containing a processor to perform any of the described methods. Such a computer-readable medium has instructions for receiving environmental data from a sensor.
[0141] While the disclosed technology has been made public with reference to the preferred embodiments and examples detailed above, it should be understood that these examples are intended to be illustrative rather than limiting. Modifications and combinations will readily occur to those skilled in the art upon consideration, and such modifications and combinations will be within the spirit of the invention and the scope of the appended claims.
Claims
1. A computer-implemented method for providing driver assistance alerts to a driver, the method comprising: Receives a series of environmental data on driving status, including at least video from a camera, return data from an optical sensor, and position data from a GNSS receiver, wherein the camera, the optical sensor, and the GNSS receiver are coupled to a processor carried by the vehicle; The environmental data is processed as input to an end-to-end neural network, wherein the end-to-end neural network is trained to generate specified steering and speed control actions in response to the current driving state; Analyze the hidden layer data and output data from the end-to-end neural network to estimate collision avoidance data, wherein the collision avoidance data includes at least: One or more detected objects within the video from the camera. Directional cues, wherein the directional cues are projections onto the head-up display based on the specified steering control action, and Risk metric, which quantifies the dissimilarity between the generated specified steering and speed control actions and the received driver steering and speed control actions; and A user interface containing driver assistance alerts based on the collision avoidance data is presented to the driver.
2. The computer implementation method of claim 1, wherein the directional prompt projected onto the head-up display within the user interface is a dynamic mustache arrow, which is based on the generated specified steering control action indication of a specified vehicle orientation relative to the current vehicle orientation.
3. The computer-implemented method according to claim 1 or claim 2, wherein obtaining the risk metric further comprises: Calculate the cross-entropy between the generated specified steering and speed control actions and the current driver's steering and speed control actions, and standardize the cross-entropy calculation to generate a risk metric output. The risk metric output is an indicator of the risk of an impending collision, and it increases proportionally as the current driver's steering and speed control actions deviate further from the generated specified steering and speed control actions.
4. The computer-implemented method according to any one of claims 1 to 3, further comprising, in response to the risk measurement output: When the risk metric output is less than a predetermined threshold, the manual driving mode is maintained, wherein the manual driving mode includes allowing the vehicle to apply the current driver's steering and speed control input actions, and When the risk measurement output is equal to or greater than the predetermined threshold, the autonomous driving mode is entered, wherein the autonomous driving mode includes causing the vehicle to apply the specified steering and speed control input actions.
5. The computer-implemented method according to any one of claims 1 to 4, further comprising classifying the risk measure as a risk level by defining a specific risk level as a risk measure output that is contained within a predetermined range between a lower limit and an upper limit.
6. The computer implementation method according to any one of claims 1 to 5, wherein the user interface presents to the driver either a quantitative risk including the risk metric output or a categorized risk including the risk level based on the risk metric value.
7. The computer implementation method according to any one of claims 1 to 6, wherein the driver assistance alarm comprises one or more of a visual display, an audio signal, or a tactile signal.
8. The computer implementation method according to any one of claims 1 to 7, wherein the user interface further presents the generated specified steering and speed control actions to the driver.
9. The computer implementation method according to any one of claims 1 to 8, wherein the environmental data is preprocessed before being provided to the end-to-end neural network, the preprocessing further comprising: The environmental data of the current driving state is tokenized to generate an environmental data token. The environmental data tokens are mapped to a reduced-dimensional vector space to generate environmental data embeddings. The environmental data embedding is added to the location embedding to generate the input embedding of the end-to-end neural network, wherein the location embedding preserves the spatial information of the environmental data token.
10. The computer implementation method according to any one of claims 1 to 9, wherein the end-to-end neural network is a transducer model trained for end-to-end autonomous driving, and the transducer model processing the environmental data further comprises: The generated input embedding is processed together with compressed embeddings from nine or more previous driving states within at least three seconds, and in response to the current driving state, a compressed embedding of the current driving state and specified steering and speed control actions is generated as output.
11. The computer-implemented method according to any one of claims 1 to 10, further comprising extracting a set of attention weights from the transformer model and generating an attention map containing a projection of the extracted attention weights using the location embedding, wherein: The magnitude of a specific attention weight increases proportionally to the importance of that specific attention weight in generating the specified steering and speed maneuvers. Implicit object detection within a region of real space surrounding the vehicle is based on: (i) a comparison of the average attention weight value within the region with another average attention weight value in one or more neighboring regions; and (ii) the location embedding.
12. The computer implementation method according to any one of claims 1 to 11, wherein presenting the one or more detected objects in the head-up display via the user interface further includes color-coding attention weights within the attention map to achieve visual recognition of the implicitly detected objects, and projecting the superimposed color-coded attention map onto the head-up display.
13. The computer-implemented method according to any one of claims 1 to 12, further comprising storing the history of the video from the camera and the driver assistance alerts presented to the driver in a driving database, wherein the driving database can be used for additional data analysis and data auditing after the driving activity is completed.
14. A computer-implemented method for training a neural network to generate driver assistance alert data, the method comprising: Receives a series of environmental data on driving status generated by human driving, which includes at least video from a camera, return data from an optical sensor, and position data from a GNSS receiver, wherein the camera, the optical sensor, and the GNSS receiver are coupled to a processor carried by the vehicle; Processing the environmental data as input for the imitation training of an end-to-end neural network includes training the end-to-end neural network to generate specified steering and speed control actions in response to the current driving state; The training described therein includes analyzing hidden layer data and output data from the end-to-end neural network to estimate collision avoidance data, wherein the collision avoidance data includes at least: Directional cues, wherein the directional cues can be projected onto the head-up display as specified steering control actions, and Speed control prompts, wherein the speed control prompts can be projected onto the head-up display as specified speed control actions; Therefore, the attention weights of the end-to-end neural network in the hidden layer data indicate the regions in the video from the camera that contribute the most to the generated specified steering and speed control actions.
15. The method according to any one of claims 1 to 14, further comprising configuring a system including the end-to-end neural network and further including a risk measurement generator, comprising: The parameters of the risk metric generator are trained to generate a normalized risk metric that quantifies the dissimilarity between the generated specified steering and speed control actions and received driver steering and speed control actions that are different from the generated specified steering and speed control actions, thereby projecting the normalized risk metric onto the head-up display.
Citation Information
Patent Citations
Multi-functional inventory storage and delivery system
US20240326258A1
Collision-avoidance system for autonomous-capable vehicles
CN109117709A
Lane changing method and device of unmanned vehicle
CN110733506A
Automatic driving method and device, electronic equipment and storage medium
CN114194211A
Providing driver feedback
US10442443B1