Autonomous alert adjustment based on driver feedback

A feedback-based system using machine learning to adjust alert thresholds and types based on driver reactions addresses alert fatigue, improving safety and efficiency in autonomous driving systems.

WO2025165897A1PCT designated stage Publication Date: 2025-08-07NETRADYNE INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/013618
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-30
Filing Date
2025-01-29
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing automated driving systems face issues with alert fatigue due to frequent and inaccurate alerts, leading to desensitization and compromised driver reaction times, which can increase the risk of accidents.

Method used

Implementing a feedback-based system using a machine learning model to adjust alert thresholds and types based on driver reactions, personalizing alert generation by detecting facial expressions and head movements, and validating event detections with remote computing devices.

Benefits of technology

Improves alert accuracy and reduces fatigue by personalizing alerts to individual drivers, enhancing safety and efficiency in autonomous driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025013618_07082025_PF_FP_ABST
    Figure US2025013618_07082025_PF_FP_ABST
Patent Text Reader

Abstract

A computing device can detect a first instance of a first type of event based on a first sequence of images using an event threshold. The computing device can generate a first alert. The computing device can receive a second sequence of images depicting a reaction of a driver of the vehicle to the first type of alert. The computing device can execute a machine learning model using the second sequence of images to detect the reaction of the driver to the first type of alert. The computing device can adjust the event threshold based on the reaction of the driver. The computing device can detect a second instance of the first type of event using the adjusted threshold. The computing device can generate a second alert within the vehicle. In doing so, the computing device can reduce alert fatigue and personalize alert generation for individual drivers.
Need to check novelty before this filing date? Find Prior Art

Description

AUTONOMOUS ALERT ADJUSTMENT BASED ON DRIVER FEEDBACKCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 626,911, filed January 30, 2024, the entirety of which is incorporated by reference herein.TECHNICAL FIELD

[0002] This application relates generally to using machine learning techniques for automated alerts.BACKGROUND

[0003] Automated driving systems have garnered significant interest and investment in recent years due to their potential to revolutionize transportation by enhancing safety, efficiency, and convenience. These systems utilize various sensors, such as cameras, LiDAR, and radar, in conjunction with sophisticated algorithms, to perceive the surrounding environment and make driving decisions autonomously.

[0004] Drivers may sometimes drive unsafely or inefficiently. Unsafe driving behavior may endanger the driver and other drivers and may risk damaging the vehicle. Unsafe driving behaviors may also lead to fines. For example, highway patrol officers may issue a citation for speeding. Unsafe driving behavior may also lead to accidents, which may cause physical harm, and which may, in turn, lead to an increase in insurance rates for operating a vehicle. Inefficient driving, which may include hard accelerations, may increase the costs associated with operating a vehicle.

[0005] Attempts to reduce unsafe driving behaviors have involved generating alerts that inform the driver of an upcoming obstacle or to alert the driver that the driver is engaging in risky driving behavior (e.g., drifting into an adjacent lane). However, even these alerts have their fair share of problems. For instance, the increasing complexity of in-vehicle alert systems has raised concerns about driver overload and distraction. Modem vehicles are equipped with a myriad of sensors and systems designed to monitor driving conditions, vehicle performance, and potential hazards. Each of these sensors and systems may generate alerts for different reasons, which can result in frequent alerts to a driver. While these alerts are intended to enhance safety by informing the driver of critical information in real-time, there is a growingbody of evidence suggesting that an excessive number of alerts can have the opposite effect. Specifically, frequent interruptions can lead to cognitive overload, where the driver's ability to process information efficiently is compromised. This phenomenon, known as "alert fatigue," can desensitize drivers to warnings, potentially leading to delayed reactions or the complete disregard of critical alerts. Consequently, the effectiveness of safety measures is diminished, and the risk of accidents may increase.SUMMARY

[0006] Systems may attempt to use image processing techniques for automatic alert generation. For example, a computing system located on a vehicle (e.g., operating within or controlling a vehicle) may continuously receive image or video data of the environment surrounding or inside a vehicle. The computing system may apply a rules engine or a machine learning algorithm to images to determine whether to generate an alert (e.g., an audible or visual alert, such as to alert a vehicle ahead is within a proximity threshold of the vehicle). However, given the amount of data that is typically in any singular image, let alone multiple images of a video, the rules engine or machine learning algorithm may not be able to account for every situation or scenario depicted in the image or video. Thus, the computing system may have gaps and not accurately identify instances to generate an alert or change operation of the vehicle that may be difficult to remove or overcome.

[0007] A computing device implementing the systems and methods described herein may overcome the aforementioned technical problems by implementing a machine learning model and a driver response feedback loop. To do so, for example, a computing device of a vehicle can receive a first sequence of images of an environment surrounding the vehicle or within the vehicle. The computing device can detect a first instance of a first type of event based on the first sequence of images using an event threshold. The computing device can generate a first alert within the vehicle. The computing device can receive a second sequence of images from the first camera or a second camera depicting a reaction of a driver of the vehicle to the first type of alert. The computing device can execute a machine learning model using the second sequence of images to detect the reaction of the driver to the first type of alert. The computing device can determine whether the reaction of the driver was positive or negative. The computing device can adjust the event threshold based on the reaction of the driver and / or the determination of whether the reaction of the driver was positive or negative (e.g., decrease or increase the threshold if the reaction was negative and / or maintain thethreshold if the reaction was positive). The computing device can detect a second instance of the first type of event using the adjusted threshold. The computing device can generate a second alert within the vehicle responsive to detecting the second instance of the first type of event. In this way, the computing device can use a feedback loop with a machine learning model to generate alerts, which can improve the accuracy of alerts generated by the computing device. Additionally, because the threshold may be adjusted based on reactions of an individual driver, the computing device can personalize alert generation for individual drivers.

[0008] In one embodiment, a method for automatic alert configuration adjustment in a vehicle can include receiving, by one or more processors from a first camera mounted to or in the vehicle, a first sequence of images of an environment surrounding the vehicle or within the vehicle; responsive to detecting, by the one or more processors, a first instance of a first type of event based on the first sequence of images using an event threshold, generating, by the one or more processors, a first alert within the vehicle; receiving, by the one or more processors from the first camera or a second camera mounted to or in the vehicle, a second sequence of images depicting a reaction of a driver of the vehicle to the first type of alert; executing, by the one or more processors, a machine learning model using the second sequence of images to detect the reaction of the driver to the first type of alert; adjusting, by the one or more processors, the event threshold based on the reaction of the driver to the first type of alert; receiving, by the one or more processors from the first camera mounted to or in the vehicle, a third sequence of images of the environment surrounding the vehicle or within the vehicle; and responsive to detecting, by the one or more processors, a second instance of the first type of event based on the third sequence of images using the adjusted threshold, generating, by the one or more processors, a second alert within the vehicle.

[0009] Executing the machine learning model can include detecting, by the one or more processors, one or more facial expressions of the driver in the second sequence of images; detecting, by the one or more processors, one or more head movements of the driver in the second sequence of images; and classifying, by the one or more processors, the reaction of the driver as positive or negative based on the detected one or more facial expressions and one or more head movements. Executing the machine learning model can include classifying, by the one or more processors, the reaction of the driver as positive or negative based on the second sequence of images; and increasing, by the one or more processors, the event threshold responsive to classifying the reaction of the driver as negative.

[0010] Detecting the first type of event can include at least one of detecting, by the one or more processors, a second vehicle in a blind spot of the vehicle; or detecting, by the one or more processors, an object within a defined distance of the vehicle. Adjusting the event threshold can include increasing, by the one or more processors, the event threshold responsive to the reaction of the driver indicating the first alert was generated too early; or decreasing, by the one or more processors, the event threshold responsive to the reaction of the driver indicating the first alert was generated too late.

[0011] The method can further include storing, by the one or more processors in a memory, a driver profile associated with the driver; and updating, by the one or more processors, the driver profile based on the reaction of the driver. Generating the second alert can include selecting, by the one or more processors, a type of the second alert as at least one of an audio alert, a visual alert, or a haptic alert based on the reaction of the driver to the first alert.

[0012] Adjusting the event threshold can include maintaining, by the one or more processors, separate thresholds for different types of driving environments; categorizing, by the one or more processors, the first instance of the first type of event as occurring in a first type of driving environment; and adjusting, by the one or more processors, only the event threshold associated with the first type of driving environment based on the reaction of the driver. Executing the machine learning model can include classifying, by the one or more processors, the reaction of the driver into one of a plurality of reaction categories. Adjusting the event threshold can include adjusting, by the one or more processors, the event threshold by an amount determined based on the classification of the reaction of the driver.

[0013] Generating the first alert can include generating, by the one or more processors, the first alert based on environmental factors of the environment surrounding the vehicle. Detecting the reaction of the driver can include detecting, by the one or more processors, movement of the driver subsequent to the first alert for a defined time duration. Generating the first alert can include generating, by the one or more processors, an audible or visual alert within the vehicle. The method can further include generating, by the one or more processors, a driver score for the driver; and determining, by the one or more processors, a magnitude of the adjustment to the event threshold based on the driver score for the driver.

[0014] The method can further include transmitting, by the one or more processors, the first sequence of images to a remote computing device responsive to detecting the first instanceof the first type of event; and receiving, by the one or more processors from the remote computing device, an indication that the detection of the first instance of the first type of event was inaccurate, wherein adjusting the event threshold based on the reaction of the driver to the first type of alert is further based on the indication that the detection of the first instance of the first type of event was inaccurate.

[0015] The method can include, responsive to detecting the reaction of the driver to the first type of alert, transmitting, by the one or more processors, the first sequence of images to a remote computing device, wherein when the remote computing device does not detect the first instance of the first type of event in the first sequence images, the remote computing device transmits an instruction to the one or more processors to adjust the event threshold for the first type of event. Transmitting, by the one or more processors, the first sequence of images to the remote computing device can be responsive to determining, by the one or more processors, the reaction of the driver to the first type of alert is negative.

[0016] In another embodiment, one or more processors configured by machine- readable instructions to receive, from a first camera mounted to or in a vehicle, a first sequence of images of an environment surrounding the vehicle or within the vehicle; responsive to detecting a first instance of a first type of event based on the first sequence of images using an event threshold, generate a first alert within the vehicle; receive, from the first camera or a second camera mounted to or in the vehicle, a second sequence of images depicting a reaction of a driver of the vehicle to the first alert; execute a machine learning model using the second sequence of images to detect the reaction of the driver to the first alert; adjust the event threshold based on the reaction of the driver to the first alert; receive, from the first camera mounted to or in the vehicle, a third sequence of images of the environment surrounding the vehicle or within the vehicle; and responsive to detecting a second instance of the first type of event based on the third sequence of images using the adjusted threshold, generate a second alert within the vehicle.

[0017] The one or more processors can be further configured to execute the machine learning model to detect one or more facial expressions of the driver in the second sequence of images; detect one or more head movements of the driver in the second sequence of images; and classify the reaction of the driver as positive or negative based on the detected one or more facial expressions and one or more head movements.

[0018] The one or more processors can be further configured to responsive to detecting the reaction of the driver to the first type of alert, transmit the first sequence of images to a remote computing device, wherein when the remote computing device does not detect the first instance of the first type of event in the first sequence images, the remote computing device transmits an instruction to the one or more processors to adjust the event threshold for the first type of event. The one or more processors can be configured to transmit the first sequence of images to the remote computing device responsive to determining the reaction of the driver to the first type of alert is negative.

[0019] In another embodiment, a method for automatic alert configuration adjustment in a vehicle includes storing, by one or more processors, a profile for a driver in memory, the profile containing an indication to generate a first type of alert in response to detecting a first type of event; receiving, by the one or more processors from a camera mounted to or in the vehicle, a first sequence of images of an environment surrounding the vehicle or within the vehicle; responsive to detecting, by the one or more processors, a first instance of the first type of event based on the first sequence of images, generating, by the one or more processors, a first alert of the first type of alert within the vehicle based on the indication in the profile of the driver; in response to detecting, by the one or more processors, a response by the driver to the first alert from a second sequence of images, adjusting, by the one or more processors, the indication in the profile to indicate a second type of alert to generate in response to detecting the first type of event; receiving, by the one or more processors from the camera mounted to or in the vehicle, a third sequence of images of the environment surrounding the vehicle or within the vehicle; and responsive to detecting, by the one or more processors, a second instance of the first type of event based on the third sequence of images, generating, by the one or more processors, a second alert within the vehicle based on the adjusted indication in the profile of the driver.

[0020] The profile of the driver can include indications of types of alerts for a plurality of types of alerts for different types of events. Generating the first alert of the first type of alert within the vehicle can include selecting, by the one or more processors, the first type of alert of the plurality of types of alerts based on the indication in the profile of the driver. Detecting the response by the driver can include detecting, by the one or more processors, a manual input on a user interface or device of the vehicle, detecting, by the one or more processors, a voice command from the driver, or detecting, by the one or more processors a gesture of the driver captured by a second camera within the vehicle.

[0021] The method can further include changing, by the one or more processors, the adjusted indication in the profile of the driver based on a user input at a user interface. The method can further include identifying the profile of the driver upon the driver entering the vehicle; and using, by the one or more processors based on the identification of the profile, the indication from the profile to generate the first type of alert.

[0022] The method can further include, responsive to detecting the response of the driver to the first alert, transmitting, by the one or more processors, the first sequence of images to a remote computing device, wherein when the remote computing device does not detect the first instance of the first type of event in the first sequence images, the remote computing device transmits an instruction to the one or more processors to adjust the type of alert for the first type of event. Transmitting the first sequence of images to the remote computing device is responsive to determining the response of the driver to the first type of alert is negative.

[0023] These and other aspects and implementations are discussed in detail below. The foregoing information and the following detailed description include illustrative examples of various aspects and implementations, and provide an overview or framework for understanding the nature and character of the claimed aspects and implementations. The drawings provide illustration and a further understanding of the various aspects and implementations, and are incorporated in and constitute a part of this specification. Aspects can be combined and it will be readily appreciated that features described in the context of one aspect of the invention can be combined with other aspects. Aspects can be implemented in any convenient form. For example, by appropriate computer programs, which may be carried on appropriate carrier media (computer-readable media), which may be tangible carrier media (e.g., disks) or intangible carrier media (e.g., communications signals). Aspects may also be implemented using suitable apparatus, which may take the form of programmable computers running computer programs arranged to implement the aspect. As used in the specification and in the claims, the singular form of ‘a’, ‘an’, and ‘the’ include plural referents unless the context clearly dictates otherwise.BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Non-limiting embodiments of the present disclosure are described by way of example with reference to the accompanying figures, which are schematic and are not intended to be drawn to scale. Unless indicated as representing the background art, the figures representaspects of the disclosure. For purposes of clarity, not every component may be labeled in every drawing. In the drawings:

[0025] FIG. 1 illustrates an example environment showing a computing system for a feedback-based driving system, according to an embodiment;

[0026] FIG. 2 illustrates a flowchart of a method for implementing a feedback-based driving system by adjusting a threshold, according to an embodiment;

[0027] FIGs. 3A and 3B illustrate sequences of example driver responses to an alert, according to an embodiment;

[0028] FIG. 4 illustrates a flowchart of a method for implementing a feedback-based driving system by adjusting a profile, according to an embodiment;

[0029] FIG. 5 illustrates a sequence diagram of a sequence for implementing a computing system for autonomous alert adjustment, according to an embodiment; and

[0030] FIG. 6 illustrates a sequence diagram of a sequence for implementing a computing system for autonomous alert adjustment using intentional feedback, according to an embodiment.DETAILED DESCRIPTION

[0031] Reference will now be made to the illustrative embodiments depicted in the drawings, and specific language will be used here to describe the same. It will nevertheless be understood that no limitation of the scope of the claims or this disclosure is thereby intended. Alterations and further modifications of the inventive features illustrated herein, and additional applications of the principles of the subject matter illustrated herein, which would occur to one skilled in the relevant art and having possession of this disclosure, are to be considered within the scope of the subject matter disclosed herein. Other embodiments may be used and / or other changes may be made without departing from the spirit or scope of the present disclosure. The illustrative embodiments described in the detailed description are not meant to be limiting of the subject matter presented.

[0032] Autonomous vehicle or assisted vehicle guidance can use image analysis to generate alerts. However, while images can play a role in the functioning of self-driving technology, real-time image analysis can be improved upon for more accurate and / or safe self-driving and / or alert generation. For example, images can be blurry for various reasons, such as because of adverse weather conditions (e.g., heavy rain, snow, or fog), which can limit the system's ability to accurately interpret the surroundings. Additionally, different drivers may have different preferences regarding the timing of alerts and / or the types of alerts that are generated regarding their driving and / or the environment surrounding the vehicles.

[0033] For example, a computing system located on a vehicle (e.g., operating within or otherwise controlling a vehicle) and / or operating remote from the vehicle may continuously receive image or video data of the environment surrounding or inside the vehicle. The computing system may apply a rules engine or a machine learning algorithm to images to determine whether to generate an alert (e.g., an audible, visual, or haptic alert, such as to alert a vehicle ahead is within a proximity threshold of the vehicle). In doing so, the computing system may automatically generate an alert in scenarios in which the driver of the vehicle does not believe is necessary or otherwise generate alerts at inaccurate times. Thus, the computing system may generate false alerts. A system that is configured to generate alerts and / or decisions based on feedback from drivers indicating the accuracy of the alerts and / or the desire for such alerts can improve the functioning of an automatic alert generation system.

[0034] To address these technical challenges, a computer or computing device implementing the systems and methods described herein may implement a feedback-based driving system. The feedback-based driving system can use reaction data of reactions that a driver has to alerts generated within a vehicle to adjust a frequency and / or type of such alerts. For example, a computing device of a vehicle may receive a first sequence of images depicting the environment surrounding the vehicle or inside of the vehicle. The computing device can analyze the first sequence of images to detect a first instance of a first type of driving event (e.g., driver drowsiness, lane drifting, speeding, upcoming stop sign or red traffic light, upcoming pedestrian, etc.). In doing so, the computing device can determine a confidence score to generate the alert or in the first type of driving event and determine the confidence score exceeds an event threshold. Responsive to detecting the first instance of the first type of driving event, the computing device can generate an alert (e.g., a haptic alert, an audio alert, and / or a visual alert) within the vehicle (e.g., within the cabin of the vehicle). The computing device can monitor the driver of the vehicle subsequent to generating the alert by receiving or otherwise processing a second sequence of images depicting the driver in a time period after (e.g., immediately after) the alert was generated. The time period after the alert can have a defined time duration.

[0035] The computing device can use a machine learning model to detect a reaction of the driver to the alert. The computing device can determine whether the reaction was positive or negative. Based on the determination, the computing device can adjust or maintain the event threshold (e.g., increase or decrease the event threshold responsive to determining the reaction was negative or maintain (e.g., not change) the event threshold responsive to determining the reaction was positive). The computing device can subsequently use the adjusted or maintained event threshold to determine whether to generate an alert, such as for the same type of driving event (e.g., the first type of driving event). The computing device can repeat this process for any number of alerts generated while the driver is driving or otherwise in the vehicle or a different vehicle. In this way, the computing device can automatically adjust the frequency and / or likelihood of generating alerts for driving events based on driver feedback, thus improving the accuracy of the alerts and / or personalizing the alerts to the driver to reduce alert fatigue.

[0036] In some cases, the computing device can further improve the effectiveness of alerts based on driver feedback by updating a profile of the driver in real-time. For example, in addition to or instead of adjusting the event threshold, the computing device can change the types of alerts that are generated based on driver feedback. For instance, a computing device can store a profile for a driver in memory. The profile can contain an indication to generate a first type of alert in response to detecting a first type of event. The computing device can receive a first sequence of images of an environment surrounding the vehicle or within the vehicle from a camera mounted to or in the vehicle. The computing device can detect a first instance of the first type of event. Responsive to the detection, the computing device can generate a first alert of the first type within the vehicle based on the indication in the profile of the driver. The computing device can receive a second sequence of images depicting a reaction (e.g., a response) of the driver to the first alert. The computing device can determine whether the reaction was positive or negative. Responsive to determining the reaction was negative, the computing device can adjust the indication in the profile to indicate to generate a different type of alert (e.g., a second type of alert) in response to future detection of the first type of event. Subsequently, the computing device can detect a second instance of the first type of event and generate an alert of the second type of alert based on the indication of the second type of alert in the profile of the driver. In this way, the computing device can automatically adjust the type of alerts that are generated for driving events based on driver feedback, thusimproving the effectiveness of the alerts and increasing the likelihood of a positive reaction to such alerts.

[0037] In some cases, the computer can validate a determined reaction prior to adjusting any thresholds and / or profiles of drivers. For example, responsive to detecting an event, generating an alert, and detecting a negative reaction from the driver to the alert, the computer can transmit the sequence of images depicting the event to a remote computing device. A user at the remote computing device, or a separately trained computer model, can determine whether the detection of the event was accurate or inaccurate. A user at the computing device or the computer can determine the detection of the event was accurate or inaccurate based on the sequence of images depicting the detected event. The remote computing device (e.g., responsive to a user input) can transmit an indication to the computer that the detection of the event was accurate or inaccurate, or otherwise whether the remote computing device detected the same event as the computing device. The computer can receive the indication. The computer can adjust the threshold or type of alert responsive to the indication indicating the event detection was inaccurate or that the remote computing device did not detect the same event. For instance, the computer may only adjust the threshold responsive to determining the event detection was inaccurate.Feedback-based Driving System

[0038] FIG. 1 depicts an example environment that includes example components of a system in which such a computer can selectively generate alerts in a feedback-based system. Various other system architectures may include more or fewer features and / or may utilize the techniques described herein to achieve the results and outputs described herein. Therefore, the system depicted in FIG. l is a non-limiting example.

[0039] FIG. 1 illustrates a system 100, which includes components of a feedback-based driving system 105 for using images captured by a camera attached to or integrated on or within a vehicle 110 for autonomous driving or alert generation. The system 100 can include the feedback-based driving system 105, the vehicle 110, and / or a cloud computing system 115. The feedback-based driving system 105 can include a computing device 120, camera(s) 125, and / or a communication interface 130. The feedback-based driving system 105 may include an alert device, such as an audio alarm, a warning light, another type of visual indicator, and / or a haptic feedback device (e.g., one or more vibrating seats or a vibrating steering wheel). The feedback-based driving system 105 can be mounted on a dashboard, windshield, or other areainside the vehicle 110. The system 100 is not confined to the components described herein and may include additional or other components, not shown for brevity, which are to be considered within the scope of the embodiments described herein.

[0040] The vehicle 110 can be any type of vehicle, such as a car, truck, van, sportutility -vehicle (SUV), motorcycle, semi-tractor trailer, or other vehicle that can be driven on a road or another environment. The vehicle 110 can be operated by a user, or, in some implementations, can include an autonomous vehicle control system (not pictured) that navigates the vehicle 110 or provides navigation assistance to an operator of the vehicle 110. The vehicle 110 can be a vehicle of a fleet of vehicles that are owned and / or operated by an entity (e.g., a business or organization) to transport goods, materials, and / or individuals.

[0041] The vehicle 110 can include the feedback-based driving system 105, which can be used to detect objects within images captured by the camera(s) 125 when the vehicle is parked and / or as the vehicle is driving down a road and generate alerts or control the vehicle 110 based on the detected objects. As outlined above, the feedback-based driving system 105 can include the computing device 120. The computing device 120 can be mounted on or in the vehicle 110. In some cases, the computing device 120 is a computing device that automatically controls the vehicle for self-driving or generates audible or visual alerts via devices within the vehicle 110.

[0042] The computing device 120 can include the storage 135, which can store images and / or video captured by the camera(s) 125, machine learning model(s) 140, and an application 150. The storage 135 can be a computer-readable memory that can store or maintain any of the information described herein that is generated, accessed, received, transmitted, or otherwise processed by the computing device 120. The storage 135 can maintain one or more data structures, which may contain, index, or otherwise store each of the values, pluralities, sets, variables, vectors, numbers, or thresholds described herein. The storage 135 can be accessed using one or more memory addresses, index values, or identifiers of any item, structure, or region of memory maintained by the storage 135.

[0043] The storage 135 may be internal to the computing device 120 or may exist externally to the computing device 120 and accessed via a suitable bus or interface. In some implementations, the storage 135 can be distributed across many different storage elements. The computing device 120 (or any components thereof) can store, in one or more regions of the memory of the storage 135, the results of any or all computations, determinations,selections, identifications, generations, constructions, or calculations in one or more data structures indexed or identified with appropriate values.

[0044] The computing device 120 can include or be in communication with a communication interface 130 that can communicate wirelessly with other devices. The communication interface 130 of the computing device 120 can include, for example, a Bluetooth communications device, a Wi-Fi communication device, or a 5G / LTE / 3G cellular data communications device. The communication interface 130 can be used, for example, to transmit any information described herein to the cloud computing system 115, including images and / or videos that the computing device 120 receives from the camera(s) 125. The communication interface 130 can also be used, for example, to receive indications of driving events that the processor detects from the images or videos, in some cases with the images or videos from which the processor detects the driving events.

[0045] The camera(s) 125 can include any type of camera or multiple cameras that are capable of capturing images or videos of the environment surrounding the vehicle 110 and / or within the vehicle 110, including a configuration having an external camera view of the surrounding environment and a cabin-facing camera view of the driver. The camera(s) 125 may periodically capture images or video while the vehicle 110 is turned on and parked and / or as the vehicle 110 is moving. For example, the camera(s) 125 can capture an image or video of a stop sign 180 as the vehicle 110 approaches or passes the stop sign 180. In another example, the camera(s) 125 can capture an image or video of the driver before, during, and / or after a driving event. In some cases, the camera(s) 125 can capture images or video when the vehicle 110 is turned off (e.g., when the vehicle is off but in a surveillance mode). The camera(s) 125 may capture images or videos and transmit the images or videos to the computing device 120. The computing device 120 may receive the images or videos and store images or videos in the storage 135 and / or transmit the images or videos to other vehicles or to the cloud computing system 115.

[0046] The cloud computing system 115 can be or include one or more computing devices that are configured to train and distribute machine learning models (e.g., neural networks, support vector machines, random forests, language models (e.g., transformer models, small language models, large language models, etc.), etc.) for object detection from images and / or decisions for alert generation or autonomous driving decisions. The cloud computing system can include a processor 170 and / or a memory 175. The cloud computingsystem 115 may have more processing resources (e.g., more cores, processors, or memory), so the cloud computing system 115 may be more accurate in event determinations than on-vehicle determinations. The cloud computing system 115 may receive training sequences of images from computing devices of multiple vehicles and train an individual machine learning model to detect events (e.g., driving events) depicted in a sequence of images. The cloud computing system 115 can train the machine learning model to detect events based further on sensor data (e.g., acceleration, speed, gyroscope data, weather data, etc.) that correspond to the individual sequences of images, identifications of types of drivers that were driving when the sequences of images were captured, and / or identifications of the types of vehicles that were being driven. The cloud computing system 115 may train one or more machine learning models using such sequences of images and / or other data collected from different vehicles (e.g., vehicles of a fleet of vehicles) and transmitted to the cloud computing system 115 over a network. The cloud computing system 115 can train the machine learning model to do so and transmit the machine learning models (e.g., copies of the machine learning models) to computing devices of vehicles (e.g., the computing device 120) once the machine learning model or models are sufficiently trained. After transmitting the machine learning models to the computing devices, the cloud computing system 115 can continue to train a local version of the machine learning models with training images to improve the machine learning models and / or account for changes in new objects that appear in the environment and / or changes in the cameras that are capturing the images. The cloud computing system 115 can provision the updated models to the computing devices of vehicles after each training iteration and / or at defined time intervals such that the computing devices can continue to use a more accurate machine learning model over time.

[0047] The cloud computing system 115 may use parallel processing techniques or have different computers perform different tasks to facilitate the operations described herein. For example, one computer of the cloud computing system 115 can establish connections with the computers of vehicles to receive sequences of images from feedback-based driving systems of the vehicles (e.g., the feedback-based driving system 105), another computer of the cloud computing system 115 can process the received sequences of images using a machine learning model to detect events in the sequences of images and train the machine learning model based on the prediction (e.g., by using backpropagation techniques), and another computer of the cloud computing system 115 can establish connections with different computing devices of vehicles and transmit the trained machine learning model to the different computing devices touse for autonomous driving and / or alert generation. Any combination of one or more of the computers of the cloud computing system 115 and / or the feedback-based driving system 105 may perform such processes.

[0048] A computing device 155 can be a computer that is configured to receive messages and / or present messages or other data on a user interface. The computing device 155 can include a processor 160 and a memory 165. The computing device 155 can be any type of computing device, such as a mobile phone, laptop computer, desktop computer, smart watch, gaming console, personal data assistant, dashboard computer, or other computing device. A user associated with the vehicle 110, such as a fleet manager of a fleet containing the vehicle no, can view detected sequences of detected events through the computing device 155. For example, responsive to detecting an event (e.g., a driving event) in a sequence of images, the feedback-based driving system 105 can transmit the sequence of images of the event to the cloud computing system 115. The computing device 155 can retrieve or request the sequence of images of the event from the cloud computing system 115 by transmitting a message or a request for the sequence of images to the cloud computing system 115. The user accessing the computing device 155 can view the sequence of images and determine whether the detection of the event was correct or accurate. The user can provide an input indicating the determination and the cloud computing system 115 can transmit the indication to the cloud computing system 115, which in some cases can forward the indication to the feedback-based driving system 105. The feedback-based driving system 105 can determine whether and / or how to adjust future event detections and / or alerts based on such indications of event detection accuracy.

[0049] In some cases, the cloud computing system 115 can operate to validate or detect events in sequences of images transmitted to the cloud computing system 115 by the feedbackbased driving system. For example, responsive to detecting an event (e.g., a driving event) in a sequence of images, the feedback-based driving system 105 can transmit the sequence of images of the event to the cloud computing system 115. The cloud computing system 115 can input the sequence of images into the machine learning model stored in the memory 175 that is trained to detect events in sequences of images. The cloud computing system 115 can execute the machine learning model based on the input to cause the machine learning model to output an indication of whether an event was detected in the sequence of images. The cloud computing system 115 can transmit the indication to the feedback-based driving system 105. The feedback-based driving system 105 can determine whether and / or how to adjust future event detections and / or alerts based on such indications of event detection accuracy.

[0050] The camera(s) 125 may include any number and any type of camera or video camera that can capture images or video of areas surrounding and / or inside the vehicle 110. The camera(s) 125 can communicate with the computing device 120 via a vehicle interface, which may include a CAN bus or an on-board diagnostics interface. The camera(s) 125 can capture images or video and transmit the captured images or video to the computing device 120. The computing device 120 can receive the captured images or video and transmit the images or video to the cloud computing system 115 to use to detect driving events for a driver of the vehicle 110.

[0051] The computing device 120 can operate to autonomously drive the vehicle 110 and / or generate audible, visual, and / or haptic alerts within the vehicle. The computing device 120 can do so using a machine learning architecture on images or video captured by the camera(s) 125. For example, the camera(s) 125 can generate a plurality of images (e.g., a sequence of images or images of a video) depicting the environment surrounding the vehicle 110 and / or within the vehicle 110. The camera(s) 125 can capture the individual images of a sequence of images at a defined interval or pseudo-randomly. The camera(s) 125 can transmit or send the images to the computing device 120 responsive to capturing the images and / or in a batch at defined time intervals. The camera(s) 125 can attach timestamps to the individual images as metadata indicating the times in which the camera(s) 125 captured the images and / or the times in which the camera(s) 125 transmitted or sent the images to the computing device 120. Any number of cameras of the vehicle 110 similar to the camera(s) 125 can transmit such images to the computing device 120 over time as the vehicle 110 is driving or is at rest.

[0052] The computing device 120 can receive the images from the camera 125 and process the images to detect events or driving events (e.g., driver drowsiness, swerving across a lane line, speeding, a missed stop sign, being too close (e.g., within a threshold distance) to another vehicle, etc.) depicted in the images. For example, the machine learning model(s) 140 can be or include a machine learning model 140 (e.g., a neural network, a support vector machine, a random forest, a language model (e.g., transformer model, small language model, large language model, etc.), etc.) that is configured to use image processing techniques (e.g., as a convolutional neural network) to detect objects and / or other features (e.g., colors, tones, contexts, etc.) in individual images and / or across images when using the individual images as input in a sequence. For instance, the computing device 120 can execute the machine learning model 140 using each image of a sequence of images as input at once or separately execute the machine learning model 140 for each image of the sequence of images. The machine learningmodel 140 can use the detected objects to detect events depicted in the sequence of images, or otherwise automatically detect the events from the sequence of images. In some cases, the machine learning model 140 can output a type of the event detected from the sequences of images.

[0053] In some cases, when generating an output indicating whether an event was detected from a sequence of images, the machine learning model 140 can use a threshold (e.g., an event threshold or a confidence threshold) to determine whether an event was detected in the sequence of images. For example, the machine learning model 140 may include one or more trained or learned weights and / or parameters. The machine learning model 140 can apply the trained or learned weights and / or parameters to the sequence of images to determine a confidence or confidence score that the machine learning model 140 has regarding whether the sequence of images depicts an event. The machine learning model 140 can generate a confidence score indicating a likelihood that the sequence of images depicts an event and / or a confidence score indicating a likelihood that the sequence of images does not depict an event. The machine learning model 140 or the application 150 can compare the confidence score to the threshold. Responsive to determining a confidence score indicating a likelihood of an event exceeds the threshold, the machine learning model 140 or the application 150 can generate or output an indication that an event was detected from the sequence of images. Otherwise, the machine learning model 140 or the application 150 can generate or output an indication that no event was detected from the sequence of images. The machine learning model 140 and / or the application 150 can repeat this process for any number of sequences of images over time as the camera(s) 125 transmit images or sequences of images to the computing device 120. In some cases, the machine learning model 140 and / or the application 150 can repeat the process at set time intervals and / or at a defined interval of images received from the camera(s) 125. The sequences of images can be of the environment surrounding the vehicle 110 and / or of the environment within the vehicle 110.

[0054] In some cases, the machine learning model(s) 140 can be configured to generate confidence scores for different types of events from individual sequences of images. For example, the machine learning model(s) 140 can generate a confidence score for each of a plurality of event types based on the input sequence of images. The machine learning model(s) 140 and / or the application 150 can compare the confidence scores to a threshold (e.g., the same threshold or a threshold specific to each event type). Responsive to determining none of the confidence scores exceeds the threshold, the machine learning model(s) 140 and / or theapplication 150 can generate an output indicating no event was detected. Otherwise, the machine learning model(s) 140 and / or the application 150 can output an identification of the type of event for which the confidence score exceeded the threshold.

[0055] The application 150 can generate alerts responsive to detecting events from sequences of images. For example, responsive to determining the confidence score for an alert exceeds a threshold (e.g., based on the comparison by the application 150 and / or the machine learning model 140), the application 150 can generate an alert within the vehicle 110. The application 150 can generate the alert by generating an audio alert (e.g., through a microphone within the vehicle 110), a visual alert (e.g., by activating a light or flashing a light), or a haptic alert (e.g., by activating a haptic device within the vehicle). The application 150 can similarly activate such alerts for any number of detected events.

[0056] In some cases, the application 150 can generate different types of alerts depending on the types of the detected events. For example, responsive to detecting a drowsy driver, the application 150 can be configured to activate a horn or a beeping sound to attempt to wake up the driver. Responsive to detecting an event indicating veering off of the road, the application 150 can generate a visual alert on a side in which the vehicle 110 is veering to show the driver the direction of the veering. Responsive to detecting another vehicle is too close to the vehicle 110, the application 150 can activate a haptic feedback device (e.g., in the driver’s seat of the vehicle 110) to cause the haptic feedback device to vibrate. The application 150 can generate any type of alert depending on the type of the detected event.

[0057] The application 150 can monitor the reaction of the driver of the vehicle 110. The application 150 can do so responsive to activating or generating an alert. For example, the application 150 can receive a sequence of images (e.g., a second sequence of images) depicting the driver of the vehicle 110 subsequent to generating the alert. The sequence of images can include images depicting the driver within a defined time period (e.g., one second or two seconds before and / or after the alert was generated). The application 150 can receive the sequence of images from the same camera that generated the sequence of images based on which the alert was generated or from a different camera (e.g., a second camera). The sequence of images can depict the movements of the driver during and / or after the time in which the alert was generated or activated. Examples of such movements can include facial movements, head movements, body movements, hand movements, arm movements, etc. In some cases, the application 150 can identify the sequence of images by filtering a plurality of received imagesto only include images with timestamps beginning or a defined time period before the generation or activation of the alert and / or ending a defined time (e.g., one second, two seconds, three seconds, etc.) after the generation of the alert.

[0058] The application 150 can receive the sequence of images and input the sequence of images into a machine learning model (e.g., a neural network, a support vector machine, a random forest, or a language model) of the machine learning model(s) 140 configured or trained to detect body movements and / or expressions in sequences of images. The application 150 can execute the machine learning model based on the sequence of images to cause the machine learning model to detect a reaction to the alert. The machine learning model can output a magnitude of a reaction and / or a positive or negative sign (e.g., a positive magnitude or a negative magnitude) (e.g., a valence) based on the movements of the driver within the sequence of images. For example, the machine learning model can analyze and / or detect the facial expression and / or head movements (e.g., a headshake) to determine the driver did not appreciate an alert (e.g., a negative reaction of a negative magnitude) and a magnitude of the unappreciation. In some cases, the facial expressions can be or include micro expressions, such as raising an eyebrow, squinting, creases around eyes, raised cheeks, lifted comers of mouth, downturned mouth corners, compressed lips, raised eyebrows, widened eyes, slightly parted lips, wrinkled nose, raised eyebrows, etc. In another example, the machine learning model may analyze and / or detect hand movements and / or hand configurations to determine the magnitude and / or a positive or negative sign of a reaction. The machine learning model can determine the sign and / or the magnitude of a reaction from a sequence of images based on movements of any combination of body part of the driver.

[0059] The application 150 can dynamically adjust the event threshold (e.g., the confidence threshold) based on reactions detected by the machine learning model. For example, responsive to detecting the reaction of the driver from the sequence of images, the application 150 can adjust the threshold used to detect the event that caused the generation of the alert based on the determined reaction. The application 150 can adjust the threshold that is specific to the event type that was detected and / or the threshold used to detect multiple types of events.

[0060] In some cases, the application 150 can adjust the event threshold based on instructions from the cloud computing device 115. For example, the machine learning model(s) 140 can detect an event from a first sequence of images using the event threshold.The application 150 can generate an alert responsive to the machine learning model(s) 140 detecting the event and then detect a negative reaction to the alert. Responsive to detecting the negative reaction to the alert, the application 150 can transmit the first sequence of images to the cloud computing system 115 with an identification of the type of event the machine learning model(s) 140 detected from the first sequence of images. The cloud computing system 115 can execute a model to determine whether the model detects the event identified in the message. Responsive to determining the model did not detect the same type of event, or any event, from the first sequence of images, the cloud computing system 115 can transmit a message or instruction back to the application 150 indicating the event or the same type of event was not detected. The application 150 can adjust the event threshold based on the message or instruction.

[0061] However, responsive to determining the model did detect the same type of event, or any event, from the first sequence of images, the cloud computing system 115 can transmit a message or instruction back to the application 150 indicating the event or the same type of event was detected. The application 150 may maintain or not adjust the event threshold based on the message or instruction.

[0062] The application 150 can adjust the threshold by increasing or decreasing the threshold. The application 150 can increase or decrease the threshold by an amount proportional to the magnitude of the detected reaction. The application 150 and / or the machine learning model(s) 140 can subsequently use the adjusted threshold to detect events, such as to detect events of the same event type when adjusting a threshold specific to the event type.

[0063] In some cases, the application 150 can additionally or instead adjust the types of alerts that are generated based on the detected reactions to generated alerts. For example, the application 150 can store a profile for the driver of the vehicle 110 in the storage 135. The profile can include the threshold or thresholds for the different event types that the application 150 and / or the machine learning models 140 can use to detect events and / or that can be adjusted to detect such events, as described above. The profile can additionally or instead include an indication of a type of alert to generate. In some cases, the profile can include multiple indications of types of alerts that correspond to different event types. Responsive to detecting an event, the application 150 can identify a type of alert (e.g., audible, visual, and / or haptic) to generate for the event, in some cases based on the type of the detected event. The application 150 can generate or activate the alert based on the identified type of alert in the profile.

[0064] The application 150 can adjust the indications in the profile based on the reactions of the driver of the vehicle 110 to the generated alerts. For example, responsive to detecting a negative reaction to a type of alert generated for a particular type of event, the application 150 can adjust the type of the alert to a different type (e.g., from audible to visual). In some cases, the application 150 can compare a magnitude of the negative reaction to a threshold (e.g., a reaction threshold), and adjust the type of alert responsive to determining the magnitude exceeds or otherwise satisfies the reaction threshold. The application 150 can similarly adjust the types of alerts for any number of types of events. Subsequently, the application 150 can use the adjusted type of alert when detecting a type of event corresponding to the adjusted type of alert.

[0065] FIG. 2 illustrates a flow of a method 200 for implementing a feedback-based driving system executed by a data processing system, according to some embodiments. The data processing system can be or include a computing system of a vehicle (e.g., the feedbackbased driving system 105) and / or a remote computing system (e.g., a cloud server or the cloud computing system 115) for using a feedback-based alert generation, in accordance with an embodiment. The method 200 is shown to include steps 202-222. However, other embodiments may include additional or alternative steps, or may omit one or more steps altogether. Different steps can be performed by different computing systems (e.g., the computing system of the vehicle can perform one or more of the steps 202-222 and / or the remote computing system can perform one or more of the steps 202-222) and / or the different computing systems and can operate together to perform individual steps of the steps 202-222.

[0066] In step 202, the data processing system can receive one or more images. The one or more images can be or include a first sequence of images or images (e.g., a first sequence of images) of a video (e.g., video data) of an environment surrounding a vehicle or the environment within the vehicle (e.g., depicting the driver). The first sequence of images can be consecutively captured images by the same camera. The images of the first sequence can be captured at a set time interval (e.g., every five seconds) or pseudo-randomly. The images can be images of a road on which the vehicle is driving or of areas adjacent to the road. The data processing system can receive the images from a camera mounted to or within the vehicle (e.g., mounted on the dashboard of the vehicle). The images can be or include a JPEG, RAW, or another image file type. The visual data can be or include one or more standalone images or one or more frames of a video. The data processing system can receive the images as a driver is driving the vehicle, when the driver has parked the vehicle, and / or when the vehicleis in a standby mode. The data processing system can receive such images at defined intervals in the case of static standalone images or as the camera streams a video to the data processing system in individual frames.

[0067] In step 204, the data processing system can generate a first alert within the vehicle. The data processing system can generate the first alert responsive to detecting a first instance of a first type of event based on the first sequence of images. For example, the data processing system can input the first sequence of images into a machine learning model. In some cases, the data processing system can collect and include sensor data (e.g., vehicle speed data or other data regarding characteristics of the vehicle and / or the environment surrounding the vehicle) generated at the time of generation of the first sequence of images in the input into the machine learning model. The machine learning model may be trained to detect events and / or types of the events based on the input first sequence of images and / or sensor data. Examples of types of events that the machine learning model may detect include driver drowsiness, swerving across a lane line, speeding, a missed stop sign, being too close (e.g., within a threshold distance) to another vehicle, changing lanes into a lane with a vehicle in the blind spot, detecting an object within a defined distance of the vehicle, etc. The data processing system can execute the machine learning model based on the input to generate a confidence score for an event or a particular type of event (e.g., event type).

[0068] The data processing system can use a threshold (e.g., a confidence score threshold or an event threshold) to detect an event from the first sequence of images. For example, based on the input, the machine learning model can generate the confidence score for an event or a particular type of event. The data processing system or the machine learning model can compare the confidence score to a threshold (e.g., a threshold associated with all or multiple event types or specific to the type of event for which the machine learning model generated the confidence score). The data processing system can detect or determine an event and / or a specific type of the event is depicted in the first sequence of images responsive to determining the confidence score exceeds or otherwise satisfies the threshold. In some cases, the machine learning model can generate confidence scores for multiple types of events based on the first sequence of images and identify a confidence score that exceeds a threshold for a specific event type to detect the event.

[0069] Responsive to detecting the event (e.g., responsive to determining the confidence score exceeds a threshold), the data processing system can generate an alert. Thedata processing system can generate the alert by activating a device within the vehicle. For example, the data processing system can activate an auditory device, a visual device, and / or haptic device responsive to detecting the event. In some cases, the data processing system can generate an alert of an alert type based on the type of event that was detected. For example, the data processing system can store mappings of event types to types of alerts. Responsive to detecting an event, the data processing system can identify an event type of the event and compare the identified event type to the mapping. Based on the comparison, the data processing system can identify the type of alert to generate and generate the identified type of alert.

[0070] In some cases, the data processing system can detect the event and / or determine the type of the alert based on a profile of the driver driving the vehicle when the event occurred. For example, the data processing system can store the profile of the driver in memory. The profile can be or include a data structure, such as a table. The profile can include thresholds for different types of events and / or types of alerts for the different types of events. When detecting the events, the data processing system can use the thresholds stored in the profile when comparing the confidence scores generated by the machine learning model to the thresholds. Responsive to detecting an event, the data processing system can identify an indication of the type of alert that corresponds to the event type of the detected event. The data processing system can activate or otherwise generate an alert of the identified alert type.

[0071] In step 206, the data processing system can receive a second sequence of images. The second sequence of images can be images of a video (e.g., video data) of an environment within the vehicle (e.g., depicting the driver). The sequence of images can be consecutively captured images by the camera that captured the first sequence of images and / or a different camera (e.g., a second camera) mounted to or in the vehicle. The images can be of the same or a similar type to the first sequence of images. The second sequence of images can depict the driver before, during, and / or subsequent to generating and / or activating the alert. Accordingly the second sequence of images can depict a reaction of the driver to the alert.

[0072] In some cases, the data processing system can identify the second sequence of images from a plurality of images captured depicting the driver. For example, the data processing system can continuously receive images (e.g., individual images and / or images of a video) depicting the driver from the camera. The individual images can each correspond to a timestamp indicating the time in which the image was generated or captured. The dataprocessing system can identify a timestamp of the time in which the data processing system generated the alert. The data processing system can identify a first image or an initial image associated with a timestamp at or immediately after (e.g., the closest timestamp to the timestamp of the alert). The data processing system can identify the image and identify a defined number of images or images for a time period (e.g., one second, two seconds, three seconds, etc.) after (e.g., a defined time duration after) the timestamp of the initial image generated by the camera depicting the driver. The identified images can be the second sequence of images. In some cases, the data processing system can identify the first initial image based on the image corresponding with a time stamp a defined time period before the alert was generated, providing more details regarding the change in position (e.g., a change from a first position or pose to a second position or pose) or movement of the driver from a time prior to the alert to a time after the alert.

[0073] In step 208, the data processing system can execute a machine learning model using the second sequence of images. The machine learning model can be trained or configured to determine a reaction or change in position (e.g., movement or change in pose) of an individual within a sequence of images. The data processing system can execute the machine learning model using the second sequence of images as input. The execution can cause the machine learning model to output a value (e.g., a valence). The value can indicate a reaction to the generated alert. The value can have a sign (e.g., positive or negative) and a magnitude. The sign can indicate whether the reaction was positive (e.g., a positive reaction may be a thumbs up and / or a head nod) or negative (e.g., a negative reaction may be a thumbs down, a headshake, and / or a specific hand configuration). The machine learning model apply learned weights and / or parameters to the second sequence of images to detect movements by different body parts and / or changes in facial expressions or micro expressions by the driver. The machine learning model may detect such features and the changes in the features across the second sequence of images. The machine learning model may apply weights and / or parameters to the features and / or changes in features to generate a value indicating the reaction of the driver to the alert.

[0074] For example, the alert can be an audible alert that is played within the cabin of the vehicle. The second sequence of images can include an initial image captured when the audible alert is played. The subsequent images of the second sequence of images can depict how the driver reacted to the audible alert and specifically a head nod by the driver. Themachine learning model can detect the head nod and output a positive value with a high magnitude indicating a high positive reaction to the audible alert.

[0075] In another example, the first sequence of images can include different environmental factors (e.g., objects in the environment surrounding the vehicle) that cause the data processing system to generate an alert. An example of detecting environmental factors can include detecting a pedestrian crossing the road in front of the vehicle to cause the vehicle to generate an alert or a different object the vehicle is approaching. The data processing system can detect movement by the driver within a defined duration subsequent to the alert from the second sequence of images, such as movement indicating surprise by the driver to the alert or a sudden movement by the driver caused by contacting the object in front of the vehicle. The data processing system can output a value indicating the sudden movement was a result of contacting the object. The data processing system may not adjust the threshold based on such a determination or may otherwise decrease the threshold to cause the data processing system to generate future alerts earlier to avoid such future collisions.

[0076] In another example, referring now to FIG. 3A, a sequence 300 of a driver’s reaction or response to an alert is shown. The sequence 300 can include a timeline 301 of images. On the timeline 301, the sequence 300 can include an initial image 302 depicting an image of the driver’s face prior to any alert being played or activated. The sequence 300 can include an image 304 depicting an image of the driver’s face at the time in which the alert (e.g., the audio alert) was played within the vehicle. The sequence 300 can include an image 306 depicting an image of the driver’s face at a time after the alert was played. As depicted in the sequence, the driver’ s face remains the same in the images 302 and 304, but changed in the image 306 to depict an angry driver. In some cases, the changes in the driver’s face can be or include micro expressions relating to anger, such as, such as lowered or drawn eyebrows, bulging or intense eyes, compressed lips, etc. The data processing system can execute the machine learning model using the sequence 300 of images, and the machine learning model can output a negative value with a high magnitude based on the change in facial expression (e.g., the micro expressions) from an ambivalent demeanor to an angry demeanor or another negative demeanor associated with the expressions.

[0077] In some cases, the driver may not visually react to the alert. For example, referring now to FIG. 3B, a sequence 308 of a driver’s reaction or response to an alert is shown. The sequence 300 can include a timeline 309 of images. On the timeline 309, an initial image310 depicting an image of the driver’s face prior to any alert being played or activated. The sequence 308 can include an image 312 depicting an image of the driver’s face at the time in which the alert (e.g., the audio alert) was played within the vehicle. The sequence 308 can include an image 314 depicting an image of the driver’ s face at a time after the alert was played. As depicted in the sequence, the driver’s face remains the same in each of the images 310-314 of the sequence 308. The data processing system can execute the machine learning model using the sequence 308 of images, and the machine learning model can output a value with a magnitude of zero or substantially zero based on the lack of any change in facial expression of the driver.

[0078] In some cases, the data processing system can detect reactions or valence values indicating the sign and magnitude of reactions of drivers to alerts using response data other than visual data. For example, the data processing system can receive, from a microphone within the vehicle, audio data that includes sound segments during the same or a similar time period to the sequence of images that is used to detect the reaction (e.g., a defined time duration before and / or after generation of the alert). The data processing system can analyze the audio data using Fourier Transform techniques and a language model or another machine learning model to determine the context or content of the audio. For instance, the audio data can include intentional speech saying an alert was incorrect or speech that the machine learning model trained to detect reactions (or another machine learning model trained to detect reactions based on audio data) can use to determine reactions. For instance, the data processing system can use the machine learning model to determine the driver spoke a word with a particular tone or that has a meaning of surprise or disgust. The machine learning model can generate a negative value with a high magnitude in response to detecting such a reaction. In another example, the machine learning model can detect intentional speech saying an alert was incorrect. In such cases, the machine learning model can output a negative value with a high magnitude in response to detecting such a reaction.

[0079] In another example, after generating an alert, the data processing system can output a set of audio data with predefined questions in a survey about the driver’s experience with the alert (e.g., a survey about whether the driver appreciated the alert or preferences for improving alerts for a similar event type in the future). The driver can respond to each of the questions by speaking. The microphone within the vehicle can capture the speech and transmit the speech data to the data processing system. The data processing system can execute the machine learning model configured to determine contexts of speech and / or to determinereaction data from speech to determine a reaction value for the answers. In some cases, the data processing system can record the answers for preferences for future alerts for the same event type in the future, and the data processing system can adjust the driver’ s profile according to the preferences so the data processing system can generate the alerts according to the preferences for future detected events of the event type.

[0080] Another example of a type of reaction can be a selection of a button, such as a physical button or a virtual button. For instance, in response to generating an alert, the driver can select a virtual button on a user interface in the vehicle or a physical button in the vehicle indicating the alert was incorrect or correct. The machine learning model can generate a negative value with a high magnitude in response to detecting the selection of the button.

[0081] In some cases, the data processing system can determine an aggregate reaction value based on the different factors or types of reaction data that the data processing system captures. For example, the data processing system can cause the machine learning models to generate separate values for the respective audio data, image data, and / or button selection data. The data processing system can execute a function on the generated values to generate an aggregated value, such as by calculating an average, a median, or a sum of the values. The data processing system can use the aggregated value to determine whether to adjust the threshold and / or type of alert to generate for the event or the event type, as described above. The data processing system can calculate the aggregate value in any manner based on any type of data.

[0082] Referring again to FIG. 2, at step 210, the data processing system can determine whether the reaction by the driver to the alert was a positive reaction or a negative reaction. The data processing system can do so based on the output value of the machine learning model indicating the reaction to the alert depicted in the second sequence of images. For example, the data processing system can identify the sign associated with the value. The data processing system can determine a positive reaction (e.g., classify the reaction as positive) responsive to identifying a positive sign and / or a negative reaction (e.g., classify the reaction as negative) responsive to identifying a negative sign.

[0083] In some cases, the data processing system may only determine a positive or a negative reaction responsive to determining the magnitude of the value output by the machine learning model exceeds or satisfies a reaction threshold. For example, the data processing system can compare the magnitude of the value to the reaction threshold. Responsive todetermining the magnitude does not exceed or satisfy the reaction threshold, the data processing system may determine not to adjust or change the event threshold. Such may be advantageous, for example, because the data processing system can avoid adjusting the likelihood of generating an alert in cases in which the driver does not hear the alert.

[0084] In some cases, the data processing system may determine a negative reaction when the magnitude does not exceed the reaction threshold. For example, the data processing system can compare the magnitude of the value to the reaction threshold. Responsive to determining the magnitude does not exceed or satisfy the reaction threshold, the data processing system may determine a negative reaction. Such may be advantageous, for example, because the data processing system can determine the alert was not effective.

[0085] Responsive to determining a positive reaction from the second sequence of images, at step 212, can maintain the event threshold. For example, the data processing system can determine the sign of the generated value for the reaction is positive. Responsive to the determination, the data processing system may not adjust the event threshold and end the method 200 or otherwise proceed to operation 220. By doing so, the data processing system can avoid changing the event threshold when the driver indicated the alert was proper (e.g., reacted positively to the alert).

[0086] For instance, a driver may not be wearing a seatbelt. The data processing system may detect an event based on images depicting the driver not wearing a seatbelt using a threshold associated with not wearing a seatbelt and generate an alert accordingly. The driver may react positively to the alert, such as by smiling or performing a thumbs up motion. The data processing system can receive images of the positive reaction and determine the reaction was positive using the machine learning model configured to determine reaction values. Responsive to determining the positive reaction, the data processing system may not transmit images based on which the event was detected and / or may not change the threshold, depending on the configuration of the data processing system. Thus, the data processing system may use positive feedback to either make it more likely to generate similar alerts for similar events in the future or otherwise not change the likelihood of generating such alerts when training the system.

[0087] However, responsive to determining a negative reaction from the second sequence of images, at step 214, can transmit a message containing the first sequence of images to a remote computing device. For example, the data processing system can determine the signof the generated value for the reaction is negative. Responsive to the determination, the data processing system can identify the first sequence of images based on which the data processing system detected the first instance of the first type of event. The data processing system can generate a message containing the first sequence of images and transmit the message and first sequence of images to the remote computing device.

[0088] At operation 216, the data processing system can determine whether to maintain or adjust the event threshold. The data processing system may do so based on an instruction or determination by the remote computing device of whether the remote computing device detected the first instance of the event detected by the data processing system. For example, the data processing system can detect a negative reaction (e.g., a reaction with a negative valence value) to an alert, such as an alert generated based on a detection of an event of being too close to another vehicle from the first sequence of images. Responsive to detecting the negative reaction, the data processing system can transmit the first sequence of images to the remote computing device. The remote computing device can execute a model (e.g., a computer model or machine learning model configured or trained to detect events from sequences of images) to determine whether any events are depicted in the first sequence of images. In some cases, the remote computing system may have more processing power than the on-vehicle data processing system, so the remote computing system may be able to more accurately detect or determine whether the event should have been detected from the first sequence of images. When the remote computing device does not detect the same type of event as determined by the data processing system, the remote computing device can transmit an instruction (e.g., a message) to the data processing system to adjust the threshold for the type of event. The data processing system can adjust (e.g., increase or decrease) the threshold responsive to receiving the instruction at operation 218.

[0089] In one example, the data processing system and the remote computing device may be configured to adjust alerts for events in which the driver is at fault (e.g., swerving without a visible reason, speeding, not wearing a seatbelt, etc.). For example, a driver may be forced onto the side of the road after another vehicle cuts in front of the driver’s vehicle. The data processing system may detect an event of crossing over a line from images captured of the incident and generate an alert accordingly. In doing so, the data processing system may not detect or give enough weight to the other vehicle causing the event. The driver may react to the alert negatively, such as by furling an eyebrow or exclaiming that it was not the driver’s fault. The data processing system can transmit the images based on which the swerving eventwas detected to the remote computing device responsive to detecting the negative reaction by the driver. The remote computing device can execute the model to process the images to determine whether the event occurred or any event occurred. In doing so, the remote computing device may determine either that the event was not the fault of the driver based on the images depicting the other vehicle cutting in front of the driver or not detect the event at all because the remote computing device may only detect events that are the fault of the driver. The remote computing device can transmit (e.g., in a message) instructions to the remote computing device to adjust the threshold based on which the event was detected to the data processing system based on the determination. Responsive to receiving the instructions, the data processing system can adjust (e.g., increase) the threshold for the event such that the model can more accurately detect events that are the fault of the driver. In this way, the data processing system can proactively train the event detection machine learning model to detect events or only events that are the fault of the driver.

[0090] In another example, the remote computing device can transmit an instruction to adjust an event threshold (e.g., of the first event type). For example, the data processing system can detect a negative reaction (e.g., a reaction with a negative valence value) to an alert, such as an alert generated based on detection of an event of being too close to another vehicle from the first sequence of images. Responsive to detecting the negative reaction, the data processing system can transmit the first sequence of images to the remote computing device. The remote computing device can execute a model to determine whether any events are depicted in the first sequence of images. When the remote computing device does not detect the event or the same type of event as determined by the data processing system, the remote computing device can transmit an instruction or instructions to the data processing system to adjust the threshold for the type of event. The data processing system can adjust the threshold responsive to or based on receiving the instruction.

[0091] In some cases, the remote computing device may transmit an indication not to adjust the threshold for a detected event. For example, a driver may drive a vehicle too close (e.g., within a threshold of) to another vehicle. The data processing system may detect an event of being too close (e.g., tailgating) to the other vehicle using a threshold associated with tailgating other vehicle events from images captured of the incident and generate an alert accordingly. The driver may react to the alert negatively, such as by waving a fist in response to hearing the alert. The data processing system can transmit images based on which the tailgating event was detected to the remote computing device responsive to detecting thenegative reaction by the driver. The remote computing device can execute the model to process the images to determine whether the event occurred or any event occurred. In doing so, the remote computing device may determine that the event was correctly detected. The remote computing device can transmit (e.g., in a message) an instruction or an indication of a correct event detection to the data processing system. The data processing system may receive the instructions or indication and determine not to adjust the threshold accordingly. Thus, the data processing system can avoid drivers intentionally manipulating the event detection to reduce the amount of alerts that are generated or otherwise reduce the number of accurate alerts that are generated.

[0092] In some cases, the remote computing device can use the model to determine an accuracy of the event (e.g., accurate or inaccurate, or a level of accuracy of the event detection) or otherwise present the first sequence of images to a user at a client device that may view the first sequence of images and provide an input indicating whether the event detection was accurate or not. The remote computing system can transmit an indication of the accuracy or inaccuracy to the data processing system.

[0093] The data processing system can determine whether to adjust the event threshold based on the indication of accuracy for the event detection. For example, the data processing system may only adjust the threshold responsive to determining the indication is that the event detection was inaccurate. Such may be the case, for example, to improve the accuracy of the event detection by using the feedback by the driver. In another example, the data processing system may only adjust the threshold responsive to determining the indication is that the event detection was accurate. Such may be the case, for example, to personalize the adjustment to the driver’s reaction to make the alerts more desirable for the driver for accurately detected events. The data processing system may be configured to adjust the threshold based on any combination or number of factors.

[0094] In some cases, responsive to receiving an indication that the event was accurately detected and determining the driver had a negative reaction (e.g., a negative reaction above a threshold), the data processing system can adjust a type of the alert to generate for the future events of the same type. The data processing system can adjust the type of the alert form a first type of alert to a second type of alert, as described herein.

[0095] To adjust the threshold at operation 218, data processing system can increase or decrease the threshold (e.g., the event threshold or the confidence score threshold). The dataprocessing system can adjust the threshold based on which the data processing system initially detected the event that caused the data processing system to generate the alert. For instance, the data processing system can adjust the threshold for the event type for which the data processing system generated the alert. In some cases, the data processing system can adjust the threshold proportional to the magnitude of the value generated by the data processing system (e.g., the higher the magnitude of the value, the higher the adjustment). When adjusting the threshold, the data processing system can adjust the threshold in the profile of the driver.

[0096] In one example, the data processing system can adjust the threshold for alerts relating to drowsiness. For instance, the data processing system can generate an alert for a drowsiness event. Such an alert may be generated responsive to detecting a driver is yawning. The data processing system can generate an incorrect alert if a driver has their mouth ajar or they are singing. Such an alert may be played to them, and then they may immediately change their behavior, such as by closing their mouth, or the driver may perform some gesture which indicates that the alert was incorrect, and the detection of the alert was a false positive. The data processing system can detect the reaction by the driver and determine the reaction was negative. Responsive to the determination that the reaction was negative, the data processing system can transmit the images and / or any other data used to detect the drowsiness event to the remote computing device. In some cases, the data processing system can include an identification of the type of the event in the message. The remote computing device can process the data and determine the drowsiness event did not occur (e.g., by identifying the identification of the drowsiness event in the message and determining the drowsiness event did not occur using the model) or otherwise not detect the drowsiness event (e.g., by determining no event was detected that matched the drowsiness event in the message). Responsive to the determination, the remote computing device can transmit an instruction to the data processing system that causes the data processing system to adjust (e.g., increase or decrease) the threshold related to detecting drowsiness events to be more sensitive to avoid future false positives.

[0097] In another example, the data processing system can adjust the threshold for alerts relating to talking on a cell phone while driving. For instance, the data processing system can generate an alert for a talking on a cell phone event. Such an alert may be generated responsive to detecting a driver is holding a bag of chips. The alert may be played to them, then the driver may say it is not a cell phone, but a bag of chips. The data processing system can detect the audible reaction and determine the reaction was negative. Responsive to the determination that the reaction was negative, the data processing system can transmit theimages and / or any other data used to detect the talking on a cell phone event to the remote computing device. In some cases, the data processing system can include an identification of the type of the event in the message. The remote computing device can process the data and determine the talking on a cell phone event did not occur (e.g., by identifying the identification of the talking on a cell phone event in the message and determining the talking on a cell phone event did not occur using the model) or otherwise not detect the talking on a cell phone event (e.g., by determining no event was detected that matched the talking on a cell phone event in the message). Responsive to the determination, the remote computing device can transmit an instruction to the data processing system that causes the data processing system to adjust (e.g., increase or decrease) the threshold related to detecting cell phone events to be more sensitive to avoid future false positives.

[0098] In another example, the data processing system can adjust the threshold for alerts relating to driving too close to another vehicle. For instance, the data processing system can generate an alert for tailgating a vehicle. However, the data processing system may actually detect a hood ornament on top of the car or a different static object that's in front of the camera. The driver may respond to the alert by saying “Hey, I'm not driving, or I'm not following too close.” The data processing system can detect the audible reaction and determine the reaction is negative. Responsive to determining the audible reaction was negative, the data processing system can transmit the images and / or any other data used to detect the tailgating event to the remote computing device. In some cases, the data processing system can include an identification of the type of the event in the message. The remote computing device can process the data and determine the tailgating event did not occur (e.g., by identifying the identification of the tailgating event in the message and determining the tailgating event did not occur using the model) or otherwise not detect the tailgating event (e.g., by determining no event was detected that matched the tailgating event in the message). Responsive to the determination, the remote computing device can transmit an instruction to the data processing system that causes the data processing system to adjust (e.g., increase or decrease) the threshold related to tailgating to be more sensitive to avoid future false positives because it's having a detrimental experience for the driver.

[0099] In another example, the data processing system can adjust the threshold for alerts relating to wearing a seatbelt. For instance, the data processing system can generate an alert for failing to wear a seatbelt. However, the driver may be wearing a seatbelts but it may not be visible in the sequence of images based on which the data processing system generatedthe alert. The driver may respond or react to the alert with an audible response saying the driver is wearing a seatbelt or by showing the camera the buckled seat belt. The data processing system can detect such feedback and determine the reaction is negative. Responsive to determining the reaction was negative, the data processing system can transmit the images and / or any other data used to detect the failure to wear a seatbelt event to the remote computing device. In some cases, the data processing system can include an identification of the type of the event in the message. The remote computing device can process the data and determine the failure to wear a seatbelt event did not occur (e.g., by identifying the identification of the failure to wear a seatbelt event in the message and determining the seatbelt event did not occur using the model) or otherwise not detect the seatbelt event (e.g., by determining no event was detected that matched the failure to wear a seatbelt event in the message). Responsive to the determination, the remote computing device can transmit an instruction to the data processing system that causes the data processing system to adjust (e.g., increase or decrease) the threshold related to wearing a seatbelt to be more sensitive to avoid future false positives.

[0100] In some cases, instead of or in addition to adjusting the threshold, the data processing system can train the machine learning model that detected the incorrect or inaccurate reaction from the first sequence of images. For example, the data processing system can identify the instruction from the remote computing device indicating the first instance of the first type of event was not detected from the first sequence of images. Responsive to the identification, the data processing system can adjust the weights of the machine learning model using the first sequence of images as input and the ground truth (e.g., using backpropagation techniques with a loss function) that the first type of event was not found in first sequence of images. By doing so, the data processing system can train the machine learning model in real time.

[0101] In some cases, the data processing system can adjust the threshold based on the reaction of the driver indicating that the first alert was generated too early or too late. For example, the execution of the machine learning model on the second sequence of images can cause the machine learning model to output an indication of a reaction indicating that the first alert was generated too early or too late. Examples of reactions that indicate the first alert was generated too early include a surprised facial expression, an angry facial expression, an annoyed facial expression, etc. Examples of reactions that indicate the first alert was generated too late include throwing hands in the air or a sudden body lurch forward (e.g., in instances of a crash), etc. The machine learning model may be trained or otherwise learn to detect suchmovements and / or the corresponding indications of an early or late alert generation. The machine learning model can output such indications in addition to or instead of the value indicating a positive or negative reaction and the magnitude of the reaction, in some cases.

[0102] The data processing system can adjust the threshold (e.g., the threshold for the first event type) based on the indication of whether the first alert was too late or too early. For example, the data processing system can increase the threshold responsive to determining the indication is that the first alert was generated too early. The data processing system can decrease the threshold responsive to determining the indication is that the first alert was generated too late. In doing so, the data processing system can increase or decrease the amount of data (e.g., number of images and / or collected sensor) that is needed to detect an event to generate an alert, increasing or reducing the amount of time it may take to detect an event to generate an alert.

[0103] In some cases, the threshold for the first type of event may correspond to a particular type of driving environment (e.g., a weather related driving environment, such as raining or snowing, and / or geographical based driving environment, such as in the city, in a rural area, driving up a mountain, and / or driving in a flat area). The data processing system may similarly store thresholds for any type of driving environment and / or any combination of driving environments. When detecting the event from the first sequence of images, the data processing system can first determine the type of driving environment in which the first sequence of images were captured, such as based on geolocation data and / or image data. The data processing system can identify the threshold that corresponds to the type of driving environment and / or the first type of event. The data processing system can adjust the identified threshold based on the reaction of the driver depicted in the second sequence of images. The data processing system can similarly store and adjust any number of thresholds for different driving environments and / or alert types based on driver reactions to alerts generated in the different driving environments. Accordingly, the data processing system can uniquely adjust different thresholds for more accurate alert determinations taking the context of the alerts into account, further improving the accuracy and effectiveness of the generated alerts.

[0104] In some cases, the data processing system can adjust the threshold based on a reaction category of the reaction depicted in the second sequence of images. For example, the execution of the machine learning model on the second sequence of images can cause the machine learning model to output an indication of a classification of a type of the reaction.Examples of types of reactions can include a head nod, micro expressions, a head shake, a thumbs up, a thumbs down, a sudden change in position or pose, etc. The data processing system can store a mapping (e.g., a table) of different types of reactions to adjustments (e.g., increases and / or decreases and / or magnitudes of the increases and / or decreases) to thresholds in memory. The mapping can be learned based on previously determined adjustments made by the data processing system based on reaction values identified by the machine learning model and / or input or adjusted by a user or administrator. The data processing system can compare the indication of the type of the reaction generated by the machine learning model based on the second sequence of images with the mapping to identify an adjustment to the threshold. The data processing system can adjust (e.g., increase or decrease) the threshold based on the adjustment. In doing so, the data processing system can adjust the threshold more objectively, reducing the risk of inaccurate adjustments to the threshold.

[0105] In some cases, the data processing system can use a driver score for the driver to adjust the threshold. For example, the data processing system can include a driver score for the driver in the driver’s profile. The driver score can be a value indicating historical driving performance by the driver. The data processing system can determine the driver score or the data processing system can request or receive the driver score from a remote computing system. In one example, the driver score can be generated and / or changed over time based on positive and / or negative driving events. For instance, the data processing system or the remote computing system can increase the driver score for positive driving events that the data processing system detects (e.g., following the speed limit, stopping at stop signs, maintaining at least a threshold distance behind a vehicle, etc.) and / or decrease the driver score based on negative driving events that the data processing system detects (e.g., accidents, tailgating, speeding, texting while driving, etc.). The data processing system can update the driver score over time as the data processing system detects events.

[0106] The data processing system can use the driver score (e.g., the current state of the driver score) to adjust the threshold. For instance, the data processing system can use the score as a weight. In doing so, the data processing system can convert the driver score to a value between a defined range (e.g., between zero and one, between zero and 10, etc.). The data processing system can multiply a determined adjustment based on the weight to determine a final adjustment. In doing so, the data processing system can make larger adjustments for better drivers or drivers with higher driver scores, which can aid in the data processing system identifying the correct threshold faster.

[0107] At step 220, the data processing system can receive a third sequence of images. The third sequence of images can be images of a video (e.g., video data) of an environment surrounding a vehicle or the environment within the vehicle (e.g., depicting the driver). The sequence of images can be consecutively captured images by the same camera (e.g., the same camera that captured the first sequence of images). The images of the sequence can be captured at a set time interval (e.g., every five seconds) or pseudo-randomly . The data processing system can receive the third sequence of images in the same or a similar manner to the manner in which the data processing system received the first sequence of images.

[0108] In step 222, the data processing system can generate a second alert within the vehicle. The data processing system can generate the second alert responsive to detecting a second instance of the first type of event based on the third sequence of images. For example, the data processing system can input the third sequence of images into the machine learning model trained or configured to detect events depicted in sequences of images. In some cases, the data processing system can collect and include sensor data (e.g., vehicle speed data or other data regarding characteristics of the vehicle and / or the environment surrounding the vehicle) generated at the time of generation of the third sequence of images in the input into the machine learning model. The data processing system can execute the machine learning model based on the input to generate a confidence score for an event or a particular type of event (e.g., event type).

[0109] The data processing system can use the maintained or adjusted threshold to detect an event from the third sequence of images. For example, based on the input, the machine learning model can generate the confidence score for an event or an event of the same event type as the first detected event. The data processing system or the machine learning model can compare the confidence score to the maintained or adjusted threshold. The data processing system can detect or determine a second event (e.g., a second event of the same event type as the first event) is depicted in the sequence of images responsive to determining the confidence score exceeds or otherwise satisfies the adjusted threshold.

[0110] Responsive to detecting the event (e.g., responsive to determining the confidence score exceeds the adjusted threshold), the data processing system can generate an alert (e.g., a second alert). The data processing system can generate the alert by activating a device within the vehicle. In some cases, the data processing system can generate the alertbased on the type of the detected event. The data processing system can generate the alert in the manner described with respect to step 204.[OHl] In some cases, the data processing system can implement the method 200 to evaluate driver responses to alerts in the context of their historical driving performance and behavior patterns. For instance, if a driver with a consistently high safety score (e.g., 950 out of 1000) displays a negative gesture in response to an alert, the data processing system may interpret this as a likely false positive alert and increase the threshold corresponding to generating an alert for the event type. If a driver with a lower safety score who typically receives many alerts rarely exhibits negative gestures, but suddenly does so, this may be weighted differently in the data processing system's evaluation of the alert's validity.

[0112] In some cases, the data processing system can implement adaptive coaching strategies based on analyzing individual driver responses and receptiveness to different coaching approaches. For example, through continuous monitoring of driver reactions, including body language and behavioral changes, the data processing system can determine optimal coaching methods for each driver - whether they respond better to frequent reminders, directive instructions, or general recommendations. The data processing system may incorporate historical driving patterns to provide more contextual coaching, such as identifying specific scenarios where a driver tends to exhibit risky behavior (e.g., accelerating through stale yellow lights) and proactively providing targeted guidance. This personalized approach allows the data processing system to maximize coaching effectiveness by tailoring both the content and delivery style of alerts based on observed driver receptiveness and response patterns.

[0113] In some cases, the data processing system can use the adjustments to the threshold for self-driving decisions. For example, the data processing systems may use the same threshold (e.g., same event threshold) to detect events for self-driving decisions or otherwise a threshold that corresponds to the maintained or adjusted threshold (e.g., adjusts responsive to any adjustment of the threshold). The data processing system can use the threshold to detect events for self driving and determine operations of the vehicle. For example, the data processing system can use the maintained or adjusted threshold to detect that the vehicle is too close to another vehicle and slow the vehicle down accordingly. The data processing system can operate the vehicle using such thresholds in any manner.

[0114] FIG. 4 illustrates a flow of a method 400 for implementing a feedback-based driving system executed by a data processing system, according to some embodiments. The data processing system can be or include a computing system of a vehicle (e.g., the feedbackbased driving system 105) and / or a remote computing system (e.g., a cloud server or the cloud computing system 115) for using a feedback-based alert generation and / or self-driving decision-making, in accordance with an embodiment. The method 400 is shown to include steps 402-412. However, other embodiments may include additional or alternative steps, or may omit one or more steps altogether. Different steps can be performed by different computing systems (e.g., the computing system of the vehicle can perform one or more of the steps 402-412 and / or the remote computing system can perform one or more of the steps 402- 412) and / or the different computing systems and can operate together to perform individual steps of the steps 402-412. One or more of the steps 402-412 can occur concurrently, be the same, or otherwise correspond to the steps 202-218.

[0115] In step 402, the data processing system can store a profile for a driver. The data processing system can store the profile for the driver in memory. The profile can be or include a data structure, such as a table. The profile can include thresholds for different types of events and / or types of alerts for the different types of events. When detecting the events, the data processing system can use the thresholds stored in the profile when comparing the confidence scores generated by the machine learning model to the thresholds. Responsive to detecting an event, the data processing system can identify an indication of the type of alert that corresponds to the event type of the detected event. The data processing system can activate or otherwise generate an alert of the identified alert type.

[0116] In step 404, the data processing system can receive a first sequence of images. The data processing system can receive the first sequence of images in the same or a similar manner to the manner described with reference to step 202 of FIG. 2.

[0117] In step 406, the data processing system can generate a first alert within the vehicle. The data processing system can generate the first alert responsive to detecting a first instance of a first type of event based on the first sequence of images. The data processing system can generate the first alert in the same manner described with respect to step 204. For example, the data processing system can input the first sequence of images into a machine learning model trained to detect events and / or types of the events based on input sequences of images. In some cases, the data processing system can collect and include sensor data (e.g.,vehicle speed data or other data regarding characteristics of the vehicle and / or the environment surrounding the vehicle) generated at the time of generation of the first sequence of images in the input into the machine learning model. The data processing system can execute the machine learning model to detect the first instance of the first type of event, such as by using a threshold for the first type of event as described above (e.g., by comparing a confidence score generated by the machine learning model to the threshold for the first type of event in the profile).

[0118] Responsive to detecting the event (e.g., responsive to determining the confidence score exceeds a threshold), the data processing system can generate an alert. The data processing system can generate the alert by activating a device within the vehicle. For example, the data processing system can activate an auditory device, a visual device, and / or a haptic device responsive to detecting the event.

[0119] The data processing system can determine a first type of the alert to generate based on the detection of the event based on the stored profile of the driver driving the vehicle when the event occurred. For example, the data processing system can identify the driver’s profile from memory. The data processing system can identify the identification of the type of alert to generate for the first type of event from the profile. For instance, the data processing system can identify an indication to generate an audible, a visual, and / or a haptic alert for the first event type as the first type of alert for events of the first event type. Responsive to identifying the indication of the first type of alert, the data processing system can select and / or generate a first alert of the first type of alert within the vehicle. The data processing system can activate or otherwise generate an alert of the identified alert type.

[0120] At step 408, the data processing system can adjust the indication of the first type of alert for the first type of event in the profile. The data processing system can do so in response to and / or based on detecting a response by the driver to the first alert. For example, the data processing system can receive a second sequence of images. The second sequence of images can be images of a video (e.g., video data) of an environment within the vehicle (e.g., depicting the driver). The sequence of images can be consecutively captured images by the camera that captured the first sequence of images and / or a different camera (e.g., a second camera) mounted to or in the vehicle. The images can be of the same or a similar type to the first sequence of images. The second sequence of images can depict the driver before, during, and / or subsequent to generating and / or activating the alert. Accordingly, the second sequenceof images can depict a reaction of the driver to the alert. In some cases, the data processing system can identify the second sequence of images in the same manner as described with respect to step 208 of FIG. 2.

[0121] The data processing system can execute a machine learning model using the second sequence of images. The machine learning model can be trained or configured to determine a reaction or change in position or pose of an individual within a sequence of images. The data processing system can execute the machine learning model using the second sequence of images as input. The execution can cause the machine learning model to output a value (e.g., a valence). The value can indicate a reaction to the generated alert. The value can have a sign (e.g., positive or negative) and a magnitude. The sign can indicate whether the reaction was positive (e.g., a positive reaction may be a thumbs up and / or a head nod) or negative (e.g., a negative reaction may be a thumbs down, a headshake, a gesture, and / or a specific hand configuration). The machine learning model can apply learned weights and / or parameters to the second sequence of images to detect movements by different body parts and / or changes in facial expressions by the driver. The machine learning model may detect such features and the change of the features across the second sequence of images. The machine learning model may apply weights and / or parameters to the features and / or changes in features to generate a value indicating the reaction of the driver to the alert of the first alert type.

[0122] For example, the alert can be a visual alert (e.g., a flashing light on a display of the vehicle or a separate light within or on the exterior of the vehicle) that is generated, activated, or presented within the cabin of the vehicle and / or on the exterior of the vehicle. The second sequence of images can include an initial image captured when the visual alert is played. The subsequent images of the second sequence of images can depict how the driver reacted to the visual alert, such as a thumbs up by the driver. The machine learning model can detect the thumbs up and output a positive value with a high magnitude indicating a high positive reaction to the audible alert. In another example, the machine learning model can detect a head shake and output a negative value with a high magnitude indicating a high negative reaction to the audible alert.

[0123] Responsive to detecting a negative reaction and / or a negative reaction with a magnitude above a threshold, the data processing system can adjust the indication of the type of alert to generate for the first type of event in the profile of the driver to a second type of alert. For example, the first type of alert can be the visual alert, and the data processing systemcan change the first type of alert to an audible alert as the second type of alert. In some cases, the data processing system can change to the second type of alert by maintaining the type of the alert (e.g., keeping the alert audible if the first type of alert was audible or visual if the first type of alert was visual), but adjust the magnitude of the alert. For example, the data processing system can decrease or increase the magnitude of the alert (e.g., increase or decrease the brightness of the alert if visual, the volume of the alert if audible, and / or strength of the alert if haptic). In some cases, in which the alert is audible, the data processing system can filter through different audio messages to generate for the alert responsive to determining a high negative reaction value (e.g., valence). The data processing system can adjust the magnitude proportional (e.g., by multiplying the magnitude by a defined value, such as 0.1) to the magnitude of the negative reaction (e.g., a higher magnitude corresponds to a higher adjustment).

[0124] In some cases, the data processing system can change the type of the alert by changing one or more characteristics of the alert. For example, the data processing system can change the sound of an audible alert (e.g., change the tone), change the color or the picture shown for a visual alert, and / or change a frequency of a haptic alert. In some cases, the data processing system can make the change by comparing the magnitude of the reaction to a plurality of thresholds (e.g., alert adjustment thresholds) that each correspond to a different change in the alert. The data processing system can identify the highest threshold that the magnitude exceeds and the type of alert that corresponds to the identified threshold. The data processing systems can adjust the first type of alert to the identified type of alert that corresponds to the identified threshold.

[0125] In some cases, the data processing system can detect a response to the first alert based on one or more other factors. For example, the driver can provide a manual input on a user interface or device in the vehicle. The manual input can be a selection of virtual or physical button or another indication of the driver’s reaction to the first alert. For instance, the manual input can indicate a positive reaction or a negative reaction to the first alert. In another example, the driver can provide an audible sound to a microphone within the vehicle. The audible sound can be an indication of a positive reaction or a negative reaction to the first alert. The data processing system can identify such responses and determine whether and / or how to adjust the threshold for the event type and / or adjust the type of alert to generate for the event type.

[0126] For instance, in one example, after generating an alert, the driver can exclaim, “What?!” The driver may do so to indicate confusion as to why the alert was generated. The data processing system can receive audio data from the microphone capturing the exclamation from the driver and process the audio data to determine a negative reaction to the alert The data processing system can determine a magnitude of the negative reaction (e.g., a magnitude of the valence of the negative reaction) exceeds the threshold, which may indicate that the alert was a false positive. Responsive to the determination, the data processing system can increase the threshold for the event type to reduce similar repeat false positives. In another example, after generating an alert for swerving into an adjacent lane, the driver may speak an audible sound of “That was not my fault!”, which the driver may say because another vehicle may have forced the driver to swerve into the adjacent lane. The data processing system can receive audio data from the microphone capturing the exclamation from the driver and process the audio data to determine a negative reaction to the alert. The data processing system can determine a magnitude of the negative reaction exceeds the threshold. Responsive to the determination, the data processing system can increase the threshold for the event based on the detection of the negative reaction to reduce similar repeat false positives. In some cases, the data processing system can adjust the threshold and / or type of alert after or responsive to transmitting the images to a remote computing device and confirming the event was incorrectly detected, as described above.

[0127] In some cases, the indication or the adjusted indication can be changed or updated based on a user input. For example, a user (e.g., the driver or a fleet manager) can view the driver’s profile using a computing device or a client device. The user can change or update the indication of the type of alert to generate in response to the first type of event via a user interface being presented on the computing device or the client device. Thus, the types of alerts to generate for the driver may be manually configurable.

[0128] At step 410, the data processing system can receive a third sequence of images. The data processing system can receive the third sequence after adjusting the indication of the type of alert to generate for the first type of event in the profile. The third sequence of images can be images of a video (e.g., video data) of an environment surrounding a vehicle or the environment within the vehicle (e.g., depicting the driver). The sequence of images can be consecutively captured images by the same camera (e.g., the same camera that captured the first sequence of images). The images of the sequence can be captured at a set time interval (e.g., every five seconds) or pseudo-randomly. The data processing system can receive thethird sequence of images in the same or a similar manner to the manner in which the data processing system received the first sequence of images.

[0129] In step 412, the data processing system can generate a second alert within the vehicle. The data processing system can generate the second alert responsive to detecting a second instance of the first type of event based on the third sequence of images. For example, the data processing system can input the third sequence of images into the machine learning model trained or configured to detect events depicted in sequences of images. In some cases, the data processing system can collect and include sensor data (e.g., vehicle speed data or other data regarding characteristics of the vehicle and / or the environment surrounding the vehicle) generated at the time of generation of the third sequence of images in the input into the machine learning model. The data processing system can execute the machine learning model to detect the second instance of the first type of event (e.g., event type).

[0130] The data processing system can use the adjusted indication of the type of alert to generate in response to detecting the first type of event. For example, responsive to detecting the first event type, the data processing system can identify the profile of the driver from memory. The data processing system can identify the indication of the type of alert to generate from the type of event from the profile that was adjusted based on the driver’s reaction to the first instance of the first type of event.

[0131] The data processing system can generate an alert of the type of alert for the first type of event indicated in the profile of the driver. The data processing system can generate the alert by activating a device within the vehicle. The data processing system can generate the alert in the manner described with respect to step 406. In some cases, the data processing system can perform the method 400 concurrently and / or for the same alerts and / or events as those described with reference to the method 200.

[0132] In some cases, the driver’s profile can follow the driver as the driver uses different vehicles for alert generation. For example, the driver can enter the vehicle or a different vehicle. Upon entering the vehicle, the data processing system of the computing system of the vehicle can identify the driver’s profile, such as based on a user input. Responsive to doing so, the data processing system and / or the computing system can use the indications of alert types and / or thresholds in the profile to detect events and / or determine the types of alerts to generate for the vehicle the driver is driving.

[0133] In a non-limiting example, the data processing system can operate as an on- vehicle device of a vehicle being driven by a driver. The data processing system can store a profile for the driver in memory. The data processing system can receive a first sequence of images of the environment in front of the vehicle and sensor data from a depth perception sensor (e.g., a reflection sensor) captured during the same time period as the first sequence of images. The data processing system can process the data using a first machine learning model to detect an event that a pedestrian is crossing the road in front of the vehicle. In doing so, the data processing system can cause the first machine learning model to generate a confidence score of .70 for the event that the pedestrian is crossing the road in front of the vehicle. The data processing system can detect the event responsive to determining the confidence score exceeds an event threshold for the event of a pedestrian crossing the road in front of the vehicle of .65 stored in the profile.

[0134] Responsive to detecting the event, the data processing system can generate an alert. The data processing system can determine or select the type of alert to generate as an audible beep from a microphone in the vehicle based on an indication of an audible alert corresponding to the event of detecting a pedestrian crossed the road. The data processing system can receive and / or process images depicting the driver captured during a time period beginning one second before the alert was generated and ending two seconds after the alert was generated. The data processing system can use a second machine learning model to detect a negative reaction by the driver based on images depicting the driver shaking his or her head.

[0135] The data processing system can use the detected reaction to adjust the threshold and / or the type of alert to generate for future events of detecting a pedestrian crossing the road. For example, the data processing system can determine the driver’s reaction was negative a magnitude of the reaction of .8 from the output of the second machine learning model. Responsive to doing so, the data processing system can transmit the first sequence of images to a remote computing device. The remote computing device can determine that the event of a pedestrian crossing the road is not depicted in the first sequence of images and transmit an instruction to adjust the event threshold based on which the pedestrian crossing the road event was detected. The data processing system can determine to adjust the event threshold for the event type of detecting the pedestrian is crossing the road responsive to determining the driver’ s reaction was negative and / or responsive to the instruction from the remote computing device. The data processing system can increase the event threshold by an amount proportional to the magnitude, such as by increasing the threshold by .08 based on the magnitude of the reactionbeing .8 (e.g., by multiplying the magnitude by a defined value of 1). Additionally or instead, in some cases responsive to the instruction from the remote computing device, the data processing system can change the type of the alert from an audible alert to a visual alert based on the negative reaction to the audible alert having a magnitude exceeding a threshold associated with visual alerts. The data processing system can make such adjustments in the driver’s profile.

[0136] Subsequently, the data processing system can detect a pedestrian in front of the vehicle. The data processing system can detect the pedestrian in front of the vehicle by inputting a sequence of images depicting the pedestrian into the machine learning model configured for event detection and to cause the machine learning model to generate a confidence score of .75 for the event of a pedestrian crossing the road. The data processing system can compare .75 to the adjusted threshold of .73 to determine the confidence score exceeds the adjusted threshold and detect that the event occurred. Based on the detection, the data processing system can generate a visual alert to alert the driver of the pedestrian. The data processing system can repeat the feedback loop to adjust the threshold and / or type of alert to generate any number of times. If the data processing system generates a confidence score for the event from another sequence of images that is below the adjusted threshold, the data processing system may not detect the event or generate any alerts for the event.Example Embodiments

[0137] Certain aspects of the present disclosure generally relate to monitoring driving behavior, particularly, monitoring a driver’s response to a safety alert and modifying parameters of a driver monitoring system based on the driver’s response.

[0138] Intelligent driver or driving monitoring systems (DMS), advanced driving assistance systems (ADAS), autonomous driving systems, camera-based Al Safety systems include machine learning models to assist a driver of a vehicle in which the systems are installed. In addition, the driver monitoring systems may provide alerts to the driver to achieve positive driving behavior or reduce unsafe driving behavior. The alerts are provided to the driver upon detection of certain driving events such as collision detection, distracted driving detection, intersection crossing detection, etc. The alerts can be provided to the driver through the output modules installed in the DMS systems (such as speakers, display screens, etc.).

[0139] In a commercial driving context, the DMS systems are configured to provide alerts based on the requirements of a fleet operator / manager. For example, the frequency and thresholds for the alerts can be set based on the requirements of the fleet manager. However, the alerts may not be effective for some drivers.

[0140] Therefore, there is a need for a system to monitor a driver’s response to safety alerts and to modify the parameters of the DMS so that safety alerts are effective for more drivers.

[0141] Driver monitoring systems (DMS) are installed in vehicles to monitor the driving behavior of a driver of a vehicle. The DMS is an edge device that may include one or more sensor modules to capture sensor data related to the driving session. Further, the edge device may include a processor to process the captured sensor data to provide inferences. The edge device may generate alerts based on the inferences and the edge device has real-time alerting mechanisms that may change driving behavior. The alerts may be provided to the driver via the output modules (such as speaker, display device, haptic device) of the edge device.

[0142] The edge device may be configured with various alerting mechanisms based on the requirements of a fleet management. In some scenarios, the drivers may respond in a way that reflects an immediate and / or long-lasting improvement in their overall driving safety. However, in some scenarios, the drivers may respond to the alerts / or the way the alerts are notified in a way that indicates that the alerting mechanism is wrong and / or unwelcome. In some scenarios, the driver may not react to the alerts and continue to show unchanged driving behavior. In order to make alerting mechanisms effective, the edge device may capture and process the reactions / responses to the alerts from the drivers. Thereafter, the edge device may modify the alerts and alerting mechanisms based on the inferred reactions of the driver and driving behavior.

[0143] The edge device may include an inward camera in the cabin of the vehicle to capture the driver’s behavior. The inward camera is configured to capture visual data that may include the reaction of the driver upon raising an alert. Further, the edge device may monitor the driving behavior upon raising an alert. The edge device may detect changes in the driving behavior based on the captured visual data and inertial sensor data. The edge device may modify the alert based on the detected change in driving behavior.

[0144] Further, the edge device can be connected to a cloud server that may send configuration settings for alerts and alerting mechanisms to the edge device. In addition, the cloud server may perform processing tasks for the edge device. In one embodiment, the edge device may communicate the detected events and / or corresponding alerts to the cloud server and related sensor data that led to generation of alerts. The sensor data can be further processed at the cloud server to verify the accuracy of the detected events.

[0145] An environment (e.g., the environment of the system 100) for monitoring a driver’s response to safety alerts may be used for a feedback-based alert and / or vehicle control system, in accordance with an embodiment of the disclosure. The environment can include an edge device (e.g., the feedback-based driving system 105) connected to a cloud server (e.g., the cloud computing system 115) via a network. The edge device is installed in a vehicle and the vehicle may be a part of a fleet. The fleet management device may be used to manage the devices installed in the vehicles of the fleet. A driver may be assigned to a vehicle and the edge device may be configured to monitor driving behavior.

[0146] Some drivers of a vehicle may exhibit unsafe driving behavior for various reasons. The edge device is configured to detect driving behavior events and raises alerts accordingly to inculcate safe driving behavior. Driver responses to alerts may vary. Some drivers may not respond to alerts and continue to drive unsafely. Other drivers may respond to alerts and tend to change their driving behavior. In some cases, drivers may provide negative feedback to the alerts either verbally or with physical gestures. For example, the driver may provide verbal feedback to the edge device or make a facial gesture indicating disagreement or preference for alert adjustment, which may be directed to the camera of the edge device as a response to the alert, or without changing the direction of gaze (as shown in FIG. 3A). In some scenarios, the driver may exhibit a negative response even though the alert was valid. In other scenarios, the negative response can be because of the belief that the driver hasn’t shown unsafe driving behavior as indicated by that the alert was invalid.

[0147] In some cases, the alerts may be generated because of false positive inferences. Processing the response from the driver and inferring or receiving the feedback from the driver may help in identifying false positive inferences. The visual data (from inward-facing and / or outward-facing cameras) that has resulted in false positive inferences can be uploaded to the cloud server, which can be later used to retrain the machine learning models deployed on the edge device. Identifying false positive cases based on the driver feedback processing therebyresults in efficient utilization of bandwidth resources to selectively upload identified false positives to the cloud server. To make effective use of driver responses as feedback, the edge device may be configured to process the visual and / or other sensor data using a trained machine learning model and calculate a feedback score, and / or rating based on the type of feedback. To that effect, the edge device can be configured to process the visual data for a time duration after an alert is raised to detect a presence and type of feedback provided by the driver with respect to the alert.

[0148] In some cases, the alerts can be valid, but the driver’s feedback may indicate that the alert is invalid. To differentiate such scenarios, the driver’s past driving behavior may also be used in addition to the driver’s responses to the alert in determining the validity of the alert, thereby resulting in improvement of the identification of false positive alerts.

[0149] In one embodiment, the edge device may communicate the alerts to the driver via audio, visual, or haptic modules connected to the edge device. In one example, these modules can be included in the edge device. In another example, these modules may be integrated into the vehicle and / or a smartphone or tablet. The frequency at which alerts are generated and communicated can be defined by fleet management based on their requirements. For example, if a driver is using a mobile phone while driving, the edge device can identify that the driver is distracted in the first inference, but the edge device can be configured to communicate the alert only if the driver continues to use the mobile phone for a time period after the first inference. The time period can be configured based on the requirements of the fleet management. An alert can be communicated to the driver upon detection of unsafe driving behavior. In some examples, alerts can be generated and communicated to the driver based on external conditions such as weather, current location of the vehicle, etc. In one example, alerts can be preventative alerts such as “wear seatbelt when beginning a driving session,” “use phone to see navigation before starting the driving session,” “approaching stop sign, please stop,” etc.

[0150] Different types of unsafe driving behavior may lead to the generation of different alerts and the means used for their communication. For example, if the driver is detected as drowsy based on the visual data, a haptic alert can be selected instead of an audio or visual alert and the haptic alert can be communicated via the haptic module installed in the seat of the vehicle. In another example, if the driver is distracted because of phone usage, an audio alert can be generated and communicated instead of other alert types. However, in some cases, the frequency of alerts and the modes of alert communication may attract negativeresponses from drivers even though the configuration settings for alerts were defined by respective fleet management. This may lead to unsafe driving behavior from the drivers. In such cases, the validity of alerts may be verified by uploading and processing visual data on the cloud server. In one embodiment, the visual data may be uploaded to the cloud server when a driver has tended to show safe driving behavior in the past, as this may indicate that the detected event was a false positive and would therefore be a valuable training data example to improve the machine learning models underlying the event detection. The visual data may not be uploaded to the cloud server when a driver has tended to show unsafe driving behavior in the past, as this may indicate that the alert was valid. However, events for which the driver’s response was negative but that are expected to be valid may be uploaded as these may be useful teaching examples, or because it may be expected that the driver will be more likely to request a review of such examples. In one example, the past driving behavior of a driver can be indicated by a score / rating provided to the driver based on prior driving sessions.

[0151] In one embodiment, the edge device may process audio and video data captured by the sensor module after an alert is communicated to the driver to detect the type of response to the alert exhibited by the driver or received from the driver. The edge device may include a machine learning model to process various sensor data related to the feedback and make inferences based on the processing.

[0152] FIG. 5 shows a sequence 500 for detecting a driver’s response to a safety alert based on captured visual data of the driver using a machine learning model. In one embodiment, the machine learning model may be trained at the cloud server prior to deployment on the edge device. The edge device may modify the alert and alerting mechanisms based on the determination of a type of response or feedback received from the driver. For example, the volume of audio alerts can be increased or decreased or the sensitivity of a haptic alert can be increased or decreased based on determined feedback. In one embodiment, the edge device may modify a subsequent alert based on a change in driving behavior of the driver (which may be considered in addition to the driver’s feedback or independently). For example, upon raising an alert if it is detected that the driver has started exhibiting safe driving behavior then a subsequent alert may be modified to be less sensitive. In another example, if the driver has continued to exhibit unsafe driving behavior after raising an alert, then the subsequent alert can be modified to be more intensive (e.g., increasing the volume of the audio alert and increasing the magnitude of the haptic alert). In one example, the alert module of the edge device may include multiple alerts for one or more driving events related to unsafe driving behavior. Analert may be selected dynamically from the multiple alerts based on the driving behavior and driver’s feedback to alerts. For example, selecting alerts may focus the driver on driving behaviors for which the safety alerts are most effective, and or for which the driver’s overall experience is most positive.

[0153] FIG. 6 shows another sequence 600 for determining a driver response to a safety alert based on voice data using a machine learning model. In another embodiment, the edge device may provide a questionnaire to receive feedback from the driver with respect to the recently communicated alert. The questionnaire may include a query and one or more answers from which the driver has to choose as a response to the query. The one or more answers may indicate the response or feedback provided to the alert by the driver. The questionnaire may be presented to the driver through an audio module or a display module, as two examples. The questionnaire may be presented to the driver after the alert is communicated. In one embodiment, the questionnaire may be included in the alert. For example, the audio module (e.g., speaker) may ask the driver whether he agrees or disagrees with a recently communicated alert. The driver may provide his inputs in response to the questionnaire either through audio (e.g., saying one of the answers) or manual input (e.g., clicking an option on a touch screen, interacting with a toggle button, etc.). The edge device may record and process audio received through the microphone for making an inference related to the feedback received from the driver. In one embodiment, the edge device may capture and process the visual data from the camera for making inferences related to any visual feedback received from the driver. If the feedback is negative, the visual data or the audio data may be sent to the cloud server in addition to the sensor data that led to alert creation. In some embodiments, the content of the data transmission to the cloud server may be predicated upon verifying the driver’s past driving behavior as tending to be safe or unsafe; and further or alternatively, upon a balancing of the driver’s or fleet’s privacy settings and accuracy requirements of the safety features.

[0154] The edge device may continue to monitor the driving behavior to detect a change after raising an alert. The edge device may modify a subsequent alert upon detection of a change in the driving behavior. In addition, the edge device may modify the subsequent alert based on the feedback received for the initial alert. The edge device may generate a report including data related to alerts raised for the driver and data related to instances for which the driver has changed driving behavior from unsafe to safe. The cloud server may collect the reports from multiple edge devices installed in vehicles of a fleet and collate the data to determine the efficacy of an alert based on the reports.

[0155] Further, the cloud server may determine the alert rate of a driver based on information related to alerts generated by an edge device and received at a cloud server. The cloud server may determine the number of alert types raised for each driver for a predefined time duration. The cloud server may further determine whether the number of alerts raised per alert type decrease or increase over a time period upon raising an in-cab alert to the driver. For example, the cloud server determines the alert rate of distracted driving for a driver to be 4 alerts per 8-hour day. Furthermore, the cloud server may determine whether the alert rate for a driver has changed upon modifying the alert type and / or alerting mechanism. The cloud server may determine an alert profile for the driver based on the alerting mechanisms that have resulted in safe driving behavior. The profile of the driver may be updated based on the change in responses or feedback received for different types of alerts raised for the driver.

[0156] The edge device may continue with the predefined alerting mechanism if the driver acknowledges an alert with positive feedback and changes the driving behavior upon raising the alert. The edge device may modify the alert and / or alerting mechanism upon receiving negative feedback from the driver upon raising an alert. In addition, the edge device may assign the driver for a coaching session when there is negative feedback for an alert and the alert is determined to be a valid alert. The edge device may be configured to provide preventive alerts when the previously raised alert is determined to be valid and when the previously raised alert received negative feedback from the driver. The preventive alerts may be notifications instead of instructions. For example, notifications may communicate a message corresponding to “This is a common area for speeding, please drive under speed limits.”

[0157] In one embodiment, the driver may be assigned coaching sessions when the driver has not responded to the alerts and continued to exhibit unsafe driving behavior. In addition, the driver may also be assigned coaching sessions when the driver has responded negatively or erratically to the alert even though the alert is valid. The edge device may track the efficacy of an alert and / or modified alert based on the driving behavior. For example, the efficacy of an alert can be low if there is no change in driving behavior upon raising that alert. The modified alert can be reverted back to a predefined parameter configuration for generating alert and / or detecting a driving event if the efficacy of the modified alert is less than a predefined threshold. In one embodiment, the modified alert can be changed to a new alert modality if the efficacy of the modified alert is less than a predefined threshold. Further, the modified alert may be changed to a new alert modality upon receiving negative feedback from the driver with respect to the modified alert. Similarly, instead of changing to a new alertmodality, a different alert type (corresponding to driving behavior in a different driving scenario) may be instead prioritized for in-cab safety alerts for the driver.

[0158] In an embodiment, a method can include detecting a driving behavior event by a driver in a driving scenario; presenting a safety alert to the driver in response to the detection of the event; and determining a response by the driver to the safety alert; and modifying a subsequent safety alert for the driver based on the response.

[0159] Determining the response can include processing visual data of the driver collected immediately after the presentation of the safety alert; computing a compliance metric for the driving scenario by the driver based on the processed visual data; and determining a change in the compliance metric before and after the presentation of the safety alert.

[0160] Modifying the safety alert can include modifying how the system alerts the driver in the future in response to a subsequent detection of the driving behavior. Modifying the safety alert can include making the safety alert louder. Modifying the safety alert can include selecting a different safety alert from a plurality of safety alert messages that are associated with the driving scenario. The different safety alert can include an instruction for the driver and the safety alert was a notification and not an instruction. Modifying the safety alert can include presenting the safety alert earlier in response to detection of the driving scenario.

[0161] The method can further include detecting the driving scenario, and wherein the safety alert is presented to the driver in response to detecting the driving scenario and before the driver has a chance to engage in unsafe driving behavior. The safety alert can be presented before a stop sign crossing, or as a reminder to wear a seatbelt when beginning a trip. The method can further include determining that the driver responded to the safety alert by correcting the unsafe behavior that was detected; and, either immediately or based on compliance trends, modifying the subsequent driving event detections a second time so that it is less sensitive and the driver receives fewer in-cab notifications in the driving scenario.

[0162] The response determination can include a valence determination, in which case the system triggers fewer alerts, but still flags the alerts that it detects for offline review. Such alerts can be collected to identify potential bias in the driver monitoring system. Labeling such examples, training the model with these examples, testing on the collected alerts, and then deploying the model can be in response to achieving a lower false positive rate. The responsedetermination can include a negative valence determination, but alerts are correct. Still reduce in-cab audio feedback, but instead prioritize human coaching session (e.g., a different feedback modality compared to in-cab audio).

[0163] The method can further include tracking risk-rates per driver; analyzing efficacy of in-cab audio feedback; analyzing efficacy by comparing driving behavior when the in-cab audio feedback is present (enabled) vs. absent (disabled) for the driving scenario; sensitive settings - if sensitivity is reduced, is driver become less safe; if driver safety is not changing, then maintain lessened sensitivity threshold. If safety metrics getting worse, then revert the sensitivity threshold. The method can further include determining a feedback modality that works for a driver; and preferentially choosing that feedback modality for a different driving behavior for the same driver.

[0164] The above-mentioned systems and methods improve the efficacy of changing driving behavior by automatically modifying the driver safety alerting mechanisms using observable and measurable responses from the system / driver as feedback.

[0165] In an example, a driver safety system may monitor a driver’ s behavior in various driving scenarios and may present a safety alert to the driver in response to detecting an unsafe driving behavior. The driver’s response to the safety alert may be inferred based on sensor data or based on other feedback from the driver. Based on the driver’s response or feedback, the parameters of the driver monitoring system may be modified. Automatically modifying the alerting mechanisms of a driver safety system using a driver’s observable and measurable responses to safety alerts may improve the overall efficacy of driver safety systems for certain drivers.

[0166] A computer program product may include a non-transitory computer-readable medium having instructions stored thereon, the instructions being executable by one or more processors configured to carry out any of the systems and / or methods described herein.

[0167] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particularapplication and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of this disclosure or the claims.

[0168] Embodiments implemented in computer software may be implemented in software, firmware, middleware, microcode, hardware description languages, or any combination thereof. A code segment or machine-executable instructions may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0169] The actual software code or specialized control hardware used to implement these systems and methods does not limit the claimed features or this disclosure. Thus, the operation and behavior of the systems and methods were described without reference to the specific software code being understood that software and control hardware can be designed to implement the systems and methods based on the description herein.

[0170] When implemented in software, the functions may be stored as one or more instructions or code on a non-transitory computer-readable or processor-readable storage medium. The steps of a method or algorithm disclosed herein may be embodied in a processorexecutable software module, which may reside on a computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable media includes both computer storage media and tangible storage media that facilitate transfer of a computer program from one place to another. A non-transitory processor-readable storage media may be any available media that may be accessed by a computer. By way of example, and not limitation, such non-transitory processor-readable media may comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other tangible storage medium that may be used to store desired program code in the form of instructions or data structures and that may be accessed by a computer or processor. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc where disks usuallyreproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media. Additionally, the operations of a method or algorithm may reside as one or any combination or set of codes and / or instructions on a non-transitory processor-readable medium and / or computer-readable medium, which may be incorporated into a computer program product.

[0171] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.

[0172] The terms “data processing apparatus”, “data processing system”, “client device”, “client computing device”, “computing platform”, “computing device”, “computing system”, “user device”, or “device” can encompass all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, a system on a chip, or multiple ones, or combinations, of the foregoing. The apparatus can include special purpose logic circuitry, e.g., an FPGA or an ASIC. The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them.

[0173] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub-programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.

[0174] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data froma read-only memory or a random access memory or both. The elements of a computer include a processor for performing actions in accordance with instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a GPS receiver, a digital camera device, a video camera device, or a portable storage device (e.g., a universal serial bus (USB) flash drive), for example. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0175] To provide for interaction with a user, implementations of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube), plasma, or LCD monitor, for displaying information to the user; a keyboard; and a pointing device, e.g., a mouse, a trackball, or a touchscreen, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can include any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user.

[0176] In certain circumstances, multitasking and parallel processing systems may be advantageous. Moreover, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. For example, the computing devices described herein can each be a single module, a logic device having one or more processing modules, one or more servers, or an embedded computing device.

[0177] Having now described some illustrative implementations and implementations, it is apparent that the foregoing is illustrative and not limiting, having been presented by way of example. In particular, although many of the examples presented herein involve specificcombinations of method acts or system elements, those acts and those elements may be combined in other ways to accomplish the same objectives. Acts, elements, and features discussed only in connection with one implementation are not intended to be excluded from a similar role in other implementations or implementations.

[0178] The phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” “having,” “containing,” “involving,” “characterized by,” “characterized in that,” and variations thereof herein, is meant to encompass the items listed thereafter, equivalents thereof, and additional items, as well as alternate implementations consisting of the items listed thereafter exclusively. In one implementation, the systems and methods described herein consist of one, each combination of more than one, or all of the described elements, acts, or components.

[0179] Any references to implementations or elements or acts of the systems and methods herein referred to in the singular may also embrace implementations including a plurality of these elements, and any references in plural to any implementation or element or act herein may also embrace implementations including only a single element. References in the singular or plural form are not intended to limit the presently disclosed systems or methods, their components, acts, or elements to single or plural configurations. References to any act or element being based on any information, act, or element may include implementations where the act or element is based at least in part on any information, act, or element.

[0180] Any implementation disclosed herein may be combined with any other implementation, and references to “an implementation,” “some implementations,” “an alternate implementation,” “various implementation,” “one implementation,” or the like are not necessarily mutually exclusive and are intended to indicate that a particular feature, structure, or characteristic described in connection with the implementation may be included in at least one implementation. Such terms as used herein are not necessarily all referring to the same implementation. Any implementation may be combined with any other implementation, inclusively or exclusively, in any manner consistent with the aspects and implementations disclosed herein.

[0181] References to “or” may be construed as inclusive so that any terms described using “or” may indicate any of a single, more than one, and all of the described terms.

[0182] Where technical features in the drawings, detailed description or any claim are followed by reference signs, the reference signs have been included for the sole purpose of increasing the intelligibility of the drawings, detailed description, and claims. Accordingly, neither the reference signs nor their absence have any limiting effect on the scope of any claim elements.

[0183] The preceding description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the embodiments described herein and variations thereof. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other embodiments without departing from the spirit or scope of the subject matter disclosed herein. Thus, the present disclosure is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the following claims and the principles and novel features disclosed herein.

[0184] While various aspects and embodiments have been disclosed, other aspects and embodiments are contemplated. The various aspects and embodiments disclosed are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the following claims.

Claims

CLAIMSWhat is claimed is:

1. A method for automatic alert configuration adjustment in a vehicle, comprising: receiving, by one or more processors from a first camera mounted to or in the vehicle, a first sequence of images of an environment surrounding the vehicle or within the vehicle; responsive to detecting, by the one or more processors, a first instance of a first type of event based on the first sequence of images using an event threshold, generating, by the one or more processors, a first alert within the vehicle; receiving, by the one or more processors from the first camera or a second camera mounted to or in the vehicle, a second sequence of images depicting a reaction of a driver of the vehicle to the first type of alert; executing, by the one or more processors, a machine learning model using the second sequence of images to detect the reaction of the driver to the first type of alert; adjusting, by the one or more processors, the event threshold based on the reaction of the driver to the first type of alert; receiving, by the one or more processors from the first camera mounted to or in the vehicle, a third sequence of images of the environment surrounding the vehicle or within the vehicle; and responsive to detecting, by the one or more processors, a second instance of the first type of event based on the third sequence of images using the adjusted threshold, generating, by the one or more processors, a second alert within the vehicle.

2. The method of claim 1, wherein executing the machine learning model comprises: detecting, by the one or more processors, one or more facial expressions of the driver in the second sequence of images; detecting, by the one or more processors, one or more head movements of the driver in the second sequence of images; and classifying, by the one or more processors, the reaction of the driver as positive or negative based on the detected one or more facial expressions and one or more head movements.

3. The method of claim 1, wherein executing the machine learning model comprises:classifying, by the one or more processors, the reaction of the driver as positive or negative based on the second sequence of images; and wherein adjusting the event threshold comprises increasing, by the one or more processors, the event threshold responsive to classifying the reaction of the driver as negative.

4. The method of claim 1, wherein detecting the first type of event comprises at least one of: detecting, by the one or more processors, a second vehicle in a blind spot of the vehicle; or detecting, by the one or more processors, an object within a defined distance of the vehicle.

5. The method of claim 1, wherein adjusting the event threshold comprises: increasing, by the one or more processors, the event threshold responsive to the reaction of the driver indicating the first alert was generated too early; or decreasing, by the one or more processors, the event threshold responsive to the reaction of the driver indicating the first alert was generated too late.

6. The method of claim 1, further comprising: storing, by the one or more processors in a memory, a driver profile associated with the driver; and updating, by the one or more processors, the driver profile based on the reaction of the driver.

7. The method of claim 1, wherein generating the second alert comprises: selecting, by the one or more processors, a type of the second alert as at least one of an audio alert, a visual alert, or a haptic alert based on the reaction of the driver to the first alert.

8. The method of claim 1, wherein adjusting the event threshold comprises: maintaining, by the one or more processors, separate thresholds for different types of driving environments; categorizing, by the one or more processors, the first instance of the first type of event as occurring in a first type of driving environment; andadjusting, by the one or more processors, only the event threshold associated with the first type of driving environment based on the reaction of the driver.

9. The method of claim 1, wherein executing the machine learning model comprises: classifying, by the one or more processors, the reaction of the driver into one of a plurality of reaction categories; and wherein adjusting the event threshold comprises adjusting, by the one or more processors, the event threshold by an amount determined based on the classification of the reaction of the driver.

10. The method of claim 1, wherein generating the first alert comprises generating, by the one or more processors, the first alert based on environmental factors of the environment surrounding the vehicle; and wherein detecting the reaction of the driver comprises detecting, by the one or more processors, movement of the driver subsequent to the first alert for a defined time duration.

11. The method of claim 1, further comprising: generating, by the one or more processors, a driver score for the driver; and determining, by the one or more processors, a magnitude of the adjustment to the event threshold based on the driver score for the driver.

12. The method of claim 1, further comprising: responsive to detecting the reaction of the driver to the first type of alert, transmitting, by the one or more processors, the first sequence of images to a remote computing device, wherein when the remote computing device does not detect the first instance of the first type of event in the first sequence images, the remote computing device transmits an instruction to the one or more processors to adjust the event threshold for the first type of event.13 The method of claim 12, wherein transmitting, by the one or more processors, the first sequence of images to the remote computing device is responsive to determining, by the one or more processors, the reaction of the driver to the first type of alert is negative.

14. A system, comprising: one or more processors configured by machine-readable instructions to:receive, from a first camera mounted to or in a vehicle, a first sequence of images of an environment surrounding the vehicle or within the vehicle; responsive to detecting a first instance of a first type of event based on the first sequence of images using an event threshold, generate a first alert within the vehicle; receive, from the first camera or a second camera mounted to or in the vehicle, a second sequence of images depicting a reaction of a driver of the vehicle to the first alert; execute a machine learning model using the second sequence of images to detect the reaction of the driver to the first alert; adjust the event threshold based on the reaction of the driver to the first alert; receive, from the first camera mounted to or in the vehicle, a third sequence of images of the environment surrounding the vehicle or within the vehicle; and responsive to detecting a second instance of the first type of event based on the third sequence of images using the adjusted threshold, generate a second alert within the vehicle.

15. The system of claim 14, wherein the one or more processors are configured to execute the machine learning model to: detect one or more facial expressions of the driver in the second sequence of images; detect one or more head movements of the driver in the second sequence of images; and classify the reaction of the driver as positive or negative based on the detected one or more facial expressions and one or more head movements.

16. The system of claim 14, wherein the one or more processors are further configured to: responsive to detecting the reaction of the driver to the first type of alert, transmit the first sequence of images to a remote computing device, wherein when the remote computing device does not detect the first instance of the first type of event in the first sequence images, the remote computing device transmits an instruction to the one or more processors to adjust the event threshold for the first type of event.

17. The method of claim 16, wherein the one or more processors are configured to transmit the first sequence of images to the remote computing device responsive to determining the reaction of the driver to the first type of alert is negative.

18. A method for automatic alert configuration adjustment in a vehicle, comprising: storing, by one or more processors, a profile for a driver in memory, the profile containing an indication to generate a first type of alert in response to detecting a first type of event; receiving, by the one or more processors from a camera mounted to or in the vehicle, a first sequence of images of an environment surrounding the vehicle or within the vehicle; responsive to detecting, by the one or more processors, a first instance of the first type of event based on the first sequence of images, generating, by the one or more processors, a first alert of the first type of alert within the vehicle based on the indication in the profile of the driver; in response to detecting, by the one or more processors, a response by the driver to the first alert from a second sequence of images, adjusting, by the one or more processors, the indication in the profile to indicate a second type of alert to generate in response to detecting the first type of event; receiving, by the one or more processors from the camera mounted to or in the vehicle, a third sequence of images of the environment surrounding the vehicle or within the vehicle; and responsive to detecting, by the one or more processors, a second instance of the first type of event based on the third sequence of images, generating, by the one or more processors, a second alert within the vehicle based on the adjusted indication in the profile of the driver.

19. The method of claim 18, wherein the profile of the driver comprises indications of types of alerts for a plurality of types of alerts for different types of events, and wherein generating the first alert of the first type of alert within the vehicle comprises: selecting, by the one or more processors, the first type of alert of the plurality of types of alerts based on the indication in the profile of the driver.

20. The method of claim 18, wherein detecting the response by the driver comprises: detecting, by the one or more processors, a manual input on a user interface or device of the vehicle, detecting, by the one or more processors, a voice command from the driver, or detecting, by the one or more processors a gesture of the driver captured by a second camera within the vehicle.

21. The method of claim 18, further comprising: changing, by the one or more processors, the adjusted indication in the profile of the driver based on a user input at a user interface.

22. The method of claim 18, further comprising: identifying the profile of the driver upon the driver entering the vehicle; and using, by the one or more processors based on the identification of the profile, the indication from the profile to generate the first type of alert.

23. The method of claim 18, further comprising: responsive to detecting the response of the driver to the first alert, transmitting, by the one or more processors, the first sequence of images to a remote computing device, wherein when the remote computing device does not detect the first instance of the first type of event in the first sequence images, the remote computing device transmits an instruction to the one or more processors to adjust the type of alert for the first type of event.

24. The method of claim 23, wherein transmitting the first sequence of images to the remote computing device is responsive to determining the response of the driver to the first type of alert is negative.

Citation Information

Patent Citations

  • Method to analyze attention margin and to prevent inattentive and unsafe driving

    US10467488B2

  • Adaptive actuator interface for active driver warning

    US20140125474A1

  • Driver and environment monitoring to predict human driving maneuvers and reduce human driving errors

    US20210309261A1

  • Systems and methods for prioritizing driver warnings in a vehicle

    US20220063495A1

  • Driver adaptive collision warning system

    US7206697B2