Device and method for assisting vehicle operation based on Situational Assessment with Exponential Risk (SAFER)
A sensor-based system with neural network processing identifies and mitigates collision and intersection risks, enhancing vehicle safety through early prediction and intervention.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NAUTO INC
- Filing Date
- 2022-05-14
- Publication Date
- 2026-05-12
AI Technical Summary
Existing vehicle systems lack effective methods for predicting and mitigating collision and intersection violation risks, as well as providing driver feedback and vehicle management based on situational assessments.
A system comprising sensors to capture environmental and vehicle data, a processing unit with neural network models to analyze time-series information from multiple sensors, and generate control signals for warnings or vehicle interventions to mitigate risks.
Early identification and mitigation of collision and intersection risks, providing driver feedback and vehicle control to enhance safety and driving quality.
Smart Images

Figure 0007857404000001 
Figure 0007857404000002 
Figure 0007857404000003
Abstract
Description
Technical Field
[0001] Related Application Data
[0001] This application claims the priority and benefit of U.S. Provisional Patent Application No. 63 / 285,073, filed on December 1, 2021, U.S. Patent Application No. 17 / 726,236, filed on April 21, 2022, and U.S. Patent Application No. 17 / 726,269, filed on April 21, 2022.
[0002]
[0002] This field relates to devices and methods for assisting in vehicle driving, scoring, and certifying driving actions, and more particularly, to devices and methods for identifying driving and situation risks.
Background Art
[0003]
[0003] Sensors (such as cameras, radars, Lidars, etc.) are used in vehicles to capture images of road conditions outside the vehicle. For example, a camera can be installed on a target vehicle to monitor the driving route of the target vehicle or the vehicles around the target vehicle.
[0004]
[0004] It is desirable to use camera images to provide collision prediction and / or intersection violation prediction. It is also desirable to give warnings to the driver or automatically operate the vehicle in response to predicted risks such as collisions and intersection violations. Furthermore, it is desirable to provide an output indicating the quality of driving, which can be used for driver training and vehicle management. Alternatively or additionally, this output can also be used to compare the actual driver's actions with those of an excellent driver in a similar situation.
[0005]
[0005] Novel techniques for determining and tracking the risk of collision and / or the risk of intersection violation are described herein. Novel techniques for providing control signals to activate warning or feedback generators to warn drivers of the risk of collision and / or the risk of intersection violation and / or to mitigate such risks are also described herein. [Overview of the project]
[0006]
[0006] The device comprises a first sensor configured to provide a first input relating to the external environment of the vehicle, a second sensor configured to provide a second input relating to the operation of the vehicle, and a processing unit configured to receive a first input from the first sensor and a second input from the second sensor, the processing unit comprising a first-stage processing system and a second-stage processing system, the first-stage processing system configured to receive a first input from the first sensor, receive a second input from the second sensor, process the first input to obtain a first time series of information, and process the second input to obtain a second time series of information, the second-stage processing system comprising a neural network model configured to receive the first time series of information and the second time series of information in parallel, the neural network configured to process the first time series and the second time series to determine the probability of predictive events relating to the operation of the vehicle.
[0007]
[0007] As a non-limiting example, the first sensor may be a camera, lidar, radar, or any combination thereof, configured to sense characteristics of the environment outside the vehicle. The first input from the first sensor may include one or more time series.
[0008]
[0008] In addition, as a non-limiting example, the second sensor may be a camera, depth sensor, radar, or any combination thereof, configured to capture images and / or detect the driver's state. Alternatively or additionally, the second sensor may include one or more sensing units configured to sense one or more states of the vehicle (e.g., speed, acceleration, deceleration, braking, direction, steering angle, brake operation, wheel traction, engine state, brake pedal position, accelerator pedal position, turn signal state, etc.) (vehicle state). The second input from the second sensor may include one or more time series.
[0009]
[0009] Optionally, the processing unit is configured to combine (e.g., merge) multiple metadata streams that represent different situational aspects or are associated with them in order to determine whether the risks are combined or combined in such a manner (e.g., nonlinear, exponentially, etc.) that the risks constitute a very dangerous situation. For example, a driver smoking a cigarette in itself may not be very dangerous. However, if the same driver is also using a cell phone and is getting too close to a preceding vehicle, the combined risk (e.g., smoking risk + cell phone use risk + close proximity risk) may constitute a very dangerous situation. Risky situations may be represented by a risk value that increases nonlinearly (e.g., exponentially) due to the combination of risk factors. In some embodiments, each metadata stream may be time-series data obtained by processing raw data from one or more sensors.
[0010]
[0010] In some embodiments, the processing unit is configured to identify peaks (or spikes) in a combination of risks, where the peaks (or spikes) represent an escalating risk situation. In some embodiments, the escalating risk situation may be represented by a risk value that increases non-linearly (e.g., exponentially) with respect to the combination of risk factors.
[0011]
[0011] In some embodiments, the processing unit is configured to identify dangerous situations early enough to intervene (e.g., give warnings, feedback, recommendations, etc., to reduce one or more risk factors). In some embodiments, the processing unit may identify what other good drivers have done in similar situations to mitigate risk and use such knowledge to give warnings, feedback, or recommendations (e.g., increase distance from the vehicle ahead, change speed, change direction, change attention state, etc.) to mitigate one or more risk factors.
[0012]
[0012] Optionally, the first input has a lower number of dimensions or less complexity compared to the first time-series information.
[0013]
[0013] Optionally, the first time series information indicates a first risk factor, and the second time series information indicates a second risk factor.
[0014]
[0014] Optionally, the processing unit is configured to package the first time series and the second time series into a data structure for supplying to a neural network model.
[0015]
[0015] Optionally, the data structure consists of a two-dimensional data matrix. In some embodiments, the data matrix may include multiple sensor streams from multiple sensors of the same or different sensor types.
[0016]
[0016] Optionally, the first time series may show the external state of the vehicle at different points in time, and the second time series may show the state of the driver and / or the state of the vehicle at different points in time.
[0017]
[0017] Optionally, the probability of the predicted event is the first probability of a first predicted event, the processing unit is configured to determine the second probability of a second predicted event, the first and second predicted events relate to the driving of a vehicle, and the processing unit is configured to calculate a risk score based on the first probability of the first predicted event and the second probability of the second predicted event.
[0018]
[0018] Optionally, the first predicted event is a collision event, the second predicted event is a non-hazardous event, and the processing unit is configured to calculate the risk score based on the first probability of a collision event and the second probability of a non-hazardous event.
[0019]
[0019] The processing unit is optionally configured to calculate the risk score by applying a first weight to the first probability to obtain a first weighted probability, applying a second weight to the second probability to obtain a second weighted probability, and adding the first weighted probability and the second weighted probability.
[0020]
[0020] Optionally, the processing unit is configured to determine a third probability of a third predicted event, and the processing unit is configured to calculate the risk score based on a first probability of the first predicted event, a second probability of the second predicted event, and a third probability of the third predicted event.
[0021]
[0021] The first predicted event may be a collision event, the second predicted event may be a near-collision event, and the third predicted event may be a non-hazardous event. The processing unit is configured to calculate the risk score based on the first probability of the collision event, the second probability of the near-collision event, and the third probability of the non-hazardous event.
[0022] Optionally, the processing unit is configured to calculate the risk score by applying a first weighting to the first probability to obtain a first weighted probability, applying a second weighting to the second probability to obtain a second weighted probability, applying a third weighting to the third probability to obtain a third weighted probability, and adding the first weighted probability, the second weighted probability, and the third weighted probability.
[0023] Optionally, the first input and the second input are acquired in the past T seconds, and the processing unit is configured to process the first input and the second input acquired in the past T seconds to determine the probability of the predicted event, where T is at least 3 seconds.
[0024] Optionally, the predicted event is at a future time at least 1 second after the current time.
[0025] Optionally, the predicted event is at a future time at least 2 seconds after the current time.
[0026] Optionally, the processing unit is configured to calculate a first risk score for a first point in time based on a probability, and the processing unit is also configured to calculate a second risk score for a second point in time and identify the difference between the first risk score and the second risk score, which indicates whether a dangerous situation is escalating or calming down.
[0027] Optionally, the processing unit is configured to determine the risk score based on the probability of the predicted event.
[0028] Optionally, the processing unit is configured to generate a control signal based on the risk score.
[0029] Optionally, the processing unit is configured to generate a control signal when the risk score meets a criterion.
[0030] Optionally, the processing unit is configured to generate a control signal for operating the device when the risk score meets a criterion.
[0031] Optionally, the device includes a speaker that emits an alarm, a display or a light-emitting device that provides a visual signal, a haptic feedback device, a collision avoidance system, or a vehicle control device for a vehicle.
[0032] Optionally, each of the first time series and the second time series includes any two or more of the distance to a preceding vehicle, the distance to an intersection stop line, the speed of the vehicle, the time until a collision, the time until an intersection violation, the estimated braking distance, information regarding road conditions, information regarding special areas, information regarding the environment (such as weather, road type, etc.), information regarding traffic conditions, the time, information regarding visibility conditions, information regarding a specified object, the position of the object, the moving direction of the object, the speed of the object, the bounding box, the operation parameters of the vehicle, information regarding the state of the driver, information regarding the driver's history, the continuous driving time, the proximity of meal times, information regarding accident history, and voice information.
[0033] Optionally, the first sensor is composed of a camera, Lidar, radar, or any combination thereof configured to sense the external environment of the vehicle.
[0034] Optionally, the second sensor is composed of a camera configured to view the driver of the vehicle.
[0035] Optionally, the second sensor is composed of one or more sensing units configured to sense one or more characteristics of the vehicle.
[0036]
[0036] Optionally, the first input includes a first image, the second input includes a second image, and the first stage processing system is configured to receive the first image and the second image, process the first image to obtain first time-series information, and process the second image to obtain second time-series information.
[0037]
[0037] The device comprises a first sensor configured to provide a first input, a second sensor configured to provide a second input, and a processing unit configured to receive a first input from the first sensor and a second input from the second sensor, wherein the processing unit is configured to identify a first probability of a first predicted event and a second probability of a second predicted event, the first and second predicted events being associated with the operation of a vehicle, and the processing unit is configured to calculate a risk score based on the first probability of the first predicted event and the second probability of the second predicted event.
[0038]
[0038] Optionally, the first predicted event is a collision event, the second predicted event is a non-hazardous event, and the processing unit is configured to calculate the risk score based on the first probability of a collision event and the second probability of a non-hazardous event.
[0039]
[0039] The processing unit is optionally configured to calculate the risk score by applying a first weight to the first probability to obtain a first weighted probability, applying a second weight to the second probability to obtain a second weighted probability, and adding the first weighted probability and the second weighted probability.
[0040]
[0040] Optionally, the processing unit is configured to determine a third probability of a third predicted event, and the processing unit is configured to calculate the risk score based on a first probability of the first predicted event, a second probability of the second predicted event, and a third probability of the third predicted event.
[0041]
[0041] Optionally, the first predicted event is a collision event, the second predicted event is a near collision event, and the third predicted event is a non-hazardous event, and the processing unit is configured to calculate the risk score based on the first probability of the collision event, the second probability of the near collision event, and the third probability of the non-hazardous event.
[0042]
[0042] The processing unit is optionally configured to calculate the risk score by applying a first weight to the first probability to obtain a first weighted probability, applying a second weight to the second probability to obtain a second weighted probability, applying a third weight to the third probability to obtain a third weighted probability, and adding the first weighted probability, the second weighted probability, and the third weighted probability.
[0043]
[0043] Optionally, the first input and the second input include data acquired over the past T seconds, and the processing unit is configured to process the data acquired over the past T seconds to determine a first probability of the first predicted event and a second probability of the second predicted event, where T is at least 3 seconds.
[0044]
[0044] Optionally, the first predicted event is at a future time at least one second after the current time.
[0045]
[0045] Optionally, the first predicted event is at a future time at least 2 seconds after the current time.
[0046]
[0046] Optionally, the processing unit is configured to calculate a risk score for a first time point, and the processing unit is also configured to calculate a further risk score for a second time point, and to identify the difference between the risk score and the further risk score, the difference indicating whether the dangerous situation is escalating or fading.
[0047]
[0047] Optionally, the processing unit is configured to generate a control signal based on the risk score.
[0048]
[0048] Optionally, the processing unit is configured to generate a control signal when the risk score meets the criteria.
[0049]
[0049] Optionally, the processing unit is configured to generate a control signal to operate the device when the risk score meets the criteria.
[0050]
[0050] Optionally, the device includes a speaker for issuing an alarm, a display or light-emitting device for providing a visual signal, a haptic feedback device, a collision avoidance system, or a vehicle control device for a vehicle.
[0051]
[0051] Optionally, the processing unit is configured to determine a first probability of the first predicted event and a second probability of the second predicted event based on the first and second inputs.
[0052]
[0052] Optionally, the processing unit comprises a first-stage processing system and a second-stage processing system, wherein the first-stage processing system is configured to acquire the first input and process the first input to provide a first output, and the second-stage processing system is configured to acquire the first output and process the first output to provide a second output, wherein the first output has lower dimensions or less complexity than the first input, and the second output has lower dimensions or less complexity than the first output.
[0053]
[0053] Optionally, the processing unit includes a neural network model.
[0054]
[0054] Optionally, the neural network model is configured to receive a first time series of information indicating a first risk factor and a second time series of information indicating a second risk factor.
[0055]
[0055] Optionally, the neural network model is configured to receive the first time series and the second time series in parallel and / or process the first time series and the second time series in parallel.
[0056]
[0056] Optionally, the processing unit is configured to package the first time series and the second time series into a data structure for supplying to a neural network model.
[0057]
[0057] Optionally, the data structure consists of a two-dimensional data matrix. In some embodiments, the data matrix may include multiple sensor streams from multiple sensors of the same or different sensor types.
[0058]
[0058] Optionally, the first time series may show the external state of the vehicle at different points in time, and the second time series may show the state of the driver and / or the state of the vehicle at different points in time.
[0059]
[0059] Optionally, the first time series may represent a first characteristic of the vehicle at each different point in time, and the second time series may represent a second characteristic of the vehicle at each different point in time.
[0060]
[0060] Optionally, the first and second time series include any one or more combinations of the following: distance to the preceding vehicle, distance to the intersection stop line, vehicle speed, time to collision, time to intersection violation, estimated braking distance, information on road conditions, information on special zones, information on the environment (weather, road type, etc.), information on traffic conditions, time, information on visibility conditions, information on identified objects, object location, direction of movement of objects, object speed, boundary box, vehicle operating parameters, information on the driver's state, information on the driver's history, continuous driving time, proximity to meal times, information on accident history, and audio information.
[0061]
[0061] Optionally, the first sensor may consist of a camera, Lidar, radar, or any combination thereof, configured to sense the environment outside the vehicle.
[0062]
[0062] Optionally, the second sensor is a camera configured to view the driver of the vehicle.
[0063]
[0063] Optionally, the second sensor comprises one or more sensing units configured to sense one or more characteristics of the vehicle.
[0064]
[0064] Optionally, the first input includes a first image, the second input includes a second image, and the first stage processing system is configured to receive the first image and the second image, process the first image to obtain first time-series information, and process the second image to obtain second time-series information.
[0065]
[0065] The device comprises a first camera for viewing the environment outside the vehicle, a second camera for viewing the driver of the vehicle, and a processing unit configured to receive a first image from the first camera and a second image from the second camera, the processing unit comprising a first-stage processing system and a second-stage processing system, the first-stage processing system configured to receive a first image from the first camera, receive a second image from the second camera, process the first image to obtain first time-series information, and process the second image to obtain second time-series information, the second-stage processing system comprising a neural network model configured to receive the first time-series information and the second time-series information in parallel, the neural network configured to process the first time-series and the second time-series to determine the probability of predictive events related to the operation of the vehicle.
[0066]
[0066] Optionally, the first image has fewer dimensions or less complexity compared to the first time-series information.
[0067]
[0067] Optionally, the first time series information indicates a first risk factor, and the second time series information indicates a second risk factor. In some embodiments, there may be two or more time series, and each time series may indicate two or more risk factors.
[0068]
[0068] Optionally, the processing unit is configured to package the first time series and the second time series into a data structure for supplying to a neural network model.
[0069]
[0069] Optionally, the data structure consists of a two-dimensional data matrix. In some embodiments, the data matrix may include multiple sensor streams from multiple sensors of the same or different sensor types.
[0070]
[0070] Optionally, the first time series shows the external state of the vehicle at different points in time, and the second time series shows the state of the driver at different points in time.
[0071]
[0071] Optionally, the probability of the predicted event is the first probability of the first predicted event, the processing unit is configured to determine the second probability of the second predicted event, the first and second predicted events relate to the driving of a vehicle, and the processing unit is configured to calculate a risk score based on the first probability of the first predicted event and the second probability of the second predicted event.
[0072]
[0072] Optionally, the first predicted event is a collision event, the second predicted event is a non-hazardous event, and the processing unit is configured to calculate the risk score based on the first probability of a collision event and the second probability of a non-hazardous event.
[0073]
[0073] The processing unit is optionally configured to calculate the risk score by applying a first weight to the first probability to obtain a first weighted probability, applying a second weight to the second probability to obtain a second weighted probability, and adding the first weighted probability and the second weighted probability.
[0074]
[0074] Optionally, the processing unit is configured to determine a third probability of a third predicted event, and the processing unit is configured to calculate the risk score based on a first probability of the first predicted event, a second probability of the second predicted event, and a third probability of the third predicted event.
[0075]
[0075] Optionally, the first predicted event is a collision event, the second predicted event is a near collision event, and the third predicted event is a non-hazardous event. The processing unit is configured to calculate the risk score based on the first probability of the collision event, the second probability of the near collision event, and the third probability of the non-hazardous event.
[0076]
[0076] The processing unit is optionally configured to calculate the risk score by applying a first weight to the first probability to obtain a first weighted probability, applying a second weight to the second probability to obtain a second weighted probability, applying a third weight to the third probability to obtain a third weighted probability, and adding the first weighted probability, the second weighted probability, and the third weighted probability.
[0077]
[0077] Optionally, the first and second images are acquired in the past T seconds, and the processing unit is configured to process the first and second images acquired in the past T seconds to determine the probability of the predicted event, where T is at least 3 seconds.
[0078]
[0078] Optionally, the predicted event is at a future time at least 1 second or at least 2 seconds from the current time.
[0079]
[0079] Optionally, the processing unit is configured to calculate a first risk score for a first time point based on probability, and the processing unit is also configured to calculate a second risk score for a second time point and to identify the difference between the first risk score and the second risk score, the difference indicating whether the dangerous situation is escalating or fading.
[0080]
[0080] Optionally, the processing unit is configured to determine the risk score based on the probability of the predicted event.
[0081]
[0081] Optionally, the processing unit is configured to generate a control signal based on the risk score.
[0082]
[0082] Optionally, the processing unit is configured to generate a control signal when the risk score meets the criteria.
[0083]
[0083] Optionally, the processing unit is configured to generate a control signal for operating the device when the risk score meets the criteria.
[0084]
[0084] Optionally, the device includes a speaker for issuing an alarm, a display or light-emitting device for providing a visual signal, a haptic feedback device, a collision avoidance system, or a vehicle control device for a vehicle.
[0085]
[0085] Optionally, the first time series and the second time series each include any two or more of the following: distance to the preceding vehicle, distance to the intersection stop line, vehicle speed, time to collision, time to intersection violation, estimated braking distance, information about road conditions, information about special zones, information about the environment (weather, road type, etc.), information about traffic conditions, time, information about visibility conditions, information about identified objects, object location, direction of movement of objects, speed of objects, boundary box, vehicle operating parameters, information about the driver's state, information about the driver's history, continuous driving time, proximity to meal times, information about accident history, and audio information.
[0086]
[0086] The device comprises a first camera configured to observe the external environment, a second camera configured to observe the driver of a vehicle, and a processing unit configured to receive a first image from the first camera and a second image from the second camera, wherein the processing unit is configured to identify a first probability of a first predicted event and a second probability of a second predicted event, the first and second predicted events being associated with the driving of a vehicle, and the processing unit is configured to calculate a risk score based on the first probability of the first predicted event and the second probability of the second predicted event.
[0087]
[0087] Optionally, the first predicted event is a collision event, the second predicted event is a non-hazardous event, and the processing unit is configured to calculate the risk score based on the first probability of a collision event and the second probability of a non-hazardous event.
[0088]
[0088] The processing unit is optionally configured to calculate the risk score by applying a first weight to the first probability to obtain a first weighted probability, applying a second weight to the second probability to obtain a second weighted probability, and adding the first weighted probability and the second weighted probability.
[0089]
[0089] Optionally, the processing unit is configured to determine a third probability of a third predicted event, and the processing unit is configured to calculate the risk score based on a first probability of the first predicted event, a second probability of the second predicted event, and a third probability of the third predicted event.
[0090]
[0090] Optionally, the first predicted event is a collision event, the second predicted event is a near collision event, and the third predicted event is a non-hazardous event. The processing unit is configured to calculate the risk score based on the first probability of the collision event, the second probability of the near collision event, and the third probability of the non-hazardous event.
[0091]
[0091] The processing unit is optionally configured to calculate the risk score by applying a first weight to the first probability to obtain a first weighted probability, applying a second weight to the second probability to obtain a second weighted probability, applying a third weight to the third probability to obtain a third weighted probability, and adding the first weighted probability, the second weighted probability, and the third weighted probability.
[0092]
[0092] Optionally, the processing unit is configured to process data acquired over the past T seconds to determine a first probability of the first predicted event and a second probability of the second predicted event, where T is at least 3 seconds.
[0093]
[0093] Optionally, the first predicted event is at a future time at least 1 second or at least 2 seconds from the current time.
[0094]
[0094] Optionally, the processing unit is configured to calculate a risk score for a first time point, and the processing unit is also configured to calculate a further risk score for a second time point, and to identify the difference between the risk score and the further risk score, the difference indicating whether the dangerous situation is escalating or fading.
[0095]
[0095] Optionally, the processing unit is configured to generate a control signal based on the risk score.
[0096]
[0096] Optionally, the processing unit is configured to generate a control signal when the risk score meets the criteria.
[0097]
[0097] Optionally, the processing unit is configured to generate a control signal to operate the device when the risk score meets the criteria.
[0098]
[0098] Optionally, the device includes a speaker for issuing an alarm, a display or light-emitting device for providing a visual signal, a haptic feedback device, a collision avoidance system, or a vehicle control device for a vehicle.
[0099]
[0099] Optionally, the processing unit is configured to acquire a first image from the first camera and a second image from the second camera, and the processing unit is configured to determine a first probability of the first predicted event and a second probability of the second predicted event based on the first and second images.
[0100]
[0100] Optionally, the processing unit comprises a first-stage processing system and a second-stage processing system, wherein the first-stage processing system is configured to acquire raw data and process the raw data to provide a first output, and the second-stage processing system is configured to acquire the first output and process the first output to provide a second output, wherein the first output has lower dimensionality or less complexity compared to the raw data, and the second output has lower dimensionality or less complexity compared to the first output.
[0101]
[0101] Optionally, the processing unit includes a neural network model.
[0102]
[0102] Optionally, the neural network model is configured to receive an input that includes at least a first time series of information indicating a first risk factor and a second time series of information indicating a second risk factor.
[0103]
[0103] Optionally, the neural network model is configured to receive the first time series and the second time series in parallel and / or process the first time series and the second time series in parallel.
[0104]
[0104] Optionally, the processing unit is configured to package the input into a data structure for supplying to a neural network model.
[0105]
[0105] Optionally, the data structure consists of a two-dimensional data matrix.
[0106]
[0106] Optionally, the first time series shows the external state of the vehicle at different points in time, and the second time series shows the state of the driver at different points in time.
[0107]
[0107] Optionally, the input includes any one or more combinations of the following: distance to the preceding vehicle, distance to the intersection stop line, vehicle speed, time to collision, time to intersection violation, estimated braking distance, information on road conditions, information on special zones, information on the environment (weather, road type, etc.), information on traffic conditions, time, information on visibility conditions, information on identified objects, object location, direction of movement of objects, speed of objects, boundary box, vehicle operating parameters, information on the driver's state, information on the driver's history, continuous driving time, proximity to meal times, information on accident history, and voice information.
[0108]
[0108] The device includes a first camera configured to view the environment outside the vehicle, a second camera configured to view the driver of the vehicle, and a processing unit configured to receive a first image from the first camera and a second image from the second camera, wherein the processing unit includes a model configured to receive a plurality of inputs and generate a metric based on at least some of these inputs, wherein the plurality of inputs include at least a first time series of information indicating a first risk factor and a second time series of information indicating a second risk factor.
[0109]
[0109] Optionally, the multiple inputs may include: distance to the preceding vehicle (may indicate distance to collision), distance to the intersection stop line, vehicle speed, time to collision, time to intersection violation, estimated braking distance, information regarding road conditions (e.g., dry, wet, snow, etc.), information regarding special zones (e.g., construction zone, school zone, etc.) (e.g., for identification), information regarding the environment (e.g., city, suburb, urban area, etc.) (e.g., for identification), information regarding traffic conditions (e.g., for identification), time of day, information regarding visibility conditions (e.g., for identification), identified object (stop sign, traffic light, pedestrian, car, etc.), object position, direction of movement of the object, object speed, bounding box, vehicle operating parameters (e.g., kinematic signals such as acceleration, velocity, centrifugal force, steering angle, brake, accelerator position, turn signal), driver Information relating to the state of the vehicle (for example, to identify it), including one or more combinations of the following: (e.g., driver's gaze direction (looking up, looking down, looking left, etc.), using a phone, holding an object, smoking, eyes closed, turning or moving the head away from the face to avoid detection, gaze direction relative to the driver, gaze direction mapped to external or internal objects (e.g., head unit, mirrors, etc.), drowsiness, fatigue, anger towards the road, stress, sudden illness (heart attack, stroke, pulse, sweating, pupil dilation, etc.), information relating to the driver's history (vehicle records, accident rate, experience, age, years of driving, years in service, route experience, etc.), continuous driving time, proximity to meal times, and information relating to accident history (fatal accidents at a given time, accidents at a given time, geospatial heatmap of location-specific risk data by time, etc.).
[0110]
[0110] Optionally, the processing unit is configured to package two or more inputs into a data structure for supplying them to the model.
[0111]
[0111] Optionally, the data structure consists of a two-dimensional data matrix.
[0112]
[0112] Optionally, the model may include a neural network model.
[0113]
[0113] Optionally, the input includes first information relating to the driver's condition and second information relating to the external condition of the vehicle.
[0114]
[0114] Optionally, the neural network model is configured to receive first information regarding the driver's condition and second information regarding the external condition of the vehicle.
[0115]
[0115] Optionally, the neural network model is configured to receive the first time series and the second time series in parallel, and / or to process the first time series and the second time series in parallel.
[0116]
[0116] Optionally, the processing unit is configured to package the input into a data structure for supplying to the model.
[0117]
[0117] Optionally, the data structure includes a two-dimensional data matrix.
[0118]
[0118] Optionally, the two-dimensional data matrix includes the value of a first input from a plurality of inputs for each different time point, and the value of a second input from a plurality of inputs for each different time point.
[0119]
[0119] Optionally, the two-dimensional data matrix includes first information indicating the driver's state at different points in time, and second information indicating the external state of the vehicle.
[0120]
[0120] Optionally, the model of the processing unit is configured to predict near collisions and determine a score indicating the probability of the predicted near collision.
[0121]
[0121] Optionally, the model of the processing unit is configured to predict collisions and determine a score indicating the probability of the predicted collisions.
[0122]
[0122] Optionally, the model of the processing unit is configured to predict non-events (e.g., non-hazardous events) and to determine a score indicating the probability of the predicted non-events.
[0123]
[0123] Optionally, the metric is a weighted sum of scores for each different classification.
[0124]
[0124] Optionally, the classification may include at least "non-events" and "related events".
[0125]
[0125] Optionally, the classification includes at least “non-event,” “predicted near collision,” and “predicted collision.”
[0126]
[0126] Optionally, at least one of the inputs is based on the output from the first camera.
[0127]
[0127] Optionally, at least one of the inputs is based on the output from the second camera.
[0128]
[0128] Optionally, the processing unit is configured to provide a control signal to operate the device when the metric meets a criterion.
[0129]
[0129] The device may optionally be equipped with a speaker for generating an alarm.
[0130]
[0130] Optionally, the device comprises a display or light-emitting device for providing visual signals.
[0131]
[0131] Optionally, the device may include a component of a vehicle.
[0132]
[0132] The device comprises a first camera configured to view the environment outside the vehicle, a second camera configured to view the driver of the vehicle, and a processing unit configured to receive a first image from the first camera and a second image from the second camera, wherein the processing unit is configured to determine first information indicating the risk of collision with the vehicle based at least partially on the first image, and the processing unit is configured to determine second information indicating the driver's condition based at least partially on the second image, and the processing unit is configured to determine whether or not to provide a control signal to operate the device based on (1) the first information indicating the risk of collision with the vehicle and (2) the second information indicating the driver's condition.
[0133]
[0133] Optionally, the processing unit is configured to predict a collision at least 3 seconds before the predicted time of the predicted collision.
[0134]
[0134] Optionally, the processing unit is configured to predict a collision with sufficient lead time for the driver's brain to process the input and for the driver to take action to mitigate the risk of collision.
[0135]
[0135] The aforementioned sufficient lead time may depend on the driver's condition.
[0136]
[0136] Optionally, the first information indicating the risk of collision includes a predicted collision, the processing unit is configured to determine an estimated time until the predicted collision occurs, and the processing unit is configured to provide the control signal if the estimated time until the predicted collision occurs is less than a threshold.
[0137]
[0137] Optionally, the device includes a warning generator, and the processing unit is configured to provide a control signal to the device to issue a warning to the driver if the estimated time until a predicted collision occurs is less than a threshold.
[0138]
[0138] Optionally, the device includes a vehicle control device, the processing unit being configured to provide a control signal to cause the device to control the vehicle when the estimated time until a predicted collision occurs is less than a threshold.
[0139]
[0139] The threshold is optionally variable based on second information indicating the driver's condition.
[0140]
[0140] Optionally, the processing unit is configured to repeatedly evaluate the estimated time with respect to a variable threshold, as the estimated time until the predicted collision occurs decreases, and the predicted collision approaches in time.
[0141]
[0141] The threshold is arbitrarily variable in real time based on the driver's condition.
[0142]
[0142] Optionally, the processing unit is configured to change the threshold (for example, decrease or increase) if the driver's state indicates that the driver is distracted or not paying attention to the driving task.
[0143]
[0143] Optionally, the processing unit is configured to at least temporarily withhold the provision of a control signal if the estimated time until a predicted collision occurs is longer than a threshold.
[0144]
[0144] Optionally, the processing unit is configured to determine a level of collision risk, and the processing unit is configured to adjust the threshold based on the determined level of collision risk.
[0145]
[0145] Optionally, the driver's state includes a state of distraction, the processing unit is configured to determine the level of the driver's state of distraction, and the processing unit is configured to adjust the threshold based on the determined level of the driver's state of distraction.
[0146]
[0146] Optionally, if the driver's state indicates that the driver is paying attention to the driving task, the threshold has a first value; if the driver's state indicates that the driver is distracted or not paying attention to the driving task, the threshold has a second value that is higher than the first value.
[0147]
[0147] Optionally, the threshold may also be based on sensor information indicating that the vehicle is being operated to reduce the risk of collision.
[0148]
[0148] The processing unit is optionally configured to decide whether or not to provide a control signal based on (1) first information indicating the risk of collision with a vehicle, (2) second information indicating the driver's condition, and (3) sensor information indicating that the vehicle is being operated to reduce the risk of collision.
[0149]
[0149] Optionally, the apparatus further includes a non-primary medium for storing a first model, and the processing unit is configured to process a first image based on the first model to determine the risk of collision.
[0150]
[0150] Optionally, the first model includes a neural network model.
[0151]
[0151] Optionally, the non-temporary medium is configured to store a second model, and the processing unit is configured to process a second image based on the second model to determine the driver's state.
[0152]
[0152] The processing unit is optionally configured to determine a metric value for each of a plurality of driver state classifications, and the processing unit is configured to determine whether or not a driver is engaged in a driving task based on one or more of the metric values.
[0153]
[0153] The driver state classification may optionally include two or more of the following: the driver's gaze direction (for example, looking down, looking up, looking left, looking right, etc.), using a mobile phone, smoking, holding an object, hands off the steering wheel, not wearing a seat belt, eyes closed, looking forward, one hand on the steering wheel, both hands on the steering wheel.
[0154]
[0154] The processing unit may optionally be configured to compare the metric value with the respective threshold for each driver state classification.
[0155]
[0155] The processing unit may optionally be configured to determine that a driver belongs to one of the driver status classifications if one of the corresponding metric values satisfies or exceeds one of the corresponding thresholds.
[0156]
[0156] Optionally, the first camera, the second camera, and the processing unit are integrated as components of an aftermarket device for a vehicle.
[0157]
[0157] The processing unit is optionally configured to determine the second information by processing the second image and determining whether the driver's image satisfies the driver state classification, and the processing unit is configured to determine whether the driver is engaged in a driving task based on whether the driver's image satisfies the driver state classification.
[0158]
[0158] Optionally, the processing unit is configured to process the second image based on a neural network model in order to determine the state of the driver.
[0159]
[0159] Optionally, the processing unit is configured to determine whether an object in the first image or the bounding box of the object overlaps with the region of interest.
[0160]
[0160] Optionally, the region of interest has a shape that changes in accordance with the shape of the road or lane on which the vehicle is traveling.
[0161]
[0161] Optionally, the processing unit is configured to determine the center line of the road or lane on which the vehicle is traveling, and the region of interest has a shape based on the center line.
[0162]
[0162] Optionally, the processing unit is configured to determine the distance between the vehicle and the physical position based on the y-coordinate of the physical position in the camera image provided by the first camera, where the y-coordinate is relative to the image coordinate frame.
[0163]
[0163] A method performed by the device includes the steps of: acquiring a first image generated by a first camera, the first camera being configured to view the environment outside the vehicle; acquiring a second image generated by a second camera, the second camera being configured to view the driver of the vehicle; determining first information indicating a risk of collision with the vehicle, at least partially based on the first image; determining second information indicating the driver's condition, at least partially based on the second image; and determining whether or not to provide a control signal to operate the device, based on (1) the first information indicating a risk of collision with the vehicle and (2) the second information indicating the driver's condition.
[0164]
[0164] Optionally, the first information is determined by predicting a collision, which is predicted at least 3 seconds before the expected time of occurrence of the predicted collision.
[0165]
[0165] Optionally, the first information is determined by predicting a collision, which is predicted with sufficient lead time for the driver's brain to process the input and for the driver to take action to mitigate the risk of the collision.
[0166]
[0166] The aforementioned sufficient lead time may depend on the driver's condition.
[0167]
[0167] Optionally, the first information indicating the risk of collision includes a predicted collision, the method further includes the step of determining an estimated time until the predicted collision occurs, and the control signal is provided to cause the device to provide a control signal if the estimated time until the predicted collision occurs is less than a threshold.
[0168]
[0168] Optionally, the device comprises a warning generator, and the control signal is provided such that the device warns the driver if the estimated time until a predicted collision occurs is less than a threshold.
[0169]
[0169] Optionally, the device includes a vehicle control device, the control signals provided to cause the device to control the vehicle when the estimated time until a predicted collision occurs is less than a threshold.
[0170]
[0170] The threshold is optionally variable based on second information indicating the driver's condition.
[0171]
[0171] The estimated time is optionally evaluated repeatedly with respect to a variable threshold, as the estimated time until the predicted collision occurs decreases, the predicted collision approaches in time.
[0172]
[0172] The threshold value is arbitrarily variable in real time based on the driver's condition.
[0173]
[0173] Optionally, the method further includes the step of changing the threshold if the driver's condition indicates that the driver is distracted or not paying attention to the driving task.
[0174]
[0174] Optionally, the method further includes the step of at least temporarily suspending the generation of a control signal if the estimated time until a predicted collision occurs is longer than a threshold.
[0175]
[0175] Optionally, the method further includes the steps of determining a level of collision risk and adjusting the threshold based on the determined level of collision risk.
[0176]
[0176] Optionally, the driver's state includes a state of distraction, and the method further includes the steps of determining the level of the driver's state of distraction and adjusting the threshold based on the determined level of the driver's state of distraction.
[0177]
[0177] Optionally, if the driver's state indicates that the driver is paying attention to the driving task, the threshold has a first value; if the driver's state indicates that the driver is distracted or not paying attention to the driving task, the threshold has a second value that is higher than the first value.
[0178]
[0178] Optionally, the threshold may also be based on sensor information indicating that the vehicle is being operated to reduce the risk of collision.
[0179]
[0179] The step of deciding whether or not to provide a control signal to operate the device is also performed based on sensor information indicating that the vehicle is being operated to reduce the risk of collision.
[0180]
[0180] Optionally, the step of determining first information indicating the risk of collision includes processing the first image based on a first model.
[0181]
[0181] Optionally, the first model includes a neural network model.
[0182]
[0182] The step of optionally determining second information indicating the state of the driver includes processing the second image based on a second model.
[0183]
[0183] Optionally, the method further includes the steps of determining a metric value for each of a plurality of pose classifications, and determining whether or not a driver is engaged in a driving task based on one or more metric values.
[0184]
[0184] Optionally, the pose classification includes two or more of the following: the driver's gaze direction (e.g., looking down, looking up, looking left, looking right, etc.), using a mobile phone, smoking, holding an object, taking hands off the steering wheel, not wearing a seatbelt, closing eyes, looking forward, having one hand on the steering wheel, or having both hands on the steering wheel.
[0185]
[0185] Optionally, the method further includes the step of comparing the metric value with the respective threshold for each pose classification.
[0186]
[0186] Optionally, the method further includes the step of determining that a driver belongs to one of the pause classifications if one of the corresponding metric values satisfies or exceeds one of the corresponding thresholds.
[0187]
[0187] Optionally, the method may be performed by an aftermarket device, in which the first camera and the second camera are integrated as components of the aftermarket device.
[0188]
[0188] Optionally, the second information is determined by processing the second image to determine whether the image of the driver satisfies the pose classification, the method further includes the step of determining whether the driver is engaged in a driving task based on whether the image of the driver satisfies the pose classification.
[0189]
[0189] The step of optionally determining second information indicating the state of the driver includes processing the second image based on a neural network model.
[0190]
[0190] Optionally, the method further includes the step of determining whether an object in the first image or the bounding box of such object overlaps with the region of interest.
[0191]
[0191] Optionally, the region of interest has a shape that changes in accordance with the shape of the road or lane on which the vehicle is traveling.
[0192]
[0192] Optionally, the method further includes the step of identifying the center line of a road or lane on which the vehicle is traveling, wherein the region of interest has a shape based on the center line.
[0193]
[0193] Optionally, the method further includes the step of determining the distance between a vehicle and a physical position based on the y-coordinate of the physical position in the camera image provided by the first camera, where the y-coordinate is relative to the image coordinate frame.
[0194]
[0194] Further aspects and features will become apparent from reading the detailed description below. [Brief explanation of the drawing]
[0195]
[0195] The drawings illustrate the design and practicality of the embodiments, and similar elements are referred to by common reference numerals. A more specific description of the embodiments is provided below with reference to the attached drawings for a better understanding of how the advantages and objectives are obtained. It should be understood that these drawings only illustrate exemplary embodiments and therefore do not limit the scope of the claimed invention. [Figure 1]
[0196] Figure 1 shows a device according to several embodiments. [Figure 2]
[0197] Figure 2A is a block diagram of the apparatus of Figure 1 according to several embodiments.
[0198] Figure 2B shows an example of the processing scheme of the apparatus shown in Figure 2A. [Figure 3]
[0199] Figure 3 shows an example of an image taken by the camera of the device shown in Figure 2. [Figure 4]
[0200] Figure 4 shows an example of a classifier. [Figure 5]
[0201] Figure 5 shows a method according to several embodiments. [Figure 6]
[0202] Figure 6 shows an example of an image captured by the camera in Figure 1, and the output of various classifiers. [Figure 7]
[0203] Figure 7 shows another example of an image captured by the camera in Figure 1, and the outputs of various classifiers. [Figure 8]
[0204] Figure 8 shows another example of an image captured by the camera in Figure 1, and the outputs of various classifiers. [Figure 9]
[0205] Figure 9 shows an example of a processing architecture having a first model and a second model coupled in series. [Figure 10]
[0206] Figure 10 shows an example of feature information received by the second model. [Figure 11]
[0207] Figure 11 shows another example of feature information received by the second model. [Figure 12]
[0208] Figure 12 shows a drowsiness detection method performed by the apparatus in Figure 2, according to several embodiments. [Figure 13]
[0209] Figure 13 shows an example of a processing architecture that can be implemented using the device shown in Figure 2A. [Figure 14]
[0210] Figure 14 shows examples of object detection according to several embodiments. [Figure 15]
[0211] Figure 15A shows another example of object identifiers, specifically indicating that each object identifier is a box representing a preceding vehicle.
[0212] Figure 15B shows another example of an object identifier, notably indicating that each object identifier is a horizontal line.
[0213] Figure 15C shows examples of preceding vehicle volume detection according to several embodiments.
[0214] Figure 15D shows a technique for determining a region of interest based on centerline detection.
[0215] Figure 15E illustrates the advantages of using the region of interest in Figure 15D to detect objects at risk of collision. [Figure 16]
[0216] Figure 16 shows three exemplary scenarios involving a collision with a preceding vehicle. [Figure 17]
[0217] Figure 17 shows another example of object detection when the detected object is a human being. [Figure 18]
[0218] Figure 18 shows an example of predicting human movement. [Figure 19]
[0219] Figures 19A-19B show other examples of object detection where the detected object is related to an intersection.
[0220] Figure 19C illustrates the concept of braking distance. [Figure 20]
[0221] Figure 20 shows a method for determining the distance between the target vehicle and the position in front of the vehicle. [Figure 21]
[0222] Figure 21 shows an example of a technique for generating control signals to control a vehicle and / or to issue warnings to the driver. [Figure 22]
[0223] Figure 22A shows a method including collision prediction according to several embodiments.
[0224] Figure 22B shows a method including the prediction of intersection violations according to several embodiments. [Figure 23]
[0225] Figure 23 shows techniques for determining the model to be used by the apparatus in Figure 2A, according to several embodiments. [Figure 24]
[0226] Figure 24 shows one technique for addressing risk factors. [Figure 25]
[0227] Figure 25 shows another technique for addressing risk factors. [Figure 26]
[0228] Figures 26A–26C show examples of neural network architectures that can be employed to process multiple inputs. [Figure 27]
[0229] Figure 27 shows the number of deaths by time of day and day of the week. [Figure 28]
[0230] Figure 28 shows an example of multiple inputs that the model processes to generate a prediction at least two seconds before the event occurs. [Figure 29]
[0231] Figure 29 shows an example of risk score calculation based on model predictions. [Figure 30]
[0232] Figure 30 shows an example of a model output based on multiple time-series inputs, where the model output exhibits a high "collision" state. [Figure 31]
[0233] Figure 31 shows an example of a model output based on multiple time-series inputs, where the model output indicates a high "near collision" state. [Figure 32]
[0234] Figure 32 shows an example of a model output based on multiple time-series inputs, where the model output exhibits a high "non-event" state. [Figure 33]
[0235] Figures 33A to 33K show a series of external images and their corresponding internal images, particularly illustrating the problems that can be addressed by the technique shown in Figure 25. [Figure 34]
[0236] Figures 34A to 34E show a series of external images and their corresponding internal images, particularly illustrating the problems that can be addressed by the technique shown in Figure 25. [Figure 35]
[0237] Figures 35A to 35D show a series of external images and their corresponding internal images, particularly illustrating the problems that can be addressed by the technique shown in Figure 25. [Figure 36]
[0238] Figures 36A to 36C show a series of external images and their corresponding internal images, particularly illustrating the problems that can be addressed by the technique shown in Figure 25. [Figure 37]
[0239] Figures 37A to 37C show a series of external images and their corresponding internal images, particularly illustrating the problems that can be addressed by the technique shown in Figure 25. [Figure 38]
[0240] Figure 38 shows a dedicated processing system for mounting one or more electronic devices described herein. [Modes for carrying out the invention]
[0196]
[0241] Various embodiments are described below with reference to the drawings. Note that the drawings may or may not be drawn to scale, and that elements having similar structure or function are represented by the same reference numerals throughout the drawings. Also note that the drawings are intended to facilitate the explanation of the embodiments. They do not exhaustively describe the claimed invention, nor do they limit the scope of the claimed invention. Furthermore, the illustrated embodiments do not necessarily have all aspects or advantages of the invention shown. Aspects or advantages described in relation to a particular embodiment are not necessarily limited to that embodiment and can be implemented in other embodiments even if they are not illustrated or explicitly described as such.
[0197]
[0242] Figure 1 shows a device 200 according to several embodiments. The device 200 is configured to be mounted on a vehicle, such as on the vehicle's windshield or rearview mirror. The device 200 includes a first camera 202 configured to view outside the vehicle and a second camera 204 configured to view inside the vehicle. In the illustrated embodiments, the device 200 is in the form of an aftermarket device that can be mounted on a vehicle (i.e., offline from the vehicle's manufacturing process). The device 200 may include a connector configured to connect the device 200 to the vehicle. In non-limiting examples, the connector may be a suction cup, adhesive, clamp, one or more screws, etc. The connector may be configured to removably secure the device 200 to the vehicle, in which case the device 200 can be selectively removed from and / or attached to the vehicle as needed. Alternatively, the connector may be configured to permanently secure the device 200 to the vehicle. In other embodiments, the device 200 may be a vehicle component that is installed during the vehicle's manufacturing process. It should be noted that the device 200 is not limited to the exemplary configuration, and the device 200 may have other configurations in other embodiments. For example, in other embodiments, the device 200 may have a different form factor. In other embodiments, the device 200 may be an end-user device such as a mobile phone or tablet having one or more cameras.
[0198]
[0243] Figure 2A is a block diagram of the apparatus 200 of Figure 1 according to several embodiments. The apparatus 200 comprises a first camera 202 and a second camera 204. As shown, the apparatus 200 also comprises a processing unit 210 coupled to the first camera 202 and the second camera 204, a non-temporary medium 230 configured to store data, a communication unit 240 connected to the processing unit 210, and a speaker 250 connected to the processing unit 210.
[0199]
[0244] In the illustrated embodiment, the first camera 202, the second camera 204, the processing unit 210, the non-temporary medium 230, the communication unit 240, and the speaker 250 may be integrated as components of an aftermarket device for the vehicle. In other embodiments, the first camera 202, the second camera 204, the processing unit 210, the non-temporary medium 230, the communication unit 240, and the speaker 250 may be integrated into the vehicle or installed in the vehicle during the vehicle manufacturing process.
[0200]
[0245] The processing unit 210 is configured to acquire images from the first camera 202 and the second camera 204, and to process the images from the first and second cameras 202 and 204. In some embodiments, the image from the first camera 202 may be processed by the processing unit 210 to monitor the external environment of the vehicle (e.g., for collision detection, collision avoidance, driving environment monitoring, etc.). Also in some embodiments, the image from the second camera 204 may be processed by the processing unit 210 to monitor the driver's driving behavior (e.g., whether the driver is distracted, drowsy, or focused, etc.). In further embodiments, the processing unit 210 may process the images from the first camera 202 and / or the second camera 204 to determine the risk of collision, predict collisions, provide warnings to the driver, etc. In other embodiments, the device 200 may not include the first camera 202. In such cases, the device 200 is configured to monitor only the environment inside the vehicle.
[0201]
[0246] The processing unit 210 of the device 200 may include hardware, software, or a combination of both. In a non-limiting example, the hardware of the processing unit 210 may include one or more processors and / or one or more integrated circuits. In some embodiments, the processing unit 210 may be implemented as a module and / or as part of any integrated circuit.
[0202]
[0247] The non-temporary medium 230 is configured to store data relating to the operation of the processing unit 210. In the illustrated embodiment, the non-temporary medium 230 is configured to store a model that the processing unit 210 can access and use to identify the driver's pose appearing in the image from camera 204 and / or to determine whether the driver is engaged in a driving task. Alternatively, the processing unit 210 may be configured so that the model has the function of identifying the driver's pose and / or determining whether the driver is engaged in a driving task. Optionally, the non-temporary medium 230 may be configured to store images from the first camera 202 and / or images from the second camera 204. Also, in some embodiments, the non-temporary medium 230 may be configured to store data generated by the processing unit 210.
[0203]
[0248] The model stored in the non-temporary medium 230 may be any computational or processing model, including but not limited to a neural network model. In some embodiments, the model includes feature extraction parameters, which the processing unit 210 can use to extract features from images provided by the camera 204 to identify objects such as the driver's head, hat, face, nose, eyes, and mobile device. In some embodiments, the model may also include program instructions, commands, scripts, etc. In one embodiment, the model may be in the form of an application that can be wirelessly received by the device 200.
[0204]
[0249] The communication unit 240 of the device 200 is configured to receive data wirelessly from a network such as a cloud, the internet, or a Bluetooth network. In some embodiments, the communication unit 240 may be configured to transmit data wirelessly. For example, images from the first camera 202, images from the second camera 204, data generated by the processing unit, or any combination thereof may be transmitted by the communication unit 240 to another device (e.g., a server, an accessory device such as a mobile phone, another device 200 in another vehicle, etc.) via a network such as a cloud, the internet, or a Bluetooth network. In some embodiments, the communication unit 240 may include one or more antennas. For example, the communication 240 may include a first antenna configured to provide long-range communication and a second antenna configured to provide short-range communication (e.g., via Bluetooth). In other embodiments, the communication unit 240 may be configured to physically transmit and / or receive data via a cable or electrical contacts. In such cases, the communication unit 240 may include one or more communication connectors configured to couple with a data transmission device. For example, the communication unit 240 may include a connector configured to connect to a cable, a USB slot configured to receive a USB drive, a memory card slot configured to receive a memory card, and so on.
[0205]
[0250] The speaker 250 of the device 200 is configured to provide the driver of the vehicle with an audible warning and / or message. For example, in some embodiments, the processing unit 210 may be configured to detect an impending collision between the vehicle and an object outside the vehicle. In such cases, in response to the detection of an impending collision, the processing unit 210 may generate a control signal to cause the speaker 250 to output an audible warning and / or message. As another example, in some embodiments, the processing unit 210 may be configured to determine whether the driver is engaged in a driving task. If the driver is not engaged in a driving task, or has not engaged in a driving task for a predetermined period of time (e.g., 2 seconds, 3 seconds, 4 seconds, 5 seconds, etc.), the processing unit 210 may generate a control signal to cause the speaker 250 to output an audible warning and / or message.
[0206]
[0251] Alternatively or additionally, the processing unit 210 may generate control signals that activate a haptic feedback device to warn the driver and / or provide input to a collision avoidance system.
[0207]
[0252] Although the device 200 is described as having a first camera 202 and a second camera 204, in other embodiments the device 200 may not have the first camera 202 and may include only the second camera (cabin camera) 204. In other embodiments the device 200 may also include multiple cameras configured to view the cabin inside the vehicle.
[0208]
[0253] As shown in Figure 2A, the processing unit 210 also includes a driver monitoring module 211, an object detection unit 216, a collision prediction unit 218, an intersection violation prediction unit 222, and a signal generation controller 224. The driver monitoring module 211 is configured to monitor the driver of the vehicle based on one or more images provided by the second camera 204. In some embodiments, the driver monitoring module 211 is configured to determine one or more poses of the driver. In some embodiments, the driver monitoring module 211 may also be configured to determine the driver's state, such as whether the driver is awake, drowsy, or paying attention to the driving task. In some cases, the driver's pose itself can also be considered the driver's state.
[0209]
[0254] The object detection unit 216 is configured to detect one or more objects in the environment outside the vehicle based on one or more images provided by the first camera 202. In non-limiting examples, the objects to be detected may be vehicles (e.g., automobiles, motorcycles, etc.), lane boundaries, people, bicycles, animals, road signs (e.g., stop signs, road signs, no U-turn signs, etc.), traffic lights, road markings (e.g., stop lines, lane dividers, letters painted on the road, etc.). In some embodiments, the detected vehicle may be a preceding vehicle, which is a vehicle traveling in the same lane as the target vehicle and ahead of the target vehicle.
[0210]
[0255] The collision prediction unit 218 is configured to determine the risk of collision based on the output from the object detection unit 216. For example, in some embodiments, the collision prediction unit 218 may determine that there is a risk of collision with a preceding vehicle and output information indicating such a risk of collision. In some embodiments, the collision prediction unit 218 may optionally acquire sensor information indicating the state of the vehicle, such as vehicle speed, vehicle acceleration, vehicle turning angle, vehicle turning direction, vehicle braking, vehicle travel direction, engine status, wheel traction, turn signal status, accelerator pedal position, brake pedal position, or any combination thereof. In such cases, the collision prediction unit 218 may be configured to determine the risk of collision based on the output from the object detection unit 216 and the acquired sensor information. In some embodiments, the collision prediction unit 218 may also be configured to identify the relative speed between the target vehicle and an object (e.g., a preceding vehicle) and determine that there is a risk of collision based on the identified relative speed. In some embodiments, the collision prediction unit 218 may be configured to identify the speed of the target vehicle, the speed of the moving object, the target vehicle's travel path, and the moving object's travel path, and to determine if there is a risk of collision based on these parameters. For example, if an object is moving along a path that intersects with the target vehicle's path, and the time until the object and the target vehicle collide based on their respective speeds is less than a time threshold, the collision prediction unit 218 may determine that there is a risk of collision. Note that the objects that may collide with the target vehicle are not limited to moving objects (cars, motorcycles, bicycles, pedestrians, animals, etc.), and the collision prediction unit 218 may be configured to determine the risk of collision with non-moving objects such as parked cars, road signs, utility poles, buildings, trees, and mailboxes.
[0211]
[0256] In some embodiments, the collision prediction unit 218 may be configured to identify the time it will take for a predicted collision to occur and compare that time to a threshold time. If this time is shorter than the threshold time, the collision prediction unit 218 can determine that there is a risk of collision with the target vehicle. In some embodiments, the threshold time for identifying the risk of collision may be at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 seconds, or more. In some embodiments, when predicting a collision, the collision prediction unit 218 may consider the speed, acceleration, direction of travel, braking operation, or any combination thereof of the target vehicle. Optionally, the collision prediction unit 218 may also consider the speed, acceleration, direction of travel, or any combination thereof of a detected object that is predicted to collide with the vehicle.
[0212]
[0257] The intersection violation prediction unit 222 is configured to determine the risk of an intersection violation based on the output from the object detection unit 216. For example, in some embodiments, the intersection violation prediction unit 222 may determine that there is a risk that the target vehicle will not be able to stop in a target area associated with a stop sign or red light, and output information indicating such an intersection violation risk. In some embodiments, the intersection violation prediction unit 222 may optionally also acquire sensor information indicating the state of the vehicle, such as the vehicle's speed, acceleration, turning angle, turning direction, braking, or any combination thereof. In such cases, the intersection violation prediction unit 222 may be configured to determine the risk of an intersection violation based on the output from the object detection unit 216 and the acquired sensor information.
[0213]
[0258] In some embodiments, the intersection violation prediction unit 222 may also be configured to identify a target area (e.g., a stop line) where the target vehicle is expected to stop, determine the distance between the target vehicle and the target area, and compare that distance with a threshold distance. If this distance is smaller than the threshold distance, the intersection violation prediction unit 222 may determine that there is a risk of an intersection violation.
[0214]
[0259] In other embodiments, the intersection violation prediction unit 222 may be configured to identify a target area (e.g., a stop line) where the target vehicle is expected to stop, determine the time it will take for the vehicle to reach the target area, and compare that time to a threshold time. If that time is shorter than the threshold time, the intersection violation prediction unit 222 may determine that there is a risk of an intersection violation. In some embodiments, the threshold time for identifying the risk of an intersection violation may be at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 seconds, or more. In some embodiments, when predicting an intersection violation, the intersection violation prediction unit 222 may consider the speed, acceleration, direction of travel, braking operation, or any combination thereof of the target vehicle.
[0215]
[0260] It should be noted that intersection violations are not limited to violations of stop signs or red lights, and the intersection violation prediction unit 222 may be configured to determine the risk of other intersection violations, such as a vehicle entering a road going the wrong way or a vehicle turning at an intersection with a "No Turning on Red Light" sign.
[0216]
[0261] In some embodiments, the signal generation controller 224 is configured to determine whether to generate a control signal based on the output from the collision prediction unit 218 and the output from the driver monitoring module. Alternatively or additionally, the signal generation controller 224 is configured to determine whether to generate a control signal based on the output from the intersection violation prediction unit 222 and optionally also based on the output from the driver monitoring module. In some embodiments, the signal generation controller 224 is also configured to determine whether to generate a control signal based on sensor information provided by one or more sensors in the vehicle.
[0217]
[0262] In some embodiments, the control signal is configured to cause a device (e.g., a warning generator) to provide a warning to the driver if the estimated time until a predicted collision occurs is less than a threshold (action threshold). For example, the warning generator may attract the driver's attention by outputting an audible signal, a visual signal, a mechanical vibration (steering wheel vibration), or any combination thereof. Alternatively or additionally, the control signal is configured to cause a device (e.g., a vehicle control system) to control the vehicle if the estimated time until a predicted collision occurs is less than or equal to a threshold (action threshold). For example, the vehicle control system may automatically apply the brakes, automatically release the accelerator pedal, automatically activate the hazard lights, automatically steer the vehicle, or any combination thereof. Thus, in some embodiments, the control signal may be supplied to a braking system (e.g., an automatic emergency braking system), a steering system, a lane control system, a Level 3 automation system, or any combination thereof. In some embodiments, the signal generation controller 224 may be configured to provide a first control signal to give a warning to the driver. If the driver does not take any action to mitigate the risk of collision, the signal generation controller 224 may then provide a second control signal to cause the vehicle control system to take control of the vehicle, such as automatically applying the vehicle's brakes.
[0218]
[0263] In some embodiments, the signal generation controller 224 may be a separate component (e.g., a module) from the collision prediction unit 218 and the intersection violation prediction unit 222. In other embodiments, the signal generation controller 224 or at least a portion of the signal generation controller 224 may be implemented as part of the collision prediction unit 218 and / or the intersection violation prediction unit 222. Also, in some embodiments, the collision prediction unit 218 and the intersection violation prediction unit 222 may be integrated.
[0219]
[0264] During operation, the device 200 is coupled to the vehicle such that the first camera 202 views the exterior of the vehicle and the second camera 204 views the driver inside the vehicle. While the driver operates the vehicle, the first camera 202 captures images of the exterior of the vehicle, and the second camera 204 captures images of the interior of the vehicle. Figure 2B shows an example of the processing scheme of the device 200. As shown in the figure, while the device 200 is in use, the second camera 204 provides images as input to the driver monitoring module 211. The driver monitoring module 211 analyzes the images to identify one or more poses of the driver of the target vehicle. As a non-limiting example, one or more poses include one or more of the following: looking down, looking up, looking left, looking right, using a mobile phone, smoking, holding an object, hands off the steering wheel, not wearing a seatbelt, eyes closed, looking forward, one hand on the steering wheel, and both hands on the steering wheel. In some embodiments, the driver monitoring module 211 may be configured to determine one or more states of the driver based on identified driver postures. For example, the driver monitoring module 211 may determine whether the driver is distracted based on one or more identified driver postures. As another example, the driver monitoring module 211 may determine whether the driver is drowsy based on one or more identified driver postures. In some embodiments, the driver monitoring module 211 may determine that the driver is distracted if the driver is in a specific posture (e.g., a posture for using a mobile phone). Also in some embodiments, the driver monitoring module 211 may analyze a series of driver posture classifications over a period of time to determine whether the driver is drowsy.
[0220]
[0265] The first camera 202 provides an image as input to the object detection unit 216, which analyzes the image to detect one or more objects in the image. As shown in the figure, the object detection unit 216 includes different detectors configured to detect different types of objects. In particular, the object detection unit 216 has a vehicle detection unit 260 configured to detect vehicles outside of a target vehicle, a vulnerable object detection unit 262 configured to detect vulnerable objects such as people, bicycles with people on them, and animals, and an intersection detection unit 264 configured to detect one or more items for identifying an intersection (e.g., stop signs, traffic lights, pedestrian crossing markings, etc.). In some embodiments, the object detection unit 216 may be configured to determine different types of objects based on different models. For example, there may be a vehicle detection model configured to detect vehicles, a human detection model configured to detect people, an animal detection model configured to detect animals, a traffic light detection model configured to detect traffic lights, a stop sign detection model configured to detect stop signs, a centerline detection model configured to detect the centerline of a road, and so on. In some embodiments, the different models may be different neural network models, each trained to detect different types of objects.
[0221]
[0266] The vehicle detection unit 260 is configured to detect vehicles outside the target vehicle and provide information about the detected vehicles (vehicle identifier, vehicle position, etc.) to module 221. Module 221 may be a collision prediction unit 218 and / or an intersection violation prediction unit 222. Module 221 includes an object tracking unit 266 configured to track one or more detected vehicles, a course prediction unit 268 configured to identify the course of a predicted collision, and a collision / crossing time (TTC) module 269 configured to estimate the time until a predicted collision occurs. In some embodiments, the object tracking unit 266 is configured to identify a preceding vehicle traveling ahead of the target vehicle. Also in some embodiments, the course prediction unit 268 is configured to determine the course of a predicted collision based on the identified preceding vehicle and sensor information from sensor 225. For example, based on the speed and direction of travel of the target vehicle, the course prediction unit 268 can identify the course of a predicted collision. The TTC module 269 is configured to calculate the predicted time to collision based on information about the predicted collision course and sensor information from sensor 225. For example, the TTC module 269 may calculate TTC (time-to-collision) based on the distance of the collision course and the relative speed of the preceding vehicle and the target vehicle. The signal generation controller 224 is configured to determine, based on the output from module 221 and the output from TTC module 269, whether to generate a control signal to activate a warning generator to warn the driver and / or to activate a vehicle control device to control the vehicle (e.g., automatically release the accelerator, apply the brakes, etc.). In some embodiments, if the TTC is less than a threshold (e.g., 3 seconds), the signal generation controller 224 generates a control signal to activate the warning generator and / or the vehicle control device. In some embodiments, the threshold may be adjustable based on the output from driver monitoring module 211.For example, if the output from the driver monitoring module 211 indicates that the driver is distracted or not paying attention to the driving task, the signal generation controller 224 may increase the threshold (for example, by setting the threshold to 5 seconds). This would cause the signal generation controller 224 to supply a control signal if the TTC (Time To Change) with the preceding vehicle is less than 5 seconds.
[0222]
[0267] It should be noted that module 221 is not limited to predicting collisions with preceding vehicles, but may be configured to predict collisions with other vehicles. For example, in some embodiments, module 221 may be configured to detect vehicles traveling toward the target vehicle's path, such as vehicles approaching an intersection or vehicles merging into the target vehicle's lane. In such situations, the course prediction unit 268 identifies the target vehicle's path and the paths of other vehicles, and further identifies the intersection of the two paths. The TTC module 269 is configured to determine the TTC based on the location of the intersection, the speed of the other vehicles, and the speed of the target vehicle.
[0223]
[0268] The vulnerable object detection unit 262 is configured to detect vulnerable objects outside the target vehicle and provide information about the detected objects (object identifier, object location, etc.) to module 221. For example, the vulnerable object detection unit 262 may detect a person outside the target vehicle and provide information about the detected person to module 221. Module 221 includes an object tracking unit 266 configured to track one or more detected objects (e.g., a person), a course prediction unit 268 configured to identify a predicted collision course, and a collision / crossing time (TTC) module 269 configured to estimate the time until an estimated collision occurs. Because certain objects, such as people, animals, and bicycles, may have unpredictable directions of movement, in some embodiments the course prediction unit 268 is configured to determine a box surrounding an image of the detected object to indicate the object's possible location. In some embodiments the course prediction unit 268 is configured to identify a predicted collision course based on a box surrounding an identified object (e.g., a person) and sensor information from sensor 225. For example, based on the speed of the target vehicle, the direction of travel of the target vehicle, and the box surrounding the identified object, the course prediction unit 268 may determine that the current travel path of the target vehicle intersects with the box. In this case, the predicted collision course will be the travel path of the target vehicle, and the predicted collision location will be the intersection of the travel path of the target vehicle and the box surrounding the object. The TTC module 269 is configured to calculate the time to the predicted collision based on information about the predicted collision course and sensor information from sensor 225. For example, the TTC module 269 may calculate the TTC (time-to-collision) based on the distance of the collision course and the relative speed of the preceding vehicle and the person. The signal generation controller 224 is configured to determine, based on the output from module 221 and the output from TTC module 269, whether to generate a control signal to activate a warning generator to warn the driver, and / or to activate a vehicle control device to control the vehicle (e.g., automatically release the accelerator, apply the brakes, etc.).In some embodiments, if the TTC is less than a threshold (e.g., 3 seconds), the signal generation controller 224 generates a control signal to activate the warning generator and / or vehicle control device. In some embodiments, the threshold may also be adjustable based on the output from the driver monitoring module 211. For example, if the output from the driver monitoring module 211 indicates that the driver is distracted or not paying attention to the driving task, the signal generation controller 224 may increase the threshold (e.g., set the threshold to 5 seconds). This would cause the signal generation controller 224 to supply a control signal if the TTC to an object is less than 5 seconds.
[0224]
[0269] It should be noted that module 221 is not limited to predicting collisions with humans and may be configured to predict collisions with other objects. For example, in some embodiments, module 221 may be configured to detect animals, cyclists, roller skaters, skateboarders, etc. In such situations, the course prediction unit 268 may be configured to identify the path of the target vehicle, as well as the path of the detected object (if the object is moving in one direction, such as a bicycle), and further to identify the intersection of the two paths. In other cases, when the movement of the object is more unpredictable (such as an animal), the course prediction unit 268 may identify the path of the target vehicle and a box encompassing a range of possible positions of the object, and similarly determine the intersection of the target vehicle's path and the box. The TTC module 269 is configured to determine the TTC based on the location of the intersection and the speed of the target vehicle.
[0225]
[0270] The intersection detection unit 264 is configured to detect one or more objects indicating an intersection outside the target vehicle and to provide information about the intersection (such as the type of intersection and the required stopping position for the vehicle) to module 221. In non-limiting examples, the one or more objects indicating an intersection may include traffic lights, stop signs, road signs, or any combination thereof. Intersections that the intersection detection unit 264 can detect may include stop sign intersections, signalized intersections, railway intersections, etc. Module 221 includes a course prediction unit 268 configured to identify the course of a predicted intersection violation and a TTC (Time to Collision / Cross) module 269 configured to estimate the time until the predicted intersection violation occurs. The TTC module 269 is configured to calculate the estimated time until the intersection violation occurs based on the required stopping position for the vehicle and sensor information from sensor 225. For example, the TTC module 269 may calculate the TTC (Time to Cross) based on the course distance (e.g., the distance between the vehicle's current position and the required stopping position) and the speed of the target vehicle. In some embodiments, the required stopping position may be determined by the object detection unit 216 detecting a stop line mark on the road. In other embodiments, there may be no stop line markings on the road. In such cases, the course prediction unit 268 may determine a virtual or graphical line indicating the required stopping position. The signal generation controller 224 is configured to determine, based on the output from module 221 and the output from TTC module 269, whether to generate a control signal to activate a warning generator to warn the driver and / or to activate a vehicle control device to control the vehicle (e.g., automatically release the accelerator, apply the brakes, etc.). In some embodiments, if the TTC is less than a threshold (e.g., 3 seconds), the signal generation controller 224 generates a control signal to activate the warning generator and / or the vehicle control device. In some embodiments, the threshold may also be adjustable based on the output from driver monitoring module 211.For example, if the output from the driver monitoring module 211 indicates that the driver is distracted or not paying attention to the driving task, the signal generation controller 224 may increase the threshold (for example, by setting the threshold to 5 seconds). This causes the signal generation controller 224 to supply a control signal if the time to pass through the intersection is less than 5 seconds.
[0226]
[0271] In some embodiments, the signal generation controller 224 may be configured to apply different threshold values for generating control signals based on the type of driver state indicated by the output of the driver monitoring module 211, in relation to a predicted collision with another vehicle, a predicted collision with an object, or a predicted intersection violation. For example, if the output of the driver monitoring module 211 indicates that the driver is looking at a mobile phone, the signal generation controller 224 may generate a control signal to activate a warning generator and / or a vehicle control device in response to the TTC being met or below a 5-second threshold. On the other hand, if the output of the driver monitoring module 211 indicates that the driver is drowsy, the signal generation controller 224 may generate a control signal to activate a warning generator and / or a vehicle control device in response to the TTC being below an 8-second threshold (for example, longer than the threshold when the driver is using a mobile phone). In some cases, depending on the driver's condition (e.g., if the driver is drowsy or falling asleep at the wheel), it may take time for the driver to react to an imminent collision. Therefore, a longer time threshold (compared to the TTC value) may be required to alert the driver and control the vehicle. Accordingly, the signal generation controller 224 may issue a warning to the driver and / or activate the vehicle control device earlier in response to a collision predicted in such situations.
[0227]
[0272] Driver status assessment
[0228]
[0273] As described herein, the second camera 204 is configured to view the driver inside the vehicle. While the driver is operating the vehicle, the first camera 202 captures an image of the outside of the vehicle, and the second camera 204 captures an image of the inside of the vehicle. Figure 3 shows an example of an image 300 captured by the second camera 204 of the device 200 of Figure 2. As shown, the image 300 from the second camera 202 may include an image of the driver 310 operating the target vehicle (the vehicle equipped with the device 200). The processing unit 210 is configured to process the image from the camera 202 (e.g., image 300) and determine whether the driver is engaged in a driving task. In non-limiting examples, a driving task may include paying attention to the road and environment ahead of the target vehicle, or having hands on the steering wheel.
[0229]
[0274] As shown in Figure 4, in some embodiments, the processing unit 210 is configured to process the driver's image 300 from the camera 202 and determine whether the driver belongs to a specific pose classification. In non-limiting examples, pose classifications include one or more of the following: looking down, looking up, looking left, looking right, using a mobile phone, smoking, holding an object, hands off the steering wheel, not wearing a seatbelt, eyes closed, looking forward, one hand on the steering wheel, and both hands on the steering wheel. In some embodiments, the processing unit 210 is also configured to determine whether the driver is engaged in a driving task based on one or more pose classifications. For example, if the driver's head is "looking down" and the driver is holding a mobile phone, the processing unit 210 may determine that the driver is not engaged in a driving task (i.e., the driver is not paying attention to the road or the environment in front of the vehicle). In another example, if the driver's head is "looking" to the right or left and the head rotation angle exceeds a certain threshold, the processing unit 210 may determine that the driver is not engaged in a driving task.
[0230]
[0275] In some embodiments, the processing unit 210 is configured to determine whether the driver is engaged in a driving task based on one or more poses of the driver displayed in the image, without needing to identify the direction of the driver's gaze. This function is advantageous because the direction of the driver's gaze may not be visible in the image or may not be accurately determined. For example, if the driver is wearing a hat, their eyes may not be visible to the onboard camera. The driver may also be wearing sunglasses that obstruct their vision. In some cases, if the driver is wearing clear glasses, the frames of the glasses may obstruct their vision or the lenses of the glasses may make eye detection inaccurate. Therefore, determining whether the driver is engaged in a driving task without needing to identify the direction of the driver's gaze is advantageous because the processing unit 210 can determine whether the driver is engaged in a driving task even if the driver's eyes cannot be detected and / or the direction of their gaze cannot be determined.
[0231]
[0276] In some embodiments, the processing unit 210 may use context-based classification to determine whether the driver is engaged in a driving task. For example, if the driver's head is facing downwards and the driver is holding a mobile phone on their lap, the processing unit 210 may determine that the driver is not engaged in a driving task. The processing unit 210 can make such a determination even if the driver's eyes cannot be detected (for example, if they are obscured by a hat as shown in Figure 3). The processing unit 210 may also use context-based classification to identify one or more poses of the driver. For example, if the driver's head is facing downwards, the processing unit 210 may determine that the driver is looking downwards even if the driver's eyes cannot be detected. As another example, if the driver's head is facing upwards, the processing unit 210 may determine that the driver is looking upwards even if the driver's eyes cannot be detected. As yet another example, if the driver's head is facing to the right, the processing unit 210 may determine that the driver is looking to the right even if the driver's eyes cannot be detected. As a further example, if the driver's head is turned to the left, the processing unit 210 may determine that the driver is looking to the left even if the driver's eyes cannot be detected.
[0232]
[0277] In one embodiment, the processing unit 210 may be configured to use a model to identify one or more poses of the driver and determine whether the driver is engaged in a driving task. This model may be used by the processing unit 210 to process images from the camera 204. In some embodiments, the model may be stored in a non-temporary medium 230. Also, in some embodiments, the model may be transmitted from a server and received by the device 200 via a communication unit 240.
[0233]
[0278] In some embodiments, the model may be a neural network model. In such cases, the neural network model may be trained on images of other drivers. For example, the neural network model may be trained using images of drivers to identify various poses such as looking down, looking up, looking left, looking right, using a mobile phone, smoking, holding an object, taking hands off the steering wheel, not wearing a seatbelt, closing eyes, looking forward, having one hand on the steering wheel, and having both hands on the steering wheel. In some embodiments, the neural network model may be trained to identify different poses even without detecting the eyes of a person in the image. This allows the neural network model to identify various poses and / or determine, based on context (for example, based on information captured in the image regarding the driver's state other than the direction of the driver's gaze), whether the driver is engaged in a driving task. In other embodiments, the model may be any other type of model other than a neural network model.
[0234]
[0279] In some embodiments, the neural network model may be trained to classify poses based on context and / or determine whether the driver is engaged in a driving task. For example, if the driver is holding a cell phone and their head is posed downwards towards the phone, the neural network model may determine that the driver is not engaged in a driving task (e.g., not looking at the road or the environment in front of the vehicle) without needing to detect the driver's eyes.
[0235]
[0280] In some embodiments, deep learning or artificial intelligence can be used to develop a model that identifies a driver's posture and / or determines whether a driver is engaged in a driving task. Such a model can distinguish between drivers who are engaged in a driving task and those who are not.
[0236]
[0281] In some embodiments, the model used by the processing unit 210 to identify the driver's posture may be a convolutional neural network model. In other embodiments, the model may simply be any mathematical model.
[0237]
[0282] Figure 5 shows an algorithm 500 for determining whether a driver is engaged in a driving task. For example, the algorithm 500 can be used to determine whether a driver is paying attention to the road and environment ahead of the vehicle. In some embodiments, the algorithm 500 may be implemented and / or executed using a processing unit 210.
[0238]
[0283] First, the processing unit 210 processes the image from the camera 204 and attempts to detect the driver's face based on the image (item 502). If the driver's face cannot be detected from the image, the processing unit 210 may determine that it is unclear whether the driver is engaged in the driving task. On the other hand, if the processing unit 210 determines that the driver's face is present in the image, it can determine whether the driver's eyes are closed or not (item 504). In one embodiment, the processing unit 210 may be configured to determine how the eyes are looking based on a model such as a neural network model. If the processing unit 210 determines that the driver's eyes are closed, the processing unit 210 can determine that the driver is not engaged in the driving task. On the other hand, if the processing unit 210 determines that the driver's eyes are not closed, the processing unit 210 can attempt to detect the driver's line of sight based on the image (item 506).
[0239]
[0284] Referring to item 510 of algorithm 500, if the processing unit 210 successfully detects the driver's line of sight, the processing unit 210 may then determine the direction of the line of sight (item 510). For example, the processing unit 210 may analyze the image to determine the pitch (e.g., vertical direction) and / or yaw (e.g., horizontal direction) of the driver's line of sight. If the pitch of the line of sight is within a predetermined pitch range and the yaw of the line of sight is within a predetermined yaw range, the processing unit 210 may determine that the user is engaged in a driving task (i.e., the user is looking at the road or environment in front of the vehicle) (item 512). On the other hand, if the pitch of the line of sight is not within a predetermined pitch range, or if the yaw of the line of sight is not within a predetermined yaw range, the processing unit 210 may determine that the user is not engaged in a driving task (item 514).
[0240]
[0285] Referring to item 520 of algorithm 500, if the processing unit 210 is unable to successfully detect the driver's line of sight, the processing unit 210 may then determine whether the driver is engaged in a driving task without requiring identification of the driver's line of sight (item 520). In some embodiments, the processing unit 210 may be configured to make such a determination using a model based on context (for example, based on information captured in the image regarding the driver's state other than the direction of the driver's line of sight). In some embodiments, the model may be a neural network model configured to perform context-based classification to determine whether the driver is engaged in a driving task. In one embodiment, the model is configured to process the image to determine whether the driver belongs to one or more pose classifications. If the driver is determined to belong to one or more pose classifications, the processing unit 210 may determine that the driver is not engaged in a driving task (item 522). If the driver is determined not to belong to one or more pose classifications, the processing unit 210 may determine that the driver is engaged in a driving task or that it is unclear whether the driver is engaged in a driving task (item 524).
[0241]
[0286] In some embodiments, items 502, 504, 506, 510, and 520 above may be repeatedly performed by the processing unit 210 to process multiple images of a sequence provided by the camera 204, thereby performing real-time monitoring of the driver while the driver is operating the vehicle.
[0242]
[0287] It should be noted that algorithm 500 is not limited to the example described, and algorithm 500 implemented using processing unit 210 may have other features and / or variations. For example, in other embodiments, algorithm 500 may not include item 502 (detection of the driver's face). As another example, in other embodiments, algorithm 500 may not include item 504 (detection of eyes closed). Furthermore, in even further embodiments, algorithm 500 may not include item 506 (attempt to detect gaze direction) and / or item 510 (determination of gaze direction).
[0243]
[0288] Furthermore, in some embodiments, even if the processing unit 210 can detect the driver's gaze direction, the processing unit 210 may perform context-based classification to determine whether the driver belongs to one or more poses. In some cases, the processing unit 210 may use pose classification to confirm the driver's gaze direction. Alternatively, the driver's gaze direction may be used by the processing unit 210 to confirm one or more pose classifications of the driver.
[0244]
[0289] As described, in some embodiments, the processing unit 210 is configured to determine, based on images from the camera 204, whether the driver belongs to one or more pose classifications, and based on one or more pose classifications, whether the driver is engaged in a driving task. In some embodiments, the processing unit 210 is configured to determine metric values for each of the multiple pose classifications, and based on one or more metric values, whether the driver is engaged in a driving task. Figure 6 shows an example of a classification output 602 provided by the processing unit 210 based on image 604a. In this example, the classification output 602 includes metric values for each of the different pose classifications, namely the "looking down" classification, the "looking up" classification, the "looking left" classification, the "looking right" classification, the "using a cell phone" classification, the "smoking" classification, the "holding an object" classification, the "eyes closed" classification, the "no face" classification, and the "not wearing a seatbelt" classification. The metric values for these different pose classifications are relatively low (e.g., 0.2 or less), indicating that the driver in image 604a does not match any of these pose classifications. Furthermore, in the illustrated example, the driver's eyes are not closed, allowing the processing unit 210 to determine the driver's gaze direction. The gaze direction is represented by a graphical object superimposed on the driver's nose in the image. The graphical object may include vectors or lines parallel to the gaze direction. Alternatively or additionally, the graphical object may include one or more vectors or one or more lines perpendicular to the gaze direction.
[0245]
[0290] Figure 7 shows another example of the classification output 602 provided by the processing unit 210 based on image 604b. In the example shown, the metric value for the "looking down" pose is relatively high (e.g., higher than 0.6), indicating that the driver is in the "looking down" pose. The metric values for the other poses are relatively low, indicating that the driver in image 604b does not meet these pose classifications.
[0246]
[0291] Figure 8 shows another example of the classification output 602 provided by the processing unit 210 based on image 604c. In the example shown, the metric value for the "facing left" pose is relatively high (e.g., higher than 0.6), indicating that the driver is in the "facing left" pose. The metric values for the other poses are relatively low, indicating that the driver in image 604c does not meet these pose classifications.
[0247]
[0292] In some embodiments, the processing unit 210 is configured to compare metric values with respective thresholds for each pose classification. In such cases, the processing unit 210 is configured to determine that the driver belongs to one of the pose classifications if one of the corresponding metric values satisfies or exceeds one of the corresponding thresholds. For example, the thresholds for different pose classifications may be set to 0.6. In such cases, if any of the metric values for any of the pose classifications exceeds 0.6, the processing unit 210 may determine that the driver has a pose belonging to that pose classification (i.e., a pose with a metric value greater than 0.6). In some embodiments, if any of the metric values for any of the pose classifications exceeds a preset threshold (e.g., 0.6), the processing unit 210 may determine that the driver is not engaged in a driving task. Following the examples above, if the metric value for the "looking down" pose, "looking up" pose, "looking left" pose, "looking right" pose, "using a mobile phone" pose, or "eyes closed" pose is higher than 0.6, the processing unit 210 may determine that the driver is not engaged in a driving task.
[0248]
[0293] In the example above, the same pre-set thresholds are implemented for each of the different pose classifications. In other embodiments, at least two of the thresholds for each of the at least two pose classifications may have different values. Also, in the example above, the metric values for pose classifications range from 0.0 to 1.0, with 1.0 being the highest. In other embodiments, the metric values for pose classifications may have other ranges. Also, in other embodiments, the rule for metric values may be reversed so that a lower metric value indicates that the driver fits a particular pose classification, and a higher metric value indicates that the driver does not fit a particular pose classification.
[0249]
[0294] In some embodiments, the thresholds for different pose classifications may be adjusted in an adjustment procedure so that each pose classification has its own adjusted threshold, enabling the processing unit 210 to determine whether or not the driver's image belongs to a particular pose classification.
[0250]
[0295] In some embodiments, a single model may be used by the processing unit 210 to provide multiple pose classifications. The processing unit 210 may output the multiple pose classifications in parallel or sequentially. In other embodiments, the model may include multiple submodels, each submodel configured to detect poses of a specific classification. For example, there may be a submodel for detecting a face, a submodel for detecting gaze direction, a submodel for detecting an upward-looking pose, a submodel for detecting a downward-looking pose, a submodel for detecting a right-looking pose, a submodel for detecting a left-looking pose, a submodel for detecting a pose using a mobile phone, a submodel for detecting a pose not holding a steering wheel, a submodel for detecting a pose not wearing a seatbelt, a submodel for detecting a pose with eyes closed, and so on.
[0251]
[0296] In the embodiments described above, the threshold for each pose classification is configured to determine whether the driver's image belongs to that pose classification. In other embodiments, the threshold for each pose classification may be configured to allow the processing unit 210 to determine whether the driver is engaged in a driving task. In such cases, if one or more metric values for one or more pose classifications meet or exceed one or more thresholds, the processing unit 210 may determine whether the driver is engaged in a driving task. In some embodiments, pose classifications may belong to the "distracted" class. In such cases, if the criteria for any pose classification are met, the processing unit 210 may determine that the driver is not engaged in a driving task (e.g., the driver is distracted). Examples of pose classifications belonging to "distracted" include the "looking left" pose, the "looking right" pose, the "looking up" pose, the "looking down" pose, and the "holding a cell phone" pose. In other embodiments, pose classifications may belong to the "attention" class. In such cases, if the criteria for any pose classification are met, the processing unit 210 may determine that the driver is engaged in a driving task (e.g., the driver is paying attention to driving). Examples of poses classified as "attention" include the "looking forward" pose and the "hands on the handlebars" pose.
[0252]
[0297] As illustrated in the example above, context-based classification is advantageous because it allows the processing unit 210 to identify drivers who are not engaged in the driving task, even when the driver's line of sight cannot be detected. In some cases, even if the device 200 is mounted at a large angle relative to the vehicle (resulting in the driver appearing at an odd angle and / or position in the camera image), context-based identification allows the processing unit 210 to identify drivers who are not engaged in the driving task. Aftermarket products may have different mounting positions, making it difficult to detect the eyes or line of sight. The features described herein are advantageous because they allow determination of whether a driver is engaged in the driving task, even if the device 200 is mounted in a way that prevents the detection of the driver's eyes or line of sight.
[0253]
[0298] It should be noted that the processing unit 210 is not limited to using a neural network model for pose classification and / or determining whether the driver is engaged in a driving task, and the processing unit 210 may utilize any processing technique, algorithm, or processing architecture for pose classification and / or determining whether the driver is engaged in a driving task. To give a non-limiting example, the processing unit 210 may use equations, regression, classification, neural networks (e.g., convolutional neural networks, deep neural networks), heuristics, selection (e.g., from libraries, graphs, or charts), instance-based methods (e.g., nearest neighbors), correlation methods, regularization methods (e.g., ridge regression), decision trees, Bayesian methods, kernel methods, probability, determinism, or a combination of two or more of these to process images from camera 204 for pose classification and / or determining whether the driver is engaged in a driving task. Pose classification can be a binary classification or binary score (e.g., looking up or not), a score (e.g., continuous or discontinuous), a classification (e.g., high, medium, low), or any other appropriate measure of pose classification.
[0254]
[0299] Note that the processing unit 210 is not limited to detecting a pose indicating that the driver is not engaged in a driving task (e.g., a pose belonging to the "inattentive" class). In other embodiments, the processing unit 210 may be configured to detect a pose indicating that the driver is engaged in a driving task (e.g., a pose belonging to the "attentive" class). In further embodiments, the processing unit 210 may be configured to detect both (1) a pose indicating that the driver is not engaged in a driving task and (2) a pose indicating that the driver is engaged in a driving task.
[0255]
[0300] In one or more embodiments described herein, the processing unit 210 may be further configured to determine a risk of collision based on whether the driver is engaged in a driving task. In some embodiments, the processing unit 210 may be configured to determine the risk of collision based only on whether the driver is engaged in a driving task. For example, the processing unit 210 may determine that the risk of collision is "high" when the driver is not engaged in a driving task and that the risk of collision is "low" when the driver is engaged in a driving task. In other embodiments, the processing unit 210 may be configured to determine the risk of collision based on additional information. For example, the processing unit 210 may be configured to track the time during which the driver is not engaged in a driving task and determine the level of the risk of collision based on the duration of the "not engaged in a driving task" state. As another example, the processing unit 210 may process an image from the first camera 202 to determine whether there is an obstacle (e.g., a vehicle, a pedestrian, etc.) in front of the target vehicle, and based on the detection of such an obstacle, determine the risk of collision in combination with the pose classification.
[0256]
[0301] In the above embodiment, camera images from camera 204 (which observes the environment inside the vehicle) are used to monitor the driver's engagement with the driving task. In other embodiments, camera images from camera 202 (which observes the environment outside the vehicle) may be used in a similar manner. For example, in some embodiments, camera images capturing the environment outside the vehicle may be processed by processing unit 210 to determine whether the vehicle is turning left, going straight, or turning right. Based on the direction of travel of the vehicle, processing unit 210 may adjust one or more thresholds for driver pose classification and / or one or more thresholds for determining whether the driver is engaged in the driving task. For example, if processing unit 210 determines (based on processing images from camera 202) that the vehicle is turning left, processing unit 210 may adjust the threshold for the "looking left" pose classification so that a driver looking left is not classified as not engaged in the driving task. In one embodiment, the threshold for the "looking left" pose classification may be 0.6 for vehicles going straight and 0.9 for vehicles turning left. In such a case, if the processing unit 210 determines (based on processing the image from camera 202) that the vehicle is moving straight and (based on processing the image from camera 204) that the metric value for the "looking left" pose is 0.7, then the processing unit 210 can determine that the driver is not engaged in the driving task (because the metric value of 0.7 exceeds the threshold of 0.6 for a vehicle moving straight). On the other hand, if the processing unit 210 determines (based on processing the image from camera 202) that the vehicle is turning left and (based on processing the image from camera 204) that the metric value for the "looking left" pose is 0.7, then the processing unit 210 can determine that the driver is engaged in the driving task (because the metric value of 0.7 does not exceed the threshold of 0.9 for a vehicle turning left). Therefore, as shown in the example above, a certain pose classification (for example, the "looking left" pose) may belong to the "distracted" class in one situation and to the "attention" class in another.In some embodiments, the processing unit 210 is configured to process an image of the external environment from the camera 202 to obtain an output and adjust one or more thresholds based on the output. As a non-limiting example, the output can be a classification of the driving state, a classification of the external environment, identified features of the environment, the context of the vehicle's operation, and the like.
[0257]
[0302] Detection of drowsiness
[0303] In some embodiments, the processing unit 210 may be configured to process an image (e.g., image 300) from the camera 204 and determine whether the driver is drowsy based on the processing of the image. In some embodiments, the processing unit 210 may also process an image from the camera 204 to determine whether the driver is distracted. In further embodiments, the processing unit 210 may process an image from the camera 202 to determine the risk of collision.
[0258]
[0304] In some embodiments, the driver monitoring module 211 of the processing unit 210 may include a first model and a second model configured to work together to detect driver drowsiness. Figure 9 shows an example of a processing architecture in which the first model 212 and the second model 214 are coupled in series. The first and second models 212 and 214 may be perceived as residing within the processing unit 210 and / or as part of the processing unit 210 (e.g., part of the driver monitoring module 211). Although models 212 and 214 are schematically shown within the processing unit 210, in some embodiments, models 212 and 214 may be stored in a non-temporary medium 230. Even in such cases, models 212 and 214 may still be considered part of the processing unit 210. As shown in this example, a series of images 400a to 400e from the camera 204 are received by the processing unit 210. The first model 212 of the processing unit 210 is configured to process the images 400a to 400e. In some embodiments, the first model 212 is configured to identify one or more poses for each corresponding image among images 400a to 400e. For example, the first model 212 may analyze image 400a and determine that the driver is in an "eyes open" pose and a "head straight" pose. The first model 212 may analyze image 400b and determine that the driver is in an "eyes closed" pose. The first model 212 may analyze image 400c and determine that the driver is in an "eyes closed" pose. The first model 212 may analyze image 400d and determine that the driver is in an "eyes closed" pose and a "head down" pose. The first model 212 may analyze image 400e and determine that the driver is in an "eyes closed" pose and a "head straight" pose. Only five images 400a to 400e are shown, but in other examples, the sequence of images received by the first model 212 may be five or more. In some embodiments, the camera 202 has a frame rate of at least 10 frames per second (e.g., 15 fps), and the first model 212 can continue to receive images from the camera 202 at that rate while the driver operates the vehicle.
[0259]
[0305] In some embodiments, the first model may be a single model used by the processing unit 210 to provide multiple pose classifications. The processing unit 210 may output the multiple pose classifications in parallel or sequentially. In other embodiments, the first model may include multiple submodels, each submodel configured to detect poses of a particular classification. For example, there may be a submodel that detects faces, a submodel that detects head-up poses, a submodel that detects head-down poses, a submodel that detects eyes-closed poses, a submodel that detects head-up poses, a submodel that detects eyes-open poses, and so on.
[0260]
[0306] In some embodiments, the first model 212 of the processing unit 210 is configured to identify metric values for each of a plurality of pause classifications. The first model 212 of the processing unit 210 is also configured to compare the metric values with the respective thresholds for each pause classification. In such cases, the processing unit 210 is configured to determine that the driver belongs to one of the pause classifications if one of the corresponding metric values satisfies or exceeds one of the corresponding thresholds. For example, the thresholds for different pause classifications may be set to 0.6. In such cases, if any of the metric values for any of the pause classifications exceeds 0.6, the processing unit 210 may determine that the driver has a pause that belongs to a pause classification (i.e., a pause with a metric value greater than 0.6).
[0261]
[0307] In the example above, the same pre-set thresholds are implemented for each of the different pose classifications. In other embodiments, at least two of the thresholds for each of the at least two pose classifications may have different values. Also, in the example above, the metric values for pose classifications range from 0.0 to 1.0, with 1.0 being the highest. In other embodiments, the metric values for pose classifications may have other ranges. Also, in other embodiments, the rule for metric values may be reversed so that a lower metric value indicates that the driver fits a particular pose classification, and a higher metric value indicates that the driver does not fit a particular pose classification.
[0262]
[0308] As described, in some embodiments, the first model 212 is configured to process images of the driver from camera 204 and determine whether the driver belongs to a particular pose classification. This pose classification belongs to the “drowsy” class, where each pose classification may indicate signs of drowsiness. In non-limiting examples, a pose classification in the “drowsy” class could be one or more of the following: a head-down pose, an eyes-closed pose, or any other pose that helps determine whether the driver is feeling drowsy. Alternatively or additionally, a pose classification belongs to the “awake” class, where each pose classification may indicate signs of wakefulness. In non-limiting examples, a pose classification could be one or more of the following: a pose using a mobile phone, or any other pose that helps determine whether the driver is feeling drowsy. In some embodiments, a particular pose may belong to both the “drowsy” and “awake” classes. For example, a head-up pose and an eyes-open pose may belong to either class.
[0263]
[0309] As shown in the figure, pose recognition (or classification) can be output as feature information by the first model 212. The second model 214 takes the feature information as input from the first model 212, processes the feature information, and determines whether or not the driver is feeling drowsy. The second model 214 also generates an output indicating whether or not the driver is feeling drowsy.
[0264]
[0310] In some embodiments, the feature information output by the first model 212 may be time-series data. The time-series data may be the pose classification of the driver in different images 400 at different time points. In particular, as images are generated sequentially one by one by the camera 204, the first model 212 processes the images sequentially one by one to identify the pose in each image. Once the pose classification of each image is identified by the first model 212, the identified pose classification of that image is output as feature information by the first model 212. Thus, as images are received one by one by the first model 212, the feature information of each image is also output sequentially one by one by the first model 212.
[0265]
[0311] Figure 10 shows an example of feature information received by the second model 214. As shown in the figure, the feature information includes a pose classification for each consecutive image, where "O" indicates that the driver is in an "eyes open" pose and "C" indicates that the driver is in an "eyes closed" pose. Once the second model 214 has acquired a series of feature information, it analyzes the feature information to determine whether the driver is drowsy or not. In one embodiment, the second model 214 may be configured (e.g., programmed, created, trained, etc.) to analyze patterns in the feature information and determine whether they are patterns associated with drowsiness (e.g., patterns indicating drowsiness). For example, the second model 214 may be configured to determine, based on a time series of feature information, any of the following metrics: blink frequency, eye-closed time, time taken to close the eyelids, PERCLOS (percentage of time with eyes closed), or other metrics that measure or indicate alertness or drowsiness.
[0266]
[0312] In some embodiments, if the blinking frequency exceeds a blinking frequency threshold associated with drowsiness, the processing unit 210 may determine that the driver is feeling drowsy.
[0267]
[0313] Alternatively or additionally, if the eye-closed time exceeds the eye-closed time threshold associated with drowsiness, the processing unit 210 may determine that the driver is feeling drowsy. A drowsy person may have a longer eye-closed time than a awake person.
[0268]
[0314] Alternatively or additionally, if the time it takes to close the eyelids exceeds a time threshold associated with drowsiness, the processing unit 210 may determine that the driver is feeling drowsy. The time it takes to close the eyelids is the time interval from when the eyes are substantially open (e.g., at least 80% open, at least 90% open, 100% open, etc.) to when the eyelids are substantially closed (e.g., at least 70% closed, at least 80% closed, at least 90% closed, 100% closed, etc.). This is an indicator of the speed at which the eyelids close. People who are drowsy tend to close their eyelids more slowly than people who are awake.
[0269]
[0315] Alternatively or additionally, if PERCLOS (percentage of time with eyes closed) exceeds the PERCLOS threshold associated with drowsiness, the processing unit 210 may determine that the driver is feeling drowsy. PERCLOS is a drowsiness index that indicates the percentage of time per minute that the eyes are closed for 80% or more of the time. PERCLOS represents the percentage of eyelid closure relative to pupil size over time, reflecting slow eyelid closure rather than blinking.
[0270]
[0316] It should be noted that the feature information provided from the first model 212 to the second model 214 is not limited to the pose classification example described in Figure 10, and the feature information used by the second model 214 to detect drowsiness may include other pose classifications. Figure 11 shows another example of feature information received by the second model 214. As shown in the figure, the feature information includes the pose classification for each consecutive image, where "S" indicates that the driver is in a "head up" pose and "D" indicates that the driver is in a "head down" pose. Once the second model 214 has acquired a series of feature information, it analyzes the feature information to determine whether the driver is drowsy. For example, if the classifications of "head up" and "head down" poses are repeated in a certain pattern associated with drowsiness, the processing unit may determine that the driver is feeling drowsy. In one embodiment, the second model 214 may be configured (e.g., by programming, creating, training, etc.) to analyze a pattern of feature information and determine whether or not it is a pattern associated with drowsiness (e.g., a pattern indicating drowsiness).
[0271]
[0317] In some embodiments, the feature information provided to the first model 212 through the second model 214 may have a data structure that allows different pose classifications to be associated with different time points. In some embodiments, such a data structure can also allow one or more pose classifications to be associated with a specific time point.
[0272]
[0318] In some embodiments, the output of the first model 212 may be a numerical vector (e.g., a low-dimensional numerical vector such as an embedding) that provides a numerical representation of the pose detected by the first model 212. The numerical vector may not be interpretable to humans, but it may provide information about the detected pose. In other embodiments, the output of the first model 212 may be any information that indicates, represents, or relates to an external scene, IMU signals, audio signals, etc. Also, in some embodiments, the embedding may represent any high-dimensional signals such as image signals, IMU signals, audio signals, etc.
[0273]
[0319] In some embodiments, the first model 212 may be a neural network model. In such cases, the neural network model may be trained based on images of other drivers. For example, the neural network model may be trained using images of drivers to identify different poses such as head down, head up, head straight, eyes closed, eyes open, and using a mobile phone. In other embodiments, the first model 212 may be a different type of model than a neural network model.
[0274]
[0320] Furthermore, in some embodiments, the second model 214 may be a neural network model. In such cases, the neural network model may be trained based on feature information. For example, the feature information may be any information indicating the driver's state, such as pose classification. In one implementation example, the neural network model may be trained using feature information output by the first model 212. In other embodiments, the second model 214 may be a model of a different type than a neural network model.
[0275]
[0321] In some embodiments, the first model 212 used by the processing unit 210 to identify the driver's posture may be a convolutional neural network model. In other embodiments, the first model 212 may simply be any mathematical model. Also, in some embodiments, the second model 214 used by the processing unit 210 to determine whether the driver is drowsy may be a convolutional neural network model. In other embodiments, the second model 214 may simply be any mathematical model.
[0276]
[0322] In some embodiments, the first model 212 may be a first neural network model trained to classify poses based on context. For example, if the driver's head is tilted downwards, the neural network model can determine that the driver is not looking ahead, even if the driver's eyes cannot be detected (e.g., if their eyes are obscured by a hat or cap). Also in some embodiments, the second model 214 may be a second neural network model trained to determine whether the driver is drowsy based on context. For example, if the blinking frequency exceeds a certain threshold, and / or if the driver alternates between a head-down pose and a head-up pose in a periodic pattern, the neural network model may determine that the driver is drowsy. As another example, if the time taken to close the eyelids exceeds a certain threshold, the neural network model may determine that the driver is drowsy.
[0277]
[0323] In some embodiments, one or more models can be developed using deep learning or artificial intelligence to identify a driver's posture and / or determine whether the driver is drowsy. Such models can distinguish between a drowsy driver and a awake driver.
[0278]
[0324] The processing unit 210 is not limited to using a neural network model to determine pose classification and / or whether the driver is sleepy. It should be noted that the processing unit 210 may utilize any processing technique, algorithm, or processing architecture to determine pose classification and / or whether the driver is sleepy. By way of non-limiting examples, the processing unit 210 processes an image from the camera 204, processes pose classification and / or time-series information, and uses equations, regression, classification, neural networks (e.g., convolutional neural networks, deep neural networks), heuristics, selection (e.g., from libraries, graphs, or charts), instance-based methods (e.g., nearest neighbor method), correlation methods, regularization methods (e.g., ridge regression), decision trees, Bayesian methods, kernel methods, probability, determinism, or a combination of two or more of these to determine whether the driver is sleepy. The pose classification can be a binary classification or binary score (e.g., whether the head is lowered), a score (e.g., continuous or discontinuous), a classification (e.g., high, medium, low), or any other suitable measure of pose classification. Similarly, the classification of sleepiness can be a binary classification or binary score (e.g., whether there is sleepiness), a score (e.g., continuous or discontinuous), a classification (e.g., high, medium, low), or any other suitable measure of sleepiness.
[0279]
[0325] In some embodiments, the determination of whether the driver is sleepy can be achieved at least by analyzing the pattern of the driver's pose classification occurring over a period such as less than 1 second, 1 second, 2 seconds, 5 seconds, 10 seconds, 12 seconds, 15 seconds, 20 seconds, 30 seconds, 1 minute, 2 minutes, 5 minutes, 10 minutes, 15 minutes, 20 minutes, 25 minutes, 30 minutes, 40 minutes, etc. This period can be any predetermined duration of a moving window or moving box (to identify the data generated in the last duration, e.g., data such as less than 1 second, 1 second, 2 seconds, 5 seconds, 10 seconds, 12 seconds, 15 seconds, 20 seconds, 30 seconds, 1 minute, 2 minutes, 5 minutes, 10 minutes, 15 minutes, 20 minutes, 25 minutes, 30 minutes, 40 minutes, etc. in the last).
[0280]
[0326] In some embodiments, the first model 212 and the second model 214 may be configured to work together to detect “microsleep” events, such as slow eyelid closures occurring over periods of less than 1 second, 1 to 1.5 seconds, or more than 2 seconds. In other embodiments, the first model 212 and the second model 214 may be configured to work together to detect early signs of drowsiness based on images taken over longer periods, such as 10 seconds, 12 seconds, 15 seconds, 20 seconds, 30 seconds, 1 minute, 2 minutes, 5 minutes, 10 minutes, 15 minutes, 20 minutes, 25 minutes, 30 minutes, and 40 minutes.
[0281]
[0327] As illustrated in the example above, using multiple sequential models to detect drowsiness is advantageous. In particular, the technique of combining (1) a first model that processes camera images one by one (as each camera image is generated) to identify the driver's posture, and (2) a second model that processes feature information obtained from the results of the camera image processing by the first model, eliminates the need for the processing unit 210 to collect and process a series of camera images (videos) in batches. This significantly saves computational resources and memory capacity. Furthermore, as explained in the example above, the second model does not process images from the camera. The second model receives feature information as output from the first model and processes that feature information to determine whether the driver is feeling drowsy. This is advantageous because processing feature information is simpler and faster than batch processing of camera images. Context-based classification is also advantageous because it allows the processing unit 210 to accurately identify various driver postures. In some cases, even if the device 200 is mounted at a significantly misaligned angle relative to the vehicle (which may result in the driver appearing at an odd angle and / or position in the camera image), context-based identification allows the processing unit 210 to accurately identify the driver's posture. Aftermarket products can be mounted in a variety of positions. The features described herein are also advantageous in that they allow for the determination of whether the driver is feeling drowsy, even when the device 200 is mounted at different angles.
[0282]
[0328] It should be noted that the processing unit 210 is not limited to detecting poses that indicate the driver is drowsy (e.g., poses belonging to the "drowsy" class). In other embodiments, the processing unit 210 may be configured to detect poses that indicate the driver is alert (e.g., poses belonging to the "alert" class). In further embodiments, the processing unit 210 may be configured to detect both (1) poses that indicate the driver is drowsy and (2) poses that indicate the driver is alert.
[0283]
[0329] In some embodiments, the processing unit 210 may acquire additional parameters (e.g., by receiving or determining) for determining whether the driver is drowsy. In non-limiting examples, the processing unit 210 may be configured to acquire information such as vehicle acceleration, vehicle deceleration, vehicle position relative to the driving lane, and driver involvement in driving. In some cases, one or more of the above parameters may be acquired by a second model 214, which, along with the output from the first model 212, may determine whether the driver is drowsy based on such parameters. Note that information on acceleration, deceleration, and driver involvement serves as an indicator of whether the driver is actively driving. For example, if the driver is changing speed or turning the steering wheel, the driver is less likely to become drowsy. In some embodiments, sensors built into the vehicle may provide acceleration and deceleration information. In such cases, the processing unit 210 may be wired to the vehicle system to receive such information. Alternatively, the processing unit 210 may be configured to receive such information wirelessly. In further embodiments, the apparatus 200 including the processing unit 210 may optionally further include an accelerometer for detecting acceleration and deceleration. In such cases, the second model 214 may be configured to acquire acceleration information and / or deceleration information from the accelerometer. The information regarding driver involvement may also be any information indicating whether or not the driver is operating the vehicle. In non-limiting examples, such information may include one or more of the following: turning or ceasing the steering wheel, operating or ceasing the turn signal lever, changing gears or ceasing the gears, operating or ceasing the brakes, pressing or ceasing the accelerator pedal. In some embodiments, the information regarding driver involvement may be information regarding driver involvement that occurred within a specific past period (e.g., the last 10 seconds or more, the last 20 seconds or more, the last 30 seconds or more, the last 1 minute or more, etc.).
[0284]
[0330] Furthermore, in some embodiments, the vehicle's position relative to the driving lane can be determined by the processing unit 210 processing images from an outward-facing camera 202. In particular, the processing unit 210 may be configured to determine whether the vehicle is traveling within a certain threshold from the lane's centerline. If the vehicle is traveling within a certain threshold from the lane's centerline, the driver is actively participating in driving. On the other hand, if the vehicle is beyond the threshold and far from the lane's centerline, the driver may not be actively participating in driving. In some embodiments, a second model 214 may be configured to receive images from the first camera 202 and determine whether the vehicle is traveling within a certain threshold from the lane's centerline. In other embodiments, another module may be configured to provide this function. In such cases, the output of the module is input to the second model 214 so that the model 214 can determine, based on the module's output, whether the driver is feeling drowsy.
[0285]
[0331] In addition, in one or more embodiments described herein, the processing unit 210 may be further configured to determine the risk of collision based on whether or not the driver is drowsy. In some embodiments, the processing unit 210 may be configured to determine the risk of collision based solely on whether or not the driver is drowsy. For example, the processing unit 210 may determine that the risk of collision is "high" if the driver is drowsy, and that the risk of collision is "low" if the driver is not drowsy (e.g., alert). In other embodiments, the processing unit 210 may be configured to determine the risk of collision based on additional information. For example, the processing unit 210 may be configured to track how long the driver has felt drowsy and determine the level of collision risk based on the duration of drowsiness.
[0286]
[0332] As another example, the processing unit 210 may process images from the first camera 202 to determine an output, and determine the risk of collision based on such output and a combination of pose classification and / or drowsiness detection. In non-limiting examples, the output may be a classification of the driving state, a classification of the external environment, identified characteristics of the environment, or the context of the vehicle's operation. For example, in some embodiments, the processing unit 210 may process camera images capturing the vehicle's external environment to determine whether the vehicle is turning left, going straight, or turning right, or whether there is an obstacle (vehicle, pedestrian, etc.) in front of the vehicle. If drowsiness is detected and the vehicle is turning left or right, and / or an obstacle is detected in the vehicle's path, the processing unit 210 may determine that the risk of collision is high.
[0287]
[0333] The second model 214 of the processing unit 210 is not limited to receiving only the output from the first model 212. The second model 214 may be configured to receive other information (as input) in addition to the output from the first model 212. For example, in other embodiments, the second model 214 may be configured to receive sensor signals from one or more sensors mounted on the vehicle, which may be configured to sense information about the vehicle's motion and / or operating characteristics. As a non-limiting example, the sensor signals acquired by the second model 214 may be acceleration signals, gyroscope signals, velocity signals, position signals (e.g., GPS signals), or any combination thereof. In further embodiments, the processing unit 210 may include a processing module that processes the sensor signals. In such cases, the second model 214 may be configured to receive the processed sensor signals from the processing module. In some embodiments, the second model 214 may be configured to process the sensor signals (provided by the sensors) or the processed sensor signals (provided by the processing module) to determine the risk of collision. The determination of collision risk may be based on drowsiness detection and sensor signals. In other embodiments, the determination of collision risk may be based on drowsiness detection, sensor signals, and images of the surrounding environment outside the vehicle captured by camera 202.
[0288]
[0334] In some embodiments, the processing unit 210 may also include a facial landmark detection module configured to detect one or more facial landmarks of the driver captured in the image from the camera 204. In such cases, the second model 214 may be configured to receive the output from the facial landmark detection module. In some cases, the output from the facial landmark detection module may be used by the second model 214 to determine drowsiness and / or alertness. Alternatively or additionally, the output from the facial landmark detection module may be used to train the second model 214.
[0289]
[0335] In some embodiments, the processing unit 210 may also include an eye landmark detection module configured to detect one or more eye landmarks of the driver captured in the image of the camera 204. In such cases, the second model 214 may be configured to receive the output from the eye landmark detection module. In some cases, the output from the eye landmark detection module may be used by the second model 214 to determine drowsiness and / or alertness. Alternatively or additionally, the output from the eye landmark detection module may be used to train the second model 214. Eye landmarks could be the pupil, eyeball, eyelid, or any other feature related to the driver's eyes.
[0290]
[0336] In some embodiments, if the second model 214 is configured to receive one or more other pieces of information in addition to the output from the first model 212, the second model 214 may be configured to receive one or more pieces of information and the output from the first model 212 in parallel. This allows the second model 214 to receive different pieces of information independently and / or simultaneously.
[0291]
[0337] Figure 12 shows a method 650 performed by the apparatus 200 of Figure 2A, according to several embodiments. The method 650 includes the steps of generating an image of the vehicle driver by a camera (item 652), processing the image by a first model of a processing unit to obtain feature information (item 654), providing the feature information by the first model (item 656), obtaining the feature information from the first model by a second model (item 658), and processing the feature information by the second model to obtain an output indicating whether or not the driver is drowsy (item 660).
[0292]
[0338] It should be noted that the poses that can be identified by the driver monitoring module 211 are not limited to the examples described, and the driver monitoring module 211 may identify other poses or behaviors of the driver. As a non-limiting example, the driver monitoring module 211 may be configured to detect the driver talking, singing, eating, daydreaming, or any combination thereof. Detecting cognitive distraction (such as talking) is beneficial because even when the driver is looking at the road, if the driver is cognitively distracted (compared to when they are paying attention to driving), the risk of intersection violations and collisions may be higher.
[0293]
[0339] Collision prediction
[0340] Figure 13 shows an example of a processing architecture 670 according to several embodiments. At least a portion of the processing architecture 670 can be implemented in some embodiments using the apparatus of Figure 2A. The processing architecture 670 includes a calibration module 671 configured to determine a region of interest for detecting an object in an image that is at risk of collision with a target vehicle, a vehicle detection unit 672 configured to detect a vehicle, and a vehicle state module 674 configured to acquire information about one or more states of the target vehicle. The processing architecture 670 also includes a collision prediction unit 675 having a tracking unit 676 and a time to collision (TTC) calculation unit 680. The processing architecture 670 further includes a driver monitoring module 678 configured to determine whether the driver of the target vehicle is distracted. The processing architecture 670 also includes an event trigger module 682 configured to generate a control signal 684 in response to the detection of a specific event based on the output provided by the collision prediction unit 675, and a context event module 686 configured to provide a context alert 688 based on the output provided by the driver monitoring module 678.
[0294]
[0341] In some embodiments, the vehicle detection unit 672 may be implemented by the object detection unit 216 and / or can be considered as an example of the object detection unit 216. The collision prediction unit 675 may, in some embodiments, be an example of the collision prediction unit 218 of the processing unit 210. The driver monitoring module 678 may, in some embodiments, be implemented by the driver monitoring module 211 of the processing unit 210. The event trigger module 682 may be implemented using the signal generation controller 224 of the processing unit 210 and / or can be considered as an example of the signal generation controller 224.
[0295]
[0342] During operation, the calibration module 671 is configured to determine the region of interest of the first camera 202 in order to detect vehicles that are at risk of colliding with the target vehicle. The calibration module 671 is further described with reference to Figures 15A to 15C. The vehicle detection unit 672 is configured to identify vehicles in the camera image provided by the first camera 202. In some embodiments, the vehicle detection unit 672 is configured to detect vehicles in the image based on a model, such as a neural network model, that has been trained to identify vehicles.
[0296]
[0343] The driver monitoring module 678 is configured to determine whether the driver of the vehicle in question is distracted. In some embodiments, the driver monitoring module 678 can identify one or more poses of the driver based on images provided by the second camera 204. Based on the driver's poses, the driver monitoring module 678 can determine whether the driver is distracted. In some cases, the driver monitoring module 678 may determine one or more poses of the driver based on a model, such as a neural network model, that has been trained to identify driver poses.
[0297]
[0344] The collision prediction unit 675 is configured to select one or more vehicles detected by the vehicle detection unit 672 as possible candidates for collision prediction. In some embodiments, the collision prediction unit 675 is configured to select a vehicle for collision prediction if its image intersects with a region of interest (determined by the calibration module 671) within the image frame. The collision prediction unit 675 is also configured to track the state of the selected vehicle (by the tracking unit 676). In non-limiting examples, the state of the selected vehicle being tracked may be the vehicle's position, vehicle speed, acceleration or deceleration, direction of movement, or any combination thereof. In some embodiments, the tracking unit 676 may be configured to determine whether the detected vehicle is on a course to collide with the target vehicle, based on the target vehicle's travel path and / or the detected vehicle's travel path. Also in some embodiments, the tracking unit 676 may be configured to determine that a vehicle is a preceding vehicle if its image appearing in the image frame from the first camera 202 intersects with a region of interest within the image frame.
[0298]
[0345] The TTC unit 680 of the collision prediction unit 675 is configured to calculate an estimated time for a predicted collision to occur, based on the tracking status of the selected vehicle and the status of the target vehicle (provided by the vehicle status module 674). For example, if the tracking status of the selected vehicle indicates that the vehicle is in the path of the target vehicle and is traveling at a slower speed than the target vehicle, the TTC unit 680 determines the estimated time until the selected vehicle collides with the target vehicle. As another example, if the tracking status of the selected vehicle indicates that the vehicle is a preceding vehicle ahead of the target vehicle, the TTC unit 680 determines the estimated time until the selected vehicle collides with the target vehicle. In some embodiments, the TTC unit 680 may determine the estimated time to the predicted collision based on the relative speed between the two vehicles and / or the distance between the two vehicles. The TTC unit 680 is configured to output the estimated time (TTC parameter).
[0299]
[0346] It should be noted that the collision prediction unit 675 is not limited to predicting collisions between a preceding vehicle and a target vehicle, and may be configured to predict other types of collisions. For example, in some embodiments, the collision prediction unit 675 may be configured to predict a collision between a target vehicle traveling on two different roads (e.g., intersecting roads) and another vehicle heading towards an intersection. As another example, in some embodiments, the collision prediction unit 675 may be configured to predict a collision between a target vehicle and another vehicle traveling in an adjacent lane merging into or entering the lane of the target vehicle.
[0300]
[0347] The event trigger module 682 is configured to provide a control signal based on the output provided by the collision prediction unit 675 and the output provided by the driver monitoring module 678. In some embodiments, the event trigger module 682 is configured to continuously or periodically monitor the driver's condition based on the output provided by the driver monitoring module 678. The event trigger module 682 also monitors the TTC parameter in parallel. If the TTC parameter indicates that the estimated time until a predicted collision occurs is below a certain threshold (e.g., 8 seconds, 7 seconds, 6 seconds, 5 seconds, 4 seconds, 3 seconds, etc.) and the output from the driver monitoring module 678 indicates that the driver is distracted or not paying attention to the driving task, the event trigger module 682 generates a control signal 682.
[0301]
[0348] In some embodiments, the control signal 684 from the event trigger module 682 may be transmitted to a warning generator configured to provide a warning to the driver. Alternatively or additionally, the control signal 684 from the event trigger module 682 may be transmitted to a vehicle control device configured to control the vehicle (e.g., automatically release the accelerator pedal, apply the brakes, etc.).
[0302]
[0349] In some embodiments, the threshold is variable based on the output from the driver monitoring module 678. For example, if the output from the driver monitoring module 678 indicates that the driver is not distracted and / or is paying attention to the driving task, the event trigger module 682 may generate a control signal 684 to activate a warning generator and / or a vehicle control device in response to the TTC meeting or falling below a first threshold (e.g., 3 seconds). On the other hand, if the output from the driver monitoring module 678 indicates that the driver is distracted or not paying attention to the driving task, the event trigger module 682 may generate a control signal 684 to activate a warning generator and / or a vehicle control device in response to the TTC meeting or falling below a second threshold (e.g., 5 seconds) that is higher than the first threshold.
[0303]
[0350] In some embodiments, the event trigger module 682 may also be configured to apply different threshold values for generating the control signal 684 based on the type of driver state indicated by the output of the driver monitoring module 678. For example, if the output of the driver monitoring module 678 indicates that the driver is looking at a mobile phone, the event trigger module 682 may generate a control signal 684 that activates a warning generator and / or a vehicle control device in response to the TTC meeting or falling below a 5-second threshold. On the other hand, if the output of the driver monitoring module 678 indicates that the driver is drowsy, the event trigger module 682 may generate a control signal 684 that activates a warning generator and / or a vehicle control device in response to the TTC meeting or falling below an 8-second threshold (for example, longer than the threshold when the driver is using a mobile phone). In some cases, depending on the driver's condition (e.g., the driver is drowsy or falling asleep at the wheel), it may take time for the driver to react to an imminent collision, and a longer time threshold (compared to the TTC value) may be required to alert the driver and take control of the vehicle. Therefore, the event trigger module 682 may warn the driver and / or activate the vehicle control system earlier in response to a collision expected in such situations.
[0304]
[0351] In some embodiments, the TTC unit 680 is configured to determine the TTC value of a predicted collision and then track the passage of time with respect to the TTC value. For example, if the TTC unit 680 determines that the TTC of a predicted collision is 10 seconds, the TTC unit 680 may perform a 10-second time countdown. While the TTC unit 680 is performing the countdown, it periodically outputs the TTC to inform the event trigger module 682 of the current TTC value. Thus, for each predicted collision, the TTC output by the TTC unit 680 at different times will be different values based on the countdown. In other embodiments, the TTC unit 680 is configured to repeatedly determine the TTC value of a predicted collision based on images from the first camera 202. In such a case, for each predicted collision, the TTC output by the TTC unit 680 at different times will be different values calculated by the TTC unit 680 based on images from the first camera 202.
[0305]
[0352] In some embodiments, the collision prediction unit 675 may continue to monitor the status of other vehicles and / or the target vehicle even after a collision has been predicted. For example, if the other vehicle deviates from the target vehicle's path and / or the distance between the two vehicles increases (e.g., because the other vehicle accelerates and / or the target vehicle decelerates), the collision prediction unit 675 may provide an output indicating that the risk of collision has been eliminated. In some embodiments, the TTC unit 680 may output a signal to the event trigger module 682 indicating that it does not need to generate a control signal 684. In other embodiments, the TTC unit 680 may output a predetermined arbitrary TTC value that is very high (e.g., 2000 seconds) or a negative TTC value so that the event trigger module 682 does not generate a control signal 684 when processing the TTC value.
[0306]
[0353] Embodiments of the collision prediction unit 675 (an example of the collision prediction unit 218) and the event trigger module 682 (an example of the signal generation controller 224) will be described further later.
[0307]
[0354] The context event module 686 is configured to provide context alerts 688 based on the output provided by the driver monitoring module 678. For example, if the output of the driver monitoring module 678 indicates that the driver is distracted for a duration exceeding a duration threshold or at a frequency exceeding a frequency threshold, the context event module 686 can generate an alert to warn the driver. Alternatively or additionally, the context event module 686 may generate messages to notify vehicle managers, insurance companies, etc. In other embodiments, the context event module 686 is optional, and the processing architecture 670 does not have to include the context event module 686.
[0308]
[0355] In other embodiments, item 672 may be a human detection unit, and the processing architecture 670 may be configured to predict a collision with a human and generate a control signal based on the predicted collision and the driver status output by the driver monitoring module 678, as similarly described herein.
[0309]
[0356] In a further embodiment, item 672 may be an object detection unit configured to detect objects related to an intersection, and the processing architecture 670 may be configured to predict an intersection violation and generate a control signal based on the predicted intersection violation and the driver status output by the driver monitoring module 678, as similarly described herein.
[0310]
[0357] In yet another embodiment, item 672 may be an object detection unit configured to detect multiple classes of objects, such as vehicles, people, and objects related to intersections. In such a case, the processing architecture 670 may predict vehicle collisions, pedestrian collisions, intersection violations, etc., and may be configured to generate control signals based on any of these predicted events, and based on the driver status output by the driver monitoring module 678, as similarly described herein.
[0311]
[0358] Figure 14 shows examples of object detection according to several embodiments. As shown, the object to be detected is a vehicle captured in an image provided by the first camera 202. Object detection may be performed by the object detection unit 216. In the illustrated example, each identified vehicle is assigned an identifier (e.g., the shape of a bounding box indicating the spatial extent of each identified vehicle). It should be noted that the object detection unit 216 is not limited to providing an identifier that is a rectangular bounding box for each identified vehicle, and the object detection unit 216 may be configured to provide other forms of identifiers for each identified vehicle. In some embodiments, the object detection unit 216 may distinguish a vehicle that is a preceding vehicle from other vehicles that are not a preceding vehicle.
[0312]
[0359] In some embodiments, the processing unit 210 may also track identified preceding vehicles and determine a region of interest based on the spatial distribution of such identified preceding vehicles. For example, as shown in Figure 15A, the processing unit 210 may use identifiers 750 (in the form of bounding boxes in some embodiments) of preceding vehicles identified over a certain period (e.g., the last 5 seconds, the last 10 seconds, the last 1 minute, the last 2 minutes, etc.) to form a region of interest based on the spatial distribution of identifiers 750. In the illustrated embodiment, the region of interest has certain dimensions and location (this location is at the bottom of the image frame, approximately in the horizontal center).
[0313]
[0360] In some embodiments, instead of using a bounding box, the processing unit 2 may use horizontal lines 752 to form a region of interest (Figure 15B). In the illustrated example, each horizontal line 752 represents a preceding vehicle identified over a period of time. The horizontal lines 752 can be thought of as an example of an identifier for an identified preceding vehicle. In some cases, the horizontal lines 752 can be obtained by extracting only the base of a bounding box (e.g., 750 shown in Figure 15A). As shown in Figure 15B, the distribution of horizontal lines 752 forms a region of interest (represented by the area filled with horizontal lines 752) that is roughly triangular or trapezoidal in shape. The processing unit 210 may use such a region of interest as a detection area to detect future preceding vehicles. For example, as shown in Figure 15C, the processing unit 210 may use the identifiers of identified preceding vehicles (e.g., lines 752) to form a region of interest 754, which in this embodiment has a triangular shape. The region of interest 754 may be used by the object detection unit 216 to identify a preceding vehicle. In the example shown in the figure, the object detection unit 216 detects the vehicle 756. Since at least a portion of the detected vehicle 756 is located in the region of interest 754, the object detection unit 216 may determine that the identified vehicle is a preceding vehicle.
[0314]
[0361] In some embodiments, the region of interest 754 may be determined by a calibration module in the processing unit 210 during the calibration process. Also, in some embodiments, the region of interest 754 may be periodically updated during operation of the device 200. Note that the region of interest 754 for detecting a preceding vehicle is not limited to the examples described and may have other configurations (e.g., size, shape, position, etc.) in other embodiments. Also, in other embodiments, the region of interest 754 may be determined using other techniques. For example, in other embodiments, the region of interest 754 for detecting a preceding vehicle may be determined in advance (e.g., programmed during manufacturing) without using a distribution of previously detected preceding vehicles.
[0315]
[0362] In the example above, the region of interest 754 has a triangular shape that can be determined during the calibration process. In other embodiments, the region of interest 754 may have a different shape and may be determined based on the detection of the lane centerline. For example, in other embodiments, the processing unit 210 may include a centerline detection module configured to identify the centerline of the lane or road in which the vehicle in question is traveling. In some embodiments, the centerline detection module may be configured to identify the centerline by processing images from a first camera 202. In one implementation example, the centerline detection module analyzes images from the first camera 202 to identify the centerline of the lane or road based on a model. This model may be a neural network model trained to identify centerlines based on images of various road conditions. Alternatively, it may be another type of model, such as a mathematical model or equation. Figure 15D shows an example of the centerline detection module identifying a centerline and an example of a region of interest based on the detected centerline. As shown in the figure, the centerline detection module determines a set of points 757a to 757e that represent the centerline of the lane or road in which the vehicle in question is traveling. Although five points 757a to 757e are shown, in other examples the centerline detection module may identify five or more points 757 representing the centerline, or fewer than five points 757. Also, as shown in the figure, the processing unit 210 may determine a left set of points 758a to 758e and a right set of points 759a to 759e based on points 757a to 757e. The processing unit 210 may also determine a first set of lines connecting the left points 758a to 758e and a second set of lines connecting the right points 759a to 759e. As shown in the figure, the first set of lines forms the left boundary of the region of interest 754, and the second set of lines forms the right boundary of the region of interest 754.
[0316]
[0363] In the illustrated example, the processing unit 210 is configured to determine the left point 758a as having the same y-coordinate as the center line point 757a, and an x-coordinate located at a distance d1 to the left of the x-coordinate of the center line point 757a. Similarly, the processing unit 210 is configured to determine the right point 759a as having the same y-coordinate as the center line point 757a, and an x-coordinate located at a distance d1 to the right of the x-coordinate of the center line point 757a. Thus, the left point 758a, the center line point 757a, and the right point 759a are aligned horizontally. Likewise, the processing unit 210 is configured to determine the left points 758b to 758e as having the same y-coordinates as their respective center line points 757b to 757e, and their respective x-coordinates located at distances d2 to d5 to the left of their respective x-coordinates. Furthermore, the processing unit 210 is configured to determine that the rightmost points 759b to 759e have the same y coordinates as the respective centerline points 757b to 757e, and that their x coordinates are located to the right of the respective x coordinates of the centerline points 757b to 757e at distances d2 to d5.
[0317]
[0364] In the illustrated example, d1>d2>d3>d4>d5, and the region of interest 754 takes on a tapered shape corresponding to the shape of the road as it appears in the camera image. As the first camera 202 repeatedly provides camera images of the road as the vehicle is traveling, the processing unit 210 repeatedly determines the centerline and the left and right boundaries of the region of interest 754 based on the centerline. In this way, the tapered shape of the region of interest 754 changes in response to changes in the shape of the road as it appears in the camera image (for example, the curvature of the tapered region of interest 754 changes). That is, since the centerline is identified based on the shape of the road, and the shape of the region of interest 754 is determined based on the identified centerline, the shape of the region of interest 754 is variable in response to the shape of the road on which the vehicle is traveling.
[0318]
[0365] Figure 15E illustrates the advantages of using the region of interest 754 in Figure 15D for detecting objects at risk of collision. In particular, the right side of the figure shows the region of interest 754 determined based on the centerline of the road or lane on which the vehicle in question is traveling, as described with reference to Figure 15D. The left side of the figure shows another region of interest 754, which is determined based on camera calibration as described with reference to Figures 15A-15C and has a shape that does not depend on the centerline of the road / lane (e.g., the curvature of the centerline). Because the region of interest 754 on the left does not depend on the curvature of the road / lane, the shape of the region of interest 754 does not necessarily coincide with the shape of the road / lane. Therefore, in the illustrated example, the processing unit 210 may incorrectly detect a pedestrian as being at risk of collision because it intersects with the region of interest 754. In another similar situation, the region of interest 754 in the left figure may incorrectly detect a parked vehicle outside the lane in question as an object that poses a collision risk. In some embodiments, the processing unit 210 may be configured to perform additional processing to address the problem of false detections (e.g., incorrectly detecting an object as being at risk of collision). On the other hand, region 754 on the right side is advantageous because it does not suffer from the false detection problem mentioned above.
[0319]
[0366] In other embodiments, other techniques may be employed to determine the region of interest 754, which has a shape that changes in accordance with the shape of the road. For example, in other embodiments, the processing unit 210 may include a road or lane boundary module configured to identify the left and right boundaries of the lane or road on which the target vehicle is traveling. Alternatively, the processing unit 210 may determine one or more lines that fit the left boundary and one or more lines that fit the right boundary, and determine the region of interest 754 based on the determined lines.
[0320]
[0367] In some embodiments, the processing unit 210 may be configured to determine both (1) a first region of interest (such as the triangular region of interest 754 described with reference to Figures 15A-15C) and (2) a second region of interest, such as the region of interest 754 described with reference to Figure 15D. The first region of interest may be used by the processing unit 210 to crop the camera image. For example, certain parts of the camera image that are far from the first region of interest, or at a certain distance from the first region of interest, may be cropped to reduce the amount of image data that needs to be processed. The second region of interest may be used by the processing unit 210 to determine whether a detected object poses a collision risk. For example, if a detected object, or the bounding box of a detected object, overlaps with the second region of interest, the processing unit 210 may determine that there is a risk of collision with the detected object. In some embodiments, the first region of interest may also be used by the processing unit 210 to detect a preceding vehicle in the camera image. The detected vehicle width and its corresponding position in the image coordinate system may be used by the processing unit 210 to determine a y-distance mapping, which will be described in more detail below with reference to Figure 20.
[0321]
[0368] In some embodiments, the collision prediction unit 218 may be configured to determine whether the region of interest 754 (e.g., a polygon created based on the centerline) intersects with the boundary box of a detected object, such as a preceding vehicle or pedestrian. If so, the collision prediction unit 218 determines that there is a risk of collision, and the object corresponding to the boundary box is considered for TTC calculation.
[0322]
[0369] In some embodiments, the collision prediction unit 218 may be configured to predict a collision with a preceding vehicle in at least three different scenarios. Figure 16 shows three exemplary scenarios involving a collision with a preceding vehicle. In the top figure (first scenario), the target vehicle (vehicle on the left) is traveling at a non-zero speed Vsv, and the preceding vehicle (vehicle on the right) is completely stopped, so its speed Vpov = 0. In the middle figure (second scenario), the target vehicle (vehicle on the left) is traveling at a non-zero speed Vsv, and the preceding vehicle (vehicle on the right) is traveling at a non-zero speed Vpov that is less than the speed Vsv. In the bottom figure (third scenario), the target vehicle (vehicle on the left) is initially traveling at a non-zero speed Vsv, and the preceding vehicle (vehicle on the right) is also initially traveling at a non-zero speed Vpov = Vsv. Subsequently, the preceding vehicle brakes and its speed Vpov decreases, causing the target vehicle's speed Vsv to become greater than the preceding vehicle's speed Vpov. In some embodiments, the collision prediction unit 218 is configured to predict a collision between the target vehicle and a preceding vehicle that may occur in any of the three scenarios shown in Figure 16. In one embodiment, the collision prediction unit 218 may determine the relative speed between the target vehicle and the preceding vehicle by analyzing an image sequence from the first camera 202. In another embodiment, the collision prediction unit 218 may acquire sensor information indicating the relative speed between the target vehicle and the preceding vehicle. For example, the collision prediction unit 218 may acquire a sequence of sensor information indicating the distance between the target vehicle and the preceding vehicle over a certain period of time. By analyzing the change in the distance between the vehicles during this period, the collision prediction unit 218 can determine the relative speed between the target vehicle and the preceding vehicle. In some embodiments, the collision prediction unit 218 may also acquire the speed of the target vehicle from the target vehicle's speed sensor, a GPS system, or a speed sensor other than the target vehicle's speed sensor.
[0323]
[0370] In some embodiments, the collision prediction unit 218 may be configured to predict a collision between the target vehicle and the preceding vehicle based on the relative speed between the target vehicle and the preceding vehicle, the speed of the target vehicle, the speed of the preceding vehicle, or any combination thereof. For example, the collision prediction unit 218 may determine that there is a risk of collision if (1) the object detection unit 216 detects a preceding vehicle, (2) the relative speed between the preceding vehicle and the target vehicle is not zero, and (3) the distance between the preceding vehicle and the target vehicle is decreasing. In some cases, criteria (2) and (3) may be combined to indicate whether the target vehicle is traveling faster than the preceding vehicle. In such cases, the collision prediction unit 218 may determine that there is a risk of collision if (1) the object detection unit 216 detects a preceding vehicle, and (2) the target vehicle is traveling faster than the preceding vehicle (the target vehicle is moving towards the preceding vehicle).
[0324]
[0371] In some embodiments, the collision prediction unit 218 may acquire other information to use in determining whether there is a risk of collision. As a non-limiting example, the collision prediction unit 218 may acquire information indicating that the preceding vehicle is braking (e.g., camera images, detected light, etc.), operating parameters of the target vehicle (e.g., information indicating acceleration, deceleration, turning, etc.), operating parameters of the preceding vehicle (e.g., information indicating acceleration, deceleration, turning, etc.), or any combination thereof.
[0325]
[0372] In some embodiments, the collision prediction unit 218 is configured to predict a collision at least 3 seconds before the predicted time of occurrence. For example, the collision prediction unit 218 may be configured to predict a collision at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 seconds before the predicted time of occurrence. In some embodiments, the collision prediction unit 218 is also configured to predict a collision with sufficient lead time for the driver's brain to process the input and for the driver to take action to mitigate the risk of collision. In some embodiments, sufficient lead time may depend on the driver's state as identified by the driver monitoring module 211.
[0326]
[0373] In some embodiments, the object detection unit 216 may be configured to detect humans. In such cases, the collision prediction unit 218 may be configured to predict collisions with humans. Figure 17 shows another example of object detection when the object to be detected is a human. As shown in the figure, the object to be detected is a human captured in the image provided by the first camera 202. Object detection can be performed by the object detection unit 216. In the illustrated example, each identified human is assigned an identifier (for example, in the form of a bounding box 760 indicating the spatial extent of each identified vehicle). It should be noted that the object detection unit 216 is not limited to providing an identifier that is a rectangular bounding box 760 for each identified human, and the object detection unit 216 may be configured to provide other forms of identifiers for each identified human. In some embodiments, the object detection unit 216 may distinguish a human who is in front of the target vehicle (for example, in the vehicle's path) from other humans who are not in the vehicle's path. In some embodiments, the same region of interest 754 described above for detecting a preceding vehicle may be used by the object detection unit 216 to detect a person in the path of the target vehicle.
[0327]
[0374] In some embodiments, the collision prediction unit 218 may be configured to determine the direction of movement of a detected person by analyzing a sequence of images of a person provided by the first camera 202. The collision prediction unit 218 may also be configured to determine the speed of movement of a detected person (e.g., walking or running speed) by analyzing a sequence of images of a person. The collision prediction unit 218 may also be configured to determine whether there is a risk of collision with a person based on the target vehicle's travel path and the detected direction of movement of the person. Such functionality is desirable to prevent collisions with people on a sidewalk that are not in the vehicle's path but are moving toward the target vehicle's path.
[0328]
[0375] In some embodiments, the collision prediction unit 218 is configured to determine an area adjacent to the detected person, indicating the person's possible position at a certain point in the future (e.g., the next 0.5 seconds, the next 1 second, the next 2 seconds, the next 3 seconds, etc.), based on the detected person's speed and direction of movement. The collision prediction unit 218 may then determine, based on the vehicle's speed, whether the vehicle will pass through the determined area (e.g., inside the box) indicating the person's predicted position. In one embodiment, the collision prediction unit 218 may determine whether the determined area intersects with the area of interest 754. In that case, the collision prediction unit 218 may determine that there is a risk of collision with the person and generate an output indicating a predicted collision.
[0329]
[0376] Figure 18 shows an example of a predicted human position based on the human's walking speed and direction. Because human movement is inherently somewhat unpredictable, in some embodiments, even if the detected human is standing (e.g., a pedestrian standing beside a roadway), the collision prediction unit 218 may determine a region relating to the human that indicates the human's possible location (e.g., for cases where the human starts walking or running). For example, the collision prediction unit 218 may determine a boundary box 760 (e.g., a rectangular box) surrounding the detected human, and then enlarge the dimensions of this boundary box 760 to account for the uncertainty of the human's future predicted position, where the enlarged box defines a region indicating the human's predicted position. The collision prediction unit 218 may then determine whether the target vehicle passes through the determined region indicating the human's predicted position. In one embodiment, the collision prediction unit 218 may determine whether the determined region of the enlarged box intersects with the region of interest 754. If so, the collision prediction unit 218 may determine that there is a risk of collision with the human and generate an output indicating a predicted collision.
[0330]
[0377] In some embodiments, the collision prediction unit 218 may be configured to predict collisions with a person in at least three different scenarios. In the first scenario, the detected person (or the bounding box 760 surrounding the detected person) intersects with the region of interest 754, indicating that the person is already within the target vehicle's travel path. In the second scenario, the detected person is not within the target vehicle's travel path but is standing beside the roadway. In such a case, the collision prediction unit 218 may use the region of the enlarged bounding box of the detected person, as described above, to determine whether there is a risk of collision. If the enlarged bounding box intersects with the region of interest 754 (for collision detection), the collision prediction unit 218 may determine that there is a risk of collision with the standing person. In the third scenario, the detected person is moving, and the image of the person (or its bounding box 760) does not intersect with the region of interest 754 (for collision detection). In such a case, the collision prediction unit 218 may use the region of the predicted position of the person, as described above, to determine whether there is a risk of collision. If the predicted location area intersects with the region of interest 754, the collision prediction unit 218 may determine that there is a risk of collision with a human.
[0331]
[0378] In some embodiments, the enlarged bounding box has dimensions that are the dimensions of the detected object plus an additional length, which is preset to account for the uncertainty of the object's movement. In other embodiments, the enlarged bounding box may be determined based on a prediction of the object's position. As shown in Figure 18, the detected object may have an initial bounding box 760. Based on the object's position in the image from the first camera 202, the processing unit 210 can predict the position of the moving object. As shown, the processing unit 210 can determine a box 762a representing the object's possible position 0.3 seconds into the future. The processing unit 210 may also determine a box 762b representing the object's possible position 0.7 seconds into the future and a box 762c representing the object's possible position 1 second into the future. In some embodiments, the collision prediction unit 218 may continue to predict the future position of the detected object (e.g., a person) at a specific future time and determine whether the target vehicle's path intersects with any of these positions. If so, the collision prediction unit 218 may determine that there is a risk of collision with the object.
[0332]
[0379] In some embodiments, the region of interest 754 may be expanded in response to the driver monitoring module 211 detecting that the driver is distracted. For example, the region of interest 754 can be widened in response to the driver monitoring module 211 detecting driver distraction. This has the advantage that objects outside the road or lane can be considered as collision risks. For example, if the driver is distracted, the processing unit 210 then widens the region of interest 754. This has the effect of relaxing the threshold for detecting overlap between detected objects and the region of interest 754. If a bicycle is traveling along the edge of the lane, it may overlap with the widened region of interest 754, so the processing unit 210 can detect the bicycle as a collision risk. On the other hand, if the driver is attentive (e.g., not distracted), the region of interest 754 becomes smaller and the bicycle no longer intersects with the region of interest 754. Therefore, in this scenario, the processing unit 210 does not consider the bicycle to be a collision risk, which makes sense because an attentive driver is more likely to avoid a collision with the bicycle.
[0333]
[0380] In some embodiments, to reduce the computational load, the collision prediction unit 218 does not need to determine the risk of collision for every person detected in the image. For example, in some embodiments, the collision prediction unit 218 may exclude people inside vehicles, people standing at bus stops, people sitting outside, etc., from detection. In other embodiments, the collision prediction unit 218 may consider all people detected for collision prediction.
[0334]
[0381] In some embodiments, the object detection unit 216 may use one or more models to detect various objects such as automobiles (illustrated), motorcycles, pedestrians, animals, lane dividers, road signs, traffic signs, and traffic lights. In some embodiments, the models utilized by the object detection unit 216 may be neural network models trained to identify various objects. In other embodiments, the models may be any other type of model, such as mathematical models configured to identify objects. The models utilized by the object detection unit 216 may be stored in a non-temporary medium 230 and / or incorporated as part of the object detection unit 216.
[0335]
[0382] Predicting intersection violations
[0336]
[0383] Figures 19A-19B show other examples of object detection related to intersections, where objects detected by the object detection unit 216 are relevant. As shown in Figure 19A, the object detection unit 216 may be configured to detect traffic signals 780. As shown in Figure 19B, the object detection unit 216 may be configured to detect stop signs 790. The object detection unit 216 may also be configured to detect other items related to intersections, such as road signs, curb corners, and ramps.
[0337]
[0384] In some embodiments, the intersection violation prediction unit 222 is configured to detect an intersection based on detected objects 216 detected by the object detection unit 216. In some cases, the object detection unit 216 may detect a stop line at the intersection indicating the intended stopping position of the vehicle in question. The intersection violation prediction unit 222 can determine the TTC (time-to-crossing) based on the position of the stop line and the speed of the vehicle in question. For example, the intersection violation prediction unit 222 can determine the distance d between the vehicle in question and the position of the stop line and calculate the TTC based on the formula TTC = d / V (where V is the speed of the vehicle in question). In some embodiments, as shown in Figure 19B, the intersection violation prediction unit 222 may also be configured to determine a line 792 corresponding to the detected stop line and perform a calculation to determine the TTC based on the line 792. In some cases, if the object detection unit 216 does not detect a stop line, the intersection violation prediction unit 222 may estimate the expected stopping position based on objects detected at the intersection. For example, the intersection violation prediction unit 222 can estimate the expected stopping position based on the expected stopping position and the known relative positions between the expected stopping position and surrounding objects such as stop signs and traffic lights.
[0338]
[0385] In some embodiments, instead of determining TTC, the intersection violation prediction unit 222 may be configured to determine time-to-brake (TTB) based on the position of the stop line and the speed of the vehicle in question. TTB measures the time remaining at the current speed for the driver to initiate braking to safely stop at or before the required stopping position associated with the intersection. For example, the intersection violation prediction unit 222 can identify the distance d between the vehicle in question and the position of the stop line and calculate TTB based on the current speed of the vehicle in question. In some embodiments, the intersection violation prediction unit 222 may be configured to determine braking distance BD, which indicates the distance required for the vehicle to come to a complete stop based on the vehicle's speed, and to determine TTB based on this braking distance. The faster the driving speed, the longer the braking distance. Also, in some embodiments, the braking distance may be based on road conditions. For example, even at the same vehicle speed, the braking distance may be longer on a wet road surface than on a dry road surface. Figure 19C shows the difference in required braking distances at various vehicle speeds and various road conditions. For example, as shown in the figure, a vehicle traveling at 40 km / h requires a braking distance of 9 m on a dry surface and 13 m on a wet surface. On the other hand, a vehicle traveling at 110 km / h requires a braking distance of 67 m on a dry surface and 97 m on a wet surface. Figure 19C also shows how far a vehicle travels when the driver's reaction time is 1.5 seconds. For example, a vehicle traveling at 40 km / h will travel 17 m in approximately 1.5 seconds (driver's reaction time) before the driver applies the brakes. Therefore, the total distance required for a vehicle traveling at 40 km / h to stop (considering the driver's reaction time) is 26 m on a dry surface and 30 m on a wet surface.
[0339]
[0386] In some embodiments, the intersection violation prediction unit 222 can determine TTB based on the formula: TTB = (d - BD) / V (where V is the vehicle speed). Since d is the distance from the current vehicle position to the stopping position (e.g., stop line) and BD is the braking distance, the term (d - BD) represents the remaining distance the vehicle will travel, during which the driver may react to the environment before applying the brakes. Thus, the term (d - BD) / V represents the time the driver must react to the environment before applying the brakes. In some embodiments, if TTB = (d - BD) / V <= threshold reaction time, the intersection violation prediction unit 222 can generate control signals to activate a device that warns the driver and / or a device that automatically controls the vehicle, as described herein.
[0340]
[0387] In some embodiments, this threshold response time may be 1 second or more, 1.5 seconds or more, 2 seconds or more, 2.5 seconds or more, 3 seconds or more, 4 seconds or more, etc.
[0341]
[0388] Furthermore, in some embodiments, the threshold reaction time may be variable based on the driver's state identified by the driver monitoring module 211. For example, in some embodiments, if the driver monitoring module 211 determines that the driver is distracted, the processing unit 210 may increase the threshold reaction time (for example, from 2 seconds for a non-distracted driver to 4 seconds for a distracted driver). Moreover, in some embodiments, the threshold reaction time may have different values for different driver states. For example, the threshold reaction time may be 4 seconds if the driver is distracted, and 6 seconds if the driver is drowsy.
[0342]
[0389] In some embodiments, the intersection violation prediction unit 222 may be configured to determine the distance d between the vehicle in question and the stopping position by analyzing images from the first camera 202. Alternatively or additionally, the intersection violation prediction unit 222 may receive information from a GPS system indicating the location of the vehicle in question and the location of the intersection. In such a case, the intersection violation prediction unit 222 can determine the distance d based on the location of the vehicle in question and the location of the intersection.
[0343]
[0390] In some embodiments, the intersection violation prediction unit 222 may determine the braking distance BD by referring to a table that associates different vehicle speeds with their respective braking distances. In other embodiments, the intersection violation prediction unit 222 may determine the braking distance BD by performing calculations based on a model (e.g., an equation) that takes vehicle speed as input and outputs braking distance. Also in some embodiments, the processing unit 210 may receive information indicating road conditions and determine the braking distance BD based on the road conditions. For example, in some embodiments, the processing unit 210 may receive an output from a moisture sensor indicating that it is raining. In such cases, the processing unit 210 can determine a higher value for the braking distance BD.
[0344]
[0391] In some embodiments, instead of determining TTB, or in addition to it, the intersection violation prediction unit 222 may be configured to determine the braking distance BD based on the speed V of the vehicle in question (and optionally, based on road conditions and / or vehicle dynamics), and may generate a control signal if the braking distance BD is less than the distance d to the intersection (e.g., the distance between the vehicle in question and the expected stopping position associated with the intersection), or if d-BD <= distance threshold. The control signal can activate a device that generates a warning to the driver and / or a device that controls the vehicle, as described herein. In some embodiments, the distance threshold may be adjusted based on the driver's state. For example, if the driver monitoring module 211 determines that the driver is distracted, the processing unit 210 may increase the distance threshold to account for the increased distance required for the driver to react.
[0345]
[0392] Distance estimation
[0393] In one or more embodiments described herein, the processing unit 210 may be configured to determine a distance d between the target vehicle and a position in front of the vehicle, which could be the position of an object captured by the image from the first camera 202 (e.g., a preceding vehicle, a pedestrian, etc.), the expected stopping position of the vehicle, etc. Various techniques can be employed in different embodiments to determine the distance d.
[0346]
[0394] In some embodiments, the processing unit 210 is configured to determine a distance d based on a Y:d mapping, where Y represents the y-coordinate in the image frame and d represents the distance between the target vehicle and the position corresponding to the Y-coordinate in the image frame. This concept is illustrated in the example in Figure 20, which shows an example of a method for determining the distance d between the target vehicle and the position in front of the vehicle. In the graph above, various widths of the bounding box of a preceding vehicle detected in the camera image are plotted against their respective y-coordinates (i.e., the y-component of each position of the bounding box of the detected object in the camera image), allowing for the determination of the best straight line relating the y-coordinate to each width of the bounding box. The y-coordinates in the graph above are based on a coordinate system with the origin y=0 at the top of the camera image. In other embodiments, the y-coordinates may be based on other coordinate systems (e.g., a coordinate system where the origin y=0 is at the bottom or center of the image). In the example in the figure, a larger y-coordinate value corresponds to a larger bounding box width. This is because vehicles detected closer to the camera appear larger (with a larger corresponding bounding box) and are displayed closer to the bottom of the camera image compared to other vehicles located further away from the camera. Also, in the illustrated example, the optimal straight line in the upper graph of Figure 20 has a linear equation with two parameters B = -693.41 and m = 1.46, where B is the value when y = 0 and m is the slope of the optimal straight line.
[0347]
[0395] It should be noted that the width (or horizontal dimension) in the coordinate system of a camera image is related to a real-world distance d based on the principle of homography. Therefore, the width parameter in the graph above Figure 20 can, in some embodiments, be converted to a real-world distance d based on perspective projection geometry. In some embodiments, the width:distance mapping may be obtained empirically by performing calculations based on perspective projection geometry. In other embodiments, the width:distance mapping may be obtained by measuring the actual distance d between the camera and an object at a certain location and determining the width of the object in the coordinate system of the camera image that captures the object at a distance d from the camera. Furthermore, in even more embodiments, instead of determining the width:distance mapping, a y:d mapping may be determined, which can be determined by measuring the actual distance d between the camera and a real-world location L and determining the y-coordinate of location L in the coordinate system of the camera image.
[0348]
[0396] The information in the graph below Figure 20 may, in some embodiments, be used by the processing unit 210 to determine the distance d. For example, in some embodiments, the information relating the y coordinate to the distance d may be stored in a non-temporary medium. This information may be the equation of a curve relating distance d to different y coordinates, a table containing different y coordinates and corresponding distances d, etc. During operation, the processing unit 210 can detect objects (e.g., vehicles) in the camera image from the first camera 202. The image of the detected object displayed in the camera image has coordinates (x, y) relative to the coordinate system of the camera image. For example, if the value of the y coordinate of the detected object is 510, then based on the curve in Figure 20, the distance of the detected object from the camera / target vehicle is approximately 25m.
[0349]
[0397] As another example, during operation, the processing unit 210 may determine a position in the camera image that represents the desired stopping position of the target vehicle. The position in the camera image has coordinates (x, y) relative to the coordinate system of the camera image. For example, if the value of the y coordinate of the position (representing the desired position of the target vehicle) is 490, then, based on the curve in Figure 20, the distance d between the camera / target vehicle and the desired stopping position (e.g., an actual stop line at an intersection, or an artificially created stop line) will be approximately 50 meters.
[0350]
[0398] It should be noted that the method for determining distance d is not limited to the example described, and the processing unit 210 may utilize other techniques for determining distance d. For example, in other embodiments, the processing unit 210 may receive distance information from a distance sensor, such as a sensor that uses the time-of-flight method for distance determination.
[0351]
[0399] Warning activation and / or automatic vehicle control system
[0352]
[0400] As described herein, the signal generation controller 224 is configured to generate control signals for activating a warning generator and / or control a target vehicle based on the output from the collision prediction unit 218 or the intersection violation prediction unit 222, and based on the output from the driver monitoring module 211 indicating the driver's status. The output from the collision prediction unit 218 or the intersection violation prediction unit 222 may be a TTC value indicating the time to collision (with another vehicle or other object) or the time to pass through a detected intersection.
[0353]
[0401] In some embodiments, the signal generation controller 224 is configured to compare the TTC value (which changes over time) with a threshold (threshold time) and determine whether to generate a control signal based on the comparison result. In some embodiments, the threshold used by the signal generation controller 224 of the processing unit 210 to determine whether to generate a control signal (in response to a predicted collision or predicted intersection violation) may have a minimum value of at least 1 second, or 2 seconds, or 3 seconds, or 4 seconds, or 5 seconds, or 6 seconds, or 7 seconds, or 8 seconds, or 9 seconds, or 10 seconds. This threshold is variable based on the driver's state as indicated by information provided by the driver monitoring module 211. For example, if the driver's state indicates that the driver is distracted, the processing unit 210 may adjust the threshold by increasing the threshold time from its minimum value (for example, if the minimum value is 3 seconds, the threshold may be adjusted to 5 seconds). On the other hand, if the driver's state indicates drowsiness, the processing unit 210 may adjust the threshold to, for example, 7 seconds (i.e., 5 seconds or more in this embodiment). This is because a drowsy driver may take longer to recognize the risk of collision and the need to stop, and to take action to mitigate the danger of a collision.
[0354]
[0402] Figure 21 shows an example of a technique for generating control signals to control a vehicle and / or a technique for generating a warning to the driver. In this example, the collision prediction unit 218 determines that the TTC is 10 seconds. The x-axis of the graph represents the elapsed time since the TTC determination. At time t=0, the collision prediction unit 218 determined an initial TTC of 10 seconds. As time progresses (represented by the x-axis), the TTC (represented by the y-axis) decreases based on the relationship TTC=10-t, where 10 is the initially determined time to collision, TTC, which is 10 seconds. As time passes, the predicted collision approaches in time, so the TTC also decreases. In the illustrated example, the processing unit 210 utilizes a first threshold TH1 of 3 seconds to provide a control signal (to warn the driver and / or to automatically activate the vehicle control device to reduce the risk of collision) when the driver status output by the driver monitoring module 211 indicates that the driver is not distracted. In the illustrated example, the processing unit 210 also utilizes a second threshold TH2 of 5 seconds to provide a control signal (to automatically activate the vehicle control system to warn the driver and / or reduce the risk of collision) when the driver status output by the driver monitoring module 211 indicates that the driver is distracted. In the illustrated example, the processing unit 210 also utilizes a third threshold TH3 of 8 seconds to provide a control signal (to automatically activate the vehicle control system to warn the driver and / or reduce the risk of collision) when the driver status output by the driver monitoring module 211 indicates that the driver is drowsy.
[0355]
[0403] As shown in Figure 21, four different scenarios are presented. In Scenario 1, the output of the driver monitoring module 211 indicates that the driver is not distracted (N) for 7 seconds from t=0 (corresponding to a TTC of 3 seconds). Therefore, the signal generation controller 224 uses a first threshold TH1 (TTC=3 seconds, corresponding to t=7 seconds) as the time to supply the control signal CS (to activate the warning device and / or vehicle control device). In other words, when the TTC decreases from the initial 10 seconds to reach the 3-second threshold, the signal generation controller 224 provides the control signal CS.
[0356]
[0404] In Scenario 2, the output of the driver monitoring module 211 indicates that the driver is not distracted from t=0 to t=1.5 seconds (N), and is distracted from t=1.5 seconds to 5 seconds thereafter (D) (corresponding to a 5-second TTC). Therefore, the signal generation controller 224 uses a second threshold TH2 (TTC=5 seconds, corresponding to t=5 seconds) as the time to supply the control signal CS (to operate the warning device and / or vehicle control device). In other words, as the TTC decreases from the first 10 seconds and reaches the 5-second threshold, the signal generation controller 224 provides the control signal CS. Therefore, in situations where the driver is distracted, the signal generation controller 224 provides the control signal earlier to warn the driver and / or operate the vehicle.
[0357]
[0405] In Scenario 3, the output of the driver monitoring module 211 indicates that the driver is not distracted from t=0 to t=2 seconds (N), distracted from t=2 seconds to 3.5 seconds (D), and not distracted again from t=3.5 seconds to 5 seconds and beyond (N). When the second threshold TH2 is reached at t=5 seconds, the driver's state is not distracted (N), but because the driver's state in this scenario changed from a distracted state to a non-distracted state immediately before the threshold TH2 at t=5 seconds, the signal generation controller 224 still uses the second threshold TH2 (for the distracted state). Thus, in some embodiments, the signal generation controller 224 may be configured to consider the driver's state within a time window prior to the threshold (e.g., 1.5 seconds before TH2, 2 seconds before, etc.) in order to determine whether to use the threshold to determine whether to generate a control signal. In other embodiments, the signal generation controller 224 may be configured to consider the driver's state at the time of the threshold to determine whether to use the threshold.
[0358]
[0406] In Scenario 4, the output of the driver monitoring module 211 indicates that the driver is drowsy (R) from t=0 to t=2 seconds and thereafter (corresponding to an 8-second TTC). Therefore, the signal generation controller 224 uses a third threshold TH3 (TTC=8 seconds, corresponding to t=2 seconds) as the time to supply the control signal CS (to activate the warning device and / or vehicle control device). In other words, as the TTC decreases from the initial 10 seconds and reaches the 8-second threshold, the signal generation controller 224 provides the control signal CS. Thus, in situations where the driver is drowsy, the signal generation controller 224 provides the control signal even earlier (i.e., earlier than when the driver is awake but distracted) to warn the driver and / or prompt them to operate the vehicle.
[0359]
[0407] Thus, as shown in the example above, in some embodiments, the threshold is variable in real time based on the driver's status identified by the driver monitoring module 211.
[0360]
[0408] In any of the above scenarios, if the signal generation controller 224 receives sensor information (for example, information provided by sensor 225) indicating that the driver is operating the vehicle to reduce the risk of collision (such as applying the brakes), the signal generation controller 224 may withhold the provision of control signals.
[0361]
[0409] The example and four scenarios in Figure 21 are described in relation to collision prediction, but can also be applied to intersection violation prediction. In the case of intersection violation prediction, the TTC value indicates the time to pass through the intersection. In some embodiments, the same thresholds TH1, TH2, and TH3 used to determine the timing of providing control signals for collision prediction (activating warning generators and / or vehicle control devices) can also be used for intersection violation prediction. In other embodiments, the thresholds TH1, TH2, and TH3 used to determine the timing of providing control signals for collision prediction may be different from the thresholds TH1, TH2, and TH3 used to determine the timing of providing control signals for intersection violation prediction.
[0362]
[0410] As shown in the examples above, in some embodiments, the collision prediction unit 218 is configured to determine an estimated time until a predicted collision occurs, and the signal generation controller 224 of the processing unit 210 is configured to provide a control signal to activate the device when the estimated time until a predicted collision occurs is below a threshold. In some embodiments, the device includes a warning generator, and the signal generation controller 224 of the processing unit 210 is configured to provide a control signal to cause the device to provide a warning to the driver when the estimated time until a predicted collision occurs is below a threshold. Alternatively or additionally, the device may include a vehicle control device, and the signal generation controller 224 of the processing unit 210 is configured to provide a control signal to cause the device to control the vehicle when the estimated time until a predicted collision occurs is below a threshold.
[0363]
[0411] Furthermore, as shown in the examples above, in some embodiments, the signal generation controller 224 of the processing unit 210 is configured to repeatedly evaluate the estimated time (TTC) with respect to a variable threshold as the estimated time until the predicted collision / intersection violation occurs decreases, and the predicted collision / intersection violation approaches in time.
[0364]
[0412] In some embodiments, the processing unit 210 (for example, the signal generation controller 224 of the processing unit 210) is configured to increase a threshold if the driver's state indicates that the driver is distracted or not paying attention to the driving task.
[0365]
[0413] Furthermore, as shown in the examples above, in some embodiments, the signal generation controller 224 of the processing unit 210 is configured to at least temporarily withhold the provision of control signals if the estimated time until a predicted collision occurs is longer than a threshold.
[0366]
[0414] In some embodiments, the threshold has a first value if the driver's state indicates that the driver is paying attention to the driving task, and a second value higher than the first value if the driver's state indicates that the driver is distracted or not paying attention to the driving task.
[0367]
[0415] Furthermore, as shown in the examples above, in some embodiments, the threshold is also based on sensor information indicating that the vehicle is being operated to reduce the risk of collision. For example, if sensor 225 provides sensor information indicating that the driver is applying the brakes to the vehicle, the processing unit 210 may raise the threshold to a higher value. In some embodiments, the signal generation controller 224 of the processing unit 210 is configured to determine whether or not to provide a control signal based on (1) first information indicating the risk of collision with the vehicle, (2) second information indicating the driver's state, and (3) sensor information indicating that the vehicle is being operated to reduce the risk of collision.
[0368]
[0416] In some embodiments, the processing unit 210 is configured to determine the level of collision risk, and the processing unit 210 (e.g., the signal generation controller 224 of the processing unit 210) is configured to adjust a threshold based on the determined level of collision risk.
[0369]
[0417] In some embodiments, the driver's state includes a distracted state, the processing unit 210 is configured to determine the level of the driver's distracted state, and the processing unit 210 (e.g., the signal generation controller 224 of the processing unit 210) is configured to adjust a threshold based on the determined level of the driver's distracted state.
[0370]
[0418] Furthermore, in some embodiments, different warnings may be provided based on whether the driver is paying attention, using multiple different thresholds. For example, in some embodiments, the processing unit 210 may control the device to provide a first warning having a first characteristic if there is a risk of collision (with a vehicle, pedestrian, etc.) and the driver is paying attention, and to provide a second warning having a second characteristic if there is a risk of collision and the driver is distracted. The first characteristic of the first warning may be a first warning volume, and the second characteristic of the second warning may be a second warning volume that is higher than the first warning volume. Also, in some embodiments, if the processing unit 210 determines that the risk of collision is higher, the processing unit 210 may control the device to provide a stronger warning (e.g., a louder warning and / or a more frequent beeping warning). Thus, in some embodiments, a milder warning may be provided when the target vehicle is approaching an object, and a stronger warning may be provided when the target vehicle is closer to the object.
[0371]
[0419] Similarly, in some embodiments, the processing unit 210 may control the device to provide a first warning having a first characteristic when there is a risk of intersection violation and the driver is paying attention, and to provide a second warning having a second characteristic when there is a risk of intersection violation and the driver is distracted. The first characteristic of the first warning may be a first warning volume, and the second characteristic of the second warning may be a second warning volume that is higher than the first warning volume. Also, in some embodiments, if the processing unit 210 determines that the risk of intersection violation is higher, the processing unit 210 may control the device to provide a stronger warning (e.g., a louder warning and / or a more frequent beeping warning). Thus, in some embodiments, a milder warning may be provided when the vehicle in question is approaching an intersection, and a stronger warning may be provided when the vehicle in question is closer to the intersection.
[0372]
[0420] As illustrated in the example above, the device 200 has an advantage in considering the driver's state when deciding whether to generate a control signal to activate a device that provides a warning and / or a control signal to activate a device that controls the vehicle. Since the driver's state can be used to adjust the monitoring threshold, the device 200 may provide a warning to the driver and / or control the vehicle to mitigate the risk of collision and / or intersection violation earlier, taking into account a specific state of the driver (e.g., the driver is distracted, drowsy, etc.). For example, in some embodiments, the device 200 may warn the driver and / or control the vehicle two seconds before an anticipated risk (e.g., the risk of collision or intersection violation), or even earlier, such as at least three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen seconds before. Also, in conventional monitoring systems that do not consider the driver's state, higher accuracy is built into the system to avoid false detections at the expense of improved sensitivity. By incorporating the driver's state, the device 200 can be configured to operate at low sensitivity (e.g., lower than or equivalent to existing solutions) and to increase its sensitivity only when the driver is inattentive. Improving sensitivity based on the driver's state can be achieved by adjusting one or more thresholds based on the driver's state, such as a threshold for determining the time to collision, a threshold for determining the time to pass through an intersection, a threshold for determining the time to apply the brakes, a threshold for determining whether an object intersects a region of interest (e.g., a camera calibration ROI, an ROI determined based on centerline detection, etc.), and a threshold for the confidence level of object detection.
[0373]
[0421] tailgating
[0374]
[0422] In some embodiments, the processing unit 210 may be configured to consider a scenario in which the target vehicle is following another vehicle. In some embodiments, the following is determined (e.g., measured) by the time-to-headway, which is defined as the distance to the preceding vehicle divided by the speed of the target vehicle (ego vehicle). In some embodiments, the speed of the target vehicle may be obtained from a vehicle speed detection system. In other embodiments, the speed of the target vehicle may be obtained from a GPS system. In further embodiments, the speed of the target vehicle may be determined by the processing unit 210 processing external images received from the first camera 202 of the device 200. Also in some embodiments, the distance to the preceding vehicle can be determined by the processing unit 210 processing external images received from the first camera 202. In other embodiments, the distance to the preceding vehicle may be obtained from a distance sensor, such as a sensor employing time-of-flight technology.
[0375]
[0423] In some embodiments, the processing unit 210 may determine that a pursuit state is in effect if the interval between vehicles is less than the pursuit threshold. As an example of non-limiting examples, the pursuit threshold may be 2 seconds or less, 1.5 seconds or less, 1 second or less, 0.8 seconds or less, 0.6 seconds or less, 0.5 seconds or less, etc.
[0376]
[0424] In some embodiments, the processing unit 210 may be configured to determine that there is a risk of collision if the target vehicle is following and the driver monitoring module 211 determines that the driver is distracted. The processing unit 210 may then generate control signals to cause a device (e.g., a warning generator) to provide a warning to the driver and / or cause a device (e.g., a vehicle control device) to control the vehicle, as described herein. For example, the vehicle control device may automatically apply the brakes, automatically release the accelerator pedal, automatically turn on the hazard lights, or any combination thereof.
[0377]
[0425] Rolling stop
[0378]
[0426] In some embodiments, the processing unit 210 may include a rolling stop module configured to detect rolling stop operations. In some embodiments, the rolling stop module may be implemented as part of the intersection violation prediction unit 222. During operation, the processing unit 210 can detect intersections where a vehicle needs to stop (for example, the processing unit 210 may identify stop signs, red lights, etc., based on processing images from the first camera 202). The rolling stop module can monitor one or more parameters indicating the vehicle's behavior to determine whether the vehicle is performing a rolling stop operation at an intersection. For example, the rolling stop module may acquire vehicle speed, vehicle braking, vehicle deceleration, or any combination thereof. In some embodiments, the rolling stop module can determine that a rolling stop operation has occurred by analyzing the vehicle's speed profile over a period of time as the vehicle approaches the intersection. For example, if the vehicle decelerates (indicating the driver is aware of the intersection) and the vehicle's speed does not decrease further within a period of time, the rolling stop module may determine that the driver is performing a rolling stop operation. As another example, if a vehicle slows down (indicating the driver is aware of the intersection) and then begins to increase in speed as it approaches the intersection, the rolling stop module may determine that the driver is performing a rolling stop. Alternatively, if the vehicle's speed decreases as it approaches the intersection, but does not decrease to a certain threshold within a certain distance from the required stopping position, the rolling stop module may determine that the driver is performing a rolling stop.
[0379]
[0427] In some embodiments, if the rolling stop module determines that the vehicle has not come to a complete stop (for example, if the driver has slowed down in response to a stop sign or red light but has not come to a complete stop), the intersection violation prediction unit 222 may determine that there is a risk of an intersection violation. In response to the determined risk of an intersection violation, the rolling stop module may generate a control signal to activate the device. For example, the control signal may operate a communication device to wirelessly send a message to a server system (for example, a cloud system). The server system can be used by a vehicle management company to instruct drivers or by an insurance company to identify dangerous drivers. Alternatively or additionally, the control signal may activate a warning system to warn the driver, which can serve as a method of instructing the vehicle. Alternatively or additionally, the control signal may activate the vehicle's braking system to control the vehicle to a complete stop.
[0380]
[0428] method
[0381]
[0429] Figure 22A shows a method 800 performed by the apparatus 200 of Figure 2A according to several embodiments. The method 800 includes the steps of: acquiring a first image generated by a first camera, the first camera being configured to view the environment outside the vehicle (item 802); acquiring a second image generated by a second camera, the second camera being configured to view the driver of the vehicle (item 804); determining first information indicating a risk of collision with the vehicle, at least partially based on the first image (item 806); determining second information indicating the driver's condition, at least partially based on the second image (item 808); and determining whether to provide a control signal to operate the device, based on (1) the first information indicating a risk of collision with the vehicle and (2) the second information indicating the driver's condition (item 810).
[0382]
[0430] Optionally, in method 800, the first information is determined by predicting a collision, which is predicted at least 3 seconds before the expected time of occurrence of the predicted collision.
[0383]
[0431] Optionally, in method 800, the first information is determined by predicting a collision, which is predicted with sufficient lead time for the driver's brain to process the input and for the driver to take action to mitigate the risk of the collision.
[0384]
[0432] Optionally, in method 800, sufficient lead time depends on the driver's condition.
[0385]
[0433] Optionally, in method 800, first information indicating the risk of collision includes a predicted collision, the method further includes a step of determining an estimated time until the predicted collision occurs, and the control signal is provided to cause a device to provide a control signal if the estimated time until the predicted collision occurs is less than a threshold.
[0386]
[0434] Optionally, in method 800, the device comprises a warning generator, and the control signal is provided such that the device warns the driver if the estimated time until a predicted collision occurs is less than a threshold.
[0387]
[0435] Optionally, in method 800, the device includes a vehicle control device, the control signal provided to cause the device to control the vehicle when the estimated time until a predicted collision occurs is less than a threshold.
[0388]
[0436] Optionally, in method 800, the threshold is variable based on second information indicating the driver's condition.
[0389]
[0437] Optionally, in method 800, the estimated time is repeatedly evaluated with respect to a variable threshold, as the estimated time until the predicted collision occurs decreases, the predicted collision approaches in time.
[0390]
[0438] Optionally, in method 800, the threshold is variable in real time based on the driver's condition.
[0391]
[0439] Optionally, in method 800, the method further includes the step of changing the threshold if the driver's condition indicates that the driver is distracted or not paying attention to the driving task.
[0392]
[0440] Optionally, method 800 further includes the step of at least temporarily suspending the generation of a control signal if the estimated time until a predicted collision occurs is longer than a threshold.
[0393]
[0441] Optionally, method 800 further includes the steps of determining a level of collision risk and adjusting the threshold based on the determined level of collision risk.
[0394]
[0442] Optionally, in method 800, the driver's state includes a distracted state, and the method further includes the steps of determining the level of the driver's distracted state and adjusting the threshold based on the determined level of the driver's distracted state.
[0395]
[0443] Optionally, in method 800, if the driver's state indicates that the driver is paying attention to the driving task, the threshold has a first value, and if the driver's state indicates that the driver is distracted or not paying attention to the driving task, the threshold has a second value that is higher than the first value.
[0396]
[0444] Optionally, in method 800, the threshold is also based on sensor information indicating that the vehicle is being operated to reduce the risk of collision.
[0397]
[0445] Optionally, in method 800, the step of determining whether or not to provide a control signal for operating the device may also be performed based on sensor information indicating that the vehicle is being operated to reduce the risk of collision.
[0398]
[0446] Optionally, in method 800, the step of determining first information indicating the risk of collision includes processing the first image based on a first model.
[0399]
[0447] Optionally, in method 800, the first model includes a neural network model.
[0400]
[0448] Optionally, in method 800, the step of determining second information indicating the driver's state includes processing the second image based on a second model.
[0401]
[0449] Optionally, method 800 further includes the steps of determining a metric value for each of a plurality of pose classifications, and determining whether a driver is engaged in a driving task based on one or more metric values.
[0402]
[0450] Optionally, in method 800, the pose classification includes two or more of the following: looking down, looking up, looking left, looking right, using a mobile phone, smoking, holding an object, taking hands off the steering wheel, not wearing a seatbelt, closing eyes, looking forward, one hand on the steering wheel, and both hands on the steering wheel.
[0403]
[0451] Optionally, method 800 further includes the step of comparing the metric value with the respective threshold for each pose classification.
[0404]
[0452] Optionally, method 800 further includes the step of determining that a driver belongs to one of the pause classifications if one of the corresponding metric values satisfies or exceeds one of the corresponding thresholds.
[0405]
[0453] Optionally, method 800 is performed by an aftermarket device, in which the first and second cameras are integrated as components of the aftermarket device.
[0406]
[0454] Optionally, in method 800, the second information is determined by processing the second image to determine whether the image of the driver satisfies a pose classification, and the method further includes the step of determining whether the driver is engaged in a driving task based on whether the image of the driver satisfies a pose classification.
[0407]
[0455] Optionally, in method 800, the step of determining second information indicating the driver's state includes processing the second image based on a neural network model.
[0408]
[0456] Figure 22B shows a method 850 performed by the apparatus 200 of Figure 2A according to several embodiments. The method 850 includes the steps of: acquiring a first image generated by a first camera, the first camera being configured to view the environment outside the vehicle (item 852); acquiring a second image generated by a second camera, the second camera being configured to view the driver of the vehicle (item 854); determining first information indicating an intersection violation risk based at least partially on the first image (item 856); determining second information indicating the driver's condition based at least partially on the second image (item 858); and determining whether or not to provide a control signal to operate the device based on (1) the first information indicating an intersection violation risk and (2) the second information indicating the driver's condition (item 860).
[0409]
[0457] Optionally, the first piece of information is determined by predicting an intersection violation, which is predicted at least 3 seconds before the expected time of occurrence of the predicted intersection violation.
[0410]
[0458] Optionally, the first piece of information is determined by predicting an intersection violation, the predicted intersection violation being predicted with sufficient lead time for the driver's brain to process the input and for the driver to take action to mitigate the risk of the intersection violation.
[0411]
[0459] The appropriate lead time depends on the driver's condition.
[0412]
[0460] Optionally, first information indicating the risk of an intersection violation includes a predicted intersection violation, the method further includes the step of determining an estimated time until the predicted intersection violation occurs, and the control signal is provided to cause the device to provide a control signal if the estimated time until the predicted intersection violation occurs is less than a threshold.
[0413]
[0461] Optionally, the device comprises a warning generator, and the control signal is provided such that the device issues a warning to the driver if the estimated time until a predicted intersection violation occurs is below a threshold.
[0414]
[0462] Optionally, the device includes a vehicle control device, the control signal provided to cause the device to control the vehicle when the estimated time until a predicted intersection violation occurs is less than a threshold.
[0415]
[0463] The threshold is optionally variable based on second information indicating the driver's condition.
[0416]
[0464] Optionally, the estimated time is repeatedly evaluated with respect to a variable threshold, as the estimated time until the predicted intersection violation occurs decreases, indicating that the predicted intersection violation is approaching in time.
[0417]
[0465] The threshold is arbitrarily variable in real time based on the driver's condition.
[0418]
[0466] Optionally, the method further includes the step of changing the threshold if the driver's condition indicates that the driver is distracted or not paying attention to the driving task.
[0419]
[0467] Optionally, the method further includes the step of at least temporarily suspending the generation of a control signal if the estimated time until a predicted intersection violation occurs is longer than a threshold.
[0420]
[0145] Optionally, the method further includes the steps of determining a level of risk of an intersection violation and adjusting the threshold based on the determined level of risk of an intersection violation.
[0421]
[0469] Optionally, the driver's state may include a state of distraction, and the method further includes the steps of determining the level of the driver's distraction and adjusting the threshold based on the determined level of the driver's distraction.
[0422]
[0470] Optionally, if the driver's state indicates that the driver is paying attention to the driving task, the threshold has a first value; if the driver's state indicates that the driver is distracted or not paying attention to the driving task, the threshold has a second value that is higher than the first value.
[0423]
[0471] Optionally, the threshold may also be based on sensor information indicating that the vehicle is being operated in a manner that reduces the risk of intersection violations.
[0424]
[0472] The step of deciding whether or not to provide a control signal to operate the device is also performed based on sensor information indicating that the vehicle is being operated in a manner that reduces the risk of intersection violations.
[0425]
[0473] The step of optionally determining first information indicating the risk of the intersection violation includes processing the first image based on a first model.
[0426]
[0474] Optionally, the first model may include a neural network model.
[0427]
[0475] The step of optionally determining second information indicating the driver's state includes processing the second image based on a second model.
[0428]
[0476] Optionally, the method further includes the steps of determining a metric value for each of a plurality of pose classifications, and determining whether or not a driver is engaged in a driving task based on one or more metric values.
[0429]
[0477] Optionally, the pose classification includes two or more of the following: looking down, looking up, looking left, looking right, using a mobile phone, smoking, holding an object, hands off the steering wheel, not wearing a seatbelt, eyes closed, looking forward, one hand on the steering wheel, or both hands on the steering wheel.
[0430]
[0478] Optionally, the method further includes the step of comparing the metric value with the respective threshold for each pose classification.
[0431]
[0479] Optionally, the method further includes the step of determining that a driver belongs to one of the pause classifications if one of the corresponding metric values satisfies or exceeds one of the corresponding thresholds.
[0432]
[0480] Optionally, the method may be performed by an aftermarket device, in which the first and second cameras are integrated as components of the aftermarket device.
[0433]
[0481] Optionally, the second information is determined by processing the second image to determine whether the image of the driver satisfies a pose classification, and the method further includes the step of determining whether the driver is engaged in a driving task based on whether the image of the driver satisfies a pose classification.
[0434]
[0482] The step of optionally determining second information indicating the driver's state includes processing the second image based on a neural network model.
[0435]
[0483] Model generation and integration
[0436]
[0484] Figure 23 shows a technique for determining the model to be used by the device 200 according to several embodiments. As shown, there may be multiple vehicles 910a to 910d, each equipped with a device 200a to 200d. Each of the devices 200a to 200d may have the configuration and features described with reference to the device 200 in Figure 2A. During operation, the cameras of devices 200b to 200d mounted on vehicles 910b to 910d (both external and internal surveillance cameras) capture images of the external environment of each vehicle 910b to 910d and images of each driver. The images are transmitted directly or indirectly to the server 920 via a network (cloud, internet, etc.). The server 920 includes a processing unit 922 configured to process images from devices 200b to 300d of vehicles 910b to 910d to determine the model 930 and one or more models 932. Model 930 may be configured to detect the driver's pose, and Model 932 may be configured to detect various types of objects in the camera image. Models 930 and 932 may be stored in a non-temporary medium 924 in the server 920. The server 920 may transmit Models 930 and 932 directly or indirectly to the device 200a in the vehicle 910a via a network (e.g., the cloud, the internet, etc.). The device 200a can then use Model 932 to process the images received by the camera of the device 200a and detect various poses of the driver of the vehicle 910a. The device 200a can also use Model 932 to process the images received by the camera of the device 200a and detect various objects outside the vehicle 910a and / or determine the region of interest of the camera of the device 200a.
[0437]
[0485] In the example shown in Figure 23, three devices 200b to 200d for providing images are installed in each of the three vehicles 910b to 910d. In other examples, there may be three or more devices 200 in each of three or more vehicles 910 for providing images to the server 920, or there may be fewer than three devices 200 in fewer than three vehicles 910 for providing images to the server 920.
[0438]
[0486] In some embodiments, Model 930 provided by Server 920 may be a neural network model. Model 932 provided by Server 920 may be one or more neural network models. In such cases, Server 920 may be a neural network or part of a neural network, and images from devices 200b-200d may be utilized by Server 920 to constitute Model 930 and / or Model 932. In particular, the processing unit 922 of Server 920 may constitute Model 930 and / or Model 932 by training Model 930 by machine learning. In some cases, images from different devices 200b-200d may form a rich dataset from different cameras mounted at different positions relative to the corresponding vehicle, which is useful for training Model 930 and / or Model 932. As used herein, the term “neural network” means a computing device, system, or module consisting of a number of interconnected processing elements that process information by dynamic state responses to inputs. In some embodiments, the neural network may have deep learning capabilities and / or artificial intelligence. In some embodiments, the neural network may be any computing element that can be trained using one or more datasets.As an unrestricted example, neural networks can be perceptrons, feedforward neural networks, radial-based neural networks, deep feedforward neural networks, recurrent neural networks, long-term / short-term memory neural networks, gated recurrent units, autoencoder neural networks, variational autoencoder neural networks, denoising autoencoder neural networks, sparse autoencoder neural networks, Markov chain neural networks, Hopfield neural networks, Boltzmann machines, restricted Boltzmann machines, deep belief networks, convolutional networks, deconvolutional networks, deep convolutional inverse graphics networks, generative adversarial networks, liquid state machines, extreme learning machines, echo state networks, deep resolver networks, Kohonen networks, support vector machines, neural Turing machines, modular neural networks, sequence-to-sequence models, or any combination thereof.
[0439]
[0487] In some embodiments, the processing unit 922 of the server 920 uses images to configure (e.g., train) a model 930 to identify specific poses of the driver. In a non-limiting example, the model 930 is configured to identify poses of the driver such as looking down, looking up, looking left, looking right, using a mobile phone, smoking, holding an object, taking hands off the steering wheel, not wearing a seatbelt, closing eyes, looking forward, having one hand on the steering wheel, having both hands on the steering wheel, etc. Also in some embodiments, the processing unit 922 of the server 920 may use images to configure a model to determine whether the driver is engaged in a driving task. In some embodiments, the determination of whether the driver is engaged in a driving task may be achieved by a processing unit that handles the driver's pose classification. In one implementation example, the pose classification may be the output provided by a neural network model. In such a case, the neural network model is passed to a processing unit, which may determine whether the driver is engaged in a driving task based on the pose classification from the neural network model. In other embodiments, the processing unit receiving the pose classification may be another (e.g., a second) neural network model. In this case, the first neural network model is configured to output the pose classification, and the second neural network model is configured to determine whether the driver is engaged in a driving task based on the pose classification output by the first neural network model. In such a case, model 930 can be thought of as having both the first and second neural network models. In further embodiments, model 930 may be a single neural network model configured to receive an image as input and provide an output indicating whether the driver is engaged in a driving task.
[0440]
[0488] In some embodiments, the processing unit 922 of the server 920 configures (for example, trains) the model 932 to detect various objects using images. In non-limiting examples, the model 932 may be configured to detect vehicles, people, animals, bicycles, traffic lights, road signs, curbs, road centerlines, and the like.
[0441]
[0489] In other embodiments, models 930 and / or 932 may not be neural network models, but may be of other types. In such cases, the configuration of models 930 and / or 932 by the processing unit 922 may not involve machine learning and / or may not require images from devices 200b-200d. Instead, the configuration of models 930 and / or 932 by the processing unit 922 may be achieved by the processing unit 922 determining (e.g., acquiring, calculating, etc.) the processing parameters (such as feature extraction parameters) of models 930 and / or 932. In some embodiments, models 930 and / or 932 may include program instructions, commands, scripts, parameters (e.g., feature extraction parameters), etc. In one embodiment, models 930 and / or 932 may be in the form of an application that can be received wirelessly by device 200.
[0442]
[0490] After models 930 and 932 are configured by server 920, models 930 and 932 become available to devices 200 in various vehicles 910 for identifying objects in camera images. As shown in the figure, models 930 and 932 may be transmitted from server 920 to device 200a in vehicle 910a. Alternatively, models 930 and 932 may be transmitted from server 920 to devices 200b to 200d in each vehicle 910b to 910d. After device 200a receives models 930 and 932, the processing unit in device 200a may process images generated by the device 200a's camera (internal camera) based on model 930 to identify the driver's pose and / or determine whether the driver is engaged in a driving task, as described herein, and may process images generated by the device 200a's camera (external camera) based on model 932 to detect objects outside vehicle 910a.
[0443]
[0491] In some embodiments, the transmission of models 930, 932 from server 920 to device 200 (e.g., device 200a) may be performed by server 920 "push-sending" models 930, 932, so that device 200 does not need to request models 930, 932. In other embodiments, the transmission of models 930, 932 from server 920 may be performed by server 920 in response to a signal generated and transmitted by device 200. For example, device 200 may generate and transmit a signal after device 200 is powered on or after a vehicle carrying device 200 is started. This signal is received by server 920, and server 920 transmits models 930, 932 so that device 200 can receive it. As another example, device 200 may include a user interface, such as a button, that allows users of device 200 to submit requests for models 930, 932. In this case, when the button is pressed, the device 200 sends a request for models 930 and 932 to the server 920. In response to this request, the server 920 sends models 930 and 932 to the device 200.
[0444]
[0492] Please note that the server 920 in Figure 23 is not limited to a single server device, but may consist of multiple server devices. Furthermore, the processing unit 922 of the server 920 may include one or more processors, one or more processing modules, etc.
[0445]
[0493] In other embodiments, the images acquired by server 920 do not have to be generated by devices 200b-200d. Instead, the images used by server 920 to determine (e.g., train, configure, etc.) models 930, 932 may be recorded using other devices such as mobile phones or cameras in other vehicles. Also in other embodiments, the images used by server 920 to determine (e.g., train, configure, etc.) models 930, 932 may be downloaded to server 920 from a database associated with server 920 or from a database owned by a third party.
[0446]
[0494] Multi-signal prediction model
[0447]
[0495] In the above embodiment, the processing unit 210 is described as being configured to individually determine various risk factors (e.g., risk of collision, risk of intersection violation, driver distraction, speeding, etc.) and to determine whether or not to generate a control signal to activate the device based on any of the individual risk factors that meet a specific criterion. Figure 24 shows an example of such a method in which various risk factors are evaluated by their respective triggers (e.g., their respective criteria). In some cases, if a criterion is met for one of the triggers, the processing unit 210 may generate a control signal to generate a warning. For example, as shown, if a distraction parameter determined by the processing unit 210 (e.g., by processing the driver's interior image) meets the distraction criterion, the processing unit 210 may then determine that the driver is distracted and generate a control signal to issue a warning. As another example, if a time to collision (TTC) parameter determined by the processing unit 210 meets the near-collision criterion, the processing unit 210 may generate a control signal to generate a warning. In the example above, the criteria may include a static or variable threshold, which may be user-configurable (e.g., for high, medium, or low risk tolerance). In either case, if it is determined that the risk parameter meets (e.g., exceeds) the threshold, the processing unit 210 may generate a control signal to generate a warning.
[0448]
[0496] In the technique shown in Figure 24, algorithms utilizing various corresponding thresholds operate independently and do not communicate with each other. For example, an algorithm that processes distraction risk signals does not communicate with an algorithm that processes time-to-collision risk signals. Furthermore, the thresholds used by each algorithm may not be adaptable to instantaneous situations or contexts. Therefore, risk combinations, risk relationships, and risk contexts are ignored. Moreover, nonlinearity is also ignored in the technique shown in Figure 24. The increase in risk in the threshold model is linear until the threshold is reached. However, in reality, risk can escalate nonlinearly (for example, if a preceding vehicle starts braking and changes lanes, or if an event occurs further down the road). In addition, the technique does not consider various risk factors in combination with driver intent, driver attention, cognitive load, and other factors (for example, expected responses to the environment such as other vehicles, overtaking maneuvers, or changes in traffic control signals).
[0449]
[0497] In other embodiments, various risk factors may be processed together by the processing unit 210 to determine whether they collectively indicate a non-event (e.g., a non-hazardous event). This function has the advantage of reducing false positives in the system because even if one risk factor indicates a hazardous situation (e.g., risk of collision), considering multiple risk factors together may collectively indicate a non-hazardous situation. In some cases, if only a single risk factor is used to trigger a warning, excessive false positives may occur, generating unnecessary warnings. This is undesirable, as it may lead drivers to ignore warnings or turn off the in-vehicle device to avoid false positives. The reverse is also true. In particular, even if a single risk factor indicates a non-hazardous situation (e.g., the risk factor does not exceed a threshold), considering multiple risk factors together may collectively indicate a hazardous situation, thereby reducing false negatives in the system. Specifically, risk factors may not individually exceed a threshold, but when considered in combination with a fusion / holistic approach, they may collectively indicate a very risky situation.
[0450]
[0498] Figure 25 shows an example of such technology, in which the processing unit 210 is configured to process multiple risk factors together and determine whether they collectively represent a non-event (non-dangerous event) or a risk event. In the illustrated example, various risk factors (e.g., driver distraction, time to intersection violation, time to collision, speed, etc.) are given as input to the model in the processing unit 210, and the processing unit 210 is configured to process two or more risk factors and determine whether they collectively represent a dangerous situation. In the illustrated embodiment, the processing unit 210 is configured to determine a risk score (e.g., instantaneous risk score) based on the probabilities of each different state (predicted event), such as the probability of a collision, the probability of a near-collision, and the probability of a non-risk state, or two or more of the aforementioned. In the illustrated example, the model in the processing unit 210 determines three probabilities for the predicted event based on the multiple risk factors received by the model: the probability of a collision occurring within the next two seconds, the probability of a near-collision occurring within the next two seconds, and the probability of a non-risk state occurring within the next two seconds. In other embodiments, the model may determine more than three or fewer probabilities of predicted events based on multiple risk factors. Also in other embodiments, the future time in which the probabilities are determined may be longer than two seconds (e.g., within 3, 4, 5, 6, 7, 8, 9, 10 seconds, etc.) or shorter than two seconds (e.g., within 1.9, 1.8, 1.7, 1.6, 1.5, 1.4, 1.3, 1.2, 1.1, 1.0, 0.9, 0.8, 0.7, 0.6, 0.5 seconds, etc.).
[0451]
[0499] The model of the processing unit 210 may be a neural network model trained using prior risk factors. The neural network model may be any neural network architecture. Figures 26A-26C show examples of neural network architectures that may be employed to process multiple risk factors as input by the neural network model. The neural network architecture may be, for example, any shallow network (e.g., SVM (support-vector-machine), logistic regression, etc.), an ensemble network (e.g., random forest), or any deep network (e.g., recurrent neural network, LSTM (long-short-term memory network), convolutional neural network (CNN)). The neural network model may be any type of neural network described herein.
[0452]
[0500] In some embodiments, the model may include submodels implemented by each processing unit. In such cases, these processing units can be considered sub-processing units of processing unit 210.
[0453]
[0501] As described, the processing unit 210 may be part of a device 10 that includes a first camera 202 configured to view the environment outside the vehicle and a second camera 204 configured to view the driver of the vehicle. The processing unit 210 is configured to receive a first image from the first camera and a second image from the second camera. The processing unit 210 includes a model configured to receive a plurality of inputs and generate metrics based on at least a portion of the inputs. The plurality of inputs include a first time series of information indicating at least a first risk factor and a second time series of information indicating a second risk factor.
[0454]
[0502] In some embodiments, the device 200 may include additional sensors for sensing motion characteristics related to the operation of the vehicle. In non-limiting examples, the sensors may include one or more sensors for sensing one or more of the following: acceleration, velocity, centripetal force, steering angle, brakes, accelerator position, turn signals, etc. In this case, the processing unit 210 (e.g., a model therein) may be configured to receive inputs from the sensors and determine a risk score based on such inputs. In some embodiments, such inputs may be used by the model to determine probabilities as described with reference to Figure 25. Also in some embodiments, the processing unit 210 (e.g., a model therein) may be configured to receive CAN / OBD signals as inputs. In such cases, the processing unit 210 may be connected to the vehicle and acquire signals from the vehicle communication system regarding brakes, pedal position, steering angle, velocity, etc. Alternatively or additionally, the processing unit 210 may be connected to the vehicle's onboard diagnostic system and acquire signals from such a system. Inputs from CAN / OBD and / or the vehicle's onboard diagnostic system may be utilized by a model within the processing unit 210 to determine the probability of various events, as described with reference to Figure 25. Furthermore, in some embodiments, information received by and / or output by the processing unit 210 may be transmitted via an in-vehicle Ethernet or other data transfer architecture (e.g., one capable of providing high-bandwidth data transfer).
[0455]
[0503] In further embodiments, the device 200 may include a plurality of sensors that collect raw data (external video, internal video, velocity, etc.) as input, and a processing unit that generates high-level, context-rich risk estimates based on the collected raw data. The processing unit 210 may be a single-stage processing unit (e.g., a single-stage processor) or may include a multi-stage processing unit. In embodiments where the processing unit 210 is implemented as a single-stage processing unit, the single-stage processing unit may include an end-to-end (E2E) deep fusion risk model that takes raw data as input and provides risk estimates based on the raw data. In embodiments where the processing unit 210 is implemented as a multi-stage processing unit, the processing unit 210 may include at least (1) a first-stage system that takes raw data as input and provides risk signals as output, and (2) a second-stage system that takes risk signals from the first-stage system and provides estimated risk. In some embodiments, the first-stage system is configured to acquire data of higher dimensionality or complexity (e.g., raw data) and process that data to provide an output of lower dimensionality or complexity (e.g., a vector). The lower-dimensionality or lower-complexity output (risk signal) is then input to the second-stage system, which processes this output to determine the probabilities of various events (predicted events) (e.g., collision events, near-collision events, non-hazardous events, etc.) and a risk score representing the estimated risk based on the probabilities of the various events. The output from the second-stage system may be of lower dimensionality or lower complexity compared to the input received by the second-stage system (e.g., a time series of the risk signal). In some embodiments, the models described herein may be implemented by the second-stage system. Since the estimated risk is based on the risk signal, the estimated risk can be thought of as a "fused" risk that combines or incorporates the risk signals (corresponding to each risk factor).
[0456]
[0504] In some embodiments, the first-stage system may include a plurality of first-stage processing units. For example, in one embodiment, the device 200 may include a plurality of sensors that collect raw data (e.g., outer video, inner video, velocity), a plurality of first-stage processing units (e.g., first-stage processors) that provide individual risk signals based on the raw data, and a second-stage processing unit (e.g., second-stage processor) that generates a high-level, context-rich estimate of risk based on the individual risk signals. For example, there may be a first-stage processing unit that determines a bounding box based on images or video from a first camera 202. Another example is a first-stage processing unit that determines the distance to a stop line based on images or video from the first camera 202. A further example is a first-stage processing unit that identifies a facial landmark based on images or video from a second camera 204. Note that the second-stage processing units are configured to provide estimated risk based on the outputs from the first-stage processing units, and the outputs from the first-stage processing units can be thought of as risk signals (e.g., metadata) for each risk factor that are different from the raw data obtained from the sensors.
[0457]
[0505] In some embodiments, the first stage processing unit may be implemented by one or more first neural network models, and the second stage processing unit may be implemented by one or more second neural network models. In some embodiments, two or more inputs (e.g., risk factors) may be determined by the first neural network model implementing the first stage processing unit and fed to the second neural network model implementing the second stage processing unit. The second neural network model considers the inputs together and determines whether the risk factors collectively lead to a dangerous situation or a non-dangerous situation.
[0458]
[0506] In some embodiments, the first-stage processing unit and / or the second-stage processing unit can be considered as part of the processing unit 210. Also, in some embodiments, the second-stage processing unit may include a plurality of sub-processing units.
[0459]
[0507] In a non-limiting example, the multiple inputs received by the model / processing unit / first-stage processing unit / second-stage processing unit include: distance to collision, distance to intersection stop line, vehicle speed, time to collision, time to intersection violation, estimated braking distance, information about road conditions (e.g., road surface conditions (dry, wet, snow, etc.) (may include traction control information for slip indication), information about (identifying) special areas (construction areas, school zones, etc.), information about (identifying) the environment (urban areas, suburbs, city areas, etc.), information about (identifying) traffic conditions, time, information about (identifying) visibility conditions (e.g., fog, snow, precipitation, sun angle, glare, etc.), information about (identifying) objects (stop signs, signals, pedestrians, cars, animals, poles, trees, lanes, curbs, taillights, lane mapping, etc.) (labels, etc.), object position, object direction of movement, and object speed. , bounding boxes (2D bounding boxes, 3D bounding boxes, etc.), vehicle operating parameters (kinematic signals such as acceleration, velocity, centrifugal force, steering angle, brakes, accelerator position, turn signals, traction control, etc.), information about (identifying) the driver's state (e.g., looking up, looking down, looking left, looking right, using a mobile phone, holding an object, smoking, eyes closed, head turned or moved away to avoid face detection, gaze direction relative to the driver, gaze direction mapped to external or internal objects (e.g., head unit, mirrors, etc.), changes in gaze, objects the driver is looking at, the driver's physiological state (e.g., drowsiness, fatigue, anger on the road, stress, sudden illness (heart attack, stroke, pulse, sweating, pupil dilation, etc.)), information about driving history (e.g., vehicle records, accident rate, experience, age, years of driving, years of vehicle ownership, route experience, VERA (Vision Enhanced Risk)Information provided from risk assessment tools that evaluate drivers based on driving behavior, such as driving behavior assessment, continuous driving time, proximity of meal times, information on accident history (e.g., fatal accidents at specific times, accidents at specific times, geospatial heatmaps of time-series risk data by location), sensor signals (e.g., LIDAR signals, radar signals, GPS signals, ultrasonic signals, signals communicated between vehicles and other vehicles (e.g., distance and relative speed signals communicated via 5G), signals communicated between vehicles and infrastructure (e.g., via 5G), or any combination of these (e.g., fusion), driver behavior This may include one or more combinations of information relating to activity or dynamic response (kinematic signals from vehicle systems (such as signals indicating acceleration, velocity, cornering, etc.), control signals (such as signals indicating steering angle, brake position, accelerator position, turn signal status, etc.), location-specific information (e.g., any information indicating risk by dangerous road, dangerous intersection, location, date, and / or time (such as information showing the number of fatalities by time of day and day of the week, as shown in Figure 27)), and audio signals (such as detected speech, detected baby crying, detected loud music inside the vehicle, and detected horns from the vehicle or outside the vehicle).
[0460]
[0508] In some embodiments, the vehicle's GPS system provides timing and position signals, which can then be used by a processing unit 210 to acquire (identify) road condition information, special zones (e.g., construction zones, school zones, etc.), environmental information (e.g., urban areas, suburbs, city areas, etc.), traffic conditions (identify), time of day, etc. (e.g., via an API).
[0461]
[0509] In some embodiments, the processing unit 210 may include or communicate with an object detection module. In such cases, whenever a new class (e.g., animals, utility poles, trees) is introduced to the object detection module, it can be directly input into the model. The model learns the relevance of the new classes during training. In this way, new risk signals can be introduced into the model without manually designing the logic.
[0462]
[0510] In some embodiments, the input supplied to the model may include at least two of the exemplary inputs described above. For example, in some embodiments, the input supplied to the model may include first information relating to the driver's condition and second information relating to the external conditions of the vehicle.
[0463]
[0511] In some embodiments, the processing unit 210 may be configured to package two or more of the exemplary inputs described above into a data structure for supplying to the model. The data structure may consist of a two-dimensional data matrix. In some embodiments, the model is a Temporal-CNN, and the data structure for the model's inputs may be constructed by encoding the input (risk signal) and time as columns and rows of a single-channel input image, or vice versa. The model's inputs can take the form of a time-series matrix of any individual or combined signals that can be considered to predict instantaneous situational risk. An example of such a matrix is described later with reference to Figure 28.
[0464]
[0512] In some embodiments, the model in the processing unit 210 of the device 200 is configured to receive at least a first time series input and a second time series input in parallel, and / or process the first time series and the second time series in parallel. The first time series may be, for example, 6 seconds (or any other arbitrary period) of TTC data, and the second time series may be, for example, 6 seconds (or any other arbitrary period) of velocity data. The first time series and the second time series may be combined to form a single-channel input image, which may be stored in a non-temporary medium. Alternatively, the first time series and the second time series may be stored separately. Even in such cases, the input image can be thought to be formed by logically relating the first time series and the second time series, and / or when the first time series and the second time series are input to the model in parallel.
[0465]
[0513] As described above, in some embodiments, the processing unit 210 of the device 200 may be configured to determine various probabilities of predicted events (e.g., probability of collision, probability of near collision, probability of non-hazardous event, etc.) based on time-series inputs. Figure 28 shows an example of multiple inputs processed by the model to generate predicted values at least 2 seconds before an event occurs. As shown, inputs from the past 6 seconds are input to the model in the processing unit 210, and the model makes predictions for one or more events that may occur in the future 2 seconds before an event occurs. The inputs from the past 6 seconds include time-series data within a 6-second window. In other embodiments, instead of using data from the past 6 seconds, the processing unit 210 may be configured to process data from less than 6 seconds (e.g., 5, 4, 3, 2, 1 second, etc.) to determine the probability of a predicted event, or it may be configured to determine the probability of a predicted event from data from 6 seconds or more (e.g., 7, 8, 9, 10, 11, 12, ... 30 seconds, etc.).
[0466]
[0514] In particular, as shown in Figure 28, the input set may include different parameters and their respective values for the corresponding time point in the immediately preceding period (e.g., 6 seconds in this example). Parameters (or risk factors) in the input set include distance, looking down, looking up, looking left, looking right, phone, smoking, holding an object, eyes closed, no face, time to collision, and speed. The "distance" parameter indicates the distance between the target vehicle and the preceding vehicle. The "looking down" parameter indicates whether the driver is looking down, or to what extent (e.g., a higher value indicates a higher probability that the driver is looking down). The "looking up" parameter indicates whether the driver is looking up, or to what extent (e.g., a higher value indicates a higher probability that the driver is looking up). The "looking left" parameter indicates whether the driver is looking left, or to what extent (e.g., a higher value indicates a higher probability that the driver is looking left). The "Looking Right" parameter indicates whether the driver is looking right, or to what extent (e.g., a higher value indicates a higher probability that the driver is looking right). The "Phone" parameter indicates whether the driver is using a phone (e.g., a higher value indicates a higher probability that the driver is using a phone). The "Smoking" parameter indicates whether the driver is smoking (e.g., a higher value indicates a higher probability that the driver is smoking). The "Object Holding" parameter indicates whether the driver is holding an object (e.g., a higher value indicates a higher probability that the driver is holding an object). The "Eyes Closed" parameter indicates whether the driver's eyes are closed, or to what extent (e.g., a higher value indicates a higher probability that the driver's eyes are closed). The "No Face" parameter indicates whether the driver's face is detected, or to what extent the driver's face is detected (e.g., a higher value indicates a lower probability that the driver's face is not detected). The "Time to Collision" (TTC) parameter indicates the predicted time to collision (e.g., a higher value indicates a shorter TTC). The "speed" parameter indicates the speed of the vehicle in question (for example, a higher value indicates a higher vehicle speed).
[0467]
[0515] It should be noted that the set of inputs (input data) is not limited to the examples of parameters described above, and the set of inputs may have other parameters that may be related to risk. Therefore, in other embodiments, the set of inputs may have additional parameters. In further embodiments, the set of inputs may not have all of the parameters described, but may have fewer parameters (for example, it may have 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 parameters). Furthermore, in other embodiments, the set of inputs may have parameters different from those described.
[0468]
[0516] As shown in Figure 28, the model is configured to generate model predictions based on a set of inputs (e.g., risk factors). In some embodiments, the model predictions may include (1) the probability of a collision within T seconds, (2) the probability of a near collision within T seconds, and (3) the probability of a non-hazardous event within T seconds. Since the model is trained to distinguish these three events, it can determine whether the set of inputs represents a collision event, a near collision event, or a non-hazardous event. The model can also determine the probability of each predicted event. In some embodiments, the model can determine the probability of each predicted event by applying a softmax function to the output vector / logits such that the final output for each type of predicted event is a probability value between 0 and 1, and the sum of all outputs (probabilities of different predicted events) is 1. In some embodiments, the model is trained on a sufficient number of similar examples of collision events so that it can predict a high probability of a collision event for a given set of inputs leading to a collision event. The model may also be trained on a sufficient number of similar examples of near collision events so that it can predict a high probability of a near collision event for a given set of inputs leading to a near collision event. The model may also be trained on a sufficient number of non-event analogous examples to enable it to predict non-events with high probability for input sets that lead to non-events. In other embodiments, the model predictions may include additional predictions or include only one or two of the three probabilities described above. For example, in other embodiments, instead of predicting the probabilities of collision events, near-collision events, and non-hazardous events, the model may be configured to predict only two probabilities: “risky events” and “non-hazardous events.”
[0469]
[0517] After the probabilities of various predicted events (e.g., collision events, near-collision events, non-hazardous events, etc.) have been determined, the processing unit 210 can then determine a metric (e.g., a risk score) based on these probabilities. In some embodiments, the metric determined by the processing unit 210 may indicate whether the risk factors collectively result in a hazardous or non-hazardous situation. In some embodiments, the metric may be a risk score indicating the degree of risk. For example, a higher risk score may indicate higher risk, and a lower risk may indicate lower risk.
[0470]
[0518] In some embodiments, the metric determined by the processing unit 210 based on the probabilities of various events may be a weighted sum of the scores (probabilities) of each different event. In other embodiments, the metric may be calculated as the sum of weighted scores (probabilities).
[0471]
[0519] Figure 29 shows an example of risk score calculation based on model predictions. In the example above, the model in processing unit 210 predicts a 6% probability of a collision within T seconds, a 22.8% probability of a near collision within T seconds, and a 71.2% probability of a non-hazardous event within T seconds. If a weight of 1.0 is applied to the collision prediction, a weight of 0.5 is applied to the near collision prediction, and a weight of 0.0 is applied to the non-hazardous event prediction, processing unit 210 may calculate the risk score as the sum of the weighted predictions as follows: 1.0 * 6% + 0.5 * 22.8% + 0 * 71.2% = 17.4.
[0472]
[0520] In the example at the bottom of Figure 29, the model of processing unit 210 predicts that the probability of a collision within T seconds is 73.4%, the probability of a near collision at T seconds is 21%, and the probability of a non-hazardous event at T seconds is 5.6%. If a weight of 1.0 is applied to the collision prediction, a weight of 0.5 is applied to the near collision prediction, and a weight of 0.0 is applied to the non-hazardous event prediction, processing unit 210 may calculate the risk score as the sum of the weighted predictions as follows: 1.0 * 73.4% + 0.5 * 21% + 0 * 5.6% = 83.9.
[0473]
[0521] In some embodiments, the processing unit 210 may be configured to compare the risk score with a threshold of 1 or more. If the threshold is met (e.g., exceeded), the processing unit 210 may then generate a control signal to generate an alarm and / or activate a vehicle control device (e.g., automatically apply the brakes, turn on the exterior lights, sound the horn, etc.). Following the two examples above in Figure 29, the risk scores are 17.4 in the upper example and 83.9 in the lower example. If the threshold for applying a warning and / or vehicle control device is set to 70, the processing unit 210 determines that the threshold is not met in the upper example (because the risk score of 17.4 is less than the threshold of 70). Therefore, in the upper example, the processing unit 210 does not generate a control signal to generate an alarm or activate a vehicle control device. On the other hand, the processing unit 210 determines that the threshold is met in the lower example (because the risk score of 83.9 is greater than the threshold of 70). In such cases, the processing unit 210 generates control signals to trigger an alarm and / or to operate the vehicle control device (e.g., automatically apply the brakes, decelerate, sound the horn, activate the exterior lights, provide haptic feedback, or any combination thereof).
[0474]
[0522] Alternatively or additionally, if the risk score is above a threshold, the processing unit 210 may generate a control signal to notify the vehicle manager or vehicle management system that the driver has performed inappropriate driving.
[0475]
[0523] In some embodiments, if the risk score falls below a threshold or another threshold, the processing unit 210 may not generate a control signal and / or activate a speaker to provide the driver with verbal praise and / or generate a control signal to notify the vehicle manager or vehicle management system that the driver has driven well.
[0476]
[0524] In some embodiments, the model may learn what is considered good driving behavior. In such cases, the model can detect specific events and determine what is considered good driving behavior for the detected event. The processing unit 210 may also compare actual driving behavior (e.g., by acquiring vehicle control device signals or analyzing in-cabin images) with good driving behavior to see how well the driver is adhering to good driving behavior. In some embodiments, the results of the comparison may be transmitted to a vehicle management system. In such cases, the vehicle management system may use such information to improve the driver's driving skills and / or to commend good drivers. Alternatively or additionally, the processing unit 210 may generate control signals that provide the driver with a message of commendation (when the driver is demonstrating good driving behavior) and a message of guidance (when the driver is not demonstrating good driving behavior). As a non-limiting example, good driving behavior includes slowing down in areas with heavy traffic, slowing down when there are pedestrians, and slowing down when approaching an intersection.
[0477]
[0525] Figure 30 shows an example of a model output based on multiple time-series inputs, where the model output exhibits a high "collision" state. The left column of the figure, in particular, shows examples of various inputs fed into the model. Each input example includes various parameters and their respective values for the corresponding time point within a preceding period (6 seconds in the example), as shown and explained with reference to Figure 28. The middle column of Figure 30 shows the intermediate representation learned by the model. Specifically, a spleness map of the neural network model is displayed, visualizing which parts of the input the model is focusing on. The model processes the input to determine the probability of each of three events: collision events, near-collision events, and non-hazardous events (shown in the right column of the figure). In all examples shown in Figure 30, the "collision event" has the highest probability (compared to the probabilities of the other two events).
[0478]
[0526] Figure 31 shows an example of a model output based on multiple time-series inputs, where the model output exhibits a high "near collision" state. The left column of the figure, in particular, shows examples of various inputs fed into the model. Each example input includes various parameters and their respective values for the corresponding time point within a preceding period (6 seconds in the example), as shown and explained with reference to Figure 28. The middle column of Figure 31 shows the intermediate representation learned by the model. In particular, a spleness map of the neural network model is displayed, visualizing which parts of the input the model is focusing on. The model processes the input to determine the probability of each of three events: collision events, near collision events, and non-hazardous events (shown in the right column of the figure). In all examples shown in Figure 31, the "near collision event" has the highest probability (compared to the probabilities of the other two events).
[0479]
[0527] Figure 32 shows an example of a model output based on multiple time-series inputs, where the model output shows a high "non-event" state. In particular, the left figure shows an example of the inputs fed into the model. Each example input includes various parameters and their respective values for the corresponding time point within a preceding period (6 seconds in the example), as shown and explained with reference to Figure 28. The center figure of Figure 32 shows the intermediate representation learned by the model. In particular, a spleness map of the neural network model is displayed, visualizing which parts of the input the model is focusing on. The model processes the input and determines the probability of each of the three events: collision events, near-collision events, and non-hazardous events (shown in the right column of the figure). In the example in Figure 32, the "non-hazardous event" has the highest probability (compared to the probabilities of the other two events).
[0480]
[0528] As shown in the examples in Figures 30-32, each set of inputs can form a certain pattern. In some cases, the model may be a neural network that can be trained to make decisions based on such patterns. For example, a neural network model may be trained using a set of inputs labeled as “collision” events (as in the example shown in Figure 30). Alternatively, a neural network model may be trained using a set of inputs labeled as “near-collision” events (as in the example shown in Figure 31). Furthermore, a neural network may be trained using a set of inputs labeled as “non-risk” events (as in the example shown in Figure 32). Through training, the model learns to pay different attention to different sets of inputs. For example, referring to the middle figures in Figures 30-32, we can see that the model pays different attention to each event type. In some embodiments, the model may be trained using a supervised approach. In such an approach, the input signals over time are encoded as a single-channel image and input to a CNN model. This model is trained to predict the probability of future events occurring at a given time interval, such as non-risk events, near-collision events, and collision events. To train the model, these prediction targets may be created by a human labeler using a “supervised” training approach. In some embodiments, the model may be trained to predict events that occur within T seconds (e.g., 1.5 seconds, 2 seconds, 3 seconds, 4 seconds, etc.). It should be noted that the model training may take any form and is not limited to supervised learning in which the model learns directly from provided labels. For example, in other embodiments, the model may be trained using reinforcement learning in which the agent learns from its environment through trial and error.
[0481]
[0529] As described above, the model described herein is advantageous because it can reduce both false positive cases (e.g., cases where a warning is generated for a non-dangerous situation) and false negative cases (e.g., cases where a warning is not generated when a dangerous situation is present). Figures 33-37 illustrate examples of these scenarios.
[0482]
[0530] Figures 33A–33K show a series of external images and their corresponding internal images, particularly illustrating the problems that can be addressed by the technology in Figure 25. This example specifically illustrates how issuing overly sensitive warnings without considering the surrounding environment hinders the purpose of assisting the driver. Figures 33A–33C are snapshots corresponding to the scenario (frames 1, 5, and 10 from a 25fps (frames / second) video). In the three illustrated frames, before the "object holding" event begins, the driver is (1) keeping their eyes on the road, (2) not speeding, and (3) maintaining a safe distance from the vehicle in front. During the "object holding" event, the driver maintains a safe distance from the vehicle in front, is not speeding, and is looking at the road as is evident from the snapshots (see Figures 33D–33G). In this example, the distraction module detected the "holding an object" pose and generated a warning because it assumed the "holding an object" pose was a distraction that could lead to a dangerous situation. However, considering the context, the "holding an object" event alone is not particularly dangerous. Therefore, generating a warning here would be a false positive. As seen in the next frame (see Figures 33H-33K), the driver appears irritated by the warning and lowers the visor that obstructs the camera. Warnings in such unimportant scenarios reduce the relevance of the warning and decrease the driver's trust in it. As this example shows, the distraction signal was triggered because the detection score for "holding an object" exceeded a predefined threshold. However, this is by no means a dangerous scenario, as the driver was looking at the road, not speeding, and maintaining a safe distance from the vehicle in front. The technique in Figure 25 is advantageous in that, by considering the combination of detected inputs, the model recognizes that the "holding an object" event alone is not dangerous when viewed in context with other inputs, and the processing unit 210 appropriately postpones generating a control signal to issue a warning.
[0483]
[0531] Figures 34A–34E show a series of external images and corresponding internal images, particularly illustrating problems that can be addressed by the technique in Figure 25. In this scenario, the warning was not triggered before the near collision. This is because, despite the driver being repeatedly distracted, the individual distraction segments did not exceed the distraction duration threshold. Thus, a false negative problem arises in this example. Figures 34A–34B are several relevant snapshots corresponding to the scenario (frames 158 and 166 from 25fps video). In these frames, the driver can be seen being repeatedly distracted for relatively short periods (looking left, talking). Ultimately, the vehicle gets too close to the preceding vehicle, resulting in a near collision scenario (frames 214, 212, and 225 from 25fps video) (see Figures 34C–34E). The combination of repeated distractions while approaching a preceding vehicle is a dangerous scenario, regardless of the length of the driver's distractions. On the other hand, the new model in Figure 25 was able to detect this “near collision” event and provided a high risk score of 65.7 for this predicted near collision event. The new model was able to do this because it had access to both the distraction stream (the first time series of input) and the TTH stream (the second time series of input). This is advantageous because the model in Figure 25 can detect situations where, even if the risk factors do not individually exceed the threshold, they combine to create a very dangerous situation.
[0484]
[0532] Figures 35A–35D show a series of external images and their corresponding internal images, particularly illustrating the problems that can be addressed by the technique in Figure 25. In a scenario similar to Example 2, the driver is temporarily distracted (looking to the left) while approaching a vehicle stopped in front at an intersection, as seen in the snapshots (frames 90, 110, and 125 from the 25fps video) (see Figures 35A–35C). Although the period of distraction is relatively short, it is a fairly dangerous scenario that ultimately leads to a collision. The frame below shows the moment of collision (frame 134 from the 25fps video) (see Figure 35D).
[0485]
[0533] Figures 36A–36C show a series of external images and their corresponding internal images, particularly illustrating problems that can be addressed by the technique in Figure 25. In this example, as seen in the snapshots (frames 29 and 65 from the 25fps video), the driver was occasionally distracted (using a cell phone and looking down), and the vehicle in front was backing up (see Figures 36A–36B). However, the existing system did not trigger a warning. Therefore, this case is also an example of a false negative. Using the model in Figure 25, processing unit 210 predicted that this event was a "collision" event and assigned a high-risk score of 83.9 based on processing the raw distraction stream (first time series of input) and the TTH / TTC stream (second time series of input). The frame shown in Figure 36C shows the moment of collision (frame 126 from the 25fps video).
[0486]
[0534] Figures 37A–37C show a series of external images and their corresponding internal images, particularly illustrating the problems that can be addressed by the technique in Figure 25. In this scenario, a warning was generated, but the warning was irrelevant because, although the driver was distracted (looking down), they were driving on an open road without speeding for a long time, as seen in the snapshots (frames 49, 143, and 213 from the 25fps video) (see Figures 37A–37C). On the other hand, the model in Figure 25 predicted this event as a non-hazardous event (calm / insignificant) and calculated a low-risk score of 17.4 using cues from distraction, speeding, and the TTH / TTC stream. Therefore, this model was able to prevent the problem of false positives.
[0487]
[0535] It should be noted that the model in Figure 25 may be implemented in the device 200 in various ways. For example, in some embodiments, a trigger-based approach may be used. In such cases, the model is executed only when one or more conditions are met (e.g., when distraction reaches a threshold). When triggered, the model is executed for the entire event period. As another example, in other embodiments, a continuous approach may be used. In such an approach, the model is executed continuously for the data of the most recent X seconds, based on a slide window.
[0488]
[0536] In the examples above, the model is described as providing model predictions that include predicted “collision” events, predicted “near-collision” events, predicted “non-risk” events, and the probability of each of them (first output). However, the first output is not limited to these examples and may be any driving condition related to risk. For example, in other embodiments, the first output may also include a metric indicating the severity of the collision (e.g., high-speed collision or minor contact), and / or characteristics of the risk event such as the presence or absence of sudden braking, sudden steering, or “good” driving behavior.
[0489]
[0537] In the example above, the probabilities of various predicted events (first output) are used by the processing unit 210 to calculate a risk score (second output) based on the sum of the weighted probabilities. In other embodiments, other techniques may be used to calculate the risk score. For example, an arbitrary continuous function may be used to map the direct model output (e.g., the probabilities of various events) to a single score.
[0490]
[0538] In a further embodiment, the processing unit 210 may identify the change in the risk score (second output) over time as a third output. This third input is useful for indicating whether the situation is becoming more dangerous or safer, and to what extent, at a particular moment.
[0491]
[0539] In further embodiments, the processing unit 210 may determine a relevant or based fourth output indicating the optimal action to be taken at any given time, particularly in dangerous scenarios. For example, in some embodiments, the processing unit 210 may determine a course of action for a "good driver" based on a set of inputs indicating a particular dangerous situation. In such cases, the model can better assess new risks by not only processing external risk factors but also evaluating how the current driver's actions compare to ideal driving behavior. These good actions may include slowing down, changing gaze, or changing lanes.
[0492]
[0540] In some embodiments, the processing unit 210 (e.g., a model within it) may be configured to process first image data from a first camera 202 and second image data from a second camera 204 to produce one or more outputs (e.g., probabilities of different event categories), where the first image data is generated during a first time window and the second image data is generated during a second time window that is at least partially identical to the first time window. For example, the first image data may be generated from t=2 to t=8 seconds and the second image data may be generated from t=3 to t=10 seconds. In other embodiments, the first and second time windows may be entirely identical. For example, the first image data may be generated from t=1 to t=5 seconds and the second image data may be generated from t=1 to t=5 seconds. In further embodiments, the first and second time windows may be non-identical. For example, a series of close-proximity distractions captured in the second image data during time window A may trigger a state of increased external risk captured in the first image data during a later time window B.
[0493]
[0541] Furthermore, in some embodiments, the processing unit 210 (e.g., a model within it) may be configured to predict future distractions based on the driver's past behavior. For example, even if a momentary distraction has ended, the processing unit 210 can predict that a future distraction (e.g., within the next few minutes) is likely to occur based on past behavior detected by the processing unit 210 (e.g., distraction events that occurred frequently in a past time window). Distraction events may be considered frequent if multiple events occur in close proximity in time (e.g., events occurring within a specific period, such as within 2 minutes, 1 minute, 30 seconds, 15 se...
Claims
1. The device comprises a first sensor configured to provide a first input relating to the external environment of a vehicle, a second sensor configured to provide a second input relating to the operation of the vehicle, and a processing unit configured to receive a first input from the first sensor and a second input from the second sensor, wherein the processing unit includes a first-stage processing system and a second-stage processing system, the first-stage processing system configured to receive a first input from the first sensor, receive a second input from the second sensor, process the first input to obtain a first time series of information, and process the second input to obtain a second time series of information, the second-stage processing system comprises a neural network model configured to receive the first time series of information and the second time series of information in parallel, the neural network configured to process the first time series and the second time series to determine the probability of predictive events relating to the operation of the vehicle. The aforementioned predicted event is for a future time at least one second after the current time. The first input contains first raw data, and the first time-series information is of lower dimensionality or less complex compared to the first input.
2. The apparatus according to claim 1, wherein the first time series information indicates a first risk factor, and the second time series information indicates a second risk factor.
3. The apparatus according to claim 1, wherein the processing unit is configured to package the first time series and the second time series into a data structure for supplying to a neural network model.
4. The apparatus according to claim 1, wherein the first time series shows the external state of the vehicle at each different point in time, and the second time series shows the state of the driver and / or the state of the vehicle at each different point in time.
5. The apparatus according to claim 1, wherein the probability of the predicted event is the first probability of a first predicted event, the processing unit is configured to determine the second probability of a second predicted event, the first and second predicted events relate to the operation of a vehicle, and the processing unit is configured to calculate a risk score based on the first probability of the first predicted event and the second probability of the second predicted event.
6. The apparatus according to claim 5, wherein the first predicted event is a collision event, the second predicted event is a non-hazardous event, and the processing unit is configured to calculate the risk score based on a first probability of a collision event and a second probability of a non-hazardous event.
7. The apparatus according to claim 5, wherein the processing unit is configured to calculate the risk score by applying a first weighting to the first probability to obtain a first weighted probability, applying a second weighting to the second probability to obtain a second weighted probability, and adding the first weighted probability and the second weighted probability.
8. The apparatus according to claim 5, wherein the processing unit is configured to determine a third probability of a third predicted event, and the processing unit is configured to calculate the risk score based on a first probability of a first predicted event, a second probability of a second predicted event, and a third probability of a third predicted event.
9. The apparatus according to claim 8, wherein the first predicted event is a collision event, the second predicted event is a near collision event, and the third predicted event is a non-hazardous event, and the processing unit is configured to calculate the risk score based on the first probability of the collision event, the second probability of the near collision event, and the third probability of the non-hazardous event.
10. The apparatus according to claim 8, wherein the processing unit is configured to calculate the risk score by applying a first weight to the first probability to obtain a first weighted probability, applying a second weight to the second probability to obtain a second weighted probability, applying a third weight to the third probability to obtain a third weighted probability, and adding the first weighted probability, the second weighted probability, and the third weighted probability.
11. The apparatus according to claim 1, wherein the first input and the second input are acquired in the past T seconds, and the processing unit is configured to process the first input and the second input acquired in the past T seconds to determine the probability of the predicted event, where T is at least 3 seconds.
12. The apparatus according to claim 1, wherein the processing unit is configured to calculate a first risk score for a first time point based on probability, and the processing unit is also configured to calculate a second risk score for a second time point and to identify the difference between the first risk score and the second risk score, the difference indicating whether the dangerous situation is escalating or fading.
13. The apparatus according to claim 1, wherein the processing unit is configured to determine a risk score based on the probability of a predicted event.
14. The processing unit is configured to generate a control signal based on the risk score, and this control signal is used to operate the device, the device is A speaker that emits a warning, A display or light-emitting device that provides visual signals. Haptic feedback device, Collision avoidance system, or The apparatus according to claim 13, comprising a vehicle control device for a vehicle.
15. The apparatus according to claim 1, wherein the first time series and the second time series each include any two or more of the following: distance to a preceding vehicle, distance to a stop line at an intersection, vehicle speed, time to collision, time to intersection violation, estimated braking distance, information on road conditions, information on special zones, information on the environment, information on traffic conditions, time, information on visibility conditions, information on identified objects, position of objects, direction of movement of objects, speed of objects, boundary boxes, vehicle operating parameters, information on the driver's state, information on the driver's history, continuous driving time, proximity to meal times, information on accident history, and voice information.
16. The first sensor includes, and / or includes a camera, lidar, radar, or any combination thereof, configured to sense the environment outside the vehicle. The apparatus according to claim 1, wherein the second sensor includes a camera configured to view the driver of the vehicle.
17. The apparatus according to claim 1, wherein the first input includes a first image, the second input includes a second image, and the first stage processing system is configured to receive the first image and the second image, process the first image to obtain first time-series information, and process the second image to obtain second time-series information.