Method for detecting passenger in vehicle by using artificial intelligence
Cameras and AI models in vehicles accurately detect passenger seating configurations, reducing errors and ensuring safe travel by preventing unsafe vehicle operation when passengers are not properly seated.
Patent Information
- Application Number
- JP2025060918
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-11
- Filing Date
- 2025-04-02
- Publication Date
- 2025-10-24
AI Technical Summary
Existing vehicles, including autonomous and human-driven ones, struggle with accurately detecting whether passengers are properly seated, leading to false positives and negatives due to weight sensors and seat belt sensors, which can result in unsafe travel conditions.
Implementing cameras that capture images and utilize artificial intelligence models to generate passenger data, determining passenger locations and seating configurations, and triggering appropriate actions if seating criteria are not met.
Reduces false positives and negatives in passenger seating detection, enhancing driving safety by ensuring passengers are correctly seated before the vehicle moves.
Smart Images

Figure 2025161757000001_ABST
Abstract
Description
[Technical Field]
[0001] This specification relates generally to vehicles, and more particularly to detecting passengers in vehicles using artificial intelligence. [Background technology]
[0002] Vehicles—whether autonomous vehicles (AVs) (including fully autonomous or partially autonomous), driven by a human driver, or other types of vehicles—often operate by sensing their environment using various sensors (e.g., radar, optical, acoustic, humidity, etc.). This environment may include other objects within the environment, some of which may be mobile. Such objects may include other vehicles, bicyclists, pedestrians, animals, etc. [Brief explanation of the drawings]
[0003] The present disclosure is presented by way of example, and not by way of limitation, and may be more fully understood by reference to the following detailed description when considered in conjunction with the figures in which:
[0004] [Figure 1] FIG. 1 illustrates a block diagram of an example autonomous vehicle (AV) that can use artificial intelligence (AI) to detect passengers within the vehicle, according to some implementations of the present disclosure. [Figure 2] FIG. 2 illustrates a flow diagram of an example method for using AI to detect passengers in a vehicle, according to some implementations of the present disclosure. [Figure 3] FIG. 3 illustrates a block diagram of an exemplary AI subsystem according to some implementations of the present disclosure. [Figure 4] FIG. 4 illustrates a top-down cutaway view of an example vehicle that uses AI to detect passengers within the vehicle, according to some implementations of the present disclosure. [Figure 5] FIG. 5 is a front view of a visual alert for detecting passengers in a vehicle using AI, according to some implementations of the present disclosure. [Figure 6]FIG. 6 illustrates a block diagram of an example computing device capable of using AI to detect passengers in a vehicle, according to some implementations of the present disclosure. Summary of the Invention
[0005] In one implementation, a method for detecting passengers in a vehicle using artificial intelligence (AI) is disclosed. The method includes obtaining one or more images captured by one or more cameras of the vehicle. The method includes generating passenger data indicating locations of one or more passengers in the vehicle using one or more AI models and the one or more images. The method includes generating vehicle area data indicating one or more areas of the vehicle in which the one or more passengers are located based on the passenger data. The method includes determining whether at least one passenger seating configuration criterion is met based on the passenger data and the vehicle area data. The method includes causing the vehicle to perform an action associated with the passenger seating configuration in response to determining that the at least one passenger seating configuration criterion is met.
[0006] In another implementation, a system for detecting passengers in a vehicle using AI is disclosed. The system includes a memory and a processing unit coupled to the memory. The processing unit is configured to perform one or more operations. The one or more operations include obtaining one or more images captured by one or more cameras of the vehicle. The one or more operations include generating passenger data indicating locations of one or more passengers in the vehicle using one or more AI models and the one or more images. The one or more operations include generating vehicle area data indicating one or more areas of the vehicle in which the one or more passengers are located based on the passenger data. The one or more operations include determining whether at least one passenger seating configuration criterion is met based on the passenger data and the vehicle area data. The one or more operations include causing the vehicle to perform an action associated with a passenger seating configuration in the vehicle in response to determining that the at least one passenger seating configuration criterion is met.
[0007] Another aspect of the present disclosure includes a non-transitory computer-readable medium storing instructions that, when executed by one or more processing devices, cause the one or more processing devices to perform one or more operations. The one or more operations include obtaining one or more images captured by one or more cameras of a vehicle. The one or more operations include generating passenger data indicating locations of one or more passengers of the vehicle using one or more AI models and the one or more images. The one or more operations include generating vehicle area data indicating one or more areas of the vehicle in which the one or more passengers are located based on the passenger data. The one or more operations include determining whether at least one passenger seating configuration criterion is met based on the passenger data and the vehicle area data. The one or more operations include causing the vehicle to perform an action associated with a passenger seating configuration in the vehicle in response to determining that the at least one passenger seating configuration criterion is met. DETAILED DESCRIPTION OF THE INVENTION
[0008] Vehicles, such as autonomous vehicles (AVs) (including vehicles deploying various driver assistance features) or vehicles driven by a human driver, can carry one or more passengers from a starting location to a destination. It is often unsafe for the vehicle to travel if one or more passengers are not properly seated. Some vehicles have weight sensors that can detect whether a passenger is seated in a particular seat of the vehicle. Some vehicles have sensors that can detect whether a particular seat belt is fastened. Using a combination of these sensors, a vehicle may be able to detect whether a passenger is seated in a seat of the vehicle without the passenger's seat belt being fastened.
[0009] One drawback of vehicles that use weight sensors and seat belt sensors is that they do not always accurately detect whether a passenger is properly seated in the vehicle. For example, the weight sensor may detect a heavy object in the seat, causing the vehicle to incorrectly determine that a person is seated in that seat (a false positive). In another example, the weight sensor may fail to detect when a very small person (e.g., a child) is seated in the vehicle seat (a false negative). In yet another example, the weight sensor may not detect when two people are sitting in the same seat in the vehicle, or if passengers are positioned in other unsafe seating configurations.
[0010] Aspects and implementations of the present disclosure address these and other challenges of existing vehicles. In one implementation, a vehicle may include one or more cameras that may be located at various locations on the vehicle (e.g., inside the vehicle, mounted on the exterior of the vehicle). The one or more cameras may capture images of various locations associated with the vehicle (e.g., inside the vehicle, outside the vehicle, etc.). The vehicle may provide the captured images as input to one or more artificial intelligence (AI) models. The AI models may generate one or more outputs, including passenger data, based on the one or more captured images. The passenger data may indicate whether a passenger is present in the image and the passenger's location. A vehicle area subsystem of the vehicle may acquire the passenger data and generate vehicle area data. The vehicle area data may indicate an area of the vehicle in which a detected passenger is located. The vehicle area data may indicate other information about the passenger (e.g., whether the passenger is a child, whether the passenger is smoking, etc.). A passenger seating subsystem of the vehicle may acquire the passenger position data and / or vehicle area data and determine whether a passenger seating configuration criterion has been met based on the passenger position data and / or vehicle area data. The passenger seating configuration criteria may be based on one or more conditions, such as multiple passengers being located in the same seat of the vehicle, passengers being located in an area of the vehicle that is not a seat, or all passengers being children. In response to the passenger seating subsystem of the vehicle determining that the passenger seating configuration criteria has been met, the passenger seating subsystem may cause the vehicle to perform one or more actions associated with the passenger seating configuration of the vehicle (e.g., generating an alert to notify passengers of the vehicle, preventing the vehicle from driving, etc.).
[0011] Advantages of the disclosed technology and system include, but are not limited to, reducing errors in vehicles detecting whether passengers are seated in a safe seating configuration. By using AI models and other computing processes to detect passenger locations, determine whether passengers are seated in a proper seating configuration, and determine whether passengers comply with other seating practices, the false positives and false negatives discussed above are reduced, which results in improved driving performance. Additionally, if the vehicle is an AV, the AV may automatically respond to one or more passengers not seated in a proper seating configuration, for example, by preventing the AV from moving or by stopping the AV.
[0012] In some implementations, the vehicle may include an AV. In those instances where the implementation description refers to an AV, it should be understood that similar technologies may be used in various driver assistance systems that do not rise to the level of a fully autonomous driving system. More specifically, the disclosed technology may be used in Society of Automotive Engineers (SAE) Level 2 driver assistance systems that implement steering, braking, acceleration, lane centering, adaptive cruise control, and other driver support. Similarly, the disclosed technology may be used in SAE Level 3 driver assistance systems that are capable of autonomous driving under limited (e.g., highway) conditions. Such systems use rapid and accurate detection and tracking of moving objects to alert the driver of approaching objects and allow the driver to make final driving decisions (e.g., in SAE Level 2 systems) or certain driving decisions such as deceleration and lane changes (e.g., in SAE Level 3 systems) without requiring driver feedback. Furthermore, while the implementation description refers to an AV, many of the subsystems, processes, and technologies are applicable to non-AV vehicles, such as human-operated vehicles. A vehicle may include a motor vehicle (such as a car, truck, bus, motorcycle, all-terrain vehicle, recreational vehicle, any specialized agricultural or construction vehicle, etc.), an aircraft (such as an airplane, helicopter, drone, etc.), a marine vessel (such as a ship, boat, yacht, submarine, etc.), or any other self-propelled vehicle (e.g., a robot, a factory or warehouse robotic vehicle, a sidewalk delivery robotic vehicle, etc.).
[0013] 1 is a diagram illustrating components of an example AV 100 capable of using AI to detect passengers within the vehicle, according to some implementations of the present disclosure. The AV 100 may include a vehicle capable of operating in an autonomous mode (with no or reduced human input).
[0014] The environment 101 around the AV 100 (sometimes referred to as the "driving environment") may include any objects (moving or non-moving) located outside the AV 100, such as roads, buildings, trees, bushes, sidewalks, bridges, mountains, other vehicles, pedestrians, animals, etc. The driving environment 101 may be urban, suburban, rural, etc. In some implementations, the driving environment 101 may be an off-road environment (e.g., a cultivated field or other agricultural land). In some implementations, the driving environment may be an indoor environment (e.g., an industrial plant environment, a shipping warehouse, a hazardous area of a building). In some implementations, the driving environment 101 may be substantially flat, with various objects moving parallel to the surface (e.g., parallel to the surface of the Earth). In other implementations, the driving environment 101 may be three-dimensional and include objects capable of moving along all three directions (e.g., balloons, leaves, etc.). Hereinafter, the term "driving environment" should be understood to include all environments in which autonomous movement of a self-driving vehicle may occur. For example, the "operating environment" may include any possible flight environment of an aircraft or a marine environment of a marine vessel. Objects in the operating environment 101 may be located at any distance from the AV 100, from a close distance of a few feet (or less) to several miles (or more).
[0015] As described herein, in a semi-autonomous or partially autonomous driving mode, the AV 100 assists with one or more driving maneuvers (e.g., steering, braking, and / or accelerating to perform lane centering, adaptive cruise control, an advanced driver assistance system (ADAS), or emergency braking), but a human driver is expected to maintain situational awareness of the AV 100's surroundings and supervise the assisted driving maneuvers. Here, while the AV 100 may perform all driving tasks in certain situations, the human driver is expected to remain responsible for assuming control as needed.
[0016] For simplicity and brevity, various systems and methods are described below in conjunction with AV100, although similar technologies may be used in various driver assistance systems that fall short of a fully autonomous driving system. In the United States, the SAE defines different levels of automated driving operation to indicate how much or how little control the vehicle has over the driving; however, different organizations in the United States, or elsewhere, may classify the levels differently. More specifically, the disclosed systems and methods may be used in SAE Level 2 (L2) driver assistance systems that implement steering, braking, acceleration, lane centering, adaptive cruise control, and other driver support. The disclosed systems and methods may be used in SAE Level 3 (L3) driver assistance systems that are capable of autonomous driving under limited (e.g., highway) conditions. Similarly, the disclosed systems and methods may be used in vehicles using SAE Level 4 (L4) automated driving systems that operate autonomously under most normal driving conditions and require only occasional attention from a human operator. In all such driver assistance systems, accurate lane estimation may be performed automatically without driver input or control (e.g., while the vehicle is moving), resulting in improved reliability of vehicle positioning and navigation, and overall safety of autonomous, semi-autonomous, and other driver assistance systems. As noted above, in addition to the manner in which the SAE classifies levels of autonomous driving operation, other organizations in the United States or other countries may classify levels of autonomous driving operation differently. Without limitation, the systems and methods disclosed herein may be used in driver assistance systems defined by the levels of autonomous driving operation of these other organizations.
[0017] An example AV 100 may include a sensing system 110. The sensing system 110 may include various electromagnetic (e.g., optical) and non-electromagnetic (e.g., audio) sensing subsystems and / or devices. The sensing system 110 may include one or more LIDARs 112, which may be laser-based units capable of determining the distance to and velocity of objects within the driving environment 101. The LIDARs 112 may include one or more light sources that generate and emit signals and one or more detectors of signals reflected back from the objects. In some implementations, the LIDARs 112 may perform 360-degree scans in the horizontal direction. In some implementations, the LIDARs 112 may be capable of spatial scanning along both the horizontal and vertical directions. In some implementations, the field of view may be up to 90 degrees vertically (e.g., at least a portion of the area above the horizon is scanned by the radar signal). In some implementations, the field of view may be spherical (consisting of two hemispheres).
[0018] The sensing system 110 may include one or more radars 113, which may be any system that utilizes radio or microwave frequency signals to detect objects within the operating environment 101 of the AV 100. The radars 113 may be configured to sense both the spatial location of objects (including their spatial dimensions) and their velocity (e.g., using Doppler shift techniques). Hereinafter, "velocity" refers to both how fast an object is moving (object speed) and the direction of the object's motion. Each of the lidar 112 and radar 113 may include a coherent sensor, such as a frequency-modulated continuous wave (FMCW) lidar or radar sensor. For example, the radar 113 may use heterodyne detection for velocity determination. In some implementations, the functionality of ToF and coherent radar is combined into a radar unit that can simultaneously determine both the distance to a reflecting object and the radial velocity of the reflecting object. Such a unit may be configured to operate in a non-coherent sensing mode (ToF mode) and / or a coherent sensing mode (e.g., a mode using heterodyne detection), or both modes simultaneously. In some implementations, multiple lidars 112 or radars 113 may be mounted on the AV 100. The sensing system 110 may further include one or more sonars 114, which in some implementations may be ultrasonic sonars.
[0019] In some implementations, the sensing system 110 may further include one or more cameras 115 configured to capture images. The cameras 115 may include one or more external cameras 116. The external cameras 116 may be mounted on the AV 100 and positioned to capture images of the driving environment 101. The cameras 115 may include one or more internal cameras 117. The internal cameras 117 may be mounted on the AV 100 and positioned to capture images of an internal portion of the AV 100. The internal portion of the AV 100 may include a portion of the AV 100 within the body of the AV 100 where one or more passengers can move around, sit, stand, stow, or perform other activities. In some implementations, the cameras 115 (either external 116 or internal 117) may be mounted on an external portion of the AV 100 (e.g., on the roof, on the side of the vehicle, on the rear of the vehicle, etc.). In some implementations, the camera 115 (either external 116 or internal 117) may be mounted on an interior portion of the AV 100 (e.g., the underside of the roof, a wall, the interior portion of the windshield, etc.).
[0020] In one or more implementations, the images captured by the camera 115 may be a two-dimensional projection of an area from the perspective of the camera's 115 lens (e.g., a portion of the driving environment 101, a portion of the interior of the AV 100) onto a projection surface (flat or non-flat) of the camera 115. Some of the cameras 115 of the sensing system 110 may be video cameras configured to capture a continuous (or quasi-continuous) stream of images. The sensing system 110 may also include one or more infrared (IR) sensors 119.
[0021] AV 100 may include data processing system 120. Data processing system 120 may include one or more computers or computing devices. Data processing system 120 may include hardware or software that receives data from sensing system 110, processes the received data, and determines how AV 100 should operate within driving environment 101. In some implementations, data processing system 120 may receive non-electromagnetic data, such as audio data (e.g., ultrasonic sensor data or data from a microphone picking up an emergency vehicle siren), temperature sensor data, humidity sensor data, pressure sensor data, weather data (e.g., wind speed and direction, precipitation data), etc.
[0022] Data processing system 120 may include a positioning subsystem 122. Positioning subsystem 122 uses positioning data (e.g., global positioning system (GPS) data, inertial measurement unit (IMU) data, or other positioning data) to help accurately determine the location of AV 100. Data processing system 120 may include a mapping subsystem 124. Mapping subsystem 124 may obtain or calculate map data (e.g., GPS data, geographic information system (GIS) data, satellite data, traffic data, or other data) that may provide map information to AV 100. In some implementations, AV 100 may receive positioning or map data from one or more servers via a data network (e.g., a cellular network). In this manner, AV 100 may store temporary positioning or map data, for example, data related to the geographic region in which AV 100 is located.
[0023] Data processing system 120 may include passenger detection subsystem 130. Passenger detection subsystem 130 may detect one or more passengers on AV 100, determine whether the one or more passengers are seated in an appropriate configuration, determine other information about the passengers, and generate outputs usable by AV control system (AVCS) 140 and other systems of AV 100, as discussed herein.
[0024] Passenger detection subsystem 130 may include a position subsystem 132. The position subsystem may determine one or more positions of one or more passengers of AV 100, as discussed herein. Passenger detection subsystem 130 may include a vehicle area subsystem 134. The vehicle area subsystem 134 may determine one or more areas of AV 100 in which one or more passengers are located, as discussed herein. Passenger detection subsystem 130 may include a passenger seating subsystem 136. Passenger seating subsystem 136 may determine whether passenger seating configuration criteria are met (e.g., passengers are not in an appropriate seating configuration) based on data generated or output by position subsystem 132 or vehicle area subsystem 134, and if so, passenger seating subsystem 136 may transmit data to AVCS 140 or other systems of 100, as discussed herein. In some implementations, passenger detection subsystem 130 may include an AI subsystem 138. The AI subsystem 138 may include one or more AI models that the position subsystem 132, the vehicle area subsystem 134, or the passenger seat subsystem 136 may use to perform various actions, as discussed herein.
[0025] Data processed or generated by the data processing system 120, including the occupant detection subsystem 130, may be used by the AVCS 140 of the AV 100. The AVCS 140 may include one or more algorithms that control how the AV 100 should behave in various driving situations and environments. For example, the AVCS 140 may include a navigation system for determining a global driving path to a destination. The AVCS 140 may also include a driving path selection system for selecting a particular path through the immediate driving environment 101, which may include selecting a lane, navigating traffic congestion, selecting a location to make a U-turn, selecting a trajectory for a parking maneuver, etc. The AVCS 140 may also include an obstacle avoidance system for safely avoiding various objects or other obstacles (e.g., stones, stranded vehicles, pedestrians crossing the road ignoring traffic rules or signals, etc.) within the driving environment 101 of the AV 100. The obstacle avoidance system may be configured to evaluate the size of an obstacle and the trajectory of the obstacle (if the obstacle is moving) and select an optimal driving strategy (e.g., braking, steering, accelerating, etc.) to avoid the obstacle. AVCS 140 may also include a system that prevents AV 100 from moving or causes AV 100 to stop in response to receiving an indication from passenger detection subsystem 130 that the passengers are not in the proper seating configuration.
[0026] The algorithms and modules of AVCS 140 may generate control outputs for use by various systems and components of AV 100, such as powertrain, braking, and steering 150, vehicle electronics 160, signaling 170, and other systems and components not explicitly shown in FIG. 1 . These systems and components may modify the operation of AV 100 based on the control outputs. Powertrain, braking, and steering 150 may include an engine (internal combustion engine, electric engine, etc.), transmission, differential, axles, wheels, steering mechanism, and other systems. Vehicle electronics 160 may include an on-board computer, engine management, ignition, communication systems, car computers, telematics, in-car entertainment systems, and other systems and components. Signaling 170 may include high and low headlights, stop lights, turn signals and taillights, horns and alarms, interior lighting systems, dashboard notification systems, passenger notification systems, radio and wireless network transmission systems, etc. Some of the commands output by AVCS 140 may be delivered directly to powertrain, braking, and steering 150 (or signaling 170), while other commands output by AVCS 140 are first delivered to vehicle electronics 160, which generates commands to powertrain, braking, and steering 150 and / or signaling 170.
[0027] In one example, AVCS 140 may determine that an obstacle identified by data processing system 120 should be avoided by slowing the vehicle until a safe speed is reached and then steering the vehicle around the obstacle. AVCS 140 may output commands to powertrain, brakes, and steering 150 (either directly or via vehicle electronics 160) to (1) reduce fuel flow to the engine and slow engine speed by changing the throttle setting, (2) downshift the drivetrain via the automatic transmission into a lower gear, (3) engage the brake unit (working in coordination with the engine and transmission) to reduce vehicle speed until a safe speed is reached, and (4) use the power steering mechanism to implement a steering maneuver until the obstacle is safely bypassed. AVCS 140 may then output commands to powertrain, brakes, and steering 150 to resume the vehicle's previous speed setting.
[0028] As used herein, the term "object" may include any body, item, device, body, or thing (moving or non-moving) located outside of AV100, such as other vehicles, bicyclists, pedestrians, animals, roads, buildings, trees, bushes, sidewalks, bridges, mountains, piers, embankments, runways, or other objects.
[0029] FIG. 2 is a flowchart illustrating one embodiment of a method 200 for detecting occupants in a vehicle using artificial intelligence, according to some implementations of the present disclosure. A processing device having one or more central processing units (CPUs), one or more graphics processing units (GPUs), and / or a memory device communicatively coupled to the CPU and / or GPU may implement method 200 and / or each of its individual functions, routines, subroutines, or operations. The processing device may include processing logic, which may include hardware, software, or a combination of both. Method 200 may be directed to systems and components of a vehicle. In some implementations, the vehicle may be an autonomous vehicle (AV), such as AV 100 of FIG. 1. In some implementations, the vehicle may be a driver-operated vehicle equipped with a driver assistance system, e.g., a Level 2 or Level 3 driver assistance system, that provides limited assistance for certain vehicle systems (e.g., systems such as steering, braking, acceleration, etc.) or under limited driving conditions (e.g., highway driving). Method 200 can be used to improve the performance of AVCS 140. In certain implementations, a single processing thread may execute method 200. Alternatively, two or more processing threads may execute method 200, with each thread executing one or more individual functions, routines, subroutines, or operations of method 200. In an exemplary embodiment, the processing threads executing method 200 may be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, the processing threads executing method 200 may execute asynchronously with respect to each other. Various operations of method 200 may be performed in a different order (e.g., reversed) compared to the order shown in FIG. 2. Some operations of method 200 may be performed simultaneously with other operations. Some operations may be optional. In some implementations, the occupant detection subsystem 130 may execute one or more operations of method 200.
[0030] At block 210, processing logic obtains one or more images captured by one or more cameras of a vehicle. The vehicle may include AV 100. The one or more cameras may include one or more cameras 115. In one implementation, the one or more images captured by the one or more cameras 115 may include one or more images of an interior portion of AV 100.
[0031] In some implementations, the one or more cameras 115 may represent a single camera. In other implementations, the one or more cameras 115 may represent multiple cameras 115. When multiple cameras 115 are used, the cameras 115 may be positioned in, on, or around the AV 100 such that the occupant detection subsystem 130 may generate a panoramic image from multiple images obtained from the multiple cameras 115. The multiple images may be captured by the cameras 115 simultaneously or nearly simultaneously. A panoramic image may include an image made up of portions of multiple images stitched together. The occupant detection subsystem 130 may use software (e.g., photography software) to combine the multiple images together into a panoramic image. In some implementations, the sensing system 110 may provide one or more images (which may include a panoramic image) to the occupant detection subsystem 130.
[0032] At block 220, processing logic generates passenger data (e.g., using location subsystem 132). The passenger data may indicate one or more locations of one or more passengers in the vehicle. The processing logic may use one or more AI models and one or more images. The one or more AI models may use the one or more images as input. The one or more AI models may include one or more AI models of AI subsystem 138. In one embodiment, the one or more AI models are trained using an AI training system, which is described in further detail below in connection with FIG. 3.
[0033] 3 illustrates one embodiment of an AI training system 300 according to an implementation of the present disclosure. As shown in FIG. 3, the AI training system 300 may include a training subsystem 310, which may include a training data engine 312, a training engine 314, a validation engine 316, a selection engine 318, or a test engine 320. The AI training system 300 may include one or more AI models 330A-N. The AI training system 300 may include an input / output component 340.
[0034] In one embodiment, the AI models 330A-N may include one or more artificial neural networks (ANNs), decision trees, random forests, support vector machines (SVMs), clustering-based models, Bayesian networks, or other types of machine learning models. ANNs generally include a feature representation component with a classifier or recurrent layer that maps features to a target output space. ANNs may include multiple nodes (neurons) arranged in one or more layers, and neurons may be connected to one or more neurons via one or more edges ("synapses"). Synapses may carry signals from one neuron to another, and weights, biases, or other configurations of neurons or synapses may adjust the value of the signal. Training an ANN may include adjusting the weights or other features of the ANN based on outputs generated by the ANN during training.
[0035] ANNs may include, for example, convolutional neural networks (CNNs), recurrent neural networks (RNNs), or deep neural networks. CNNs, a particular type of ANN, host multiple layers of convolutional filters. Pooling may be performed to address nonlinearities in lower layers, and a multilayer perceptron is typically added on top to map the features extracted by the convolutional layers to a decision (e.g., classification output). Deep networks may include ANNs with multiple hidden layers or shallow networks with zero or few (e.g., one or two) hidden layers. Deep learning is a class of machine learning algorithms that uses a cascade of multiple layers of nonlinear processing units for feature extraction and transformation. Each successive layer uses the output from the previous layer as input. RNNs are a type of ANN that contain memory, allowing the ANN to capture time dependencies. RNNs can learn input-output mappings that depend on both current and past inputs. RNNs address past and future measurements and make predictions based on this continuous measurement information. One type of RNN that can be used is a long short-term memory (LSTM) neural network.
[0036] ANNs may learn in a supervised (e.g., classification) or unsupervised (e.g., pattern analysis) manner. Some ANNs (e.g., deep neural networks) may include a hierarchy of layers, with different layers learning different levels of representations corresponding to different levels of abstraction. In deep learning, each level learns to transform its input data into slightly more abstract and complex representations.
[0037] In one or more embodiments, the AI models 330A-N may include multi-modal generative AI models 330A-N, transformer-based AI models 330A-N, or another type of AI model 330A-N. The AI models 330A-N may include generative capabilities, which may include the ability to generate new original data. The AI models 330A-N may include discriminative capabilities, which may include the ability to make predictions based on existing data patterns. The multi-model AI models 330A-N may include AI models 330A-N that can accept multiple forms of data as input and / or generate multiple forms of data as output (e.g., text data, image data, video data, audio data, etc.).
[0038] In some embodiments, large, multi-model generative AI models 330A-N can leverage their world knowledge in zero-shot detection. Because the generative AI models 330A-N may receive input data via prompts containing text, the detection set may be arbitrarily large (e.g., the input data is not limited to a set of fixed objects), which can add to the discriminatory capabilities of the AI models 330A-N.
[0039] In one embodiment, a generative AI model can deviate from a machine learning model based on its ability to generate new, original data rather than making predictions based on existing data patterns. Generative AI models may include generative adversarial networks (GANs), variational autoencoders (VAEs), or large-scale language models (LLMs). In some instances, generative AI models may employ different approaches to train or learn the underlying probability distributions of training data compared to some machine learning models. For example, a GAN may include a generative network and a discriminative network. The generative network attempts to generate synthetic data samples that are indistinguishable from real data, while the discriminative network attempts to correctly classify between real and fake samples. Through this iterative adversarial process, the generative network can gradually improve its ability to generate increasingly realistic and diverse data.
[0040] Generative AI models also have the ability to capture and learn complex, high-dimensional structures in data. One goal of generative AI models is to model the underlying data distribution and generate new data points that have the same properties as the training data. Some machine learning models (e.g., non-generative AI models) focus on optimizing task-specific predictions.
[0041] In some embodiments, the AI models 330A-N may be trained on a corpus of data. In some embodiments, the AI models 330A-N may be models that are first pre-trained on a corpus of data to create a base model and then fine-tuned with more data related to a particular task set to create a more task-specific or targeted model. The base model may first be pre-trained using a corpus of data, which may include data in the public domain, licensed content, and / or proprietary content. Such pre-training can be used by the AI models 330A-N to learn a wide range of elements, including image or speech recognition, general sentence structure, common phrases, terminology, natural language structure, and other elements. In some embodiments, this first base model may be trained using self-supervision or unsupervised training on such datasets.
[0042] In some embodiments, the second portion of training, which includes fine-tuning, may be unsupervised, supervised, reinforcement, or any other type of training. In some embodiments, this second portion of training may include some elements of supervision, including learning techniques incorporating human- or machine-generated feedback, training according to a set of guidelines, or training on a previously labeled dataset. In a non-limiting example related to reinforcement learning, the outputs of the AI models 330A-N during training may be ranked by a user according to various factors, including accuracy, usefulness, veracity, acceptability, or any other metric useful in the fine-tuning portion of training. In this way, the AI models 330A-N can learn to favor these and any other factors relevant to the user when generating responses. More details about training are provided below.
[0043] In some embodiments, the AI models 330A-N may include one or more pre-trained or fine-tuned models. By way of non-limiting example, in some embodiments, the goal of "fine-tuning" may be achieved with a second, or third, or any number of additional models. For example, the output of a pre-trained model may be input into a second AI model trained in a manner similar to the "fine-tuned" portion of the training described above. In this way, two more AI models may accomplish a similar task as one model that was pre-trained and then fine-tuned.
[0044] In one implementation, the training subsystem 310 may manage the training and testing of the AI models 330A-N. The training data engine 312 may generate training data (e.g., a set of training inputs and a set of target outputs) for training the AI models 330A-N. In an exemplary embodiment, the training data engine 312 may initialize a training set T to null (e.g., Rockwell). The training data engine 312 may obtain data to be added to the training set T. In the present disclosure, in some implementations, one of the training data may include an image and ground truth. The image may include an image of a portion (e.g., an interior portion) of the AV 100, which may or may not include a passenger in the image. The ground truth associated with the image may include data indicating whether the image includes one or more passengers, data indicating one or more positions of the one or more passengers in the image (if the image includes at least one passenger), or other data. The training data engine 312 may add training data to the training set T and may determine whether the training set T is sufficient for training the AI models 330A-N. The training set T may, in some embodiments, be sufficient for training the AI models 330A-N if the training set T includes a threshold amount of training data. In response to determining that the training set T is not sufficient for training, the training data engine 312 may identify or obtain additional portions of training data. In response to determining that the training set T is sufficient for training, the training data engine 312 may provide the training set T to the training engine 314.
[0045] The training engine 314 can use training data (e.g., training set T) to train the AI models 330A-N. AI models 330A-N may refer to model artifacts generated by the training engine 314 using the training data, which may include training inputs and, in some implementations, corresponding target outputs (e.g., correct answers for each training input). The training engine 314 can input the training data to the AI models 330A-N so that the AI models 330A-N can find patterns in the training data and configure themselves based on those patterns.
[0046] If the AI models 330A-N use supervised learning, the training engine 314 may assist the AI models 330A-N in determining whether the AI models 330A-N map training inputs to target outputs (predicted answers). If the AI models 330A-N use unsupervised learning, the training engine 314 may input training data to the AI models 330A-N. The AI models 330A-N may configure themselves based on the input training data, but because the training data may not include the target outputs, the training engine 314 may not assist the AI models 330A-N in determining whether the AI models 330A-N provided the correct outputs during the training process.
[0047] The validation engine 316 may validate the trained AI models 330A-N using the corresponding feature sets from the validation set from the training data engine 312. The validation engine 316 may determine the accuracy of each of the trained AI models 330A-N based on the corresponding feature sets from the validation set. When the training data may not include the target output, validating the trained AI models 330A-N may include obtaining the output from the AI model 330A-N and providing the output to another entity for evaluation. The other entity may include another AI model configured to evaluate the output of the AI model being trained. The other entity may include a human. The validation engine 316 may discard trained AI models 330A-N that have an accuracy that does not meet a threshold accuracy or otherwise fails evaluation. In some embodiments, the selection engine 318 may select trained AI models 330A-N that have an accuracy that meets the threshold accuracy. In some embodiments, the selection engine 318 may be able to select the trained AI model with the highest accuracy of multiple trained AI models 330A-N. In some implementations, the selection engine 318 may receive input from another AI model or a human and may select a trained AI model based on the input.
[0048] The test engine 320 may be capable of testing the trained AI models 330A-N using corresponding feature sets from a test set from the training data engine 312. For example, a first trained AI model 330A-N trained using a first feature set from the training set may be tested using a first feature set from the test set. The test engine 320 may determine the trained AI model 330A-N that has the highest accuracy or other evaluation of all of the trained AI models 330A-N based on the test set.
[0049] The input / output component 340 of the AI training system 300 may be configured to provide data as input to the AI models 330A-N and obtain one or more outputs. For example, the input / output component 340 may provide training data from the training engine 314 to one or more AI models 330A-N and obtain respective AI model 330A-N outputs. In another embodiment, the input / output component 340 may provide a test data set to one or more AI models 330A-N and obtain respective AI model 330A-N outputs.
[0050] As described above, in some embodiments, the AI models 330A-N may include multi-modal generative AI models 330A-N. The AI models 330A-N can generate new content based on input data provided to them. The generative AI models 330A-N may be supported by a prompt subsystem (not shown), which may reside on the passenger detection subsystem 130 of FIG. 1. The prompt subsystem may enable components of the passenger detection subsystem 130 to access the generative AI models 330A-N. The prompt subsystem may be configured to perform automatic identification and facilitate the acquisition of relevant and timely contextual information for efficient and accurate processing of prompts by the generative AI models 330A-N. Communication between the prompt subsystem and the generative AI models 330A-N of the AI subsystem 138 may, in some embodiments, be facilitated by a generative model application programming interface (API). In additional or alternative embodiments, the generative model API may convert prompts generated by the prompt subsystem into an unstructured natural language format, and conversely, convert responses received from the generative AI models 330A-N into any suitable format (including, for example, any structured, proprietary format that may be used by the prompt subsystem). Similarly, the data management API may support instructions that may be used to communicate data requests to components of the passenger detection subsystem 130 and the format of data received from such components.
[0051] The prompt interface may support any suitable type of input (e.g., text input, audio input, image input, etc.). The prompt interface may further support any suitable type of output (e.g., text output, audio output, image output, etc.). In some embodiments, the prompt subsystem may include a prompt analyzer that supports various operations of the present disclosure. For example, the prompt analyzer may receive an input (e.g., an image received from one or more cameras 115 of the sensing system 110) and generate one or more intermediate prompts to the generative AI models 330A-N to determine what type of data the generative AI models need to successfully respond to the input. Upon receiving a response from the generative AI models 330A-N, the prompt analyzer may analyze the response and form a request for relevant contextual data from the passenger detection subsystem 130. The prompt analyzer may then generate a prompt to the generative AI models 330A-N that includes the original prompt and the contextual data. In some embodiments, the prompt analyzer may itself include a lightweight generative AI model that processes the intermediate prompt and determines what type of contextual data is needed by the generative AI models 330A-N along with the original prompt to ensure a meaningful response from the generative AI models 330A-N.
[0052] The prompt subsystem may contain (or be accessible to) instructions stored on one or more tangible, machine-readable storage media of a computing device and executable by one or more processing units of the computing device. In one embodiment, the prompt subsystem may be implemented on a single machine. In some embodiments, the prompt subsystem may be a combination of client and server components.
[0053] In some implementations, the trained AI models 330A-N may be provided to the AI subsystem 138 of the AV 100. For example, the AI training system 300 may be in data communication with the AV 100 via a data network. The passenger detection subsystem 130 may detect one or more passengers of the AV 100 using one or more AI models 330A-N, as discussed herein. The AI subsystem 138 may include an input / output component 340 configured to provide data as input to the AI models 330A-N and obtain one or more outputs. As described herein, for example, in the case of the AI models 330A-N used by the position subsystem 132, the input / output component 340 may provide one or more images from the camera 115 as input to the AI models 330A-N and obtain one or more outputs from the AI models 330A-N, which may indicate whether the one or more images include an image of a passenger, and if so, where the passenger is located. The input / output component 340 may be further configured to provide one or more outputs to components of the AV 100 (e.g., components of the occupant detection subsystem 130).
[0054] Returning to FIG. 2 , in one implementation, the passenger detection subsystem 130 may acquire one or more images from a camera of the sensing system 110. The passenger detection subsystem 130 may provide one or more images from one or more cameras 115 to the position subsystem 132. The position subsystem 132 may provide one or more images to an input / output component 340 of the AI subsystem 138. The input / output component 340 may provide the one or more images as input to one or more AI models 330A-N. The one or more AI models 330A-N may generate passenger data based on the one or more images. The passenger data may indicate one or more locations of one or more passengers of the AV 100. The input / output component 340 may acquire the passenger data and provide it to the position subsystem 132.
[0055] In some implementations, using one or more AI models 330A-N and one or more images may include using a panoramic image as input to the one or more AI models. In one implementation, using one or more AI models 330A-N and one or more images may include generating one or more embeddings based on the one or more images and using the one or more embeddings as input to the one or more AI models 330A-N. The embeddings may include digital representations of the corresponding image(s). The digital representations may include vectors, e.g., vectors of floats. In some embodiments, the occupant detection subsystem 130, the position subsystem 132, the AI subsystem 138, or some other system of AV100 may perform a compression operation on the vectors. In some embodiments, the vectors may include a large number of floats (e.g., hundreds or thousands of floats). Compressing the vectors to reduce the number of floats or otherwise reduce the size of the vectors may result in the AI models 330A-N processing the embeddings using fewer computing resources than using uncompressed vectors.
[0056] In some implementations, generating an embedding based on the one or more images may include generating an embedding for each image of the one or more images. In one implementation, generating an embedding based on the one or more images may include generating a single embedding based on all images of the one or more images. Generating an embedding based on the one or more images may include generating an embedding based on a panoramic image.
[0057] In one or more implementations, one or more AI models 330A-N may represent a single AI model 330A. If the one or more images include a single image, the AI model 330A may receive the single image from the input / output component 340, process the image to generate passenger data, and provide the passenger data to the input / output component 340. In implementations where the one or more images include multiple images, the AI model 330A may receive the images simultaneously and generate passenger data for each image. The input / output component 340 may combine the different passenger data into passenger data and provide the passenger data to the position subsystem 132.
[0058] In some implementations, one or more AI models 330A-N may represent multiple AI models 330A-N. In one implementation, each AI model 330A-N may correspond to a camera 115 of one or more cameras 115. Each AI model 330A-N may be trained with training data including images captured from the perspective or position of the camera 115 corresponding to the respective AI model 330A-N. As mentioned above, the cameras 115 may include one or more external cameras 116 and one or more internal cameras 117, which will be discussed in more detail below in connection with FIG. 4.
[0059] 4 illustrates an exemplary AV 100 according to some implementations of the present disclosure. The AV 100 may include a first row of seats 402 (e.g., driver's seat and passenger seat), a second row of seats 404 (e.g., captain's chair row), a third row of seats 406 (e.g., a rear seat bench), and a storage area 408 (e.g., a trunk). The AV 100 may include a first interior camera 117A mounted in an upper center portion of the windshield and facing toward the rear of the AV 100. The AV 100 may include a second interior camera 117B mounted on the interior roof of the right side of the AV 100 between the first row 402 and the second row 404. The AV 100 may include a third interior camera 117C mounted on the interior roof of the left side of the AV 100 between the second row 404 and the third row 406. The AV 100 may include a fourth interior camera 117D mounted on an interior wall of the AV 100 within the storage area 408 and facing the interior of the storage area 408.
[0060] In some implementations, first AI model 330A may correspond to first internal camera 117A. First AI model 330A may be trained with training data including images captured from a camera at a position and orientation similar to that of first internal camera 117A. First AI model 330A may receive images captured by first internal camera 117A. Second AI model 330B may correspond to second internal camera 117B. Second AI model 330B may be trained with training data including images captured from a camera at a position and orientation similar to that of second internal camera 117B. Second AI model 330B may receive images captured by second internal camera 117B. Third AI model 330C may correspond to third internal camera 117C. Third AI model 330C may be trained on training data including images captured from a camera at a position and orientation similar to that of third internal camera 117C. Third AI model 330C may receive images captured by third internal camera 117C. Fourth AI model 330D may correspond to fourth internal camera 117D. Fourth AI model 330D may be trained with training data including images captured from a camera at a position and orientation similar to fourth internal camera 117D. Fourth AI model 330D may receive images captured by fourth internal camera 117D.
[0061] Returning to FIG. 2 , as described above, the AI models 330A-N may generate output in response to processing images from the cameras 115, and the location subsystem 132 may generate passenger data based on the output. In one implementation, the passenger data may indicate whether the image includes one or more images of one or more passengers. In some implementations where the image includes one or more images of one or more passengers, the passenger data may indicate one or more locations of the one or more passengers. The data indicating the passenger locations may include data describing a bounding box (e.g., data indicating the dimensions of the box, data indicating the location of the box, data indicating the size of the box, etc.) that contains at least a portion of the passenger's image within the input image. The bounding box may include a two-dimensional box or a three-dimensional box. In one or more implementations, for each passenger location, the passenger data may include a confidence score. The confidence score may include an indicator generated by the AI models 330A-N that indicates the AI model's 330A-N confidence level that the location contains a passenger.
[0062] In one embodiment, if the AI models 330A-N include generative AI models 330A-N, the input to the AI models 330A-N may include a prompt including one or more images from one or more cameras 115. The prompt may include a command to the AI models 330A-N to determine whether an image of the one or more images from the cameras 115 includes one or more passengers. The prompt may include an image of a vehicle (e.g., an image similar to FIG. 4 ), and the prompt may further instruct the AI models 330A-N to determine locations within the vehicle where the one or more passengers are located. The generative AI models 330A-N may process the prompt and output passenger data, which may include the locations of the one or more passengers within the vehicle.
[0063] At block 230, processing logic generates vehicle area data. The vehicle area data may indicate one or more areas of the vehicle in which one or more passengers are located. The generation of the vehicle area data may be based on the passenger data of block 220. In one implementation, vehicle area subsystem 134 may perform block 230.
[0064] In some implementations, the one or more areas of the vehicle may include one or more seats of the vehicle (e.g., first seat row 402, second seat row 404, or third seat row 406 of AV100 in FIG. 4). The seats of the vehicle may include areas of the vehicle designated as locations where passengers can sit while the vehicle is operating. The one or more areas of the vehicle may include a floor of the vehicle. The floor may include areas of the vehicle where passengers walk or stand (e.g., between the seats of the vehicle). The one or more areas of the vehicle may include a storage area of the vehicle (e.g., storage area 408 of AV100 in FIG. 4). A storage area may include an area of the vehicle designated to hold luggage or other objects but not as a location where passengers can be located while the vehicle is operating. A storage area may include a trunk of the vehicle, a truck bed of the vehicle, etc. The one or more areas of the vehicle may include the exterior of the vehicle. The exterior of the vehicle may include the sides of the vehicle, the rear of the vehicle, the roof of the vehicle, etc.
[0065] In some embodiments, the vehicle area subsystem 134 may receive passenger data from the position subsystem 132. The vehicle area subsystem 134 may provide the passenger data to the AI subsystem 138, which may generate output based on the passenger data using one or more AI models 330A-N. The AI subsystem 138 may provide output to the vehicle area subsystem 134, which may generate vehicle area data.
[0066] In some implementations, the AI models 330A-N may use passenger data as input and may output one or more areas of the vehicle where one or more passengers are located. The AI models 330A-N may be trained on training data. The training data may include data similar to the passenger data and corresponding ground truth including data indicating one or more areas where one or more passengers indicated by the passenger data are located. In some implementations, if the passenger data indicates the absence of a passenger, the AI models 330A-N may output vehicle area data that does not indicate an area of the vehicle, or the AI models 330A-N may not execute such passenger data. In one or more implementations, the passenger data may further include one or more images from the camera 115. The one or more AI models 330A-N may use the data and one or more images indicated in the passenger data to generate an output.
[0067] In one embodiment, generating vehicle area data based on the passenger data may include determining passenger locations of one or more passengers. Determining the passenger locations may include determining a location of the vehicle in which the passengers are located. Generating vehicle area data based on the passenger data may further include determining whether the passenger locations correspond to an area of the vehicle. Each area of the vehicle may include location data indicating where the area within the vehicle is located. If the passenger locations correspond to the locations of the areas of the vehicle, the vehicle area data may indicate that the passenger is located in the area of the vehicle. In one embodiment, if the AI models 330A-N include generative AI models 330A-N, the vehicle area subsystem 134 may provide a prompt to the generative AI models 330A-N, where the prompt may include an image of the passenger, an image of the location of the areas of the vehicle, and text asking the generative AI models 330A-N whether the passenger is located at the location. An output of the generative AI models 330A-N may indicate whether the passenger is located at the location.
[0068] Determining that the passenger location corresponds to a vehicle area location may include determining whether the passenger location overlaps with the vehicle area location by more than a threshold amount. For example, the AI models 330A-N may determine, based on the passenger data, that 80% of the passenger location overlaps with the first seat and 20% of the passenger location overlaps with the floor. The threshold amount may include 75%. Because the passenger location exceeds the threshold amount, the vehicle area subsystem 134 may determine that the passenger is located in the first seat and may generate vehicle area data based on the determination.
[0069] In some implementations, determining that the passenger's position corresponds to a location in the area of the vehicle may include determining whether the passenger is positioned in a predetermined position or orientation. For example, the vehicle area subsystem 134 may determine that the passenger's buttocks are not located within the area of the vehicle and are not directly above the seating area of the vehicle seat (which may indicate the passenger is not properly seated in the seat). The vehicle area subsystem 134 may determine that the passenger is sitting, standing, kneeling, lying down, or in some other position. The vehicle area subsystem 134 may determine that the passenger is not facing the front of the vehicle (e.g., the passenger's body is turned 90 degrees in their chair) or that the passenger is away from the area of the vehicle. In response to the passenger being in a predetermined position, the vehicle area subsystem 134 may determine that the passenger's position does not correspond to a location in the area of the vehicle. If the AI models 330A-N include generative AI models 330A-N, the vehicle area subsystem 134 may provide a prompt to the generative AI models 330A-N, which may include an image of a passenger and text asking whether the passenger is in a predetermined position or orientation. The output of the generative AI models 330A-N may include data indicating whether the passenger is in a predetermined position or orientation.
[0070] In some implementations, the location subsystem 132 or the vehicle area subsystem 134 (or the AI models 330A-N used by one or more of these subsystems 132, 134) may determine whether a passenger is a child or an adult. Determining whether a passenger is a child may include determining whether the passenger is an infant, toddler, young child, pre-teen, teenager, or some other age group. If the AI models 330A-N include generative AI models 330A-N, the vehicle area subsystem 134 may provide a prompt to the generative AI models 330A-N, where the prompt may include an image of the passenger and text asking whether the passenger is a child or an adult or text instructing the generative AI model to determine the passenger's age. In one or more implementations, the location subsystem 132 or the vehicle area subsystem 134 (or the AI models 330A-N used by one or more of these subsystems 132, 134) may determine whether a passenger is engaged in a predetermined activity. The predetermined activity may include smoking, using an e-cigarette, or engaging in some other activity. For example, the AI models 330A-N may be trained on images of passengers of various ages or images of passengers engaging in various activities. If the AI models 330A-N include generative AI models 330A-N, the vehicle area subsystem 134 may provide a prompt to the generative AI models 330A-N, which may include an image of the passenger and text asking whether the passenger is engaging in the predetermined activity.
[0071] At block 240, processing logic determines whether at least one passenger seating configuration criterion is met. The determination may be based on passenger data or vehicle area data. In some implementations, passenger seating subsystem 136 may perform block 240.
[0072] In one or more implementations, the passenger seating subsystem 136 may receive passenger data or vehicle area data from the position subsystem 132 or the vehicle area subsystem 134. The passenger seating subsystem 136 may determine whether the passenger data or vehicle area data meets one or more passenger seating configuration criteria.
[0073] In one implementation, the passenger seat configuration criteria may include a first passenger of one or more passengers positioned in a first seat of the vehicle and vehicle seat belt data indicating that a seat belt for the first seat is not fastened. The vehicle seat belt data may include data generated by the vehicle that indicates the state of a particular seat belt (e.g., not fastened, fastened, etc.). A seat belt assembly of the vehicle may include one or more sensors that can detect whether a seat belt buckle is in a seat belt receptacle, and the seat belt assembly may provide data indicative of the state of the seat belt to the vehicle (e.g., data processing system 120).
[0074] In some implementations, the passenger seating configuration criteria may include a plurality of passengers, one or more passengers, positioned in a first seat of the vehicle. For example, the vehicle area data may indicate that a first passenger is positioned in a first seat, and the vehicle area data may indicate that a second passenger is positioned in the same first seat.
[0075] In one or more implementations, the passenger seating configuration criteria may include that the first passenger is located in a first area of the vehicle, and the first area is not a seat of the vehicle. For example, the vehicle area data may indicate that the first passenger is located in the storage area 408 of the AV 100 or in the truck bed of the AV 100. Such areas are not seats of the AV 100.
[0076] In some implementations, passenger seating configuration criteria may include one or more passengers all being children. For example, the vehicle area data may indicate that the vehicle includes three passengers and that all three passengers are children. In one or more implementations, passenger seating configuration criteria may include a first passenger being a child and the first passenger not being seated in a child safety seat. In some implementations, the vehicle area data may include data indicating whether a passenger is seated in a child safety seat. In some implementations, passenger seating configuration criteria may include a first passenger of the vehicle being a child and the first passenger being located in a predetermined seat of the vehicle. The predetermined seat may include a seat in the first row 402 of the vehicle, the driver's seat, the passenger seat, or some other seat.
[0077] In one implementation, passenger seating configuration criteria may include one or more passengers smoking or using e-cigarettes. For example, vehicle area data may indicate that a passenger is smoking or using an e-cigarette. Passenger seating configuration criteria may also include passengers engaging in another predetermined activity.
[0078] In some embodiments, if the AI models 330A-N include generative AI models 330A-N, the passenger seating subsystem 136 may provide a prompt to the generative AI models 330A-N. The prompt may include text asking whether at least one passenger seating configuration criterion is met based on one or more passenger seating configuration criteria, passenger data or vehicle area data, and passenger data or vehicle data. The output of the generative AI models 330A-N may indicate whether the generative AI models 330A-N have determined that at least one passenger seating configuration criterion is met.
[0079] At block 250, in response to determining that at least one passenger seating configuration criterion is met at block 240, processing logic causes the vehicle to perform an action associated with the passenger seating configuration within the vehicle. The action associated with the passenger seating configuration within the vehicle may include generating a passenger alert. The passenger alert may include one or more visual or audible indications configured to provide information to one or more passengers indicating that the passenger seating configuration is invalid.
[0080] In some implementations, the passenger alert may include a visual alert. The visual alert may include an icon. The icon may include a symbol corresponding to the met criterion(s). The visual alert may include a displayed string of text data corresponding to the met criterion(s). The illuminated icon may include a visual representation of the vehicle and an indication of which criterion(s) within the vehicle are met. In one or more implementations, the passenger alert may include an audible alert. The audible alert may include a sound produced by a speaker in the vehicle. The sound may include a beep, a tone, or some other sound. The audible alert may include spoken utterances providing information about the met criterion(s).
[0081] FIG. 5 illustrates an example visual alert 500 according to some implementations of the present disclosure. The visual alert 500 may be displayed on a screen on the dashboard of the AV 100. In response to the passenger seating subsystem 136 (1) determining from vehicle area data that a passenger is located in the left seat of the second row 404 of the AV 100 and (2) determining that the seat belt of the left seat of the second row 404 is not fastened, the passenger seating subsystem 136 may generate a passenger alert. The passenger alert may include a seat belt icon 502, text 504 that may read "Seat belt not fastened," and a visual representation 506 of the AV 100 (e.g., similar to FIG. 4 ) in the visual alert 500, with another icon 508 appearing above the left seat of the second row 404. The passenger alert may include a speaker on the AV 100 emitting the audio message "Seat belt not fastened."
[0082] As another example, in response to the passenger seating subsystem 136 determining from the vehicle area data that a passenger is located within the storage area 408 of the AV 100, the passenger seating subsystem 136 may generate a passenger alert. The passenger alert may include text 504 stating "Passenger not in vehicle seat" and a visual representation 506 of the AV 100, with an icon 508 appearing in the storage area 408. The passenger alert may include a speaker on the AV 100 stating "Passenger not in vehicle seat."
[0083] Returning to FIG. 2 , performing an action associated with a passenger seating configuration in the vehicle at block 250 may include autonomously modifying the operation of the AV 100. Autonomously modifying the operation of the AV 100 may include causing the AVCS 140 to prevent the AV 100 from driving. For example, before the AV 100 begins autonomous driving, the passenger seating subsystem 136 may determine that a passenger seating configuration criterion is met and may provide an indication to the AVCS 140, which may not operate the powertrain, brakes, or steering 150. One or more passengers in the AV 100 may reconfigure such that the passenger seating configuration criterion is no longer met. The passenger seating subsystem 136 may determine that none of the one or more passenger seating configuration criteria are met and may provide another indication to the AVCS 140. The AVCS 140 may then operate the powertrain, brakes, or steering 150 to allow the AV 100 to drive.
[0084] In one implementation, the AV 100 may already be moving when the passenger seating subsystem 136 determines that the passenger seating configuration criteria have been met. Therefore, autonomously modifying the operation of the AV 100 may include stopping the AV 100. Stopping the AV 100 may include the AV 100 pulling off the road and stopping. For example, a passenger may unbuckle their seat belt while the AV 100 is moving. The passenger seating subsystem 136 may determine from the vehicle area data and seat belt data that a passenger is located in a particular seat and that the seat belt in that seat is not fastened. The passenger seating subsystem 136 may send an indication to the AVCS 140, which may operate the powertrain, brakes, or steering 150 to pull the AV 100 over to the side of the road, slow down, and stop it.
[0085] In some implementations, AVCS 140 may not immediately prevent or stop AV 100 from driving in response to receiving an indication from passenger seating subsystem 136. AVCS 140 may wait a predetermined amount of time before preventing AV 100 from driving or stopping AV 100. This may allow one or more passengers to modify their behavior to meet passenger seating configuration criteria before AVCS 140 takes action.
[0086] In some embodiments, in response to determining in block 240 that at least one passenger seating configuration criterion is not met, processing logic may return to block 210 and method 200 may be executed again using one or more different images captured by one or more cameras 115. The one or more different images may include one or more images captured by one or more cameras 115 at a later time than the one or more images of the first iteration of block 210. In some implementations, execution of method 200 may be repeated at predetermined intervals (e.g., every second, every five seconds, every ten seconds, etc.). In some implementations, execution of method 200 may be repeated while the vehicle is on, moving, or otherwise moving.
[0087] As described above, in some implementations, the position subsystem 132 may use multiple AI models 330A-N to determine one or more positions of one or more passengers in a vehicle. The position subsystem 132 may provide each image of one or more images to a different AI model 330A-N, and each AI model 330A-N may generate an output that may indicate the passenger's position. In some implementations, different images may show portions of the same location in a vehicle. For example, in FIG. 4, an image captured by the first internal camera 117A may show, among other areas, the center seat in the third row 406, and the third internal camera 117C may also show, among other areas, the center seat in the third row 406. In one or more implementations, the outputs of the different AI models 330A-N processing these images may conflict. For example, the first AI model 330A output may indicate that the passenger is located in the center seat of the third row 406, and the second AI model 330B output may indicate that the passenger is not located in the center seat of the third row 406. In some implementations, the vehicle area subsystem 134 may determine which passenger data to use.
[0088] In some implementations, the passenger data may include first passenger data generated by the first AI model 330A. The first passenger data may indicate that a first location of the one or more locations of the vehicle contains one or more passengers. The passenger data may further include second passenger data generated by the second AI model 330B. The second passenger data may indicate that the first location does not contain any passengers. The first passenger data and the second passenger data may each include a respective confidence score generated by the respective AI model 330A, 330B. The confidence score may include an indicator indicating the AI model's level of confidence in determining whether the first location contains a passenger. The location subsystem 132 may provide the passenger data to the vehicle area subsystem 134. The vehicle area subsystem 134 may determine that the confidence score of the first passenger data is higher than the confidence score of the second passenger data. In response, the vehicle area subsystem 134 may use the first passenger data to generate vehicle area data associated with the first location. In some implementations, the vehicle area subsystem 134 may ignore the second passenger data associated with the first location, but may use the second passenger data associated with other locations in the vehicle.
[0089] FIG. 6 shows a block diagram of an example computing device 600 capable of using AI to detect passengers in a vehicle, according to some implementations of the present disclosure. The example computing device 600 may be connected to other computing devices in a local area network (LAN), an intranet, an extranet, and / or the Internet. The computing device 600 may operate in the capacity of a server in a client-server network environment. The computing device 600 may be a personal computer (PC), a set-top box (STB), a server, a network router, switch or bridge, or any device capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that device. Furthermore, while only a single example computing device is shown, the term “computer” shall also be considered to include any group of computers that, individually or jointly, execute a set (or sets) of instructions to perform any one or more of the methodologies discussed herein.
[0090] An embodiment of a computing device 600 may include a processing unit 602 (also referred to as a processor or CPU), a main memory 604 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM), etc.), a static memory 606 (e.g., flash memory, static random access memory (SRAM), etc.), and a secondary memory (e.g., data storage device 618), which may communicate with each other via a bus 630.
[0091] Processing unit 602 (which may include logic processing 603) represents one or more general-purpose processing units, such as a microprocessor, a CPU, or the like. More specifically, processing unit 602 may be a complex instruction set computer (CISC) microprocessor, a reduced instruction set computer (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor executing other instruction sets, or a processor executing a combination of instruction sets. Processing unit 602 may also be one or more special-purpose processing units, such as a GPU, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or the like. According to one or more aspects of the present disclosure, processing unit 602 may be configured to execute instructions to perform a method, such as method 200 for detecting occupants in a vehicle using AI.
[0092] An example of computing device 600 may further include a network interface device 608, which may be communicatively coupled to a network 620. The network interface device 608 may include a network card, a network interface controller, or some other network interface. The network 620 may include a local area network (LAN), an intranet, an extranet, the Internet, a modem, a router, a switch, or some other network or network device. In some embodiments, computing device 600 may communicate data with other systems or devices over network 620. An example of computing device 600 may further include a video display 610 (e.g., a liquid crystal display (LCD), a touch screen, or a cathode ray tube (CRT)), an alphanumeric input device 612 (e.g., a keyboard), a cursor control device 614 (e.g., a mouse), and an audio signal generating device 616 (e.g., a speaker).
[0093] Data storage device 618 may include a computer-readable storage medium (or more specifically, a non-transitory computer-readable storage medium) 628 having stored thereon one or more sets of executable instructions 622. According to one or more aspects of the present disclosure, executable instructions 622 may include executable instructions for performing method 200.
[0094] The executable instructions 622 may also reside, for example, completely or at least partially within the main memory 604 and / or within the processing unit 602 during their execution on the computing device 600, with the main memory 604 and the processing unit 602 also constituting computer-readable storage media. The executable instructions 622 may further be transmitted or received over a network via the network interface device 608.
[0095] While computer-readable storage medium 628 is illustrated in FIG. 6 as a single medium, the term "computer-readable storage medium" should be considered to include a single medium or multiple media (e.g., centralized or distributed databases, and / or associated caches and servers) that store one or more sets of operating instructions. The term "computer-readable storage medium" should also be considered to include any medium capable of storing or encoding a set of instructions for execution by a machine that cause the machine to perform any one or more of the methods described herein. Thus, the term "computer-readable storage medium" should be considered to include, but not be limited to, solid-state memory, and optical and magnetic media.
[0096] In some cases, a particular component of AV 100 (e.g., sensing system 110, data processing system 120, AVCS 140, or other component) may include computing device 600.
[0097] In some implementations, certain components of the AV 100—such as the occupant detection subsystem 130, the position subsystem 132, the vehicle area subsystem 134, the passenger seating subsystem, or the AI subsystem 138—may be implemented on a computing device 600 external to the AV 100. The AV 100 may provide data to the computing device 600 over a network 620, and the components may perform method 200, and the computing device 600 may provide data to the AV 100 over the network 620. The computing device 600 may include a server. In one implementation, obtaining one or more images captured by the one or more cameras 115 in block 210 of method 200 may include the sensing system 110 or the data processing system 120 of the AV 100 providing the one or more images to a server external to the AV 100 over the network 620. The computing device 600 may receive the one or more images and provide them to the occupant detection subsystem 130 executing on the server. The occupant detection subsystem 130 executing on the server may perform method 200 discussed herein. In block 250, causing AV100 to perform an action associated with the passenger seating configuration in AV100 may include the server sending data to AV100, and in response to processing the data, data processing system 120 may send a command to AVCS140 to generate a passenger alert, autonomously modify the operation of AV100, or perform some other action, as discussed herein.
[0098] Some portions of the above detailed descriptions are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, understood to be a self-consistent sequence of steps leading to a desired result. The steps require physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.
[0099] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise indicated, and as will be apparent from the discussion that follows, throughout the description, discussions using terms such as "identify," "determine," "modify," "store," "adjust," "produce," "return," "compare," "generate," "stop," "cause," "load," "copy," "replace," "perform," or the like, will be understood to refer to the actions and processes of a computer system, or similar electronic computing device, that manipulate and transform data represented as physical (electronic) quantities in the computer system's registers and memory into other data similarly represented as physical quantities in the computer system's memory or registers or other such information storage, transmission, or display device.
[0100] Examples of the present disclosure also relate to an apparatus for carrying out the methods described herein. This apparatus may be specially constructed for the required purposes, or it may be a general-purpose computer system selectively programmed by a computer program stored within the computer system. Such a computer program may be stored on a computer-readable storage medium, such as, but not limited to, any type of disk, including optical disks, CD-ROMs, and magneto-optical disks, read-only memory (ROM), random-access memory (RAM), EPROM, EEPROM, magnetic disk storage media, optical storage media, flash memory devices, other types of machine-accessible storage media, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.
[0101] The methods and displays presented herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform the required method steps. The required structure for a variety of these systems appears as set forth in the description below. Additionally, the scope of the present disclosure is not limited to any particular programming language. It will be understood that a variety of programming languages can be used to implement the teachings of the present disclosure.
[0102] It is to be understood that the above description is intended to be illustrative, and not restrictive. Many other implementations will be apparent to those skilled in the art upon reading and understanding the above description. While the present disclosure describes particular examples, it will be recognized that the systems and methods of the present disclosure are not limited to the examples described herein, but may be modified and practiced within the scope of the appended claims. Accordingly, the specification and drawings should be considered in an illustrative, and not a restrictive, sense. The scope of the present disclosure should, therefore, be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.
Claims
1. 1. A method comprising: obtaining a plurality of images captured by a plurality of cameras of a vehicle; generating passenger data indicative of a location of one or more passengers in the vehicle using one or more artificial intelligence (AI) models and the plurality of images; generating vehicle area data based on the passenger data indicative of one or more areas of the vehicle in which the one or more passengers are located; determining whether at least one passenger seating configuration criterion is met based on the passenger data and the vehicle area data; and in response to determining that the at least one passenger seating configuration criterion is met, causing the vehicle to perform an action associated with a passenger seating configuration within the vehicle.
2. the vehicle comprises an autonomous vehicle (AV); The action associated with the passenger seating configuration in the vehicle is: generating a passenger alert based on the at least one passenger seating configuration criterion being met; or and autonomously modifying operation of the AV based on the at least one passenger seating configuration criterion being met.
3. Autonomously modifying the operation of the AV includes: Preventing the AV from running; or 3. The method of claim 2, further comprising performing at least one of: stopping the AV;
4. using the one or more AI models and the plurality of images, generating a panoramic image from the plurality of images and using the panoramic image as an input to the one or more AI models; or 10. The method of claim 1, comprising at least one of: generating embeddings based on the plurality of images; and using the embeddings as inputs to the one or more AI models.
5. The one or more areas of the vehicle are: one or more seats in the vehicle; a floor of the vehicle; a storage area of the vehicle.
6. generating the vehicle area data based on the passenger data; determining a passenger location of the one or more passengers; and determining whether the location of the passenger overlaps with a location in one of the one or more areas of the vehicle by more than a threshold amount.
7. The at least one passenger seating configuration criterion comprises: a first passenger of the one or more passengers is located in a first seat of the vehicle, and seat belt data for the vehicle indicates that the seat belt for the first seat is not fastened; a plurality of passengers of the one or more passengers are located in the first seat of the vehicle; or 2. The method of claim 1, comprising at least one of: the first passenger is located within a first area of the one or more areas of the vehicle, and the first area is not a seat of the vehicle.
8. The at least one passenger seating configuration criterion comprises: all of said one or more passengers are children; a first passenger of the one or more passengers is a child, and the first passenger is not seated in a child safety seat; or The method of claim 1 , including at least one of: a second passenger of the one or more passengers is smoking.
9. 1. A system comprising: Memory and a processing unit, coupled to the memory, obtaining a plurality of images captured by a plurality of cameras of a vehicle; generating passenger data indicative of a location of one or more passengers in the vehicle using one or more artificial intelligence (AI) models and the plurality of images; generating vehicle area data based on the passenger data indicative of one or more areas of the vehicle in which the one or more passengers are located; determining whether at least one passenger seating configuration criterion is met based on the passenger data and the vehicle area data; and a processing device configured to perform operations including causing the vehicle to perform an operation associated with a passenger seating configuration within the vehicle in response to determining that the at least one passenger seating configuration criterion is met.
10. the vehicle comprises an autonomous vehicle (AV); The action associated with the passenger seating configuration in the vehicle is: generating a passenger alert based on the at least one passenger seating configuration criterion being met; or and autonomously modifying operation of the AV based on the at least one passenger seating configuration criterion being met.
11. Autonomously modifying the operation of the AV includes: Preventing the AV from running; or and stopping the AV.
12. using the one or more AI models and the plurality of images, generating a panoramic image from said plurality of images; or and generating an embedding based on the plurality of images.
13. The one or more areas of the vehicle are: one or more seats in the vehicle; a floor of the vehicle; a storage area for the vehicle.
14. generating the vehicle area data based on the passenger data; determining a passenger location of the one or more passengers; and determining whether the location of the passenger overlaps with a location in one of the one or more areas of the vehicle by more than a threshold amount.
15. The at least one passenger seating configuration criterion comprises: a first passenger of the one or more passengers is located in a first seat of the vehicle, and seat belt data from the vehicle indicates that the seat belt for the first seat is not fastened; a plurality of passengers of the one or more passengers are located in the first seat of the vehicle; or 10. The system of claim 9, comprising at least one of: the first passenger is located within a first area of the one or more areas of the vehicle, and the first area is not a seat of the vehicle.
16. The at least one passenger seating configuration criterion comprises: all of said one or more passengers are children; a first passenger of the one or more passengers is a child, and the first passenger is not positioned in a child safety seat; or 10. The system of claim 9, including at least one of: a second passenger of the one or more passengers is smoking.
17. A non-transitory computer-readable medium that, when executed by one or more processing devices, causes the one or more processing devices to: obtaining a plurality of images captured by a plurality of cameras of a vehicle; generating passenger data indicative of a location of one or more passengers in the vehicle using one or more artificial intelligence (AI) models and the plurality of images; generating vehicle area data based on the passenger data indicative of one or more areas of the vehicle in which the one or more passengers are located; determining whether at least one passenger seating configuration criterion is met based on the passenger data and the vehicle area data; and in response to determining that the at least one passenger seating configuration criterion is met, causing the vehicle to perform an action associated with a passenger seating configuration within the vehicle.
18. The computer-readable medium of claim 17 , wherein the passenger data includes a confidence score for each location of the one or more locations of the one or more passengers.
19. the passenger data includes first passenger data generated by a first AI model of the one or more AI models, the first passenger data indicating that a first location of the one or more locations of the vehicle includes a passenger of the one or more passengers; the passenger data further includes second passenger data generated by a second AI model of the one or more AI models, the second passenger data indicating that the first location does not include a passenger of the one or more passengers; the processing further includes determining that the confidence score of the first passenger data is higher than the confidence score of the second passenger data; The computer-readable medium of claim 18 , wherein generating the vehicle area data is based on the first passenger data.
20. the vehicle comprises an autonomous vehicle (AV); The action associated with the passenger seating configuration in the vehicle is: generating a passenger alert based on the at least one passenger seating configuration criterion being met; or and autonomously modifying operation of the AV based on the at least one passenger seating configuration criterion being met.