Detection of vulnerable road users who have lost control in the automotive environment
The system uses sensor data to identify and respond to VRUs at risk of losing control by analyzing reference points and tendency states, enhancing collision prevention and compliance in autonomous vehicles.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- WAYMO LLC
- Filing Date
- 2025-11-19
- Publication Date
- 2026-06-01
Smart Images

Figure 2026089686000001_ABST
Abstract
Description
Technical Field
[0001] This specification generally relates to autonomous vehicles. More specifically, this specification relates to the automatic detection of objects in an automotive environment that are at risk of losing control of their operation.
Background Art
[0002] Autonomous (fully and partially self-driving) vehicles (AVs) operate by sensing the external environment with various electromagnetic (e.g., radar and optical) and non-electromagnetic (e.g., voice and humidity) sensors. Some autonomous vehicles chart a driving route within the environment based on the sensed data. The driving route can be determined based on Global Positioning System (GPS) data and roadmap data. GPS and roadmap data can provide information about the static aspects of the environment (buildings, street layouts, road closures, etc.), while dynamic information (information about other vehicles, pedestrians, streetlights, etc.) is obtained from simultaneously collected sensed data. The accuracy and safety of the driving route, as well as the accuracy and safety of the speed regime selected by the autonomous vehicle, depend on the ability of the driving algorithm to timely and accurately identify various objects present in the external environment and process information about the environment to provide correct instructions to vehicle control and the drivetrain.
Brief Description of the Drawings
[0003] This disclosure is shown by way of example and not limitation and can be more fully understood when considered in connection with the following detailed description, with reference to the figures.
[0004] [Figure 1] FIG. 1 is a diagram showing the components of an exemplary autonomous vehicle (AV) that can detect a loss of control (LoC) of a vulnerable road user (VRU) in a driving environment according to some implementations of the present disclosure. [Figure 2A]Figure 2A shows a first stage of exemplary architecture of a portion of a vehicle perception and planning system that can identify VRUs at risk of LoC in a driving environment, as can be seen in several implementations of the present disclosure. [Figure 2B] Figure 2B shows a second stage of exemplary architecture of a portion of a vehicle perception system that can identify a VRU at risk of LoC in a driving environment, as demonstrated by several implementations of the present disclosure. [Figure 3A] Figure 3A shows one example of a late-stage fusion architecture for the computer vision backbone of the VRU perception model of Figure 2A, based on several implementations of the present disclosure. [Figure 3B] Figure 3B shows one example of an exemplary early fusion architecture of the computer vision backbone for the VRU perception model of Figure 2A, based on several implementations of the present disclosure. [Figure 4A-4C] Figures 4A–4C show tracking of a reference point for a pedestrian walking, as part of VRU control loss monitoring in a driving environment, as demonstrated by several implementations of this disclosure. [Figures 5A-5C] Figures 5A to 5C illustrate tendency state detection for pedestrians in motion, as part of VRU control loss monitoring in a driving environment, as demonstrated by several implementations of this disclosure. [Figure 6] Figure 6 illustrates an exemplary method for identifying vulnerable road users at risk of losing control in a driving environment, based on several implementations of this disclosure. [Figure 7] Figure 7 shows a block diagram of an exemplary computer device that can identify and respond to the presence of a VRM in an operating environment that is at risk of losing control of its operation, as demonstrated by some implementations of the present disclosure. [Overview of the project]
[0005] One implementation discloses a system including an autonomous vehicle sensing system. The sensing system is for collecting sensing data of the autonomous vehicle's environment. The system further includes an autonomous vehicle perception system. The perception system processes the sensing data using one or more machine learning models to identify multiple reference points associated with vulnerable road users (VRUs) in the environment, identify one or more height differences to the multiple reference points, and determine that the VRU is at risk of losing control based on a change in at least one height difference, causing the autonomous vehicle's control system to perform evasive action.
[0006] In another implementation, a system is disclosed that includes an autonomous vehicle sensing system. The sensing system is for collecting sensing data of the autonomous vehicle's environment. The system further includes an autonomous vehicle perception system. The perception system processes the sensing data using one or more machine learning models to identify multiple reference points associated with vulnerable road users (VRUs) in the environment and to determine the VRU's tendency state score. The tendency state score characterizes whether the VRU is in a tendency position or is likely to transition to a tendency position. The perception system further identifies one or more height differences relative to the multiple reference points and determines that the VRU is at risk of losing control based on the VRU's acceleration and at least one of one or more height differences or changes in the tendency state score. The perception system further causes the autonomous vehicle's control system to perform evasive maneuvers.
[0007] Another implementation disclosed is a method comprising: using the autonomous vehicle's sensing system to collect sensing data of the autonomous vehicle's environment; processing the sensing data with one or more machine learning models to identify a plurality of reference points associated with vulnerable road users (VRUs) in the environment; identifying one or more height differences to the plurality of reference points; determining that the VRU is at risk of loss of control based on a change in at least one of the one or more height differences; and causing the autonomous vehicle's control system to perform an evasive action. [Modes for carrying out the invention]
[0008] Autonomous vehicles, or vehicles deploying various driver assistance features, can use multiple sensor modalities to facilitate the detection and identification of objects in the driving environment and track the future trajectories of such objects. Sensors may include radio detection and ranging (radar) sensors, optical detection and ranging (lidar) sensors, multiple digital cameras, ultrasonic sensors, position sensors, and other similar sensors. Different types of sensors may offer different complementary benefits. For example, radar and lidar emit electromagnetic signals (radio signals or optical signals) that are reflected from objects and transmit information about the distance to the object (e.g., from the time of flight of the signal) and the velocity of the object (e.g., from the Doppler shift of the frequency of the reflected signal). Radar and lidar can scan the entire 360-degree field of view by using a series of continuous sensing frames. Sensing frames may include numerous reflections covering the external environment in a high-density grid of return points. Each return point may be associated with the distance to the corresponding reflective object and the line-of-sight velocity of the reflective object (the component of velocity along the line of sight).
[0009] LiDAR possesses high spatial resolution due to its submicron optical wavelengths, which makes it easy to obtain many close-up return points from the same object. This enables accurate detection and tracking of objects once they enter the range of the LiDAR sensor. Depending on the specific LiDAR model, LiDAR has an operating range of 150 to 350 m, with higher ranges typically achieved by more powerful and expensive systems.
[0010] Radar sensors are inexpensive, require less maintenance than lidar sensors, have a wider operating range, and are more resistant to adverse weather conditions. As a result of using much longer wavelengths (radio), the resolution of radar data is much lower than that of lidar. In particular, while radar can accurately determine the speed of moving objects (rather than those moving too slowly for the radar receiver), detecting the precise location of an object can often be problematic.
[0011] Cameras (e.g., still cameras or video cameras) can acquire high-resolution images at both close range (where LiDAR operates) and long range (where LiDAR cannot reach). A camera captures a two-dimensional projection of the three-dimensional external space onto the image plane (or some other non-planar imaging surface). Cameras have a longer operating range than LiDAR, but have a higher error margin along the radial direction compared to the lateral direction when determining the position of an object.
[0012] Camera and LiDAR images (and radar images in some applications) can be processed by various object detection models, including deep learning neural network models. Such models can determine the position and orientation of objects, as well as the evolution of their position and orientation over time. These models can further classify objects by type (e.g., trucks, cars, school buses, motorcyclists, pedestrians, and / or similar), manufacturer, model, and / or similar.
[0013] Driving environments can change very rapidly and can create situations where various objects and road users are at risk of colliding with other objects. Among the objects at highest risk in such situations are pedestrians, bikers, scooter riders, motorcyclists, and / or other users who, unlike drivers or passengers of other vehicles, are not protected from collisions with the body of a vehicle. Such high-risk users are referred to herein as vulnerable road users (VRUs). For example, a collision of a VRU with another object, such as a vehicle body or another VRU, can cause serious injury to the VRU, even at low speeds. In some situations, the risk of an accident in which a VRU may be involved can be further increased. For example, a biker or motorcyclist may lose control due to road defects, ice, equipment failure, carelessness, lack of experience, contact with an external object (vehicle or another VRU), and / or any other reason. Such situations are referred to herein as loss of control (LoC) situations. LoC events can cause a VRU to alter its operating patterns in many unpredictable ways. For example, a pedestrian walking on a sidewalk might slip or trip and land on the road. A biker or scooter attempting to regain control might quickly change gradient across the road. A motorcyclist might miscalculate the cornering angle or surface traction while turning, resulting in a fall and slide across multiple lanes. Early detection of situations prone to developing into loss-of-control events by autonomous vehicles equipped with driver assistance technologies, and by the vehicle itself, is crucial for the safety of all road users, especially VRUs that are not protected by a rigid vehicle body. Automatic detection of potential LoC events is difficult because collecting large amounts of data related to the occurrence of such situations is challenging (due to the relatively low percentage of driving missions in which LoC events occur).
[0014] The aspects and implementations of this disclosure address these and other challenges of existing object detection and tracking technologies by providing systems and technologies for efficiently and timely identifying VRUs at risk of loss of control in the process of experiencing loss of control (e.g., sliding, tipping over) or post-LoC (e.g., lying on the ground, sitting, getting up, and / or similar). Timely detection of imminent or likely LoC situations or post-LoC states allows a vehicle (e.g., an autonomous vehicle) to take appropriate and timely response measures to eliminate or reduce the risk of a secondary collision with the VRU. Furthermore, in some cases, a vehicle may cause a VRU to experience LoC even without direct physical contact with the VRU. For example, a vehicle may, by action, or sometimes by stopping, cause a VRU to attempt evasive maneuvers, altering its planned trajectory in a way that causes the VRU to lose control and tip over, such as by swerving, braking, or jumping. The driver or owner of a vehicle that caused such an event (e.g., an autonomous vehicle owner) may have a legal (and / or moral) obligation to remain at the scene of the incident and respond to the VRU and / or any official investigation that may be initiated. Accordingly, the vehicle may be required to have the capability to detect such an event, even if direct contact with the VRU (or any other object) has not been detected.
[0015] In some implementations, the disclosed technology includes a VRU monitoring system that uses sensing data (e.g., LiDAR, radar, camera data, and / or similar) to identify various VRUs in the environment, namely pedestrians (including joggers), bike cyclists, scooter riders, wheelchair users, skateboarders, and / or other road or sidewalk users not protected by the rigid walls of vehicles, and determines the state of operation of the identified VRU, including location, speed, acceleration, history, or motion, and / or similar. A trained reference point detection model may use sensing data associated with individual objects to determine the locations of multiple reference points within the VRU. In one embodiment of a pedestrian, such reference points may include one or more upper body reference points, such as the nose, chin, top of the head, ears, and / or similar, as well as one or more lower body reference points, such as the ankles, knees, contact points between the feet and the ground, and / or similar. The LoC detection engine uses, for example, the height difference (difference) between the upper and lower body reference points, and D(t)=H upper (t) New H lowerOne or more LoC indices can be calculated using, for example, the height difference between the nose and ankle(s), and this difference D(t) can be tracked over different time frames of sensing data, e.g., t1, t2, ... The LoC detection engine may determine the onset of LoC by identifying a change in height difference D(t) that exceeds a threshold, representing a person who is losing balance and beginning to fall to the ground. In some implementations, the selection of the reference point may depend on a particular type of object. For example, since a motorcyclist's facial features may be obscured by a helmet, the reference point detection model may use the geometric center of the helmet, the top of the helmet, the helmet visor, and / or similar. In some implementations, the selection of the reference point may be associated with the type of object, identified by an additional object type detection model (or an additional classification head of the model that determines the reference point). In some implementations, alternative reference points with some equivalence may be selected based on the VRU's field of view. For example, the nose or chin may be used as a reference point for pedestrians facing the sensing system, the ears may be selected for pedestrians positioned laterally to the sensing system, the center of the back of the head may be selected for pedestrians facing away from the sensing system, and / or similar points may be used. Some implementations may track additional reference points (and corresponding height differences), such as the right and / or left shoulders, right and / or left elbows, right and / or left wrists, right and / or left hips, right and / or left knees, etc.
[0016] When the LoC detection engine determines that an LoC event has begun, the vehicle may respond by decelerating (e.g., reducing the amount of throttle and / or braking) or by moving to a driving lane (if not occupied) that is further away from the VRU than the vehicle's current lane of movement. In some implementations, the LoC module may use additional inputs to reliably determine the beginning of an LoC event and eliminate false positive detections. For example, a change in D(t) may be caused by benign activity, such as a pedestrian sitting down to fix a loose shoe or bending down to pick up an object they dropped. To eliminate such false positives, the LoC module may further track the VRU's acceleration. An LoC event can be detected if a change in D(t) has occurred (or is in the process of occurring) and the VRU is experiencing (negative) threshold acceleration, which is likely to be related to an interruption of the VRU's normal movement, such as a fall, collision with another object, hard braking, and / or similar.
[0017] An additional trend state detection model may process the sensing data to determine the likelihood that a VRU captured within the sensing data is in a trend position (e.g., lying on the ground), for example, by determining a trend state score S. A trend state can be a continuous value within a specific, predetermined (training) range, e.g., S=0 (perfectly upright) and S=1 (perfectly trended). The trend state score S(t) can be tracked, as can the height difference D(t), and the rate of change of S(t) can be used to indicate the onset of a LoC event. In some implementations, the determination of an LoC event is conditional on at least one of the trend state score S(t) or the rate of change of the height difference D(t) exceeding its respective threshold. In other implementations, the determination of an LoC event is conditional on both the trend state score S(t) and the height difference D(t) exceeding their respective thresholds.
[0018] Numerous other implementations are disclosed herein. Advantages of the disclosed technologies and systems include, but are not limited to, the timely and efficient identification of VRUs, which have a high probability of losing trajectory control and a risk of colliding with or being in collision with other objects and road users, and which take appropriate protective measures to mitigate the risk of such accidents. Further advantages include facilitating an appropriate response to an event in accordance with legal requirements, when a vehicle (e.g., an autonomous vehicle) may be involved in the circumstances in which the event occurs without direct physical interaction between the vehicle and the VRU ("non-contact collision").
[0019] In these cases where the implementation description refers to autonomous vehicles, it should be understood that similar technologies may be used in various driver assistance systems that do not reach the level of fully autonomous driving systems. More specifically, the disclosed technologies may be used in Level 2 driver assistance systems that implement steering, braking, acceleration, lane centering, adaptive cruise control, and other driver support. Similarly, the disclosed technologies may be used in Level 3 driver assistance systems that enable autonomous driving under limited conditions (e.g., highways). Such systems may use high-speed, accurate detection and tracking of objects to inform the driver of approaching vehicles and / or other objects, allowing the driver to make final driving decisions (e.g., in a Level 2 system) or certain driving decisions (e.g., in a Level 3 system), such as deceleration or lane changes, without requiring driver feedback.
[0020] Figure 1 shows components of an exemplary autonomous vehicle (AV) 100 that can detect a Vulnerable Road User (VRU) Loss of Control (LoC) in a driving environment, as implemented in several implementations of the present disclosure. An autonomous vehicle could be an automobile (such as a car, truck, bus, motorcycle, all-terrain vehicle, recreational vehicle, any special agricultural, or construction vehicle), or any other self-propelled vehicle capable of operating in an automated driving mode (without or with reduced human input) (e.g., a robot, a factory, or warehouse robotic vehicle, a sidewalk delivery robotic vehicle, etc.). "Object" can include any body, item, device, vehicle, or article (with or without movement) located outside the autonomous vehicle, such as a road, building, tree, bush, sidewalk, bridge, mountain, other vehicle, pier, dike, runway, animal, bird, or other thing, as used herein.
[0021] The driving environment 101 may include any (moving or stationary) objects located outside the AV, such as roads, buildings, trees, bushes, sidewalks, bridges, mountains, other vehicles, and pedestrians. The driving environment 101 may be an urban area, suburban area, or rural area. In some implementations, the driving environment 101 may be an off-road environment (e.g., cultivated land or farmland). In some implementations, the driving environment may be an indoor environment, such as an industrial plant environment, a shipping warehouse, or a hazardous area of a building. In some implementations, the driving environment 101 may be substantially flat, with various objects moving parallel to the surface (e.g., parallel to the Earth's surface). In other implementations, the driving environment may be three-dimensional and may include objects capable of moving along all three directions (e.g., balloons, leaves). Hereafter, the term “driving environment” should be understood to include all environments in which autonomous operation of a self-propelled vehicle may occur. For example, “driving environment” may include any possible flight environment for an aircraft or a marine environment for a marine vessel. Objects in the driving environment 101 may be located at any distance from AV, from a few feet (or less) to several miles (or more).
[0022] As described herein, in semi-autonomous or partially autonomous driving modes, the vehicle assists with one or more driving operations (e.g., steering, braking, and / or acceleration for lane centering, adaptive cruise control, advanced driver-assistance systems (ADAS), or emergency braking), but the human driver is expected to situationally perceive the vehicle's surroundings and supervise the assisted driving operations. In such driving modes(s), the vehicle may perform all driving tasks in certain situations, but the human driver is expected to take control as needed.
[0023] For the sake of simplification and brevity, various systems and methods will be described later in conjunction with autonomous vehicles, but similar technologies may be used in various driver assistance systems that do not reach the level of fully autonomous driving systems. In the United States, the Society of Automotive Engineers (SAE) defines different levels of autonomous driving operations to indicate how much or how little a vehicle controls the driving, although different organizations in the United States or other countries may classify the levels differently. More specifically, the disclosed systems and methods may be used in SAE Level 2 (L2) driver assistance systems that implement steering, braking, acceleration, lane centering, adaptive cruise control, and other driver support. The disclosed systems and methods may be used in SAE Level 3 (L3) driver assistance systems that enable autonomous driving under limited conditions (e.g., highways). Similarly, the disclosed systems and methods may be used in vehicles using SAE Level 4 (L4) autonomous driving systems that operate autonomously under most normal driving conditions and require only occasional attention from a human operator. In all such driver assistance systems, an accurate assessment of the driving environment can be performed automatically without driver input or control (e.g., while the vehicle is moving), resulting in improved reliability of vehicle positioning and navigation, as well as overall safety of autonomous, semi-autonomous, and other driver assistance systems. As stated above, in addition to the way the SAE classifies levels of autonomous driving operations, other organizations in the United States or other countries may classify levels of autonomous driving operations differently. The systems and methods disclosed herein may be used in driver assistance systems defined by the levels of autonomous driving operations of these other organizations, but are not limited to these.
[0024] The AV100 in this embodiment may include a sensing system 110. The sensing system 110 may include various electromagnetic (e.g., optical) and non-electromagnetic (e.g., audio) sensing subsystems and / or devices. The sensing system 110 may include a radar 114 (or more radars 114), which may be any system that utilizes radio or microwave frequency signals to sense objects within the driving environment 101 of the AV100. The radar(s) 114 may be configured to sense both the spatial position of objects (including their spatial dimensions) and their velocity (e.g., using Doppler shift technology). Hereinafter, “velocity” refers to both how fast an object is moving (object speed) and the direction of the object’s motion. The sensing system 110 may include a lidar 112, which may be a laser-based unit capable of determining the distance to an object and the velocity of an object within the driving environment 101. Each of the lidar 112 and radar 114 may include a coherent sensor, such as a frequency-modulated continuous-wave (FMCW) lidar or radar sensor. For example, radar 114 may use heterodyne detection for velocity determination. In some implementations, the functionality of ToF and coherent radar is combined into a radar unit that can simultaneously determine both the distance to a reflecting object and the line-of-sight velocity of the reflecting object. Such a unit may be configured to operate in a non-coherent sensing mode (ToF mode) and / or a coherent sensing mode (e.g., a mode using heterodyne detection), or both modes simultaneously. In some implementations, multiple lidar 112 or radar 114 may be mounted on the AV100.
[0025] Lidar 112 may include one or more light sources that generate and emit signals, and one or more detectors that receive signals reflected back from objects. In some implementations, Lidar 112 may perform a 360-degree scan in the horizontal direction. In some implementations, Lidar 112 may have the ability to perform spatial scans along both the horizontal and vertical directions. In some implementations, the field of view may be up to 90 degrees vertically (e.g., at least a portion of the area above the horizon is scanned by the radar signal). In some implementations, the field of view may be spherical (consisting of two hemispheres).
[0026] The sensing system 110 may further include one or more cameras 118 to capture images of the driving environment 101. The images may be two-dimensional projections of the driving environment 101 (or a portion of the driving environment 101) onto the projection surface (flat or non-flat) of the camera(s). Some of the cameras 118 of the sensing system 110 may be video cameras configured to capture a continuous (or quasi-continuous) stream of images of the driving environment 101. The sensing system 110 may also include one or more infrared (IR) sensors 119. The sensing system 110 may further include one or more sonars 116, which in some implementations may be ultrasonic sonars.
[0027] The sensing data acquired by the sensing system 110 may be processed by the data processing system 120 of the AV 100. For example, the data processing system 120 may include a perception and planning system 130. The perception and planning system 130 may be configured to detect and track objects in the driving environment 101 and recognize the detected objects. For example, the perception and planning system 130 may be able to analyze images captured by the camera 118 and have the ability to detect traffic signals, road signs, road layout (e.g., traffic lane boundaries, intersection topology, parking space designations, etc.), the presence of obstacles, etc. The perception and planning system 130 may also further receive radar sensing data (Doppler data and ToF data) to determine the distance to various objects in the environment 101 and the velocity of such objects (radially and, in some implementations, laterally, as described below). In some implementations, the perception and planning system 130 may use radar data in combination with data captured by the camera(s) 118, as described in more detail below.
[0028] The perception and planning system 130 may include an object detection model 132 that deploys one or more suitable computer vision models to identify objects of interest, including VRUs and VRU types 254, such as pedestrians, bikers, motorcyclists, scooter riders, skateboarders, wheelchair users, and / or similar, which identify areas of the driving environment 101 that depict individual objects, classify the objects by type (e.g., vehicles, pedestrians, pedestrians, motorcyclists, scooter riders, skateboard riders, wheelchair users, and / or similar). The object detection model 132 may crop camera / lidar / radar images to portions of images (also referred to herein as patches) associated with these individual VRUs.
[0029] The perception and planning system 130 may further include a tracking and prediction component 134 for monitoring how the driving environment 101 evolves over time by determining and monitoring the positions and velocities of various objects, for example, identified by an object detection model 132. In some implementations, the tracking and prediction component 134 may track the changing appearance of the environment due to the AV's behavior in relation to the environment. In some implementations, the tracking and prediction component 134 may predict how various tracked objects in the driving environment 101 will be positioned within a predicted time frame. The prediction may be based on the current position and velocity of the tracked object, as well as the previous position and velocity (and in some cases, acceleration) of the tracked object. For example, based on stored data for object 1 (referred to herein as “tracking”) showing the position / velocity of object 1 in the tracking and prediction component 134 for the previous 3-second period, it can be concluded that object 1 is maintaining a constant velocity. Therefore, the tracking and prediction component 134 can predict where object 1 is likely to be within the next 3 or 5 seconds of its motion. As another example, based on the trajectory of object 2 showing its decelerating motion as it approaches a road crossing over the previous 2 seconds, the tracking and prediction component 134 can conclude that object 2 is approaching a stop sign before turning onto a side road. Therefore, the tracking and prediction component 134 can predict where object 2 is likely to be within the next 1 or 3 seconds. The tracking and prediction component 134 can periodically check the accuracy of its predictions and modify them based on new data acquired from the sensing system 110.
[0030] The perception and planning system 130 may further include a reference point detection model 135 that determines the locations of various reference points of the VRU detected by the object detection model 132 (and tracked by the tracking and prediction component 134). Exemplary reference points may include various skeletal points of the VRU, facial features of the VRU, and / or similar. The perception and planning system 130 may further include LoC detection 138 that determines, based on the reference points, the dynamics of various LoC indices, e.g., the height difference D(t) between two or more reference points of the VRU. LoC detection 138 may determine whether the dynamics of the height difference D(t) satisfy one or more threshold conditions to estimate the likelihood of an LoC event. Furthermore, LoC detection 138 may use the input of a trend state detection 136 model, which processes cropped images of the VRUs and outputs a trend state score S vor for each VRU, indicating that the VRU is likely in a trend position, for example, lying on the ground or sliding, falling, and / or similar. Based on the dynamics of additional information such as height difference Δ(t), trend state score S(t), and / or acceleration a(t), LoC detection 138 may determine whether the VRU is out of control, in the process of losing control, or at risk of losing control of its movement.
[0031] The perception and planning system 130 may also receive further information from a positioning subsystem 122, which may include a GPS transceiver and / or an inertial measuring unit (IMU) configured to acquire information about the AV's position relative to the Earth and its surroundings. The positioning subsystem can use positioning data, e.g., GPS data and IMU data, in conjunction with the perception data to help accurately determine the AV's position relative to fixed objects in the driving environment 101 (e.g., roads, lane boundaries, intersections, sidewalks, crosswalks, road signs, curbs, surrounding buildings, etc.), whose positions may be provided by road graph information 124. In some implementations, the data processing system 120 may receive non-electromagnetic data, such as audio data (e.g., ultrasonic sensor data or data from a microphone picking up an emergency vehicle siren), temperature sensor data, humidity sensor data, pressure sensor data, and weather data (e.g., wind speed and direction, precipitation data).
[0032] Various systems and subsystems of the data processing system 120 may have software stored in one or more system memory devices 126. The system memory 126 may include any volatile or non-volatile memory devices, such as read-only memory (ROM), random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), flash memory, flip-flop memory, or any other device capable of storing data. RAM may be static memory, such as dynamic random access memory (DRAM), synchronous DRAM (SDRAM), or static random access memory (SRAM). In some implementations, the system memory 126 may be on-chip memory.
[0033] The operation of the data processing system 120 may be carried out by one or more processors 128, which may include CPUs, GPUs, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc. As used herein, “processor” means a device capable of executing instructions that code arithmetic, logical, or I / O operations, for example, stored in the system memory 126. In some implementations, the processors 128 and the system memory 126 may be implemented as a single controller, for example, as an FPGA.
[0034] Data generated by the perception and planning system 130, including the positioning subsystem 122, object detection model 132, tracking and prediction component 134, reference point detection 135, trend state detection 136, LoC detection 138, and / or other systems and components, can be used by an autonomous driving system, such as a vehicle control system (VCS) 140. The VCS 140 may include one or more algorithms that control how the AV should behave in various driving situations and environments. For example, the VCS 140 may include a navigation system for determining a global driving route to a destination point. The VCS 140 may also include a driving route selection system for selecting a specific route through the immediate driving environment, which may include selecting a traffic lane, navigating through traffic congestion, selecting a place to make a U-turn, selecting a trajectory for parking operations, etc. The VCS 140 may also include an obstacle avoidance system for safely avoiding various obstacles in the AV's driving environment, such as stones, stalled vehicles, and pedestrians crossing in disregard of traffic rules and signals. An obstacle avoidance system may be configured to evaluate the size of an obstacle and its trajectory (if the obstacle is moving) and to select the optimal driving strategy (e.g., braking, steering, acceleration) to avoid the obstacle.
[0035] The VCS140 algorithm and modules can generate instructions for various systems and components of a vehicle, including the powertrain, brakes, and steering 150, vehicle electronics 160, signaling 170, and other systems and components not explicitly shown in Figure 1. The powertrain, brakes, and steering 150 may include the engine (internal combustion engine, electric engine, etc.), transmission, differential, axles, wheels, steering mechanism, and other systems. The vehicle electronics 160 may include the onboard computer, engine management, ignition, communication systems, car computer, telematics, in-car entertainment system, and other systems and components. The signaling 170 may include high headlights, low headlights, stop lights, turn signals, and taillights, horn and alarms, interior lighting systems, dashboard notification systems, passenger notification systems, radio, and wireless network transmission systems. Some of the commands output by the VCS140 can be delivered directly to the powertrain, brakes, and steering 150 (or signaling 170), while other commands output by the VCS140 are first delivered to the vehicle electronics 160, which generates commands for the powertrain, brakes, and steering 150, and / or signaling 170.
[0036] For example, the perception and planning system 130 may determine that a VRU identified by the data processing system 120 has lost control of its movement (e.g., skidding, falling, losing balance, and / or similar) and that it needs to be avoided, for example by slowing down the AV until it reaches a safe speed, and / or, if the VRU is a motorcyclist, cyclist, or scooter rider, by steering the AV vehicle to move away from a sidewalk, pedestrian crossing, or the lane the VRU is traveling in. The VCS140 is configured to output commands to the powertrain, brakes, and steering 150 (directly or via the vehicle electronics 160) to (1) reduce the engine speed by modifying the throttle setting, thereby reducing the fuel flow to the engine; (2) downshift the drivetrain to a lower gear via the automatic transmission; (3) engage the brake unit to reduce the vehicle's speed to a safe speed (working in conjunction with the engine and transmission); and (4) use the power steering mechanism to perform steering operations to move away from the VRU. The VCS140 may then output commands to the powertrain, brakes, and steering 150 to resume the vehicle's previous speed setting.
[0037] Figure 2A shows a first stage 200 of an exemplary architecture of a portion of a vehicle perception and planning system that can identify a VRU at risk of LoC in a driving environment, according to several implementations of the present disclosure. As illustrated, input to the perception and planning system (e.g., perception and planning system 130 in Figure 1) may include data acquired by the sensing system 110, for example, by lidar 112, radar 114, and / or camera(s) 118. The acquired data may be provided to the perception and planning system 130 by a camera image acquisition module 210, a lidar data acquisition module 220, and / or a radar data acquisition module 230. More specifically, the camera image acquisition module 210 can obtain a series of camera images, for example, a two-dimensional projection of the driving environment (or a portion thereof) on an array of sensing detectors (e.g., charge-coupled devices, or CCD detectors, complementary metal-oxide-semiconductor, or CMOS detectors, and / or similar). Individual camera images may have pixels of varying intensities, either in one color (for black and white images) or multiple colors (for color images). Camera images can be panoramic images or images showing specific parts of a driving environment. Camera images may contain a large number of pixels. The number of pixels may depend on the image resolution. Each pixel may be characterized by one or more intensity values. Black and white pixels may be characterized, for example, by a single intensity value representing the brightness of the pixel, where a value of 1 corresponds to a white pixel, a value of 0 to a black pixel (and vice versa). Intensity values can be continuous (or discrete) values between 0 and 1 (or between any other selected limits, e.g., 0 and 255). Similarly, color pixels may be represented by multiple intensity values, e.g., three intensity values (e.g., when using the RGB color encoding scheme) or four intensity values (e.g., when using the CMYK color encoding scheme). Camera images can be preprocessed, for example, by downscaling (combining multiple pixel intensity values into a single pixel value), upsampling, filtering, and denoising.Camera images(s) can be in any suitable digital format (JPEG, TIFF, GIG, BMP, CGM, SVG, etc.).
[0038] The LiDAR image acquisition module 220 (and similarly, the radar image acquisition module 230) can provide LiDAR (radar) images, which may include a set of return points (point clouds) corresponding to laser (radar) beam reflections from various objects in the driving environment. Each return point can be understood as a data unit (pixel), including the coordinates of the reflective surface, line-of-sight velocity data, intensity data, and / or similar. For example, the LiDAR image acquisition module 220 (and radar image acquisition module 230) can provide an image including an intensity map I(R,θ,φ), where R, θ, and φ are sets of polar coordinates. In some implementations, Cartesian, elliptic, parabolic, or other preferred coordinates may be used instead. The intensity map identifies the intensity of radar (LiDAR) reflections for various points in the field of view. The coordinates of an object (or the surface of an object) reflecting a lidar (radar) signal can be determined from directional data (e.g., polar angle θ and azimuth angle φ in the lidar transmission direction) and distance data (e.g., radial distance R, determined from the time of flight of the lidar signal). The lidar and / or radar image may further include velocity data of various reflecting objects, identified based on the detected Doppler shift of the reflected signal. Figure 2A illustrates an implementation in which three data acquisition modules are deployed, and one or more data acquisition modules may not be present (or may be disabled) in some implementations. For example, the camera image acquisition module 210 and the lidar (or radar) image acquisition module 220 may be deployed while the radar image acquisition module 230 (or lidar image acquisition module 220) is not deployed.
[0039] Camera (C) images, LiDAR (L) images, and / or radar (R) images may be large images of the entire driving environment or images of important parts of the driving environment (e.g., camera images acquired by the front camera(s) of the vehicle's sensing system). The acquired camera, LiDAR, and / or radar images may be processed by an object detection model 132, for example, a model (or more models) trained to identify individual objects 232 in the driving environment and to crop the camera / LiDAR / radar images to portions of the image (also called patches, as herein) associated with the individual objects 232. The object detection model 132 may be (or include) any suitable computer vision model, a machine learning model trained to identify areas containing objects of interest, such as vehicles, pedestrians, animals, road signs, buildings, structures, overpasses, and / or similar.
[0040] Objects identified by the object detection model 132 can be tracked by the tracking and prediction component 134, which uses different timestamps t j ,for example,
number
number
number
number
[0041] Camera (C), LiDAR (L), and / or radar (R) image patches cropped using object detection model 132 may be provided to a VRU perception model 235, which uses the provided patches to identify one or more LoC indices associated with a given VRU state. The VRU state refers to a set of detected reference points of the VRU, a score indicating the likelihood that the VRU is in a trend state, and / or other indices that may indicate an ongoing or imminent loss of control by the VRU. In some implementations, the VRU perception model 235 may include an input patch that processes intermediate features 252 and outputs LoC indices for the VRU, and a computer vision backbone 250 that generates one or more intermediate features (feature vectors, embeddings, etc.) 252 representing the contents of multiple classification networks (heads), e.g., a reference point detection network 135, a trend state detection network 136, and / or similar. In some implementations, the computer vision backbone 250 can process additional inputs, such as VRU track 254 (provided by tracking and prediction component 134), VRU type 256 (determined by object detection model 132), and / or other preferred inputs.
[0042] In some implementations, as shown by the dotted arrow in Figure 2A, the VRU type 256 may be determined by an additional classifier of the VRU perception model 235 that processes intermediate features 252 generated by the computer vision backbone 250.
[0043] In some implementations, the VRU perception model 235 may use decision tree algorithms, support vector machines, deep neural networks, etc. Deep neural networks may include convolutional neural networks, recurrent neural networks (RNNs) with one or more hidden layers, fully connected neural networks, long short-term memory neural networks, Boltzmann machines, attention mechanism networks, transformer networks, and / or similar.
[0044] The object detection model 132 and / or VRU perception model 235 can be trained using real camera images, lidar images, and / or radar images that show VRUs present in various driving environments, e.g., urban driving environments, highway driving environments, rural driving environments, off-road driving environments, and / or similar. Training can be carried out by a training engine 242 hosted by a training server 240, which may be an external server that introduces one or more processing devices, e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or similar. In some implementations, the object detection model 132 and / or VRU perception model 235 can be trained by the training engine 242 and then downloaded onto the AV's perception system. The object detection model 132 and / or VRU perception model 235 can be trained using training data, including training inputs 244 and corresponding target outputs 246 (correct matches for each training input 244), as shown in Figure 2A. During training of the object detection model 132 and / or the VRU perception model 235 (including the training computer vision backbone 250 and various classification heads), the training engine 242 can find patterns in the training data that map the training input 244 to the target output 246.
[0045] The training engine 242 may have access to a data store 241 that stores multiple camera images, lidar images, and / or radar images for actual driving conditions in various environments. The training input 244 may be annotated with labels or some other preferred mapping data 248 (ground truth annotation) that maps the training input 244 to the corresponding target output 246.
[0046] The stored training inputs 244 may include a large dataset (e.g., hundreds or thousands of images or more) containing cropped camera images / lidar / radar patches. In some implementations, the annotated training inputs can be ground truth annotated by the developer before being stored in the data store 241. During training, the training server 240 can retrieve the annotated training data from the data store 241, which includes one or more training inputs 244 and one or more target outputs 246 mapped by mapping data 248. For example, training of an object detection model 132 may be performed using images of pedestrians, vehicles, traffic signs, road markings (e.g., lane markings), and / or other objects as training inputs 244, and labels identifying the type of object and its bounding box as target outputs 246 (ground truth). Similarly, training of the reference point detection 135 model (or classifier) can be performed using a cropped image of the VRU as the training input 244, and a reference point (e.g., skeleton, face, etc.) marked by various developers as the target output 246. Training of the tendency state detection 136 model (or classifier) can be performed using a cropped image of the VRU as the training input 244, and a tendency state score S assigned by various developers as the target output 246. The assigned tendency state score S may range from 0 for an upright VRU (in one non-limiting example) to 1 for a VRU lying on the ground, with the intermediate value indicating that the VRU begins to lose control in the falling process (e.g., S=0.5), contact with the ground, e.g., a sitting state score (S=0.75), and / or any other suitable set of tendency state scores (e.g., S=0.25).
[0047] During training of the object detection model 132 and / or VRU perception model 235, the training engine 242 may modify the parameters (e.g., weights and biases) of the object detection model 132 and / or VRU perception model 235 until the models successfully learn how to predict the correct target output 246. In some implementations, the object detection model 132 and / or VRU perception model 235 may be trained separately. In various implementations, two or more VRU perception models 235 may be trained for use under different conditions and in different driving environments; for example, separate VRU perception models 235 may be trained for detecting loss of control of pedestrians, bicycle riders, motorcyclists, and / or similar. Different VRU perception models 235 may have different architectures (e.g., different number of neuron layers and different neural topologies), different configurations (e.g., activation functions), and may be trained using different hyperparameters.
[0048] The datastore 241 can be persistent storage capable of storing Lidar data, camera images, and data structures configured to facilitate accurate and rapid identification and verification of mark detection, according to various implementations of this disclosure. The datastore 241 may be hosted by one or more storage devices, such as main memory, magnetic or optical storage disks, tapes, or hard drives, network-attached storage (NAS), or storage area networks (SANs). Although shown separately from the training server 240, in some implementations the datastore 241 may be part of the training server 240. In some implementations the datastore 241 may be a network-attached file server, while in other implementations the datastore 241 may be some other type of persistent storage, such as an object-oriented database or a relational database (not shown in Figure 2A), hosted by a server machine or one or more different machines that can access the training server 240 over a network.
[0049] Figure 3A shows an exemplary late-fusion architecture 300 of the computer vision backbone 250 of the VRU perception model 235 of Figure 2A, according to several implementations of the present disclosure. As shown, the camera network 310, lidar network 320, and radar network 330 may be configured and trained to process input data of the corresponding modalities. For example, the camera network 310 may process camera patches 302 associated with individual VRUs (e.g., pedestrians) and generate camera embeddings 312 that constitute digital representations of various appearance features of the VRUs within the patches. During training, the camera network 310 learns how to efficiently encode appearance features via the camera embeddings 312. The camera embeddings 312 may have 256 bits, 512 bits, 1024 bits, or several other numbers of bits, which can be empirically set, for example, along with the architecture of the camera network 310, based on experimentation, to determine the optimal value of bits for a given target environment. In some implementations, the camera network 310 may be (or include) a neural network of artificial neurons. Neurons may be associated with learnable weights and biases. Neurons may be arranged in layers. Some layers may be hidden layers. The camera network 310 may include multiple hidden neuron layers and may be configured to perform computations that enable the detection of the VRU state. In some implementations, the camera network 310 may include multiple convolutional layers with suitable learning parameters, including kernel / mask size, kernel / mask weights, slide step size, etc. Convolutional layers may be arranged alternately with padding layers and may be followed by one or more pooling layers, e.g., a max pooling layer, an average pooling layer, etc. Some layers of the camera network 310 may be fully connected layers.In some implementations, the camera network 310 can be a fully connected layer network, a convolutional neural network, a recurrent neural network (RNN), a long-term short-term memory model (LSTM), an attention-grabbing network, a transformer network, and / or similar, or a combination thereof.
[0050] Similarly, the lidar network 320 may process a lidar patch 304 for the same VRU and generate a lidar embedding 314. The radar network 330 may similarly process a radar patch 306 and generate a radar embedding 316 for the VRU. The lidar embedding 314 / radar embedding 316 constitutes a digital representation of each portion of the lidar / radar point cloud captured by the lidar / radar patch. Through training, the lidar network 320 (and / or radar network 330) is trained to generate lidar embeddings 314 (radar embeddings 316) that efficiently represent the data for each captured object. The lidar embedding 314 (radar embeddings 316) may have the same number of bits as the camera embedding 312. In some implementations, the number of bits in the lidar embedding 314 (radar embeddings 316) may differ from the number of bits in the camera embedding 312. In some implementations, the LiDAR network 320 (and / or radar network 330) may have a U-network architecture in which a convolutional subnetwork (encoder) shrinks the features of a LiDAR patch (and / or radar patch) along its height and width dimensions and increases their size along its feature dimensions. The deconvolutional network (decoder) then shrinks the feature dimensions while simultaneously expanding the features along its width and height dimensions.
[0051] In some implementations, various additional network architectures or variations of network architectures may be used to implement camera networks 310, lidar networks 320, and / or radar networks 330, such as networks with residual connectivity, networks with multiple paths, attention-focused networks (self-attention and cross-attention), transformer networks, convolutional neural networks with low-density convolutions, and / or similar.
[0052] The camera embedding 312 can be combined with the lidar embedding 314, and further combined with the radar embedding 316 (e.g., concatenated or otherwise aggregated), and the combined embeddings can be processed by an aggregation network 340 that outputs intermediate features 252.
[0053] Figure 3B shows an exemplary initial fusion architecture 301 of a computer vision backbone 250 of the VRU perception model 235 of Figure 2A, according to several implementations of the present disclosure. Having the initial fusion architecture 301, the computer vision backbone 250 processes camera patches 302 using a self-attention camera network 310. Having the initial fusion architecture 301, the computer vision backbone 250 also processes lidar patches 304 using a self-attention lidar network 320. Each of the self-attention camera network 310 and the self-attention lidar network 320 transforms a corresponding input into an embedding by identifying the image, not only by the individual pixels (or groups of pixels) of the corresponding image, but also by extracting information present in the correlations of the pixels (or groups of pixels). For example, the relative arrangement of pixels depicting the cyclist's arm and the bicycle wheel carries important information about the cyclist's state (e.g., whether the cyclist maintains or loses control of the bicycle). In addition to the self-attention mechanism used for processing camera patch 302 and lidar patch 304, the cross-attention mechanism 315 learns the interrelationships between camera pixels and lidar pixels, starting from the early stages of embedding generation. In some implementations, self-attention and cross-attention in the early fusion network may be implemented using one or more transformer blocks that treat pixels (or groups of pixels) as query key value groups. For brevity, the early fusion architecture 301 is exemplified, for example, by the computer vision backbone 250, which processes camera and lidar inputs, but in some implementations, the early fusion backbone can also process radar input or other additional inputs and has a cross-attention mechanism 315 that learns the cross-correlations between various input modality pairs (e.g., camera-lidar cross-correlations, camera-radar cross-correlations, lidar-radar cross-correlations, and / or similar).
[0054] Furthermore, referring to Figure 2A, in some implementations, the computer vision backbone 250 processes each frame of the input data, for example, a given timestamp t j This can be done individually for camera / lidar / radar patches collected for. In some implementations, the computer vision backbone 250 receives input data frames t1, t2, ... t N It is possible to process sliding windows simultaneously, and the windows slide by a specific number of (stride) frames (e.g., M ≤ N) in each processing iteration. In one exemplary non-restrictive implementation, the sliding window is processed when the stride of the sliding window is M = 2 frames, t j+1 -t j It can contain N=10 frames spaced at 0.1-second intervals, and as a result, the updated set of processed frames is processed every 0.2 seconds, with two new frames added to the previous set of frames and the two oldest frames removed from the set of processed frames.
[0055] The reference point detection 135 can be (or include) a network having two or more fully connected layers that uses, for example, the intermediate feature 252 as an input and outputs the coordinates of various reference points for the VRU. In some implementations, the reference points can be represented by their coordinates and a predicted label, for example, output = {x, y, z; label}. The label can include "nose", "above the head", "jaw", "left ear", "right ear", "left shoulder", "right shoulder", "left hip", "right hip", "left knee", "right knee", "left ankle", "right ankle", and / or any label that the reference point detection 135 is trained to recognize. In some implementations, the coordinates x, y, z can be Cartesian coordinates. In other implementations, the coordinates x, y, z can be spherical coordinates, cylindrical coordinates, or any other suitable coordinates. The coordinates x, y, z can be determined using various position recognition data, such as the exact distance to a given reference point and the bearing angle to a given reference point, which are used as inputs to the reference point detection 135 and / or the computer vision backbone 250, for example, lidar data, or camera data that can include an exact bearing angle with distance determined using a suitable mapping from the perspective camera field to the global system of coordinates. The reference point detection 135 can select one or more upper body reference points (e.g., height z upper ), and one or more lower body reference points (e.g., height z lower ), and calculate the height difference(s), D = z upper - z lower . For example, in the case of a pedestrian walking towards (or laterally to) a sensing system (e.g., an autonomous vehicle sensing system), the upper body reference points can include the nose, jaw, upper part of the head, etc., and the lower body reference points can include the ankle(s), knee(s), etc. In the case of a pedestrian walking away from the sensing system, the upper body reference points can include the ear, upper part of the head, center of the head, etc., and the lower body reference points can include the heel(s), ankle(s), knee(s), etc.
[0056] Figures 4A–4C illustrate the tracking of a reference point of a walking pedestrian as part of VRU control loss monitoring in a driving environment, as demonstrated by several implementations of this disclosure. Figure 4A shows a walking pedestrian 400 maintaining control of its movement. For example, a height difference D can be identified by the reference point detection 135 in Figure 2A. L This represents the vertical distance between the pedestrian's nose (upper body reference point) and the pedestrian's left ankle (lower body reference point). Similarly, the height difference D R This indicates the vertical distance between the pedestrian's nose and the pedestrian's right ankle. Figure 4B shows pedestrian 402 at the start of the LoC event (slip and fall), with height difference D. L , and D R Compared to pedestrian 400 who did not experience a Line of Climax (LoC) during walking, pedestrian 404 experiences a sudden decrease. As shown in Figure 4B, the onset of the LoC event is accompanied by a substantial deceleration of the pedestrian, at the start of the fall, with forward movement ceasing. Figure 4C shows pedestrian 404 in the later stages of LoC. Pedestrian 404 has a height difference D between his nose and left ankle. L The number of limbs decreased further, and the right ankle was no longer visible from the sensory system's perspective as the person sat on the ground.
[0057] Referring further to Figure 2A, the intermediate features 252, which are patches used as input to the VRU perception model 235, may be processed by a tendency state detection 136, which may be another classifier network that outputs (or includes) a tendency state score S of the VRU. The tendency state score S can be a continuous value within a specific predetermined range (set as part of the training of the tendency state detection 136), for example, between S=0, which corresponds to the VRU's perfectly upright state, and S=1, which corresponds to the VRU's perfectly trendy state.
[0058] Figures 5A–5C illustrate tendency state detection for a moving pedestrian as part of VRU control loss monitoring in a driving environment, as demonstrated by several implementations of this disclosure. Figure 5A shows a moving pedestrian 500 maintaining control of its motion in an upright position, associated with a tendency state score close to S=0. Figure 5B shows a moving pedestrian 502 beginning to lose control (e.g., due to mispositioning of feet in contact with the pavement) and leaning excessively forward. The tendency state score of the moving pedestrian 504 may be substantially above the upright score S=0, but may also be substantially less than the entire tendency score S=0. Figure 5C shows a pedestrian 504 in a substantially tendency position with a tendency state score close to S=1, for example, after a fall by pedestrian 502 in Figure 5B.
[0059] Figure 2B shows a second stage 201 of an exemplary architecture of a part of a vehicle perception system that can identify a VRU at risk of LoC in a driving environment, according to several implementations of the present disclosure. The second stage 201 may include LoC detection 138. In some implementations, LoC detection 138 may be a component that unfolds one or more heuristics, formulas, tables, and / or similar. In some implementations, LoC detection 138 may include one or more decision tree algorithms and / or other machine learning techniques, including but not limited to supporting vector machines.
[0060] In various implementations, LoC detection 138 may use one or more of the following as input: reference point detection 135, trend state detection 136, detected VRU track 254, VRU type 256, and / or similar. In some implementations, LoC detection uses one or more height differences {D} between various pairs of reference points output by reference point detection 135. j}=D1, D2, ... D N For example, the height difference D between some upper body reference features (e.g., nose, ears, chin, top of head, shoulders, etc.) and some lower body reference features (e.g., knees, hips, etc.). L , and D RThis may include, but is not limited to, one or more LoC metrics. Some implementations use multiple differences, e.g., D1, D2, ... D N → D agg This can be aggregated into fewer (e.g., one) aggregated differences to reduce the changes in reference features that occur during normal motion. For example, the average of the nose-ankle differences may be calculated, D agg =(D L +D R ) / 2 can be a more stable representation of a cyclist's or pedestrian's footing. In some implementations, multiple aggregated differences can be calculated; for example, one aggregated difference can be calculated to represent the nose-ankle difference, and another aggregated difference can be used to represent the shoulder-knee difference. In some implementations, differences of different mean sizes (e.g., ear-ankle and shoulder-hip differences) can nevertheless be calculated over a specific observation time, e.g., several seconds, and the difference and the corresponding mean <D j >The ratio is aggregated by rescaling,
number
[0061] In some implementations, LoC detection 138 is, for example, t K =t1, t2, ... for multiple sensing data frames, the difference {D(t K )}=D1(t k ), D2(t k ), ...D N (t K) is monitored (and aggregated as described above, where applicable). In some implementations, LoC detection 138 also averages time-dependent differences over a specific (e.g., empirically selected) number of frames / hour, for example, over the last 2 seconds, half a second, and / or any other time interval. In some implementations, LoC detection 138 uses the rate of change over time (including differential and discrete derivatives) dD j Based on the fact that (t) / dt increases beyond the normal range of variation represented by the threshold rate of change, it can be determined that the onset of LoC has occurred:
number
[0062] In some implementations, the rate of change,
number
[0063] In some implementations, in addition to satisfying threshold velocity or threshold change conditions, VRUs may need to be observed for at least a threshold time that allows for LoC determination to be made in order to eliminate or reduce false positive LoC predictions. In some implementations, the selection of reference points used for LoC prediction and / or specific threshold conditions may depend on a specific VRU type 256 determined by the object detection model 132 and / or computer vision backbone 250 (see Figure 2A). For example, facial or head features may be hidden under a motorcyclist's helmet, and specific points associated with the helmet itself (e.g., geometric center, top, visor, and / or similar) may be used as sources for one or more upper body reference points. Similarly, threshold conditions may be set differently for different types of VRUs or depending on a particular driving environment. For example, larger changes in height difference for motorcyclists rounding curbs may be more common on rural highways than on urban roads. As a result, the threshold for motorcyclists in such rural environments may be set higher than in urban environments.
[0064] In some implementations, the LoC detection 138 may use a track 254 of the VRU as an additional input. Track 254 may be provided by a tracking and prediction component 134 that monitors the behavior of various VRUs from the moment they enter the field of view of the sensing system until they leave the field of view. In one exemplary implementation, track 234 is the coordinate X(t) of the VRU. j )(For example, along multiple spatial dimensions), the velocity V(t) of VRU j ), various times j+1 This may include the acceleration a(tj) of the VRU in (for example, corresponding to different sensing frames), and / or similar. In one non-exclusive implementation, the tracking and prediction component 134 may include, for example, X(t j+1 )=X(tj)+V( tj )(t j+1 -t j )+a(t j )(t j+1 -tj ) 2 / 2, V(t j+1 ) = V(t j )+a(t j )(t j+1 -t j The future state of the VRU can be predicted as follows: The new output of object detection model 132 is obtained over the subsequent time t j+1 , t j+2 When available for ..., the VRU track 254 may be updated using a Kalman filter, which computes a weighted combination of predicted motion states (based on the VRU motion model as described above) and observed motion states (generated by the object detection model 132).
[0065] In some implementations, to reliably determine the onset of an LoC event and eliminate false positive detections, the LoC detection 138 may use track 254 to exclude several benign activities, such as a pedestrian sitting on a curb, bending to pick up a dropped item, loading / unloading a vehicle, and / or similar. For example, to exclude false positives, the LoC detection 138 may use a change in the reference point difference (e.g., as disclosed above) and VRU a(t j A combination of accelerations can be used. An LoC event can be determined to have occurred if a substantial change in the difference(s) ΔD has occurred (or is in the process of occurring) and the VRU is experiencing a (negative) threshold acceleration, which is likely to be related to an interruption of normal VRU movement as a result of, for example, a fall, collision with another object, hard braking, and / or similar. In one non-limiting embodiment, the reference point index M is a weighted average of the difference(s) ΔD and the negative acceleration -a. RP It is possible to calculate, for example,
number
[0066] In some implementations, LoC detection 138 may also use a trend state score S(t) and / or the rate of change of the trend state score dS(t) / dt. In some implementations, an LoC event can be determined if the trend state score S(t) exceeds an empirical threshold ST, or if the rate of change of the LoC score dS(t) / dt exceeds another empirical threshold S'T. Condition dS(t) / dt ≥ S' T This may indicate the onset of a VRU in a LoC event, such as a loss of control and a rollover process. State S(t)≧S T This may indicate the aftermath of an LoC event that just occurred in a VRU lying (or sitting) on the ground. In other implementations, the determination of the start of an LoC event may be based on both the rate of change dS(t) / dt and the negative acceleration experienced by the VRU. In one non-restrictive embodiment, the propensity score index M is a weighted average of the rate of change dS(t) / dt of the propensity score. PS , and negative acceleration -a can be calculated, for example,
number
[0067] In some implementations, the determination of an LoC event may be subjected to one or more filters in the filtering stage 260 to determine whether the LoC detection process is appropriate or likely to lead to a false positive detection. Although shown as being performed after LoC detection 138, in some implementations, the filtering stage 260 may precede LoC detection 138. For example, the filtering stage 260 may include a tracking history filter 261 that determines the duration T for tracking a VRU. If the duration T is less than a set threshold time T0 (e.g., 1 to several seconds, depending on the driving environment and the type of VRU), the VRU may not be a real object but an artifact, such as a mirror-like surface (e.g., a window or side panel of another vehicle) reflecting another VRU. If the VRU has been observed for at least time T0, the filtering stage 260 can accept the VRU as a real object.
[0068] The filtering stage 260 may further include a field of view filter 262 that determines whether the vehicle's sensing system has a clear field of view of the VRU. For example, if the VRU's field of view is acquired through the window of another object (car, bus, and / or similar), the field of view filter 262 may determine that various calculated LoC indices are not reliable enough to detect the occurrence (or risk of occurrence) of an LoC event. Similarly, if a portion of the VRU (e.g., 30%, 40%, 50%, etc., at least a certain threshold portion of the VRU's body) is obstructed by another (closer-positioned) object, the field of view filter 262 may determine that the sensing data is not reliable enough and may, for example, postpone the LoC determination until a clearer field of view of the VRU becomes available.
[0069] The filtering stage 260 may further include a VRU-type filter 263 for determining whether the object is of a suitable type for LoC detection 138. For example, LoC detection may be unreliable for a child in a stroller or a rider on a recumbent bike, where the reference point (and / or trend state score) has relatively low reliability and can lead to a large possibility of false positive determinations.
[0070] The filtering stage 260 may further include a distance / speed filter 264 that determines whether the VRU is too far from the vehicle to be considered. For example, if an object is at such a distance d and traffic (including vehicles) is moving at a speed u relative to the vehicle, the distance / speed filter 264 may determine that at least time d / u has elapsed before the vehicle reaches the VRU, even if the VRU's behavior changes dramatically as a result of, for example, a rollover, a slip, a collision with another object, and / or something similar. If time d / u exceeds the vehicle's stopping time (considering the vehicle's current speed), the distance / speed filter 264 may disable the LoC detection 138 processing of the VRU at the current point in time, but may enable such processing at a later point in time if, for example, the distance d and / or the vehicle's speed changes.
[0071] When the LoC detection 138 determines that an LoC event has begun, the vehicle may perform an LoC response 270. For example, the LoC detection 138 may assume that the VRU has immediately lost control and outputs a command to the VCS 140 to perform a driving action that maximizes the isolation between the vehicle and the VRU, such as braking, reducing the amount of throttle, nudging, moving to a different available lane, and / or similar actions.
[0072] In some implementations, the LoC response 270 may depend on the reliability of LoC detection. Reliability is measured by the reference point index M. KP , and / or trend state index M PSThe value of the indicator, for example, can refer to the extent to which one or both indicators exceed a threshold indicator(s). If reliability is high, the LoC response 270 may include further observation and tracking of the VRU until reliability improves. The LoC response 270 can be performed by the VCS140, which includes, but is not limited to, immediate braking, delayed braking, performing nudging within the same travel lane, moving to a different lane, and / or otherwise increasing the separation between the vehicle and the object.
[0073] In some implementations, if the distance from the autonomous vehicle to the VRU at the time of LoC detection is less than a specific empirically defined distance, such as 6-10 meters, the LoC response 270 may, by default, assume that the vehicle's driving path likely caused or influenced the loss of control by the VRU (even if direct contact with the VRU is not detected). The LoC response 270 may then stop the vehicle on the side of the road until the situation is resolved, for example, until the VRU gets up and resumes operation, or until further development occurs, for example, until police arrive to address the situation. In some implementations, the vehicle may stop a remote assistant 282 and request commands from it, which may be deployed by a dispatch server 280 for the autonomous vehicle fleet. For example, the remote assistant 282 (which may be a human, computing software, or both) may review sensing data (including camera feeds and / or lidar / radar data logs) to determine whether the autonomous vehicle caused the situation that occurred and whether or not the vehicle should resume operation in the situation. The remote assistant 282 can then communicate this decision to the autonomous vehicle.
[0074] Figure 6 shows an exemplary method 600 for identifying vulnerable road users at risk of loss of control in a driving environment, as demonstrated by several implementations of the present disclosure. A processing device having one or more processing units (CPUs), one or more graphics processing units (GPUs), one or more parallel processing units (PPUs), and memory devices communicably coupled to the CPUs, GPUs, and / or PPUs may perform method 600 and / or each of its individual functions, routines, subroutines, or operations. Method 600 may apply to vehicle systems and components. In some implementations, the vehicle may be an autonomous vehicle. In some implementations, the vehicle may be a driver-operated vehicle that provides limited assistance through specific vehicle functions (e.g., systems such as steering, braking, and acceleration) or has driver assistance systems, e.g., Level 2 or Level 3 driver assistance systems, under limited driving conditions (e.g., highway driving). A processing device that implements Method 600 (e.g., processor 128 in Figure 1) may execute instructions issued by the perception and planning system 130 in Figure 1, more specifically by the object detection model 132, tracking and prediction component 134, reference point detection 135, trend state detection 136, LoC detection 138, and / or similar, during the driving operation of a vehicle. The operation of Method 600 may be performed in response to instructions stored in non-temporary computer-readable memory (e.g., system memory 126 in Figure 1). In a particular implementation, a single processing thread may implement Method 600. Alternatively, two or more processing threads may implement Method 600, each thread may perform one or more individual functions, routines, subroutines, or operations of the Method. In exemplary embodiments, the processing threads implementing Method 600 may be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, the processing threads implementing Method 600 may be executed asynchronously with respect to each other. Some of the operations of Method 600 may be performed in a different order than the order shown in Figure 5. Some of the operations of Method 600 may be performed simultaneously with other operations.Some actions can be optional.
[0075] In block 610, method 600 may include collecting sensing data of the autonomous vehicle's environment using the autonomous vehicle's sensing system (e.g., sensing system 110, see Figure 1). The sensing data may include camera data (e.g., one or more camera images), LiDAR data (e.g., LiDAR point cloud), radar data (e.g., radar returns), or any combination thereof. In some implementations, either the data / image may be a cropped patch / part of the sensing data associated with a particular VRU in the environment (e.g., detected and cropped by object detection model 132). In some implementations, the sensing system may include a camera for collecting camera data. In some implementations, the sensing system may include a LiDAR sensor for collecting LiDAR data and / or a radar sensor for collecting radar data. In some implementations, VRUs may include pedestrians, bicyclists, motorcyclists, scooter riders, skateboarders, wheelchair users, and / or similar.
[0076] In block 620, method 600 continues processing the sensing data by one or more machine learning models (e.g., VRU perception model 235, see Figure 2A) to identify multiple reference points associated with the VRU. In some implementations, the multiple reference points may include at least one reference point associated with the upper body of the VRU and at least one reference point associated with the lower body of the VRU (e.g., as shown in Figures 4A-4C). In some implementations, by identifying the multiple reference points, one or more machine learning models may be trained to process a portion of the camera data associated with the VRU, and at least a portion of the lidar data associated with the VRU, or a portion of the radar data associated with the VRU.
[0077] In some implementations, identifying a reference point may involve the actions shown in the left callout portion of Figure 6. More specifically, in block 622, method 600 may use a backbone (e.g., computer vision backbone 250 in Figure 2A) to process the sensing data and generate one or more intermediate features (e.g., intermediate feature 252). In block 624, method 600 may use a first classifier (e.g., reference point detection 135 in Figure 2A) to continue processing one or more intermediate features and output multiple reference points.
[0078] In some implementations, in block 625, method 600 may further include processing the perceived data by one or more machine learning models to determine a trend state score for the VRU. The trend state score may characterize the likelihood that the VRU is in or transitioning to a trend position (e.g., as shown in Figures 5A-5C). In some implementations, determining the trend state score may include the operation shown in the right callout portion of Figure 6. More specifically, in block 626, method 600 may include processing one or more intermediate features using a second classifier (e.g., trend state detection 136 in Figure 2A) to output a trend state score.
[0079] In block 630, method 600 may continue to identify one or more height differences relative to multiple reference points (e.g., as shown in Figures 4A-4C). In some implementations, one or more height differences may be identified for multiple time periods (or frames) during which the VRU is observed (tracked).
[0080] In block 640, method 600 may include determining that a VRU is at risk of loss of control based on one or more height differences and / or changes in the trend state score. For example, determining that a VRU is at risk of loss of control may include detecting that changes in one or more height differences satisfy a threshold condition. In some implementations, the threshold condition may include one or more height differences that experience at least a threshold change and / or at least a threshold rate of change (e.g., over time or frames).
[0081] In some implementations, determining that a VRU is at risk of loss of control may be based more on the VRU's acceleration (e.g., determined by the tracking and prediction component 134). In some implementations, determining that a VRU is at risk of loss of control may be based more on the propensity state score.
[0082] In some implementations, determining that a VRU is at risk of loss of control may be based on one or more height differences that satisfy a first threshold condition, and a change (or rate of change) in the propensity state score that satisfies a second threshold condition. For example, for LoC decisions, the height difference (or aggregated height difference) or the rate of change in the height difference may be based on a specific empirical value (e.g., empirical value D). T , and V T Regarding ΔD>D T , or dD(t) / dt>V T ) must be greater than or equal to the trend state score, or the rate of change of the trend state score, (experience S T , and U T Regarding another experience point (S(t)>S T , or dS(t) / dt>U T It must be greater than or equal to )
[0083] In some implementations, determining that a VRU is at risk of loss of control involves one or more height differences and a third threshold condition (e.g., empirical weights W1 and W2, and value X). T Regarding the weighted sum W 1· ΔD+W 2· S(t)>XT This can be based on a combination with a propensity state score that satisfies the following conditions.
[0084] In some implementations, determining that a VRU is at risk of loss of control is based on a fourth threshold condition (e.g., empirical value a). T Regarding a>a T The acceleration of the VRU that satisfies , and a fifth threshold condition (e.g., empirical value DD) T Regarding ΔD>DD, T One or more height differences that satisfy the fourth threshold condition, or the acceleration of the VRU that satisfies the fourth threshold condition, and the sixth threshold condition (Experience SS) T , and UU T Regarding S(t)>SS T , or dS(t) / dt>UU T This can be based on a trend state score that satisfies the following conditions. In some implementations, either the threshold condition (and value) may differ for different types of VRUs (for example, as detected by object detection model 132, see Figure 2A).
[0085] In block 650, method 600 may cause the autonomous vehicle's control system (e.g., VCS140 in Figure 2B) to continue performing an evasive maneuver. The evasive maneuver may include a change in the autonomous vehicle's speed, e.g., braking or acceleration. The evasive maneuver may also include a lateral shift of the autonomous vehicle, e.g., a nudge within the same traffic lane, a change in traffic lane, and / or similar, to move away from the VRU. In some implementations, causing the autonomous vehicle's control system to perform an evasive maneuver is in response to the confidence level of the LoC decision exceeding a threshold confidence level.
[0086] Figure 7 shows a block diagram of an exemplary computer device 700 capable of identifying and responding to the presence of a VRM in an operating environment that is at risk of loss of operational control, as in some implementations of this disclosure. The exemplary computer device 700 may be connected to other computer devices in a LAN, intranet, extranet, and / or the Internet. The computer device 700 may operate as a server in a client-server network environment. The computer device 700 may be a personal computer (PC), a set-top box (STB), a server, a network router, a switch or bridge, or any device capable of executing a set of instructions (sequentially or otherwise) that specify actions to be taken by such a device. Furthermore, although only a single embodiment of a computer device is shown, the term “computer” shall also be considered to include any group of computers that individually or collectively execute a set of instructions (or sets of instructions) to carry out any one or more of the methods discussed herein.
[0087] An exemplary computer device 700 may include a processing device 702 (also referred to as a processor or CPU), main memory 704 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM), etc.), static memory 706 (e.g., flash memory, static random access memory (SRAM), etc.), and secondary memory (e.g., a data storage device 718), which may communicate with each other via bus 730. In some implementations, the processing device 702 may be or include the processor 128 in Figure 1, and the main memory 704 may be or include the system memory 126 in Figure 1.
[0088] The processing device 702 (which may include a logic processor 703) represents one or more general-purpose processing devices, such as a microprocessor, a central processing unit, or similar. More specifically, the processing device 702 may be a composite instruction set computer (CISC) microprocessor, a reduced instruction set computer (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor that executes other instruction sets, or a processor that executes combinations of instruction sets. The processing device 702 may also be one or more application-specific processing devices, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or similar. According to one or more aspects of this disclosure, the processing device 702 may be configured to execute instructions that perform method 600 for identifying vulnerable road users at risk of loss of control in a driving environment.
[0089] An embodiment of the computer device 700 may further include a network interface device 708 that can be communicatively coupled to a network 720. An embodiment of the computer device 700 may further include a video display 710 (e.g., a liquid crystal display (LCD), a touchscreen, or a cathode ray tube (CRT)), an alphanumeric input device 712 (e.g., a keyboard), a cursor control device 714 (e.g., a mouse), and an audio signal generator 716 (e.g., a speaker).
[0090] The data storage device 718 may include a computer-readable storage medium (or more specifically, a non-temporary computer-readable storage medium) 728 on which one or more sets of executable instructions 722 are stored. According to one or more aspects of the present disclosure, the executable instructions 722 may include executable instructions that implement a method 600 for identifying vulnerable road users at risk of losing control in a driving environment.
[0091] The executable instructions 722 may also reside, entirely or at least partially, in main memory 704 and / or processing device 702 during their execution, for example, in an exemplary computer device 700, and main memory 704 and processing device 702 also constitute computer-readable storage media. The executable instructions 722 may also be transmitted or received over a network via network interface device 708.
[0092] Although the computer-readable storage medium 728 is shown as a single medium in Figure 7, the term “computer-readable storage medium” should be understood to include a single medium or multiple media (e.g., centralized or distributed databases, and / or associated caches and servers) that store one or more sets of operational instructions. The term “computer-readable storage medium” should also be understood to include any medium that has the ability to store or encode a set of instructions for machine execution that causes a machine to perform one or more of the methods described herein. Thus, the term “computer-readable storage medium” should be understood to include, but not be limited to, solid-state memory, as well as optical media and magnetic media.
[0093] Some parts of the detailed description above are presented in terms of algorithms of operation and symbolic representations of data bits in computer memory. These algorithmic descriptions and representations are means used by those skilled in the field of data processing to most effectively communicate the content of the work to others skilled in the art. An algorithm is understood here, and generally, to be a step-self-consistent sequence that produces a desired result. A step is one that requires the physical manipulation of a physical quantity. Usually, but not always, these quantities take the form of electrical or magnetic signals that can be stored, moved, combined, compared, and otherwise manipulated. Referring to these signals as bits, values, elements, symbols, characters, terms, numbers, or similar has sometimes proven convenient, primarily for reasons of general use.
[0094] However, it should be noted that all these terms, and similar terms, are associated with appropriate physical quantities and are merely convenient labels applied to those quantities. Otherwise, unless specifically stated, as will be evident from the following considerations, any considerations throughout the explanation using terms such as “specify,” “determine,” “store,” “adjust,” “produce,” “return,” “compare,” “generate,” “stop,” “load,” “copy,” “insert,” “replace,” “implement,” or similar terms will be understood to refer to the operation and processes of a computer system or similar electronic computing device that manipulates and converts data represented as physical (electronic) quantities in computer system registers and memory to other data similarly represented as physical quantities in computer system memory or registers or other such information storage, transmission, or display devices.
[0095] Examples of this disclosure also relate to apparatus for carrying out the methods described herein. Such apparatus may be specifically constructed for a required purpose or may be a general-purpose computer system selectively programmed by computer programs stored within the computer system. Such computer programs may be stored on computer-readable storage media, including, but not limited to, any type of disk, including optical disks, CD-ROMs, and magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic disk storage media, optical storage media, flash memory devices, other types of machine-accessible storage media, or any type of media suitable for storing electronic instructions, each coupled to a computer system bus.
[0096] The methods and representations presented herein are not inherently related to any particular computer or other device. Various general-purpose systems may be used with the programs taught herein, or it may be convenient to construct more specialized devices to carry out the necessary method steps. The necessary structures for various such systems will appear as described below. In addition, the scope of this disclosure is not limited to any particular programming language. It will be understood that various programming languages can be used to carry out the teachings of this disclosure.
[0097] It should be understood that the above description is intended to be illustrative and not restrictive. Many other embodiments will become apparent to those skilled in the art upon reading and understanding the above description. While this disclosure describes specific embodiments, it will be recognized that the systems and methods of this disclosure are not limited to the examples described herein and can be modified and implemented within the scope of the appended claims. Therefore, the specification and drawings should be considered illustrative, not restrictive. Accordingly, the scope of this disclosure should be determined by reference to the appended claims, along with the entire scope of equivalents to which such claims are entitled.
Claims
1. It is a system, A sensing system for an autonomous vehicle, comprising a sensing system that collects sensing data of the environment of the autonomous vehicle, The perception system of the autonomous vehicle, wherein the perception system is One or more machine learning models process the sensing data to identify multiple reference points associated with vulnerable road users (VRUs) in the environment. Identifying one or more height differences relative to the aforementioned multiple reference points, Based on a change in at least one of the aforementioned height differences, it is determined that the VRU is at risk of loss of control. The control system of the aforementioned autonomous vehicle includes a perception system that causes it to perform evasive maneuvers, A system equipped with these features.
2. The system according to claim 1, wherein the plurality of reference points include at least one reference point associated with the upper body of the VRU and at least one reference point associated with the lower body of the VRU.
3. The system according to claim 2, wherein the perceptual system detects that the change in one or more height differences satisfies a threshold condition in order to determine that the VRU is at risk of loss of control.
4. The aforementioned threshold condition is, The one or more of the aforementioned height differences experience at least a threshold change, or The system according to claim 3, comprising at least one of the following: one or more height differences experience at least a threshold change rate.
5. The system according to claim 1, wherein the VRU is determined to be at risk of loss of control from a further viewpoint of the acceleration of the VRU.
6. The system according to claim 1, wherein one or more machine learning models further determine a trend state score for the VRU, the trend state score characterizes the possibility that the VRU is in a trend position or is transitioning to a trend position, and the VRU is determined to be at risk of loss of control from a further perspective of the trend state score.
7. The aforementioned VRU, The change in one or more height differences that satisfies the first threshold condition, and the trend state score that satisfies the second threshold condition, A combination of one or more height differences and the trend state score that satisfies the third threshold condition, or The acceleration of the VRU that satisfies the fourth threshold condition, and at least one of the one or more height differences that satisfy the fifth threshold condition, or the trend state score that satisfies the sixth threshold condition, The system according to claim 6, which is determined to be at risk of loss of control based on at least one of the following.
8. The one or more machine learning models mentioned above are as follows: A backbone trained to process the aforementioned sensing data and generate one or more intermediate features, A first classifier trained to process one or more intermediate features and output the plurality of reference points, The system according to claim 6, comprising a second classifier trained to process one or more intermediate features and output the propensity state score.
9. The sensing data includes camera data and at least one of lidar data or radar data, and the sensing system A camera for collecting the aforementioned camera data, The system comprises at least one of the following: a lidar sensor for collecting lidar data, or a radar sensor for collecting radar data. In order to identify the plurality of reference points, one or more machine learning models, A portion of the camera data associated with the VRU, The system according to claim 1, trained to process at least a portion of the lidar data associated with the VRU, or a portion of the radar data associated with the VRU.
10. The aforementioned VRU is as follows: Pedestrians, Bicycle cyclist, Motorcyclist, Scooter rider, Skateboard rider, or The system according to claim 1, comprising at least one of the following: a wheelchair user.
11. It is a system, A sensing system for an autonomous vehicle, comprising a sensing system that collects sensing data of the environment of the autonomous vehicle, The perception system of the autonomous vehicle, wherein the perception system is The sensing data is processed by one or more machine learning models, Identify multiple reference points associated with vulnerable road users (VRUs) within the aforementioned environment, Determining the trend state score of the VRU, wherein the trend state score characterizes the possibility that the VRU is in a trend position or has transitioned to a trend position. Identifying one or more height differences relative to the aforementioned multiple reference points, The acceleration of the VRU, and The change in one or more height differences, or Based on at least one of the aforementioned trend status scores, it is determined that the VRU is at risk of loss of control. A system comprising a perception system that causes the control system of the autonomous vehicle to perform an avoidance maneuver.
12. The aforementioned VRU, The acceleration of the VRU that satisfies the first threshold condition, and the one or more height differences that satisfy the second threshold condition, The system according to claim 11, wherein the VRU is determined to be at risk of loss of control based on at least one of the combinations of the acceleration of the VRU and one or more height differences that satisfy a third threshold condition.
13. The aforementioned VRU, The acceleration of the VRU that satisfies the first threshold condition, and the trend state score that satisfies the second threshold condition, or The system according to claim 11, wherein the VRU is determined to be at risk of loss of control based on at least one of the combinations of the acceleration of the VRU and the trend state score that satisfies a third threshold condition.
14. It is a method, Using the autonomous vehicle's sensing system, collect sensing data about the autonomous vehicle's environment, One or more machine learning models process the sensing data to identify multiple reference points associated with vulnerable road users (VRUs) in the environment. Identifying one or more height differences relative to the aforementioned multiple reference points, Based on a change in at least one of the aforementioned height differences, it is determined that the VRU is at risk of losing control. A method comprising causing the control system of the autonomous vehicle to perform an avoidance maneuver.
15. Determining that the VRU is at risk of loss of control includes detecting that the change in one or more height differences satisfies a threshold condition. The aforementioned threshold condition is, The one or more of the aforementioned height differences experience at least a threshold change, or The method according to claim 14, comprising at least one of the following: one or more height differences experience at least a threshold change rate.
16. The method according to claim 14, wherein the determination that the VRU is at risk of loss of control is further based on the acceleration of the VRU.
17. The method according to claim 14, wherein one or more machine learning models further determine a propensity state score for the VRU, the propensity state score characterizes the likelihood that the VRU is in or transitioning to a propensity position, and the determination that the VRU is at risk of loss of control is further based on the propensity state score.
18. The determination that the VRU is at risk of loss of control is The change in one or more height differences that satisfies the first threshold condition, and the trend state score that satisfies the second threshold condition, A combination of one or more height differences and the trend state score that satisfies the third threshold condition, or The acceleration of the VRU that satisfies the fourth threshold condition, and at least one of the one or more height differences that satisfy the fifth threshold condition, or the trend state score that satisfies the sixth threshold condition, The method according to claim 17, based on at least one of the following.
19. Processing the sensing data using one or more machine learning models is Using the backbone, the sensing data is processed to generate one or more intermediate features, Using the first classifier, process one or more intermediate features and output the multiple reference points, The method according to claim 17, comprising using a second classifier to process one or more intermediate features and output the propensity state score.
20. The sensing data includes camera data and at least one of lidar data or radar data, and the sensing system A camera for collecting the aforementioned camera data, The system comprises at least one of the following: a lidar sensor for collecting lidar data, or a radar sensor for collecting radar data. In order to identify the plurality of reference points, one or more machine learning models, A portion of the camera data associated with the VRU, The method according to claim 14, trained to process at least a portion of the lidar data associated with the VRU, or a portion of the radar data associated with the VRU.