System and method for emergency vehicle detection

Through two-stage detection technology, combined with machine learning and smoother model, autonomous vehicles can accurately identify and respond to activate emergency vehicles, solving the problem of inaccurate identification in the prior art and improving the environmental perception and response capabilities of autonomous vehicles.

CN120303175APending Publication Date: 2025-07-11AURORA OPERATIONS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380078233.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-09
Filing Date
2023-11-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently identify and distinguish between activated emergency vehicles in autonomous vehicles, especially in complex environments, resulting in insufficient timely and accurate responses of autonomous vehicles.

Method used

Using two-stage detection technology, firstly identify the emergency vehicle through a machine learning model and determine its status, then confirm the activation status in the multi-frame image data through a smoother model, and optimize decisions based on the signal indicator model.

Benefits of technology

It improves the identification accuracy and response speed of autonomous vehicles for emergency vehicles, reduces the complexity of model training, and enhances the environmental perception ability of autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303175A_ABST
    Figure CN120303175A_ABST
Patent Text Reader

Abstract

Systems and methods for emergency vehicle detection are provided. An example method includes obtaining sensor data including image frames indicative of an actor in an environment of an autonomous vehicle. For respective image frames, the example method includes determining, using a machine learning model, that an actor is an emergency vehicle and generating output data indicating that the emergency vehicle is active or inactive. An example method includes storing attribute data for a respective image frame (e.g., in a buffer). The attribute data includes output data and times associated with respective image frames. Once the buffer reaches a particular threshold, the example method includes determining, based on the attribute data (e.g., using a second model), that the emergency vehicle is an active emergency vehicle. An example method includes performing an action of an autonomous vehicle based on activating an emergency vehicle within an environment of the autonomous vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 423,997, filed on November 9, 2022. U.S. Provisional Patent Application No. 63 / 423,997 is hereby incorporated by reference in its entirety. Background Art

[0003] Autonomous platforms can process data to sense the environment through which the autonomous platform can travel. For example, an autonomous vehicle can use various sensors to sense its environment and identify objects around the autonomous vehicle. The autonomous vehicle can identify an appropriate path through the sensed surrounding environment and navigate along the path with minimal or no human input. Summary of the Invention

[0004] The present disclosure relates to techniques for detecting an activated emergency vehicle within the environment of an autonomous vehicle. The detection techniques according to the present disclosure can use a combination of models to provide an improved image-based assessment of traffic actors for a granular assessment of actors at each frame level, while also detecting whether an actor is an activated emergency vehicle given a historical scenario across multiple frames.

[0005] For example, the perception system of an autonomous vehicle can sense its environment by obtaining sensor data indicative of the vehicle's surrounding environment. The sensor data can be captured over time and can include multiple image frames depicting actors within the vehicle's environment.

[0006] The autonomous vehicle can utilize the sensor data to determine that an actor is an emergency vehicle. For example, a machine learning model (e.g., a convolutional neural network with a ResNet-18 backbone) can analyze the corresponding image frame to detect whether the actor depicted in the image frame is an emergency vehicle. Example emergency vehicles can include: police cars, ambulances, fire trucks, tow trucks, etc. Actors identified as not representing emergency vehicles can be classified as "non-emergency vehicles" and can be filtered out of downstream analysis.

[0007] If the actor is an emergency vehicle, the machine learning model can further determine whether the emergency vehicle is in an activated state or a non-activated state. The machine learning model can be trained to determine whether an emergency vehicle is activated or non-activated by detecting whether a particular light bulb (e.g., a beacon mounted on the roof) on the emergency vehicle in the corresponding image frame is in an "on" state or an "off" state. An emergency vehicle with a light bulb in the "on" state can be considered activated, while an emergency vehicle with a light bulb in the "off" state can be considered non-activated. Based on this analysis, the machine learning model can generate output data indicating whether the actor within the corresponding image frame is an emergency vehicle in an activated or non-activated state.

[0008] The buffer can store attribute data including output data from a machine learning model. For example, the attribute data can include output data having an associated time (e.g., when the corresponding image frame is captured). When an image frame is processed by the machine learning model, the buffer can continue to store the attribute data for each image frame.

[0009] Once the buffer has accumulated a threshold amount of image frames across multiple times, a second model can analyze the number of image frames. For example, the second model (e.g., a rule-based smoother model) can process the image frames to determine whether more than 50% of the processed image frames indicate that an emergency vehicle is in an activated state. If so, the second model can be configured to determine that the emergency vehicle is an activated emergency vehicle and generate an output regarding it. In some examples, the second model can evaluate patterns, colors, light intensities, etc. within the image frames to help improve its confidence that an activated emergency vehicle is present.

[0010] Additionally or alternatively, according to the present disclosure, an autonomous vehicle can include a machine learning signal indicator model. The signal indicator model can be trained to analyze multiple image frames to inform its determination as to whether a signal indicator (e.g., a turn signal, a brake light, etc.) of an actor is in an activated or non-activated state. For example, this can include processing a current image frame with historical image frames at a previous time step to detect patterns indicating that the signal indicator is flashing, etc. As will be further described herein, the output from the signal indicator can be post-processed in a manner similar or different to that of the emergency vehicle model.

[0011] The autonomous vehicle can perform various actions based on the detection of an activated emergency vehicle and / or the detection of an activated signal indicator of an actor within the surrounding environment of the vehicle. For example, the autonomous vehicle can predict the movement of the activated emergency vehicle or other actor (e.g., using a activated left turn signal) to predict its future trajectory. Additionally, the vehicle's motion planning system can formulate a strategy regarding how to interact with and traverse the environment by considering its decision-level options for movement (e.g., avoid / don't avoid the emergency vehicle / actor, etc.). If needed, the autonomous vehicle can be controlled to physically maneuver in response to the activated emergency vehicle (e.g., pull over to the shoulder) or the actor (e.g., allow the actor to merge).

[0012] The two-stage detection technique of the present disclosure can provide many technical improvements to the performance of autonomous vehicles. For example, by using the described two-stage method to evaluate actors for emergency vehicle status, an autonomous vehicle can appropriately dedicate its on-vehicle computing resources to more discrete tasks at each stage. For example, a machine learning model can focus on the tasks of image processing and classification without worrying about time analysis, while a second / smoother model can focus on the aggregated results across multiple time steps. This helps reduce the complexity of training the model or constructing heuristics. In addition, the described detection technique allows an autonomous vehicle to identify / classify activated emergency vehicles in its surrounding environment with higher accuracy while also improving the motion response of the autonomous vehicle.

[0013] For example, in one aspect, the present disclosure provides an example method for detecting emergency vehicles of an autonomous vehicle. In some embodiments, the example computer-implemented method includes (a) obtaining sensor data including a plurality of image frames indicating actors in the environment of the autonomous vehicle. In some embodiments, the example method includes (b) for each corresponding image frame: (i) using a machine learning model, determining that the actor is an emergency vehicle; (ii) using a machine learning model, generating output data indicating whether the emergency vehicle is in an activated state or a non-activated state in the corresponding image frame; and (iii) storing attribute data of the corresponding image frame. The attribute data includes the output data of the machine learning model and the time associated with the corresponding image frame. In some embodiments, the example method includes (c) determining that the emergency vehicle is an activated emergency vehicle based on the output data associated with the corresponding image frame. In some embodiments, the example method includes (d) performing an action of the autonomous vehicle based on the activated emergency vehicle being within the environment of the autonomous vehicle.

[0014] In some embodiments of the example method, (b)(ii) includes: using a machine learning model, determining the state of the lights of the emergency vehicle in the corresponding image frame, where the state of the lights includes an on state or an off state; and using a machine learning model, based on the state of the lights, determining whether the emergency vehicle is in an activated state or a non-activated state for the corresponding image frame.

[0015] In some embodiments of the example method, (b)(i) includes using a machine learning model to determine the category of the emergency vehicle from one of the following categories: (1) police vehicle; (2) ambulance; (3) fire truck; or (4) tow truck.

[0016] In some embodiments of the example method, the output data indicates the category of the emergency vehicle.

[0017] In some embodiments of the example method, (b)(i) includes: obtaining tracking data of an actor within the environment of the autonomous vehicle, the tracking data indicating the boundary shape of the actor; determining that the centroid of the boundary shape is within the projection field of view of the autonomous vehicle; and in response to determining that the centroid of the boundary shape is within the projection field of view, generating input data for a machine learning model based on sensor data.

[0018] In some embodiments of the example method, the output data further includes tracking data associated with an emergency vehicle.

[0019] In some embodiments of the example method, (b)(iii) includes storing the attribute data in a buffer, and wherein the method further includes: determining that the buffer includes a threshold amount of attribute data of multiple image frames at multiple times; and using a second model to determine that the emergency vehicle is an activated emergency vehicle based on the attribute data of at least a subset of the multiple image frames.

[0020] In some embodiments of the example method, (c) includes: using a second model to determine at least one of the following: (1) the pattern of the lights of the emergency vehicle; (2) the color of the lights of the emergency vehicle; or (3) the light intensity of the lights of the emergency vehicle.

[0021] In some embodiments of the example method, the machine learning model is trained based on labeled training data, wherein the labeled training data is based on point cloud data and image data, and wherein the labeled training data indicates multiple training actors, and the corresponding training actors are labeled with emergency vehicle labels.

[0022] In some embodiments of the example method, the emergency vehicle label indicates the type of the emergency vehicle of the corresponding training actor or that the training actor is not an emergency vehicle.

[0023] In some embodiments of the example method, the corresponding training actor includes an activity label, and the activity label indicates that the corresponding training actor is in an activated state or a non-activated state based on the lights of the training actor.

[0024] In some embodiments of the example method, the machine learning model is a convolutional neural network.

[0025] In some embodiments of the example method, the second model is a rule-based smoothing model.

[0026] In some embodiments of the example method, the actions of the autonomous vehicle include at least one of the following: (1) predicting the movement of the activated emergency vehicle; (2) generating a motion plan for the autonomous vehicle; or (3) controlling the movement of the autonomous vehicle.

[0027] In some embodiments of the example method, (b)(iii) includes storing, in a buffer, attribute data for a corresponding image frame, the attribute data including output data of a machine learning model and a time associated with the corresponding image frame, and (c) includes determining that the buffer includes a threshold amount of attribute data for a plurality of image frames at a plurality of times.

[0028] In some embodiments of the example method, the machine learning model is further configured to process one or more historical image frames to generate output data, the one or more historical image frames being associated with one or more time steps prior to a time step associated with the corresponding time frame.

[0029] For example, in one aspect, the present disclosure provides one or more example non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations. In some embodiments, the operations include (a) obtaining sensor data including a plurality of image frames indicative of actors in an environment of an autonomous vehicle. The operations include: (b) for a corresponding image frame, (i) using a machine learning model to determine that the actor is an emergency vehicle, (ii) using the machine learning model to generate output data indicative of whether the emergency vehicle is in an active state or a non-active state in the corresponding image frame, and (iii) storing attribute data for the corresponding image frame, the attribute data including the output data of the machine learning model and a time associated with the corresponding image frame. In some embodiments, the operations include determining that the emergency vehicle is an active emergency vehicle based on the output data associated with the corresponding image frame. In some embodiments, the operations include performing an action of the autonomous vehicle based on the active emergency vehicle being within the environment of the autonomous vehicle.

[0030] In some embodiments of the example one or more non-transitory computer-readable media, (b)(ii) includes: using a machine learning model to determine a state of lights of the emergency vehicle in the corresponding image frame, wherein the state of the lights includes an on state or an off state; and using the machine learning model to determine, based on the state of the lights, whether the emergency vehicle is in an active state or a non-active state for the corresponding image frame.

[0031] In some embodiments of the example one or more non-transitory computer-readable media, the output data indicates a category of the emergency vehicle.

[0032] In some embodiments of the example one or more non-transitory computer-readable media, (b)(i) includes: obtaining tracking data of an actor within the environment of the autonomous vehicle, the tracking data indicative of a boundary shape of the actor; determining that a centroid of the boundary shape is within a projection field of view of the autonomous vehicle; and in response to determining that the centroid of the boundary shape is within the projection field of view, generating input data for the machine learning model based on the sensor data.

[0033] For example, in one aspect, the present disclosure provides an example autonomous vehicle control system for controlling an autonomous vehicle. In some embodiments, the example autonomous vehicle control system includes one or more processors and one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the autonomous vehicle control system to control the movement of the autonomous vehicle using an operating system. In some embodiments, the operating system detects an emergency vehicle by: (a) obtaining sensor data that includes a plurality of image frames indicative of actors in the environment of the autonomous vehicle; (b) for a respective image frame, (i) using a machine learning model, determining that the actor is an emergency vehicle, (ii) using a machine learning model, generating output data indicative of whether the emergency vehicle is in an active or non-active state in the respective image frame, and (iii) storing attribute data for the respective image frame, the attribute data including the output data of the machine learning model and the time associated with the respective image frame; and (c) determining that the emergency vehicle is an active emergency vehicle based on the output data associated with the respective image frame.

[0034] In some embodiments of the example autonomous vehicle control system, (b)(iii) includes storing the attribute data in a buffer. In some embodiments, the operating system further detects an emergency vehicle by: determining that the buffer includes a threshold amount of attribute data for a plurality of image frames at a plurality of times; and using a second model, determining that the emergency vehicle is an active emergency vehicle based on the attribute data for at least a subset of the plurality of image frames.

[0035] Other example aspects of the present disclosure relate to other systems, methods, vehicles, devices, tangible non-transitory computer-readable media, and apparatuses for performing the functions described herein. These and other features, aspects, and advantages of the various embodiments will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] A detailed description of embodiments for a person of ordinary skill in the art is set forth in the specification, with reference to the accompanying drawings, in which:

[0037] Figure 1 is a block diagram of an example operating scenario according to some embodiments of the present disclosure;

[0038] Figure 2 is a block diagram of an example system according to some embodiments of the present disclosure;

[0039] Figure 3A is a representation of an example operating environment according to some embodiments of the present disclosure;

[0040] Figure 3B is a representation of an example map of an operating environment in accordance with some embodiments of the present disclosure;

[0041] Figure 3C is a representation of an example operating environment in accordance with some embodiments of the present disclosure;

[0042] Figure 3D is a representation of an example map of an operating environment in accordance with some embodiments of the present disclosure;

[0043] Figure 4 is a block diagram of an example computing system for detecting emergency vehicles in accordance with some embodiments of the present disclosure;

[0044] Figure 5A is a block diagram of an example computing system for preprocessing image frames in accordance with some embodiments of the present disclosure.

[0045] Figure 5B presents a block diagram of an example model architecture for training and analyzing actor signal indicators in accordance with some embodiments of the present disclosure.

[0046] Figure 6-A-1-4 is an example vehicle in accordance with some embodiments of the present disclosure.

[0047] Figure 6B is a representation of an example vehicle maneuver in accordance with some embodiments of the present disclosure.

[0048] Figure 6C is a representation of an example training data in accordance with some embodiments of the present disclosure.

[0049] Figure 7 is a flowchart of an example method for detecting emergency vehicles in accordance with some embodiments of the present disclosure.

[0050] Figure 8A is a flowchart of an example method for detecting emergency vehicles using a preprocessing module in accordance with some embodiments of the present disclosure.

[0051] Figure 8B is a flowchart of an example method for detecting emergency vehicles using a machine learning model in accordance with some embodiments of the present disclosure.

[0052] Figure 9 is a flowchart of an example method for training and validating one or more models in accordance with some embodiments of the present disclosure;

[0053] Figure 10 is a flowchart of an example method for detecting signal indicators and controlling an autonomous vehicle in accordance with some embodiments of the present disclosure; and

[0054] Figure 11A block diagram of an example computing system for performing system verification in accordance with some embodiments of the present disclosure. Detailed Description

[0055] For example purposes only, the techniques of the present disclosure are described herein in the context of autonomous vehicles. As described herein, the techniques described herein are not limited to autonomous vehicles and can be implemented for or within other autonomous platforms and other computing systems.

[0056] Reference Figures 1 to 11 is discussed in more detail with respect to example embodiments of the present disclosure. Figure 1 A block diagram of an example operating scenario in accordance with example embodiments of the present disclosure. In the example operating scenario, environment 100 includes an autonomous platform 110 and a plurality of objects, the plurality of objects including a first actor 120, a second actor 130, and a third actor 140. In the example operating scenario, the autonomous platform 110 is capable of moving through the environment 100 and interacting with objects (e.g., the first actor 120, the second actor 130, the third actor 140, etc.) located within the environment 100. The autonomous platform 110 can optionally be configured to communicate with a remote system 160 via a network 170.

[0057] The environment 100 can be or include an indoor environment (e.g., within one or more facilities, etc.) or an outdoor environment. The indoor environment can be, for example, an environment enclosed by a structure such as a building (e.g., a service facility, a maintenance location, a manufacturing facility, etc.). The outdoor environment can be, for example, one or more regions in the external world, such as, for example, one or more rural areas (e.g., having one or more rural driving roads, etc.), one or more urban areas (e.g., having one or more urban driving roads, highways, etc.), one or more suburban areas (e.g., having one or more suburban driving roads, etc.), or other outdoor environments.

[0058] The autonomous platform 110 can be any type of platform configured to operate within the environment 100. For example, the autonomous platform 110 can be a vehicle configured to autonomously sense and operate within the environment 100. The vehicle can be a ground-based autonomous vehicle, such as, for example, an autonomous car, a truck, a van, etc. The autonomous platform 110 can be an autonomous vehicle capable of controlling, connecting to, or otherwise associating with implements, attachments, and / or accessories for transporting people or goods. This can include, for example, an autonomous tractor optionally coupled to a cargo trailer. Additionally or alternatively, the autonomous platform 110 can be any other type of transportation vehicle, such as one or more aircraft, water-based vehicles, space-based vehicles, other ground-based vehicles, etc.

[0059] The autonomous platform 110 can be configured to communicate with the remote system 160. For example, the remote system 160 can communicate with the autonomous platform 110 for assistance (e.g., navigation assistance, situation response assistance, etc.), control (e.g., fleet management, remote operation, etc.), maintenance (e.g., updates, monitoring, etc.), or other local or remote tasks. In some embodiments, the remote system 160 can provide data indicating tasks that the autonomous platform 110 should perform. For example, as further described herein, the remote system 160 can provide data indicating a trip / service for the autonomous platform 110 to perform, such as a user transportation trip / service, a delivery trip / service (e.g., for goods, merchandise, items), etc.

[0060] The autonomous platform 110 can communicate with the remote system 160 using the network 170. The network 170 can facilitate the transmission of signals (e.g., electrical signals, etc.) or data (e.g., data from computing devices, etc.), and can include any combination of various wired (e.g., twisted pair cables, etc.) or wireless communication mechanisms (e.g., cellular, wireless, satellite, microwave, radio frequency, etc.) or any desired network topology (or topologies). For example, the network 170 can include a local area network (e.g., an intranet, etc.), a wide area network (e.g., the Internet, etc.), a wireless LAN network (e.g., via Wi-Fi, etc.), a cellular network, a SATCOM network, a VHF network, an HF network, a WiMAX-based network, or any other suitable communication network (or combination thereof) for transmitting data to or from the autonomous platform 110.

[0061] For example, as Figure 1 shown, the environment 100 can include one or more objects. An object can be an object that is not in motion or not predicted to be in motion ("static object") or an object in motion or predicted to be in motion ("dynamic object" or "actor"). In some embodiments, the environment 100 can include any number of actors, such as, for example, one or more pedestrians, animals, vehicles, etc. Actors can move within the environment according to one or more actor trajectories. For example, the first actor 120 can move along any one of the first actor trajectories 122A-C, the second actor 130 can move along any one of the second actor trajectories 132, the third actor 140 can move along any one of the third actor trajectories 142, etc.

[0062] As further described herein, the autonomous platform 110 can utilize its autonomous system to detect these actors (and their movements) and plan its motion to navigate through the environment 100 based on one or more platform trajectories 112A-C. The autonomous platform 110 can include an on-vehicle computing system 180. The on-vehicle computing system 180 can include one or more processors and one or more memory devices. The one or more memory devices can store instructions executable by the one or more processors to cause the one or more processors to perform operations or functions associated with the autonomous platform 110, including implementing its autonomous system.

[0063] Figure 2 FIG. 4 is a block diagram of an example autonomous system 200 for an autonomous platform according to some embodiments of the present disclosure. In some embodiments, the autonomous system 200 can be implemented by a computing system of the autonomous platform (e.g., the on-vehicle computing system 180 of the autonomous platform 110). The autonomous system 200 can operate to obtain inputs from sensors 202 or other input devices. In some embodiments, the autonomous system 200 can additionally obtain platform data 208 (e.g., map data 210) from local or remote storage. The autonomous system 200 can generate control outputs for controlling the autonomous platform (e.g., via a platform control device 212, etc.) based on sensor data 204, map data 210, or other data. The autonomous system 200 can include different subsystems for performing various autonomous operations. The subsystems can include a positioning system 230, a perception system 240, a planning system 250, and a control system 260. The positioning system 230 can determine the position of the autonomous platform within its environment; the perception system 240 can detect, classify, and track objects and actors in the environment; the planning system 250 can determine the trajectory of the autonomous platform; and the control system 260 can convert the trajectory into vehicle control for controlling the autonomous platform. The autonomous system 200 can be implemented by one or more on-vehicle computing systems. The subsystems can include one or more processors and one or more memory devices. The one or more memory devices can store instructions executable by the one or more processors to cause the one or more processors to perform operations or functions associated with the subsystems. The computing resources of the autonomous system 200 can be shared among its subsystems, or the subsystems can have a set of dedicated computing resources.

[0064] In some embodiments, the autonomous system 200 can be implemented for an autonomous vehicle (e.g., a ground-based autonomous vehicle) or by an autonomous vehicle (e.g., a ground-based autonomous vehicle). The autonomous system 200 can perform various processing techniques on inputs (e.g., sensor data 204, map data 210) to sense and understand the environment around the vehicle and generate an appropriate set of control outputs to achieve vehicle motion planning (e.g., including one or more trajectories) for traversing the environment around the vehicle (e.g., Figure 1environment 100, etc.). In some embodiments, an autonomous vehicle implementing the autonomous system 200 can drive, navigate, operate, etc. with minimal interaction or no interaction from a human operator (e.g., a driver, a pilot, etc.).

[0065] In some embodiments, the autonomous platform can be configured to operate in multiple operation modes. For example, the autonomous platform can be configured to operate in a fully autonomous (e.g., self-driving, etc.) operation mode, where the autonomous platform is controllable without user input (e.g., can drive and navigate without input from a human operator present in or away from the autonomous vehicle, etc.). The autonomous platform can operate in a semi-autonomous operation mode, where the autonomous platform can operate with some input from a human operator present in the autonomous platform (or a human operator away from the autonomous platform). In some embodiments, the autonomous platform can enter a manual operation mode, where the autonomous platform is fully controllable by a human operator (e.g., a human driver, etc.) and can be prohibited or disabled (e.g., temporarily, permanently, etc.) from performing autonomous navigation (e.g., self-driving, etc.). The autonomous platform can be configured to operate in other modes, such as a parking or sleep mode (e.g., for use between tasks such as waiting to provide a trip / service, recharging, etc.). In some embodiments, the autonomous platform can implement vehicle operation assistance technologies (e.g., a collision mitigation system, power-assisted steering, etc.), e.g., to assist a human operator of the autonomous platform (e.g., when in the manual mode, etc.).

[0066] The autonomous system 200 can be located on (e.g., on or within) the autonomous platform and can be configured to operate the autonomous platform in various environments. The environment can be a real-world environment or a simulated environment. In some embodiments, one or more simulation computing devices can simulate one or more of the following: sensors 202, sensor data 204, communication interface 206, platform data 208, or platform control device 212 for simulating the operation of the autonomous system 200.

[0067] In some embodiments, the autonomous system 200 can communicate with one or more networks or other systems having a communication interface 206. The communication interface 206 can include any suitable components for docking with one or more networks (e.g., Figure 1 network 170, etc.), including, for example, a transmitter, a receiver, a port, a controller, an antenna, or other suitable components that can help facilitate communication. In some embodiments, the communication interface 206 can include multiple components (e.g., antennas, transmitters, or receivers, etc.) that allow it to implement and utilize various communication technologies (e.g., multiple-input multiple-output (MIMO) technology, etc.).

[0068] In some embodiments, the autonomous system 200 can communicate with one or more computing devices remote from the autonomous platform (e.g., remote system 160) via one or more networks (e.g., network 170) using the communication interface 206. For example, in some examples, one or more inputs, data, or functions of the autonomous system 200 can be supplemented or replaced by a remote system communicating via the communication interface 206. For example, in some embodiments, map data 210 can be downloaded to the remote system via the network using the communication interface 206. In some examples, one or more of the positioning system 230, the perception system 240, the planning system 250, or the control system 260 can be updated, influenced, nudged, transmitted, etc. by a remote system for assistance, maintenance, situational response coverage, management, etc.

[0069] The sensor 202 can be located on the autonomous platform. In some embodiments, the sensor 202 can include one or more types of sensors. For example, one or more sensors can include image capture devices (e.g., visible spectrum cameras, infrared cameras, etc.). Additionally or alternatively, the sensor 202 can include one or more depth capture devices. For example, the sensor 202 can include one or more light detection and ranging (LIDAR) sensors or radio detection and ranging (RADAR) sensors. The sensor 202 can be configured to generate point data that describes at least a portion of a three-hundred-sixty-degree view of the surrounding environment. The point data can be point cloud data (e.g., three-dimensional LIDAR point cloud data, RADAR point cloud data). In some embodiments, one or more sensors 202 for capturing depth information can be fixed to a rotating device so that the sensor 202 rotates about an axis. The sensor 202 can rotate about the axis while capturing data in spaced-apart sector groupings that describe different portions of a three-hundred-sixty-degree view of the surrounding environment of the autonomous platform. In some embodiments, one or more sensors 202 for capturing depth information can be solid state.

[0070] The sensor 202 can be configured to capture sensor data 204 indicative of or otherwise associated with at least a portion of the environment of the autonomous platform. The sensor data 204 can include image data (e.g., 2D camera data, video data, etc.), RADAR data, LIDAR data (e.g., 3D point cloud data, etc.), audio data, or other types of data. In some embodiments, the autonomous system 200 can obtain inputs from additional types of sensors such as an inertial measurement unit (IMU), altimeter, inclinometer, odometer device, position or location device (e.g., GPS, compass), wheel encoder, or other types of sensors. In some implementations, the autonomous system 200 can obtain sensor data 204 associated with a particular component or system of the autonomous platform. The sensor data 204 can indicate, for example, wheel speed, component temperature, steering angle, cargo or passenger status, etc. In some implementations, the autonomous system 200 can obtain sensor data 204 associated with environmental conditions such as ambient or weather conditions. In some implementations, the sensor data 204 can include multimodal sensor data. Multimodal sensor data can be obtained by at least two different types of sensors (e.g., of sensor 202) and can indicate static objects or actors within the environment of the autonomous platform. Multimodal sensor data can include at least two types of sensor data (e.g., camera and LIDAR data). In some embodiments, the autonomous platform can use the sensor data 204 for sensors remote from (e.g., non-vehicle-mounted) the autonomous platform. This can include, for example, sensor data 204 captured by different autonomous platforms.

[0071] The autonomous system 200 is capable of obtaining map data 210 associated with an environment in which the autonomous platform has been, is currently, or will be located. The map data 210 can provide information about the environment or geographical area. For example, the map data 210 can provide information about the identity and location of different driving roads (e.g., traffic roads, etc.), driving road segments (e.g., road sections, etc.), buildings or other items or objects (e.g., lamp posts, crosswalks, curbs, etc.); the position and orientation of boundaries or boundary markers (e.g., the position and orientation of traffic lanes, parking lanes, turning lanes, bicycle lanes, other lanes, etc.); traffic control data (e.g., the position and instructions of signs, traffic lights, other traffic control devices, etc.); obstacle information (e.g., temporary or permanent roadblocks, etc.); event data (e.g., road closures / traffic rule changes due to parades, concerts, sports events, etc.); nominal vehicle path data (e.g., indicating an ideal vehicle path such as along the center of a certain lane, etc.); or any other map data that provides information to assist the autonomous platform in understanding its surrounding environment and its relationship with the surrounding environment. In some embodiments, the map data 210 can include high-definition map information. Additionally or alternatively, the map data 210 can include sparse map data (e.g., lane maps, etc.). In some embodiments, the sensor data 204 can be fused with the map data 210 or used to update the map data 210 in real time.

[0072] The autonomous system 200 can include a positioning system 230, which can provide the autonomous platform with an understanding of its position and orientation in the environment. In some examples, the positioning system 230 can support one or more other subsystems of the autonomous system 200, such as by providing a unified local reference frame for performing, for example, perception operations, planning operations, or control operations.

[0073] In some embodiments, the positioning system 230 can determine the current position of the autonomous platform. The current position can include a global position (e.g., with respect to a georeferenced anchor point, etc.) or a relative position (e.g., with respect to an object in the environment, etc.). The positioning system 230 generally can include any device or circuit for analyzing the position or position change of the autonomous platform (e.g., an autonomous ground-based vehicle, etc.) or interfacing therewith. For example, the positioning system 230 can determine the position by using one or more of the following: inertial sensors (e.g., inertial measurement units, etc.), satellite positioning systems, radio receivers, networked devices (e.g., based on IP addresses, etc.), triangulation or proximity to network access points or other network components (e.g., cell towers, Wi-Fi access points, etc.), or other suitable technologies. The position of the autonomous platform can be used by various subsystems of the autonomous system 200 or provided to a remote computing system (e.g., using the communication interface 206).

[0074] In some embodiments, the positioning system 230 is capable of registering the relative positions of elements of the surrounding environment of the autonomous platform with the recorded positions in the map data 210. For example, the positioning system 230 is capable of processing sensor data 204 (e.g., LIDAR data, RADAR data, camera data, etc.) for alignment or otherwise registration to a map of the surrounding environment (e.g., from the map data 210) to understand the position of the autonomous platform within the environment. Thus, in some embodiments, the autonomous platform is able to identify its position within the surrounding environment (e.g., across six axes, etc.) based on a search on the map data 210. In some embodiments, given an initial position, the positioning system 230 is able to update the position of the autonomous platform with incremental re-alignment based on the recorded or estimated deviation from the initial position. In some embodiments, the position can be directly registered within the map data 210.

[0075] In some embodiments, the map data 210 can include a large amount of data subdivided into geographical blocks such that a desired region of the map stored in the map data 210 can be reconstructed from one or more blocks. For example, a plurality of blocks selected from the map data 210 can be stitched together by the autonomous system 200 based on the position obtained by the positioning system 230 (e.g., a plurality of blocks selected near the position).

[0076] In some embodiments, the positioning system 230 is able to determine the (e.g., relative or absolute) position of one or more attachments or accessories of the autonomous platform. For example, the autonomous platform can be associated with a cargo platform, and the positioning system 230 can provide the position of one or more points on the cargo platform. For example, the cargo platform can include a trailer or other device towed or otherwise attached to or manipulated by the autonomous platform, and the positioning system 230 can provide data describing the position of the autonomous platform as well as the cargo platform (e.g., absolute, relative, etc.). Such information can be obtained by other autonomous systems to assist in operating the autonomous platform.

[0077] The autonomous system 200 can include a perception system 240, which can allow the autonomous platform to detect, classify, and track objects and actors in its environment. Environmental features or objects perceived within the environment can be those within the field of view of the sensor 202 or predicted to be occluded relative to the sensor 202. This can include objects that are not in motion or not predicted to move (static objects) or objects that are in motion or predicted to be in motion (dynamic objects / actors).

[0078] The perception system 240 is capable of determining one or more states (e.g., current or past states, etc.) of one or more objects within the surrounding environment of the autonomous platform. For example, the state can (e.g., for a given time, time period, etc.) describe the estimated or past position (also referred to as location) of the object; the current or past speed / velocity; the current or past acceleration; the current or past direction of travel; the current or past orientation; the size / footprint (e.g., as represented by a bounding shape, object highlighting, etc.); the classification (e.g., pedestrian class versus vehicle class versus bicycle class, etc.); the uncertainty associated therewith; or other state information. In some embodiments, the perception system 240 can use one or more algorithms or machine learning models configured to identify / classify objects based on inputs from the sensors 202 to determine the state. The perception system can use different modalities of the sensor data 204 to generate a representation of the environment to be processed by one or more algorithms or machine learning models. In some embodiments, as the autonomous platform continues to perceive objects or interact with objects (e.g., maneuver or bypass, yield, etc.), the state of one or more identified or unidentified objects can be maintained and updated over time. In this way, the perception system 240 can provide an understanding of the current state of the environment (e.g., including the objects therein) informed by a record of the previous state of the environment (e.g., including the movement history of the objects therein). Such information can assist the autonomous platform in planning its movement through the environment.

[0079] The autonomous system 200 can include a planning system 250, which can be configured to determine how the autonomous platform interacts with and moves within its environment. The planning system 250 can determine one or more motion plans for the autonomous platform. A motion plan can include one or more trajectories (e.g., motion trajectories) indicating the path that the autonomous platform is to follow. A trajectory can have a certain length or time extent. The length or time extent can be defined by the computational planning horizon of the planning system 250. A motion trajectory can be defined by one or more waypoints (with associated coordinates). A waypoint can be a future position of the autonomous platform. The motion plan can be continuously generated, updated, and considered by the planning system 250.

[0080] The motion planning system 250 can determine a strategy for the autonomous platform. A strategy can be a set of discrete decisions made by the autonomous platform (e.g., yield to an actor, reverse yield to an actor, merge, lane change). The strategy can be selected from multiple potential strategies. The selected strategy can be the lowest-cost strategy determined by one or more cost functions. The cost function can, for example, evaluate the probability of a collision with another actor or object.

[0081] The planning system 250 can determine the desired trajectory for executing a policy. For example, the planning system 250 can obtain one or more trajectories for executing one or more policies. The planning system 250 can evaluate the trajectories or policies (e.g., using scores, costs, rewards, constraints, etc.) and rank them. For example, the planning system 250 can use the prediction output indicating the interaction (e.g., proximity, intersection points, etc.) between the trajectory of the autonomous platform and one or more objects to evaluate the candidate trajectories or policies of the autonomous platform. In some embodiments, the planning system 250 can use static costs to evaluate the trajectory of the autonomous platform (e.g., "avoid lane boundaries", "minimize jerk", etc.). Additionally or alternatively, the planning system 250 can use dynamic costs to evaluate the trajectory or policy for the autonomous platform based on the prediction results of the current operating scenario (e.g., the predicted trajectories or policies that result in interactions between actors, the predicted trajectories or policies that result in interactions between an actor and the autonomous platform, etc.). The planning system 250 can rank the trajectories based on one or more static costs, one or more dynamic costs, or a combination thereof. The planning system 250 can select a motion plan (and the corresponding trajectory) based on the ranking of multiple candidate trajectories. In some embodiments, the planning system 250 can select the highest-ranked candidate or the highest-ranked feasible candidate.

[0082] Then, the planning system 250 can verify the selected trajectory against one or more constraints before the autonomous platform executes the trajectory.

[0083] To assist in its motion planning decisions, the planning system 250 can be configured to perform a prediction function. The planning system 250 can predict the future state of the environment. This can include predicting the future state of other actors in the environment. In some embodiments, the planning system 250 can predict the future state based on the current or past state (e.g., as developed or maintained by the perception system 240). In some embodiments, the future state can be or include the predicted trajectory (e.g., position over time) of an object (such as another actor) in the environment. In some embodiments, one or more future states can include one or more probabilities associated therewith (e.g., marginal probability, conditional probability). For example, one or more probabilities can include one or more probabilities conditioned on the available policy or trajectory options of the autonomous platform. Additionally or alternatively, the probabilities can include probabilities conditioned on the available trajectory options of one or more other actors.

[0084] In some embodiments, the planning system 250 can perform an interactive prediction. The planning system 250 can use an understanding of how the predicted future state of the environment can be affected by the execution of one or more candidate motion plans to determine the motion plan of the autonomous platform. As an example, referring again to Figure 1, the autonomous platform 110 can determine candidate motion plans corresponding to a set of platform trajectories 112A-C that correspond to first actor trajectories 122A-C of the first actor 120, trajectory 132 of the second actor 130, and trajectory 142 of the third actor 140, respectively (e.g., having respective trajectory correspondences indicated by matching line patterns). For example, the autonomous platform 110 (e.g., using its autonomous system 200) can predict that a platform trajectory 112A that more quickly moves the autonomous platform 110 into an area in front of the first actor 120 may be associated with reducing forward speed and more quickly avoiding the first actor 120 of the autonomous platform 110 according to the first actor trajectory 122A. Additionally or alternatively, the autonomous platform 110 can predict that a platform trajectory 112B that gently moves the autonomous platform 110 into an area in front of the first actor 120 may be associated with slightly reducing speed and slowly avoiding the first actor 120 of the autonomous platform 110 according to the first actor trajectory 122B. Additionally or alternatively, the autonomous platform 110 can predict that a platform trajectory 112C that maintains parallel alignment with the first actor 120 may be associated with the first actor 120 not avoiding the autonomous platform 110 by any distance according to the first actor trajectory 122 C. Based on a comparison of the predicted scenarios to a set of desired outcomes (e.g., by scoring the scenarios based on costs or rewards), the planning system 250 can select a motion plan (and its associated trajectory) in light of the autonomous platform's interaction with the environment 100. In this manner, for example, the autonomous platform 110 can interweave its prediction and motion planning functions.

[0085] To implement the selected motion plan, the autonomous system 200 can include a control system 260 (e.g., a vehicle control system). Generally, the control system 260 can provide an interface between the autonomous system 200 and the platform control device 212 for implementing the strategies and motion plans generated by the planning system 250. For example, the control system 260 can implement the selected motion plan / trajectory by following the selected trajectory (e.g., including waypoints therein) to control the movement of the autonomous platform through its environment. The control system 260 can, for example, convert the motion plan into instructions for the appropriate platform control device 212 (e.g., acceleration control, braking control, steering control, etc.). As an example, the control system 260 can convert the selected motion plan into instructions to adjust a steering component (e.g., steering angle) by a specific degree, apply a braking force of a certain magnitude, increase / decrease speed, etc. In some embodiments, the control system 260 can communicate with the platform control device 212 via a communication channel, which can include, for example, one or more data buses (e.g., Controller Area Network (CAN), etc.), on-board diagnostic connectors (e.g., OBD-II, etc.), or a combination of wired or wireless communication links. The platform control device 212 can send or (or vice versa) obtain data, messages, signals, etc. from the autonomous system 200 via the communication channel.

[0086] The autonomous system 200 can receive an assistance signal from the remote assistance system 270 via the communication interface 206. The remote assistance system 270 can communicate with the autonomous system 200 via a network (e.g., as a remote system 160 on the network 170). In some embodiments, the autonomous system 200 can initiate a communication session with the remote assistance system 270. For example, the autonomous system 200 can initiate the session based on or in response to a trigger. In some embodiments, the trigger can be an alarm, an error signal, a map feature, a request, a location, a traffic condition, a road condition, etc.

[0087] After initiating the session, the autonomous system 200 can provide the remote assistance system 270 with the context data. The context data can include the sensor data 204 and the status data of the autonomous platform. For example, the context data can include a real-time camera feed from a camera of the autonomous platform and the current speed of the autonomous platform. An operator of the remote assistance system 270 (e.g., a human operator) can use the context data to select the assistance signal. The assistance signal can provide values or adjustments for various operating parameters or characteristics of the autonomous system 200. For example, the assistance signal can include road points (e.g., a path around an obstacle, a lane change, etc.), speed or acceleration curves (e.g., speed limits, etc.), relative motion instructions (e.g., platoon formation, etc.), operating characteristics (e.g., use of the assistance system, a reduced energy processing mode, etc.), or other signals that assist the autonomous system 200.

[0088] The autonomous system 200 can use auxiliary signals as inputs to one or more autonomous subsystems for performing autonomous functions. For example, the planning subsystem 250 can receive an auxiliary signal as an input for generating a motion plan. For example, the auxiliary signal can include constraints for generating a motion plan. Additionally or alternatively, the auxiliary signal can include cost or reward adjustments for influencing the motion plan made by the planning subsystem 250. Additionally or alternatively, the auxiliary signal can be regarded by the autonomous system 200 as a suggested input to be considered in addition to other received data (e.g., sensor inputs, etc.).

[0089] The autonomous system 200 can be platform agnostic, and the control system 260 can provide control instructions for various different platforms (e.g., multiple different autonomous platforms adapted to the autonomous control system) for autonomous movement to the platform control device 212. This can include various different types of autonomous vehicles (e.g., sedans, vans, SUVs, trucks, electric vehicles, internal combustion engine-powered vehicles, etc.) from various different manufacturers / developers operating in various different environments, and in some embodiments, perform one or more vehicle services.

[0090] For example, referring to Figure 3A , the operating environment can include a dense environment 300. The autonomous platform can include an autonomous vehicle 310 controlled by the autonomous system 200. In some embodiments, the autonomous vehicle 310 can be configured for maneuverability in a dense environment such as having a configured wheelbase or other specifications. In some embodiments, the autonomous vehicle 310 can be configured to transport goods or passengers. In some embodiments, the autonomous vehicle 310 can be configured to transport a number of passengers (e.g., passenger cars, shuttles, buses, etc.). In some embodiments, the autonomous vehicle 310 can be configured to transport goods, such as large quantities of goods (e.g., trucks, vans, walk-in vans, etc.) or smaller goods (e.g., food, personal packages, etc.).

[0091] Referring to Figure 3B, a selected top view 302 of a dense environment 300 is shown, which is covered with an example trip / service between a first location 304 and a second location 306. The example trip / service can be assigned to an autonomous vehicle 320 by, for example, a remote computing system. The autonomous vehicle 320 can be, for example, a vehicle of the same type as the autonomous vehicle 310. The example trip / service can include transporting passengers or goods between the first location 304 and the second location 306. In some embodiments, the example trip / service can include traveling to or through one or more intermediate locations, such as to load or unload passengers or goods. In some embodiments, the example trip / service can be pre-scheduled (e.g., for regular traversals, such as with respect to transportation scheduling). In some embodiments, the example trip / service can be on-demand (e.g., as requested by or for a taxi, ride-sharing, car-hailing, courier, delivery service, etc.).

[0092] Reference Figure 3C , in another example, the operating environment can include an open road driving environment 330. The autonomous platform can include an autonomous vehicle 350 controlled by the autonomous system 200. This can include an autonomous tractor for an autonomous truck. In some embodiments, the autonomous vehicle 350 can be configured for high payload transportation (e.g., transporting a large amount of goods or other cargo or passengers), such as long-distance, high payload transportation. For example, the autonomous vehicle 350 can include one or more cargo platform attachments, such as a trailer 352. Although depicted in FIG. 3 as a tow attachment, in some embodiments, one or more cargo platforms can be integrated into the autonomous vehicle 350 (e.g., attached to the chassis, etc.) (e.g., as in a van, walk-in van, etc.).

[0093] Reference Figure 3D, shows a selected top view of an open driving road environment 330, including a driving road 332, an overpass 334, transfer hubs 336 and 338, an access driving road 340, and locations 342 and 344. In some embodiments, an autonomous vehicle (e.g., autonomous vehicle 310 or autonomous vehicle 350) can be assigned an exemplary trip / service to traverse one or more driving roads 332 (optionally connected by an overpass 334) to transport goods between transfer hub 336 and transfer hub 338. For example, in some embodiments, the exemplary trip / service includes a goods delivery / transport service, such as a merchandise delivery / transport service. The exemplary trip / service can be assigned by a remote computing system. In some embodiments, transfer hub 336 can be a starting point for goods (e.g., a storage depot, a warehouse, a facility, etc.), and transfer hub 338 can be a destination location for the goods (e.g., a retailer, etc.). However, in some embodiments, transfer hub 336 can be an intermediate point in the final journey of a goods item between its respective origin and respective destination. For example, a goods item can be located at location 342 along access driving road 340. The goods item can be transported accordingly (e.g., by a human-driven vehicle, by autonomous vehicle 310, etc.) to transfer hub 336 for staging operations. At transfer hub 336, various goods items can be grouped or staged for longer-distance transportation on driving road 332.

[0094] In some embodiments of the exemplary trip / service, a group of staged goods items can be loaded onto an autonomous vehicle (e.g., autonomous vehicle 350) for transportation to one or more other transfer hubs, such as transfer hub 338. For example, although not depicted, it should be understood that open driving road environment 330 can include more transfer hubs than transfer hubs 336 and 338, and can include more driving roads 332 interconnected by more overpasses 334. A simplified map is presented here only for clarity purposes. In some embodiments, one or more goods items transported to transfer hub 338 can be assigned to one or more local destinations (e.g., by a human-driven vehicle, by autonomous vehicle 310, etc.), such as to location 344 along access driving road 340. In some embodiments, the exemplary trip / service can be pre-scheduled (e.g., for regular traversal, such as in a transportation schedule). In some embodiments, the exemplary trip / service can be on-demand (e.g., as requested by or for performing a chartered passenger transportation or merchandise delivery service).

[0095] To improve the performance of an autonomous platform, such as an autonomous vehicle that is at least partially controlled using autonomous system 200 (e.g., autonomous vehicle 310 or 350), the perception system 240 can detect emergency vehicles in accordance with example aspects of the present disclosure.

[0096] Figure 4 is a block diagram of a detection system 407 according to some embodiments of the present disclosure. The detection system 407 can be included in an emergency vehicle detection system and / or a vehicle light detection system within a perception system 240 of, for example, an autonomous vehicle. Although Figure 4 illustrates an example embodiment of a detection system 407 having various components, it should be understood that the components can be rearranged, combined, omitted, etc. within the scope of and consistent with the present disclosure.

[0097] The detection system 407 can include a preprocessing module 400, an inference module 401, a buffer 403, and a postprocessing module 404. In some examples, the inference module 401 can include a machine learning emergency vehicle model 405. In some examples, the postprocessing module 404 can include a smoother model 406.

[0098] To assist in detecting an emergency vehicle or activating a vehicle signal indicator, the detection system 407 can obtain sensor data 204. As described herein, the sensor data 204 can include data captured by one or more sensors 202 on an autonomous vehicle. This can include radar data, LIDAR data, image data, etc. For example, the sensor data 204 can include image frames captured during an instance of real-world driving and the associated times at which objects in the environment are perceived. The sensor data 204 can include data collected from other sources (e.g., roadside cameras, aircraft, etc.).

[0099] The sensor data 204 can be associated with multiple times. For example, the sensor data 204 can include multiple image frames indicating actors in the environment of the autonomous vehicle. Each respective image frame can be associated with the time / timestamp at which the image frame was captured. For example, the multiple image frames can include a series of image frames taken across multiple times and depicting actors in the environment. The actors can include, for example, another vehicle. The environment can be, for example, the external and surrounding environment of the autonomous vehicle (e.g., within the sensor field of view). In some embodiments, the sensor data 204 can include video data. Additionally or alternatively, the sensor data 204 can include multiple individual static images.

[0100] The detection system 407 can preprocess the sensor data 204. The preprocessing can be performed by the preprocessing module 400.

[0101] FIG. 5 is a block diagram of an example data flow for preprocessing sensor data according to some embodiments of the present disclosure. In FIG. 5, at a time after the sensor 202 has captured the sensor data 204, the sensor data 204 can be processed by the preprocessing module 400.

[0102] For example, the preprocessing module 400 can obtain an image frame 505 depicting a portion of the environment of the autonomous vehicle and an actor 500. The preprocessing module 400 can obtain tracking data 504 of an actor within the environment of the autonomous vehicle. The tracking data 504 can be generated by another system of the autonomous vehicle. The tracking data 504 can include a trace for the actor depicted in the corresponding image frame 505. The trace can include the state data and boundary shape 501 of the actor. The state data can include the position, speed, acceleration, etc. of the actor when the actor is sensed.

[0103] The boundary shape 501 can be a shape (e.g., a polygon) that includes the actor 500 depicted in the corresponding image frame. For example, as shown in FIG. 5, the boundary shape 501 can include a square that encapsulates the actor 500 (e.g., a bounding box). Those of ordinary skill in the art will understand that other shapes, such as circles, etc., can be used. In some embodiments, the boundary shape 501 can include a shape that matches the outermost boundary / perimeter of the actor 500 and the contour of those boundaries. The boundary shape can be generated at the per-pixel level. The tracking data 504 can include the x, y, z coordinates of the center of the boundary shape and the length, width, and height of the boundary shape. In some examples, the state of the trace can fit a multivariate normal distribution.

[0104] To help determine the relevant data set, the preprocessing module 400 can analyze the actors based on the centroids of their associated boundary shapes. For example, the preprocessing module 400 can determine that the centroid of the boundary shape 501 is within the projection field of view of the autonomous vehicle. In response to determining that the centroid of the boundary shape 500 is within the projection field of view, the preprocessing module 400 can generate input data 502 for the emergency vehicle model 405 for machine learning based on the sensor data 204.

[0105] In some examples, the preprocessing module 400 identifies the corresponding coordinates of the corners of each trace of the boundary shape 501. The preprocessing model 400 can project the three-dimensional coordinates of the corners of the boundary shape 501 of each trace into the forward camera image. For each actor whose projected centroid is within the image, the preprocessing module 400 can crop the image using the smallest shape (e.g., a square, etc.) that encapsulates the projected corners and resize it to a predetermined length and width (e.g., 224x224 in length and width).

[0106] In some examples, actors whose centroids are outside the field of view of sensor 202 can be ignored. In other examples, actors whose projected width is less than a threshold amount (e.g., L = 18.6 pixels) can be ignored. By ignoring actors / tracks whose centroids are outside the field of view of sensor 200 or whose projected width is less than the threshold, actors that are too far from the autonomous vehicle will not be evaluated. This can help focus the vehicle on in-vehicle computing resources (e.g., power, processing, memory, bandwidth) and avoid unnecessary use on actors that may not be recognized in a given frame with sufficient confidence.

[0107] The preprocessing module 400 can generate valid image patches from the cropped image frame 502. Image patches that have not been ignored can be considered valid. The image patches can be batch processed into valid batch crops. In some examples, the valid batch crops can be ranked from most important to least important. For example, batch processing the valid image crops and ranking the batch crops improves the system latency by first processing the most important batch crops.

[0108] The preprocessing module 400 can provide the valid batch crops as input data 503 to the inference module 401 for an emergency vehicle model 405 for machine learning. The inference module 401 can run the input data 503 through the emergency vehicle model 405 and output the probability of determining whether the actor 500 in the corresponding image frame 505 is an activated emergency vehicle.

[0109] The emergency vehicle model 405 can include one or more machine learning models that are trained to determine whether an actor is an emergency vehicle and whether the emergency vehicle is in an activated state. The emergency vehicle model 405 can be or otherwise include various machine learning models, such as, for example, regression networks, generative adversarial networks, neural networks (e.g., deep neural networks), support vector machines, decision trees, ensemble models, k-nearest neighbor models, Bayesian networks, or other types of models that include linear or non-linear models. Example neural networks include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks.

[0110] An emergency vehicle model 405 can be trained by using one or more model trainers and training data. One or more training or learning algorithms can be used to train the model trainer. An example training technique is backpropagation of errors. In some examples, simulations can be implemented for obtaining training data or for implementing model trainers for training or testing the model. In some examples, the model trainer can use labeled training data to perform supervised training techniques. As further described herein, the training data can include labeled image frames that have labels indicating whether an actor is an emergency vehicle, the type of the emergency vehicle, and the bulb status of the headlights of the emergency vehicle. In some examples, the training data can include simulated training data (e.g., training data obtained from simulated scenarios, inputs, configurations, environments, etc.).

[0111] Additionally or alternatively, the model trainer can use unlabeled training data to perform unsupervised training techniques. As an example, the model trainer can train one or more components of a machine learning model to perform emergency vehicle detection by using unsupervised training techniques of an objective function (e.g., cost, reward, heuristic, constraint, etc.). In some embodiments, the model trainer can perform multiple generalization techniques to improve the generalization ability of the trained model. Generalization techniques include weight decay, dropout, or other techniques.

[0112] In some examples, the emergency vehicle model 405 can be a convolutional neural network. The convolutional neural network can include, for example, ResNet-18 as a backbone to extract image features, followed by an average pooling layer and a linear layer with a single output channel, and then a sigmoid function to output a probability as to whether the actor is an emergency vehicle. Example model input dimensions can be represented as [B, C, W, H], where B is the batch size, C is the RGB channels, and W & H represent the image size (width and height). In one example, the batch size (B) can be 32, the number of RBG channels (C) can be 3, and the image size (W = H) can be 224.

[0113] The ResNet backbone can use pre-trained weights on an image dataset that is not frozen during training. The image dataset can be organized according to hierarchies, where each node of the hierarchy is depicted by hundreds / thousands of images and grouped into multiple sets of synchronized sets, and each synchronized set expresses a different concept. The synchronized sets can be interconnected by means of concept semantics and lexical relationships.

[0114] The emergency vehicle model 405 can use focal loss, which can be good for unbalanced datasets. The Adam optimizer with 0 weight decay and an initial learning rate of 10-4 can be used. The model trainer can use a custom optimizer wrapper that will reduce the learning rate when the loss has not improved for a certain number of iterations and will stop training when the learning rate has dropped to a given low value.

[0115] The emergency vehicle model 405 can merge the labels into binary targets. In an example embodiment, when an image frame contains an emergency vehicle and has an activated (or bulb-on) state when the image is captured, a positive target can be determined. In an example embodiment, when the image contains a non-emergency vehicle, a negative target can be determined. In an example embodiment, when the image contains an emergency vehicle with a non-activated (or bulb-off) state, a negative target can be determined. The emergency vehicle model 405 can ignore emergency vehicles with a non-activated (or bulb-off) state. In this way, the emergency vehicle model 405 can effectively identify activated emergency vehicles of interest to the motion planning system 250.

[0116] For a corresponding image frame, the detection system 407 can use the emergency vehicle model 405 to determine that an actor (depicted in the image frame) is an emergency vehicle. In some examples, the emergency vehicle model 405 can determine (and output) the class of the emergency vehicle from one of the following classes: police car, ambulance, fire truck, tow truck, or other types of emergency vehicles. The emergency vehicle model 405 can determine that the actor is not an emergency vehicle. For example, the emergency vehicle model 405 can determine the class of a non-emergency vehicle for a vehicle that is not an emergency vehicle.

[0117] As an example, the emergency vehicle model 405 can analyze the actor in the cropped image frame to determine the probability that the actor is an emergency vehicle. This can include analyzing the shape or position of the actor to determine if the actor is a police car, ambulance, fire truck, tow truck, or other type of emergency vehicle. In some examples, the emergency vehicle model 405 can analyze the shape or position of the lights (e.g., roof-mounted lights) of the actor to help determine if the actor is an emergency vehicle. The presence of long roof-mounted lights on a sedan can increase the probability that the actor is a police car.

[0118] The probability can reflect the confidence of the model that the actor is an emergency vehicle. In some examples, a probability above a threshold probability (e.g., 50%, 75%, 90%, etc.) can result in a positive detection of an emergency vehicle. In some examples, a probability below the threshold probability can result in determining that the actor is not an emergency vehicle in the corresponding image frame. Figure 6A-4 Examples of actors that are not emergency vehicles in the corresponding image frame are shown.

[0119] The emergency vehicle model 405 can predict whether an emergency vehicle is activated or deactivated based on a corresponding image frame. For example, the emergency vehicle model 405 can determine the state of the lights of the emergency vehicle in the corresponding image frame. The state can include an "on" state where the bulbs of the indicator lights are on / illuminated or an "off" state where the bulbs of the indicator lights are off / not illuminated.

[0120] In some examples, the emergency vehicle model 405 can determine whether a bulb is on or off based on the characteristics of the pixels associated with the lights (e.g., color, light intensity, brightness, etc.). In some examples, the characteristics can be compared with the characteristics of other pixels in the corresponding image frame. The characteristics can include, for example, color, pattern, intensity, brightness, etc. The pattern, color, light intensity, brightness, etc. can indicate whether the bulb in the corresponding image frame is on or off.

[0121] The emergency vehicle model 405 can determine whether the emergency vehicle is in an activated state or a deactivated state for the corresponding image frame based on the state of the lights. The emergency vehicle 405 can determine that the emergency vehicle is activated when the lights are in the on state. The emergency vehicle can determine that the emergency vehicle is deactivated when the lights are in the off state. In some examples, the color can indicate the type of the emergency vehicle. For example, the bulb label can indicate that the emergency vehicle is an activated police car. Example image frames with activated emergency vehicles are shown in Figure 6A-1 , Figure 6A-2 and Figure 6A-3 are shown.

[0122] As Figure 4 depicted, the emergency vehicle model 405 can generate output data 408 indicating whether the emergency vehicle is in an activated state or a deactivated state in the corresponding image frame. The output data 408 can include the image frame and the probability (as determined by the emergency vehicle model 405) that the emergency vehicle is in an activated state or a deactivated state in the corresponding image frame.

[0123] In some embodiments, the output data 408 is stored in a buffer 403. The buffer 403 can include, for example, a circular buffer of a distributed ledger (ledger) having multiple valid outputs (e.g., the last 25 valid outputs) of the emergency vehicle model 405 with a history from each vehicle trace. In some examples, the buffer 403 can be included in a post-processing module 404. In some examples, the buffer 403 can be implemented as an intermediary between the inference module 401 and the post-processing module 404.

[0124] In some embodiments, as further described herein, the buffer 403 is omitted from the detection system 407.

[0125] The output data 408 can be stored as attribute data 409. The attribute data 409 can include the output data 408 and the time associated with the corresponding image frame. The time associated with the corresponding image frame can be the time when the corresponding image frame is captured. In some examples, the attribute data 409 can include the tracking data 504 associated with the emergency vehicle. This can include associating the trace with the corresponding image frame of the detected emergency vehicle. In some examples, the output data 408 / attribute data 409 can indicate the category of the emergency vehicle (e.g., police car, etc.).

[0126] The detection system 407 can determine that there is a threshold amount of attribute data for multiple image frames at multiple times. In an example, the detection system 407 can determine that the buffer 403 includes a threshold amount of attribute data 409 for multiple image frames at multiple times. The threshold amount of attribute data can include a threshold amount of image frames (e.g., 6 frames) that have been processed by the emergency vehicle model 405.

[0127] The detection system 407 can make a final determination that the emergency vehicle is an activated emergency vehicle based on the attribute data 409 using a second model. As Figure 4 depicted, the post-processing module 404 can obtain the image frames from the buffer 403. The post-processing module 404 can include a smoother model 406. The smoother model 406 can include a downstream smoother that processes the bulb switch cycle and infers the final overall state of the emergency vehicle.

[0128] For example, the smoother model 406 can analyze multiple image frames (across multiple times) and their activation / inactivation labels in the attribute data 409, and make a final determination as to whether the depicted emergency vehicle is in an activated state or an inactivated state. The smoother model 406 can determine whether the emergency vehicle is an activated emergency vehicle by calculating that more than 50% of the multiple image frames indicate that the emergency vehicle is in an activated state.

[0129] In some examples, the smoother model 406 can determine at least one of the following: (1) the pattern of the lights of the emergency vehicle; (2) the color of the lights of the emergency vehicle; or (3) the light intensity of the lights of the emergency vehicle. Additionally or alternatively, the smoother model can determine the type of the emergency vehicle.

[0130] As an example, the smoother model 406 can determine the color of the activated light bulb based on multiple image frames stored in the buffer 403. The smoother model 406 can determine that the flashing light bulb is blue by calculating the color tags containing blue in more than 50% of the image frames stored in the buffer 403. In some embodiments, the smoother model 406 can determine the category of the activated emergency vehicle. For example, the smoother model 406 can determine that the activated emergency vehicle is a police car by calculating that more than 50% of the image frames stored in the buffer 403 have been classified as police car tags.

[0131] In some examples, the smoother model 406 can include a rule-based model that includes a heuristic set of rules. A rule set can be developed to evaluate multiple image frames as described herein.

[0132] In some examples, the smoother model 406 can include one or more machine learning models. This can include one or more machine learning models that are trained to determine whether an actor is an activated emergency vehicle given multiple image frames. The emergency vehicle model 405 can be or otherwise include various machine learning models such as, for example, regression networks, generative adversarial networks, neural networks (e.g., deep neural networks), support vector machines, decision trees, ensemble models, k-nearest neighbor models, Bayesian networks, or other types of models including linear or non-linear models. Example neural networks include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. The one or more models can be trained by using one or more model trainers and training data. One or more training or learning algorithms can be used to train the model trainers to train the models based on multiple image frames to determine whether an emergency vehicle is activated or not, the type of the emergency vehicle, characteristics, etc.

[0133] Actions of the autonomous vehicle can be performed based on the activated emergency vehicle within the environment of the autonomous vehicle. For example, data indicating the emergency vehicle within the vehicle's environment can be provided to the planning system 250. The motion planning system 250 can predict the motion of the activated emergency vehicle in the same manner as an actor perceived by the autonomous vehicle as described herein.

[0134] The motion planning system 250 can generate a motion plan for the autonomous vehicle based on the activated emergency vehicle. This can include, for example, generating trajectories for the autonomous vehicle to decelerate to provide more distance between the autonomous vehicle and the activated emergency vehicle, allowing the activated emergency vehicle to pass, allowing the activated emergency vehicle to merge into the road, etc. In some examples, the trajectory can include the autonomous vehicle changing lanes or the activated emergency vehicle pulling over.

[0135] In some examples, the motion planning system 250 is capable of considering an activated emergency vehicle in its trajectory generation and determining that the autonomous vehicle does not need to change acceleration, speed, direction, etc., because the autonomous vehicle is already appropriately positioned relative to the activated emergency vehicle. This can include scenarios where the activated emergency vehicle is already sufficiently positioned in front of the autonomous vehicle.

[0136] The actions of the autonomous vehicle can include controlling the movement of the autonomous vehicle based on the activated emergency vehicle. The motion planning system 250 can provide data indicative of a trajectory generated in the environment based on the activated emergency vehicle. The control system 260 can control the maneuvers of the autonomous vehicle based on the trajectory, as described herein. Figure 6B An example vehicle maneuver in the case of an activated emergency vehicle within the environment of the autonomous vehicle is illustrated.

[0137] Return Figure 4 , the detection system 407 can include a signal indicator model 410. The signal indicator model 410 can include one or more models configured to detect signals of objects in the surrounding environment of the autonomous vehicle. This can include, for example, the signal lights of vehicles in the surrounding environment.

[0138] As described herein, the detection system 407 can obtain sensor data 204. The sensor data 204 can include multiple image frames indicative of actors in the environment of the autonomous vehicle. The multiple image frames can include a current image frame (e.g., at the current time step t) and one or more historical image frames. The historical image frames can be associated with one or more previous time steps (e.g., t - 1, t - 2, etc.) from the current time step t of the current image frame.

[0139] As informed by the historical image frames, the signal indicator model 410 can include one or more models trained to determine characteristics of signal indicators for actors within the current image frame. For example, the signal indicator model 410 can be a convolutional neural network trained using one or more training techniques.

[0140] The signal indicator model 410 can be trained by using one or more model trainers and training data. One or more training or learning algorithms can be used to train the model trainer. An example training technique is backpropagation of errors. In some examples, simulations can be implemented for obtaining training data or for implementing the model trainer for training or testing the model. In some examples, the model trainer can use labeled training data to perform supervised training techniques. The training data can include labeled image frames with labels indicating the characteristics of the signal indicator. The characteristics can include: the signal indicator of the object, the type of the signal indicator (e.g., turn signal, brake signal, hazard signal), the position of the indicator on / relative to the actor (e.g., left, right, front, back), the bulb state of the signal indicator (e.g., whether the light is on or off), color, or other characteristics. In some examples, the training data can include simulated training data (e.g., training data obtained from simulated scenarios, inputs, configurations, environments, etc.). Figure 6C Depicts an example training data 600 including multiple image frames with labeled characteristics of signal indicators depicted in the image frames. The image frames can also include metadata indicating their corresponding time steps (e.g., t, t-1, t-2, etc.).

[0141] Additionally or alternatively, the model trainer can use unlabeled training data to perform unsupervised training techniques. As an example, the model trainer can train one or more components of a machine learning model to perform signal indicator detection by using unsupervised training techniques of an objective function (e.g., cost, reward, heuristic, constraint, etc.). In some embodiments, the model trainer can perform multiple generalization techniques to improve the generalization ability of the trained model. The generalization techniques include weight decay, dropout, or other techniques.

[0142] Figure 5B Depicts a training architecture 510 for the signal indicator model 410. When executing a training instance, the training data can include a first image crop 512A taken at time t. The first image crop 512A can be considered as the current image frame associated with the current time step t. The training data can also include historical image frames. For example, the training data can include a second image crop 514A taken at time t-1 and a third image crop 516A taken at time t-2.

[0143] Capable of processing image crops 512A, 514A, 516A to generate intermediate outputs, such as embeddings 512B, 514B, 516B. To this end, the training architecture 510 can include a model trunk. The trunk can include a common network for various tasks, such as, for example, processing various image frames. For example, the first image crop 512A can be processed using the trunk to generate a first embedding 512B associated with time step t. The second image crop 514A can be processed using the trunk to generate a second embedding 514B associated with time step t-1. The third image crop 516A can be processed using the trunk to generate a third embedding 516B associated with time step t-2. Each embedding can capture, for example, the state of the signal light (e.g., turn signal) of a vehicle that appears in the image crops 512A, 514A, 516A at their respective time steps. The state can indicate whether the light is in an active state (e.g., on) or in an inactive state (e.g., off).

[0144] The embeddings 512A, 512B, 512C can be processed by a model head to generate a training output 518. The training output 518 can indicate the state of the signal indicator of the first image crop 512A (e.g., the current image frame), which is informed by the states of the signal indicators in the second and third image crops 514A, 516A (e.g., historical image frames). For example, the signal indicator model 410 can determine whether the turn signal shown in the image crops 512A, 514A, 516A is active at the current time step t based on the previous time steps t-1 and t-2. The embeddings 512A, 512B, 512C can indicate that the light of the turn signal can be lit / turned on at time step t-2, off / not lit at time step t-1, and lit / turned on at time step t. This can indicate a blinking pattern and thus indicate that the turn signal is in an active state at the current time step t.

[0145] The training output 518 can indicate the prediction of the signal indicator model 410 during a training instance. The training output 518 can be compared with the training data to determine the progress of training and the accuracy of the model. This can include comparing the prediction of the signal indicator model 410 in the training output 518 (e.g., indicating that the turn signal is active at the current time step) with the ground truth. Based on this comparison, one or more loss metrics or objectives can be generated, and if necessary, at least one parameter of at least a portion of the signal indicator model 410 can be modified based on the loss metric or at least one objective.

[0146] The signal indicator model 410 can be trained to determine other characteristics of the signal indicator. For example, the signal indicator model 410 can be trained to determine the type of the signal indicator (e.g., turn signal, brake signal, hazard signal, other light), the position of the signal indicator relative to the actor (e.g., left turn signal, right turn signal, etc.), the color, or other characteristics. The training data can include labels indicating these characteristics, and the training output 518 of the signal indicator model 410 can indicate the predictions of the model for the type of the signal indicator, the position of the signal indicator, the color, etc. Similar to what is described above, if needed, these predictions can be compared with the ground truth to evaluate the progress of the model and modify the model parameters.

[0147] The inference architecture 520 of the signal indicator model 410 can diverge from the training architecture 510. More specifically, when evaluating the current image frame, the signal indicator model 410 may have processed historical image frames. In an example, the signal indicator model 410 may have generated a first previous embedding 524 based on the image frame at time step t - 1 and a second previous embedding 526 based on the image frame at time step t - 2. These previous embeddings can be stored in a memory (e.g., local cache) such that they can be accessed and used to inform the analysis of the image crop 522A based on the current image frame (e.g., RBG image) at the current time step t.

[0148] In the case where the signal indicator model 410 has not processed historical image frames relative to the current image frame, the signal indicator model 410 can perform its analysis based on a single current image frame without being informed by historical time frames.

[0149] The image crop 522A can be generated using the preprocessing module 400 as described herein and provided as input data to the signal indicator model 410.

[0150] The signal indicator model 410 can determine the signal indicator of the actor within the image crop 522A. For the current image frame, the detection system 407 can use the signal indicator model 410 to determine that a part of the actor (depicting the current image frame) is a signal indicator. In some examples, the signal indicator model 405 can determine (and output) the type of the signal indicator from one of the following types: turn signal indicator, hazard signal indicator, brake signal indicator, or other types of signal indicators.

[0151] As an example, the signal indicator model 410 can analyze the actor in the image crop 522A to determine the probability that a portion of the depicted actor is a signal indicator. This can include analyzing the shape or position of a subset of pixels representing the actor to determine if the subset is a turn signal light, brake light, etc. In some examples, the signal indicator model 410 can analyze the shape or position of a subset of pixels of the actor to help identify the type of signal indicator (e.g., left turn signal).

[0152] The probability can reflect the confidence of the model that the actor contains a signal indicator. In some examples, a probability above a threshold probability (e.g., 95%, etc.) can result in a positive detection. A probability below the threshold probability can result in a determination that no signal indicator was captured in the image crop 522A.

[0153] The detector system 407 can use the signal indicator model 410 and generate output data 530 based on one or more historical image frames. As an example, the vehicle indicator model 410 can predict whether the right turn signal of the actor is activated or deactivated based on the current image frame as informed by the previous embeddings 524, 526. The signal indicator model 410 can pass the image crop 522A through the backbone to generate a current embedding 522B associated with the current time step t. The current embedding 522B can be saved for analyzing the next time frame associated with time step t+1. Then, the signal indicator model 410 can pass the current embedding 522B through the model head and concatenate the current embedding 522B (e.g., at time step t) with the previous embeddings 524, 526 (e.g., at time steps t-1 and t-2) to generate output data 530.

[0154] The output data 530 can indicate one or more characteristics of the signal indicator as determined by the signal indicator model 410. The characteristics can indicate the type of signal indicator or the location of the signal indicator. The characteristics can indicate the activation state or deactivation state of the signal indicator in the current image frame. For example, as described herein, the signal indicator model 410 can determine that there is a right turn signal of the actor depicted in the current image crop 522A. The current image crop 522A (and current embedding 522B) at time step t and the first previous embedding 524 at t-1 can indicate the right turn signal as lit / on. The second previous embedding 526 at t-2 can indicate the right turn signal as unlit / off. Based on the concatenation of the current embedding 522A and the previous embeddings 524, 526, the signal indicator model 410 can determine that at time step t, for the current image frame, the right turn signal is in an activated state. Thus, the output data 530 can indicate that the signal indicator depicted in the current image frame is a valid right turn signal at time step t. In some embodiments, the output data 530 indicates, for example, that the color of the turn signal is red.

[0155] Return Figure 4 , the detector system 407 is capable of determining whether the signal indicator of the actor is activated or deactivated based at least on the current image frame. For example, the output data 530 can be provided to the post-processing module 404. The output data 530 can include the current image frame, the time step associated with the current time step t, and the characteristics associated with the signal indicator determined by the signal indicator model 410. In an example, the output data 530 can include the probability that the signal indicator is in an activated or deactivated state in the corresponding image frame (as determined by the signal indicator model 410).

[0156] In some embodiments, the output data 530 is stored in the buffer 403. The output data 530 can be stored as the attribute data 409. The attribute data 409 can include the output data 530 and the time associated with the corresponding image frame.

[0157] The detection system 407 can determine that there is a threshold amount of attribute data for multiple image frames at multiple times. In an example, the detection system 407 can determine that the buffer 403 includes a threshold amount of the attribute data 409 for multiple image frames at multiple times. The threshold amount of the attribute data can include a threshold amount of image frames (e.g., 6 frames) that have been processed by the signal indicator model 410. As previously described herein, the smoother model 409 can be used to analyze the attribute data 409 to generate the output of the detector system 407.

[0158] In some embodiments, the post-processing module 404 is not used for the output data 530. This may occur when the confidence associated with the output data 513 is sufficient because it has been tentatively inferred by the signal indicator model 410 based on the analysis of the current image frame with respect to the historical image frames.

[0159] The actions of the autonomous vehicle can be performed based on the signal indicator of the activated or deactivated actor. For example, the data indicating the signal indicator, its type, status, location, etc. can be provided to the planning system 250. The motion planning system 250 can predict the actor's motion based on the signal indicator in the manner described herein for the actors perceived by the autonomous vehicle. The status of the signal indicator can help determine the actor's intention. This can be particularly advantageous for actors whose dynamic motion parameters (e.g., direction / speed changes) may not show the intention of the actor (e.g., turning left).

[0160] The motion planning system 250 is capable of generating a motion plan for an autonomous vehicle based on the status of signal indicators. This can include, for example, generating a trajectory for the autonomous vehicle to decelerate to allow a vehicle with an activated left turn signal to turn left in front of the autonomous vehicle, making a fine adjustment within the lane to provide more distance between the autonomous vehicle and an actor on the shoulder with a flashing hazard signal light, decelerating to allow an actor with an activated right turn signal to merge into the same lane as the autonomous vehicle, decelerating in response to an actor with an activated brake light, etc. In some examples, the trajectory can include the autonomous vehicle changing lanes in response to the signal indicators of actors.

[0161] In some examples, the motion planning system 250 can take the status of signal indicators into account in its trajectory generation and determine that the autonomous vehicle does not need to change acceleration, speed, direction, etc., because the autonomous vehicle is already properly positioned relative to the actor. This can include scenarios when the actor is located behind the autonomous vehicle.

[0162] The actions of the autonomous vehicle can include controlling the movement of the autonomous vehicle based on the signal indicators of actors. The motion planning system 250 can provide data indicating the trajectory generated based on the signal indicators of actors. The control system 260 can control the maneuvering of the autonomous vehicle based on the trajectory, as described herein.

[0163] The signal indicator model 410 can run concurrently with the emergency vehicle model 405. For example, the signal indicator model 410 and the emergency vehicle model 405 can evaluate the same image frame (e.g., its image crop). The outputs from the models can be combined, stored in relation to each other, or otherwise processed in a way that provides more robust information about the actor. For example, the emergency vehicle model 405 can determine that the actor is an emergency vehicle, while the signal indicator model 410 can determine that the actor has an activated left turn signal. Thus, the combination of the outputs can inform the autonomous vehicle of the presence of an activated emergency vehicle intending to turn left.

[0164] In some embodiments, the signal indicator model 410 and the emergency vehicle model 405 can run in series, where the output from one model is used as the input to the other model.

[0165] In some embodiments, the functions of the signal indicator model 410 and the emergency vehicle model 405 can be performed by one model. This can allow the corresponding image frame processed by the emergency vehicle model 405 to be informed by historical image frames associated with previous time steps. For example, for a corresponding time frame associated with time step t, the machine learning model can be configured to process one or more historical image frames to generate output data 408. The one or more historical image frames can be associated with one or more time steps t - 1, t - 2 prior to the time step associated with the corresponding time frame.

[0166] Figure 7 Depicts a flowchart of a method 700 for detecting an activated emergency vehicle and controlling an autonomous vehicle in accordance with aspects of the present disclosure. One or more portions of method 700 can be implemented by a computing system that includes one or more computing devices, such as, for example, the computing systems described with reference to other figures (e.g., autonomous platform 110, vehicle computing system 180, remote system 160, Figure 4 , FIG. 5, Figure 10 , etc.). Each corresponding portion of method 700 can be performed by any one (or any combination) of one or more computing devices. Additionally, one or more portions of method 700 can be implemented on the hardware components of the devices described herein (e.g., as in Figure 1 to FIG. 5, Figure 10 , etc.), for example, to detect an activated emergency vehicle and control the autonomous vehicle relative thereto.

[0167] Figure 7 Elements are depicted in a particular order for purposes of illustration and discussion. Using the disclosure provided herein, one of ordinary skill in the art will understand that, without departing from the scope of the present disclosure, the elements of any method discussed herein can be adapted, rearranged, extended, omitted, combined, or modified in various ways. Elements / terms are described with reference to elements / terms described with respect to other systems and figures for purposes of exemplary illustration Figure 7 , and Figure 7 is not meant to be limiting. One or more portions of method 700 can additionally or alternatively be performed by other systems.

[0168] At 702, method 700 includes obtaining sensor data that includes a plurality of image frames indicative of actors in the environment of the autonomous vehicle. For example, a (e.g., on-vehicle) computing system of the autonomous vehicle can obtain image data from one or more cameras on the autonomous vehicle. The image data can include a plurality of image frames at a plurality of times. This can include a first image frame captured at a first time.

[0169] At 704, method 700 includes: for a corresponding image frame, using a machine learning model to determine that an actor is an emergency vehicle. For example, a computing system can access an emergency vehicle model of machine learning from an accessible memory (e.g., on-board an autonomous vehicle). As described herein, the emergency vehicle model can be trained based on labeled training data. The labeled training data can be based on point cloud data and image data. In addition, the labeled training data can indicate multiple training actors (e.g., vehicles), and at least one corresponding training actor is labeled with an emergency vehicle label. In some examples, the emergency vehicle label can indicate the type of the emergency vehicle of the corresponding training actor or that the training actor is not an emergency vehicle. In some examples, the corresponding training actor includes an active label that indicates, based on the lights of the training actor, that the training actor is in an activated state or a non-activated state.

[0170] As described herein, the emergency vehicle model can process a first image frame to predict whether it includes an activated emergency vehicle.

[0171] In some examples, the first image frame can be preprocessed according to method 800 of FIG. 8. At 802, method 800 includes obtaining tracking data of a first actor in the environment of an autonomous vehicle (e.g., depicted in the first image frame). The tracking data can indicate the boundary shape of the first actor. At 804, method 800 includes determining that the centroid of the boundary shape is within the projection field of view of the autonomous vehicle. At 806, method 800 includes, in response to determining that the centroid of the boundary shape is within the projection field of view, generating input data for the emergency vehicle model based on sensor data. The input data can include a cropped version of the first image frame depicting the first actor.

[0172] Returning to Figure 7 , at 706, method 700 includes: for a corresponding image frame, using a machine learning model to generate output data indicating whether the emergency vehicle in the corresponding image frame is in an activated state or a non-activated state. As described herein, the emergency vehicle model can process a first image frame to determine that the first actor depicted in the first image frame is a first emergency vehicle, such as a police car.

[0173] The emergency vehicle model can determine whether the first emergency vehicle is in an activated state or a non-activated state. To this end, method 850 of FIG. 8 can be used.

[0174] At 852, method 850 includes using a machine learning model to determine the state of the lights of a first emergency vehicle in a corresponding image frame, where the state of the lights includes an on state or an off state. As described herein, in some examples, this includes the emergency vehicle model analyzing the pixels of the first image frame to determine whether one or more characteristics (e.g., brightness, color, etc.) of a first light of the first emergency vehicle indicate an illuminated lighting element. As an example, the emergency vehicle model can determine that the roof-mounted lights of a police car are illuminated blue in the first image frame.

[0175] At 854, method 850 includes using a machine learning model to determine, based on the state of the lights, whether the emergency vehicle is in an active state or a non-active state for the corresponding image frame. In the case where the emergency vehicle model detects an on state (e.g., illuminated blue lights), the first emergency vehicle can be considered to be in an active state. In the case where the emergency vehicle model detects an off state, the first emergency vehicle can be considered to be in a non-active state.

[0176] Return Figure 7 , method 700 can include storing attribute data for the corresponding image frame. The attribute data can include the output data of the machine learning model and the time associated with the corresponding image frame. In an example, at 708, method 700 includes storing the attribute data for the corresponding image frame in a buffer, the attribute data including the output data of the machine learning model and the time associated with the corresponding image frame. This can include the first image frame for which the first emergency vehicle is the probability of an active emergency vehicle associated with the time (e.g., when the first image frame was captured) and the trace of the first emergency vehicle.

[0177] Method 700 can include determining that the emergency vehicle is an active emergency vehicle based on the output data associated with the corresponding image frame. In an example, at 710, method 700 includes determining that the buffer includes a threshold amount of attribute data for multiple image frames at multiple times. For example, the buffer can store the multiple second image frames as attribute data when processed and output by the emergency vehicle model. Each second image frame is stored with an indication of whether the first emergency vehicle is active or non-active, the associated time, and the trace. Once the buffer includes a threshold amount of processed image frames for the first emergency vehicle that cover a threshold number of time frames, a second model can process at least a subset of the image frames.

[0178] At 712, method 700 includes determining that an emergency vehicle is an activated emergency vehicle using a second model based on attribute data of at least a subset of multiple image frames. As described herein, the second model can include a downstream smoother model configured to confirm the presence of an activated emergency vehicle in the environment by analyzing the outputs of a machine learning emergency vehicle model over a given subset of times. As an example, the smoother model can determine that more than 50% of a first image frame and a second image frame indicate that the depicted police vehicle is activated. Thus, the smoother model can output data indicating that the first actor is an activated emergency vehicle to one or more systems on-board the autonomous vehicle.

[0179] At 714, method 700 includes performing an action of the autonomous vehicle based on the activated emergency vehicle being within the environment of the autonomous vehicle. This can include, for example, at least one of the following: (1) predicting the movement of the activated emergency vehicle; (2) generating a motion plan for the autonomous vehicle; or (3) controlling the movement of the autonomous vehicle, as described herein.

[0180] Figure 9 A flowchart of a method 900 for training one or more models in accordance with aspects of the present disclosure is depicted. For example, the model can include an emergency vehicle model or a smoother model, as described herein.

[0181] One or more portions of method 900 can be implemented by a computing system that includes one or more computing devices, such as, for example, the computing system described with reference to other figures (e.g., Figure 11 the system, etc.). Each respective portion of method 900 can be executed by any one (or any combination) of one or more computing devices. Additionally, one or more portions of method 900 can be implemented on hardware components of a device (e.g., as Figure 11 etc.) described herein, for example, to train example models of the present disclosure.

[0182] Figure 9 Elements are depicted in a particular order for purposes of illustration and discussion. Using the disclosure provided herein, one of ordinary skill in the art will understand that, without departing from the scope of the present disclosure, the elements of any method discussed herein can be adapted, rearranged, extended, omitted, combined, or modified in various ways. Elements / terms are described for purposes of exemplary illustration with reference to elements / terms described with respect to other systems and figures Figure 9 and Figure 9 is not meant to be limiting. One or more portions of method 900 can additionally or alternatively be executed by other systems.

[0183] At 902, method 900 can include obtaining training data for an emergency vehicle model for machine learning. The training data can include sensor data, perception output data, log data, simulation data, and the like. The training data can include vehicle state data, trajectories, image frames captured during real-world or simulated driving instances, the associated times at which an autonomous vehicle perceives actors / objects in the environment, and other information.

[0184] For example, sensor data that can be used as a basis for training data can be collected using one or more autonomous platforms (e.g., autonomous platform 110) or their sensors when the autonomous platform is within its environment. As an example, when a vehicle is operating along one or more driving roads, one or more autonomous vehicles or sensors can be used to collect training data. In some example methods, other sensors (such as mobile device-based sensors, ground-based sensors, aerial-based sensors, satellite-based sensors, or substantially any sensor interface configured to obtain and / or record measurement data) can be used to collect training data. In some example methods, training data can be collected from a public source that is not specific to emergency vehicles. For example, training data can be collected from an emergency vehicle dedicated channel or other publicly available online sources.

[0185] Perception output data can include data output from an autonomous vehicle's perception system. In some examples, the perception output data can include specific metadata generated by the perception system (or its functionality). For example, the perception output data can include metadata associated with the characteristics of objects / actors in an image frame captured of the environment. In some example methods, the perception output data can include vehicle trajectories. The trajectories can include the boundary shapes and state data of the actors. The state data can include the position, speed, acceleration, etc. of the actor when the actor is perceived.

[0186] Log data can include data obtained from one or more autonomous vehicles and downloaded to an offline system. The log data can be a recorded version of sensor data, perception output data, etc. The log data can be stored in an accessible memory and can be extracted to generate a specific combination of attributes for the training data.

[0187] In some examples, the training data can include simulation data. Simulation data can be collected during one or more simulation instances / runs. The simulation instances can simulate scenarios in which a simulated autonomous vehicle traverses a simulated environment and captures simulated perception output data of the simulated environment. Simulated emergency vehicles and non-emergency vehicles can be placed within the scenario such that the synthesized simulated log data reflects the simulated perception output data. In this way, the simulated log data can include emergency vehicles and non-emergency vehicles, which can then be used for training data generation.

[0188] The training data can cover vehicles and emergency vehicles from different aspects. For example, the training data can cover many non-emergency vehicle types and appearances. The training data can be biased towards close-range emergency vehicles that are easy to classify. The training data can cover rich scenarios involving emergency vehicles, including day and night, highway and urban environments, and various other traffic conditions.

[0189] In some examples, the training data can include augmented training data. Data augmentation can be applied to the training data by applying transformations to the original image data, such as cropping, flipping, rotating, resizing, color jitting, etc. Data augmentation can include statistically adjusting the vehicle trace boundary shape by sampling the state distribution to generate new trace boundary shapes. For example, the training data can contain the state coordinates x, y, z of the trace at the center of the boundary shape, as well as the length, width, and height of the boundary shape. The state of the trace can be fitted to a sampled multivariate normal distribution, and such variations can affect the image cropping position to enhance the dataset. The augmented training data can ensure that the augmented dataset is natural and likely to occur in the real world. In some example methods, the augmented training data can use sampling rate multipliers on positive and negative targets.

[0190] In some example methods, the training data can be processed by a data engine. The data engine can be used to mine data (e.g., log data) to find events of positive emergency vehicle detection. In some examples, positive emergency vehicle events can be added to the training dataset for further training of the emergency vehicle model. In some example methods, false positive emergency vehicle events can be added to the training dataset for further training. For example, the false positive event rate can be measured to achieve improvements and changes in recall compared to a baseline.

[0191] The training data can include labeled training data. For example, the training data can include label data indicating whether the actor in the corresponding image frame is an emergency vehicle or a non-emergency vehicle, the vehicle type, the activation status (e.g., indicating the activated state or non-activated state), the bulb / lamp status, etc. In some examples, the training data can include labels indicating various other details about the lights of the vehicle in the corresponding image frame (e.g., bulb color, bulb light intensity, etc.).

[0192] The labeling can include four-dimensional (4D) labeling (e.g., 3D bounding boxes around LIDAR points on an object over time) and two-dimensional (2D) labeling (e.g., 2D bounding boxes on an object within a forward camera image). The 4D and 2D labels can be associated and used to generate an image sequence (e.g., a tiled video) for each individual actor. Actor metadata, including vehicle type, bulb status, activation status, etc., can also be labeled.

[0193] If the actor is marked as an emergency vehicle, a second marking phase for activation and bulb status can be initiated. For example, at each time frame of the sequence (e.g., the frame of a tiled video), if the beacon / bulb light is flashing, the image frame can be marked as activated, and otherwise marked as non-activated. If the light is on, the activated image frame can be marked as "bulb on" or "on state", and otherwise marked as "bulb off" or "off state". The activation marking can depend on the time context, while the bulb status marking can depend on each individual frame.

[0194] Data extraction for training purposes can be similar to the cropping and filtering implemented online. For example, the data engine can mine log data to find two-dimensional (2D) labels linked to a given four-dimensional (4D) label, and extract data on emergency vehicles, bulb status (e.g., bulb on as positive, bulb off as negative), non-activated emergency vehicles, non-emergency vehicles, or other information.

[0195] Table 1 provides an overview of the example training dataset distribution.

[0196] Table 1

[0197]

[0198] The training data can include multiple training sequences divided among multiple datasets (e.g., training dataset, validation dataset, or test dataset). Each training sequence can include multiple pre-recorded perception data points, point clouds, images, etc.

[0199] At 904, method 900 includes selecting training instances based on the training data. For example, the model trainer can select a labeled training dataset to train a machine learning emergency vehicle model. The labeled training data can include actors or scenarios that are typically viewed by the emergency vehicle model or edge cases for which the model should be trained.

[0200] Training instances can also be selected based on specific goals. The goals can include true positive and false positive targets. This can help improve the model regardless of whether they are true positive or false positive events. For example, the goals can include positive targets that can indicate an activated emergency vehicle. In some example methods, the goals can include negative targets that can indicate a non-activated emergency vehicle. In some examples, the negative targets can include non-emergency vehicles.

[0201] The goals can include positive and negative targets in various contexts including day and night, highways and cities, and various other traffic conditions. In some examples, the targets can be generated from the real world or simulated driving. In some example methods, the targets can be generated from public sources that are not specific to emergency vehicles.

[0202] At 906, method 900 can include inputting training instances into a machine - learning emergency vehicle model. For example, the machine - learning model can receive training data and extract labels to determine positive and negative emergency vehicle detections. The machine - learning model can process the training data and generate machine - learning output data. In some examples, the machine - learning output data can include a baseline. In some examples, the machine - learning output data can include oversampling within the training set.

[0203] For a given training instance, a smoother model can also be evaluated. For example, the output data generated from training the machine - learning emergency vehicle model can include image frames (at multiple times) with indicators of emergency vehicles and activation states. This can be input into the smoother model, which can output a final training determination of whether the emergency vehicle is activated based on analyzing multiple training image frames across multiple times.

[0204] At 908, method 900 can include generating one or more loss metrics or one or more objectives for the machine - learning emergency model based on the output of at least a portion of the model and the labels associated with the training instance. For example, the output can be compared with the training data to determine the progress of training and the accuracy of the model.

[0205] At 910, method 900 can include modifying at least one parameter of at least a portion of the machine - learning emergency vehicle model based on at least one of the loss metric or the objective. For example, a computing system can modify at least one hyperparameter of the machine - learning emergency vehicle model. The hyperparameters of the emergency vehicle model can be adjusted to improve the maximum - F1 score or other metrics. The data engine can continuously improve the model by adding more and more data over time during training and retraining.

[0206] In some examples, system - level metrics can also be used to refine and evaluate the downstream smoother model. For example, since the smoother model can be designed to determine whether a vehicle is an emergency vehicle in an activated state (e.g., flashing lights), activation labels can be used to compare with the output to calculate metrics. When the emergency vehicle model is trained, the absolute - majority threshold of the smoother model can be updated by, for example, refining the minimum score of positive outputs in a buffer to output a final positive activated emergency vehicle detection. This can be done by evaluating recall, F1 - score, and the level of tracking false negatives. Table 3 provides the smoother results.

[0207] Table 3

[0208]

[0209] In some example methods, a machine-learned emergency vehicle model can be trained in an end-to-end manner. For example, in some implementations, the machine-learned emergency vehicle model can be fully differentiable.

[0210] After being updated, an emergency vehicle model or operating system that includes the model can be provided for verification. In some implementations, a smoother model can evaluate or verify the operating system to identify areas of performance deficiency compared to an example corpus. The smoother model can trigger retraining, debugging, etc. of the operating system based on, for example, failure to meet a verification threshold in one or more areas.

[0211] Figure 10 A flowchart of a method 1000 for detecting a signal indicator state and controlling an autonomous vehicle in accordance with aspects of the present disclosure is depicted. One or more portions of method 1000 can be implemented by a computing system that includes one or more computing devices (such as the computing systems described with reference to other figures). Each respective portion of method 1000 can be executed by any (or any combination) of one or more computing devices. Additionally, one or more portions of method 1000 can be implemented on hardware components of the devices described herein, for example, to detect and determine the state of a signal indicator.

[0212] Figure 10 Elements are depicted as being executed in a particular order for purposes of illustration and discussion. Using the disclosure provided herein, one of ordinary skill in the art will understand that, without departing from the scope of the present disclosure, the elements of any method discussed herein can be adapted, rearranged, extended, omitted, combined, or modified in various ways. Elements / terms are described for purposes of exemplary illustration with reference to elements / terms described with respect to other systems and figures Figure 10 and Figure 10 is not meant to be limiting. One or more portions of method 1000 can additionally or alternatively be executed by other systems.

[0213] At 1002, method 1000 can include obtaining data that includes a plurality of image frames indicative of actors in the environment of an autonomous vehicle. For example, a computing system (such as an on-vehicle computing system of the autonomous vehicle) can obtain data that includes a plurality of image frames indicative of a plurality of image frames including a current image frame and one or more historical image frames. As described herein, the current image frame can be associated with a current time step t, while the historical image frames can be associated with past time steps t-1, t-2, etc. The image frames can be RGB image frames captured via an on-vehicle camera of the autonomous vehicle.

[0214] At 1004, method 1000 can include: for a current image frame, using a machine learning model to identify a signal indicator of an actor. As described herein, a computing system can utilize a machine learning signal indicator model to determine that a signal indicator is depicted in an image crop of the current image frame. The signal indicator can be a turn signal, a hazard signal, a brake signal, etc. The signal indicator model can determine the type, location, color, etc. of the signal indicator.

[0215] At 1006, method 1000 can include: for a current image frame, using a machine learning model and based on one or more historical image frames to generate output data indicating one or more characteristics of the signal indicator. For example, the signal indicator model can utilize embeddings of historical image frames (at previous time steps) to process the current image frame. As described herein, this can allow the signal indicator model to cascade the data to inform its analysis of the status of the signal indicator in the current image frame. For example, the signal indicator model can process frames across multiple time steps to identify that the right turn signal light of a vehicle is lit in a flashing pattern across time steps including the current time step. Thus, the status of the signal indicator in the current image frame (at the current time step) can be the activation status of the right turn signal light.

[0216] At 1008, method 1000 can include determining that the signal indicator of the actor is activated or deactivated based at least on the current image frame. This can include storing the output data of the signal indicator model as attribute data. The attribute data can indicate the model's determination of the signal indicator status for each image frame across multiple time steps. As described herein, the status can be evaluated to determine that the actor's signal indicator (e.g., its right turn signal) is activated / on.

[0217] At 1010, method 1000 can include performing an action of an autonomous vehicle based on the signal indicator of the actor being activated. As described herein, this can include determining the intent of the actor (e.g., using an intent model), predicting the movement of the actor, generating a motion plan / trajectory for the autonomous vehicle, controlling the autonomous vehicle, etc. These actions can be performed to avoid interfering with the actor.

[0218] Figure 11 is a block diagram of an example computing ecosystem 10 according to an example implementation of the present disclosure. The example computing ecosystem 10 can include a first computing system 20 and a second computing system 40 communicatively coupled via one or more networks 60. In some implementations, the first computing system 20 or the second computing 40 can implement one or more of the systems, operations, or functions described herein for validating one or more systems or operating systems (e.g., remote system 160, in-vehicle computing system 180, autonomous system 200, etc.).

[0219] In some embodiments, the first computing system 20 can be included in an autonomous platform and used to perform the functions of the autonomous platform as described herein. For example, the first computing system 20 can be located on-board an autonomous vehicle and implement an autonomous system for autonomously operating the autonomous vehicle. In some embodiments, the first computing system 20 can represent the entire on-board computing system or a portion thereof (e.g., a positioning system 230, a sensing system 240, a planning system 250, a control system 260, a detection system 407, or a combination thereof, etc.). In other embodiments, the first computing system 20 may not be located on-board the autonomous platform. The first computing system 20 can include one or more different physical computing devices 21.

[0220] The first computing system 20 (e.g., its computing device 21) can include one or more processors 22 and a memory 23. The one or more processors 22 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be a single processor or multiple processors operably connected. The memory 23 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, one or more memory devices, flash memory devices, etc., and combinations thereof.

[0221] The memory 23 can store information accessible by the one or more processors 22. For example, the memory 23 (e.g., one or more non-transitory computer-readable storage media, memory devices, etc.) can store data 24 that can be obtained (e.g., received, accessed, written, manipulated, created, generated, stored, pulled, downloaded, etc.). The data 24 can include, for example, sensor data, map data, data associated with autonomous functions (e.g., data associated with sensing, planning, or control functions), simulation data, or any data or information described herein. In some embodiments, the first computing system 20 can obtain data from one or more memory devices remote from the first computing system 20.

[0222] The memory 23 can store computer-readable instructions 25 executable by the one or more processors 22. The instructions 25 can be software written in any suitable programming language or can be implemented in hardware. Additionally or alternatively, the instructions 25 can be executed in logically or virtually separate threads on the processor 22.

[0223] For example, the memory 23 can store instructions 25 that can be executed by one or more processors (e.g., by one or more processors 22, by one or more other processors, etc.) to perform any operations, functions, or methods / processes (or portions thereof) described herein (e.g., using the computing device 21, the first computing system 20, or other systems having a processor that executes instructions). For example, the operations can include implementing system verification (e.g., as described herein).

[0224] In some embodiments, the first computing system 20 can store or include one or more models 26. In some embodiments, the model 26 can be or can otherwise include one or more machine learning models (e.g., a machine learning emergency vehicle detection model, a machine learning operating system, etc.). As an example, the model 26 can be or can otherwise include various machine learning models, such as, for example, regression networks, generative adversarial networks, neural networks (e.g., deep neural networks), support vector machines, decision trees, ensemble models, k-nearest neighbor models, Bayesian networks, or other types of models that include linear or non-linear models. Example neural networks include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. For example, the first computing system 20 can include one or more models for subsystems that implement the autonomous system 200, including any one of the following: a localization system 230, a perception system 240, a planning system 250, or a control system 260.

[0225] In some embodiments, the first computing system 20 can obtain one or more models 26 using the communication interface 27 to communicate with the second computing system 40 via the network 60. For example, the first computing system 20 can store the model 26 (e.g., one or more machine learning models) in the memory 23. Then, the first computing system 20 can use or otherwise (e.g., by the processor 22) implement the model 26. As an example, the first computing system 20 can implement the model 26 to locate an autonomous platform in an environment, perceive the environment of the autonomous platform or objects therein, plan one or more future states for the autonomous platform to move through the environment, control the autonomous platform for interacting with the environment, detect emergency vehicles, etc.

[0226] The second computing system 40 can include one or more computing devices 41. The second computing system 40 can include one or more processors 42 and a memory 43. The one or more processors 42 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be one processor or multiple processors operably connected. The memory 43 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, one or more memory devices, flash memory devices, etc., and combinations thereof.

[0227] The memory 43 can store information that can be accessed by the one or more processors 42. For example, the memory 43 (e.g., one or more non-transitory computer-readable storage media, memory devices, etc.) can store data 44 that can be obtained. The data 44 can include, for example, sensor data, model parameters, map data, simulation data, simulated environment scenes, simulated sensor data, data associated with a vehicle trip / service, or any data or information described herein. In some embodiments, the second computing system 40 can obtain data from one or more memory devices remote from the second computing system 40.

[0228] The memory 43 can also store computer-readable instructions 45 that can be executed by the one or more processors 42. The instructions 45 can be software written in any suitable programming language or can be implemented in hardware. Additionally or alternatively, the instructions 45 can be executed in logically or virtually separate threads on the processor 42.

[0229] For example, the memory 43 can store instructions 45 that can be executed (e.g., by the one or more processors 42, by the one or more processors 22, by one or more other processors, etc.) to perform any of the operations, functions, or methods / procedures described herein (e.g., using the computing device 41, the second computing system 40, or other systems having a processor for executing the instructions, such as the computing device 21 or the first computing system 20). This can include, for example, the functions of the autonomous system 200 (e.g., localization, perception, planning, control, etc.) or other functions associated with an autonomous platform (e.g., remote assistance, mapping, fleet management, trip / service allocation and matching, etc.). This can also include, for example, validating an operating system for machine learning.

[0230] In some embodiments, the second computing system 40 can include one or more server computing devices. In the case where the second computing system 40 includes multiple server computing devices, such server computing devices can operate according to various computing architectures, including, for example, a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0231] In addition to or in place of the model 26 of the first computing system 20, the second computing system 40 can include one or more models 46. As an example, the model 46 can be or can otherwise include various machine learning models (e.g., machine-learned operating systems, etc.), such as, for example, recurrent networks, generative adversarial networks, neural networks (e.g., deep neural networks), support vector machines, decision trees, ensemble models, k-nearest neighbor models, Bayesian networks, or other types of models that include linear or non-linear models. Example neural networks include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. For example, the second computing system 40 can include one or more models of one or more autonomous systems 200.

[0232] In some embodiments, the second computing system 40 or the first computing system 20 can train one or more machine learning models of the model 26 or the model 46 by using one or more model trainers 47 and training data 48. The model trainer 47 can use one or more training or learning algorithms to train any of the model 26 or the model 46. An example training technique is backpropagation of error. In some embodiments, the model trainer 47 can use labeled training data to perform supervised training techniques. In other embodiments, the model trainer 47 can use unlabeled training data to perform unsupervised training techniques. In some embodiments, the training data 48 can include simulated training data (e.g., training data obtained from simulated scenarios, inputs, configurations, environments, etc.). In some embodiments, the second computing system 40 can implement a simulation for obtaining the training data 48 or for implementing the model trainer 47 for training or testing the model 26 or the model 46. As an example, the model trainer 47 can use an objective function (e.g., cost, reward, heuristic, constraint, etc.) to train one or more components of the machine learning model of the autonomous system 200 by unsupervised training techniques. In some embodiments, the model trainer 47 can perform multiple generalization techniques to improve the generalization ability of the trained model. Generalization techniques include weight decay, dropout, or other techniques.

[0233] For example, in some embodiments, the second computing system 40 can generate training data 48 in accordance with the example aspects of the present disclosure. For example, the second computing system 40 can generate the training data 48. For example, the second computing system 40 can implement a method in accordance with the example aspects of the present disclosure. The second computing system 40 can use the training data 48 to train the model 26. For example, in some embodiments, the first computing system 20 can include a computing system that is vehicle-mounted or otherwise associated with a real or simulated autonomous vehicle. In some embodiments, the model 26 can include a perception or machine vision model configured for in-vehicle deployment or deployment in service of a real or simulated autonomous vehicle. In this way, for example, the second computing system 40 can provide a training pipeline for training the model 26.

[0234] The first computing system 20 and the second computing system 40 can each separately include a communication interface 27 and 49. The communication interface 27, the communication interface 49 can be used to communicate with each other or with one or more other systems or devices, including systems or devices located remotely from the first computing system 20 or the second computing system 40. The communication interface 27, the communication interface 49 can include any circuitry, components, software, etc. for communicating with one or more networks (e.g., network 60). In some embodiments, the communication interface 27, the communication interface 49 can include, for example, one or more of a communication controller, a receiver, a transceiver, a transmitter, a port, a conductor, software, or hardware for transmitting data.

[0235] The network 60 can be any type of network or combination of networks that permits communication between devices. In some embodiments, the network can include one or more of a local area network, a wide area network, the Internet, a secure network, a cellular network, a mesh network, a peer-to-peer communication link, or some combination thereof, and can include any number of wired or wireless links. Communication over the network 60 can be implemented, for example, via a network interface using any type of protocol, protection scheme, encoding, format, encapsulation, etc.

[0236] Figure 11 An example computing ecosystem 10 that can be used to implement the present disclosure is illustrated. Other systems can also be used. For example, in some embodiments, the first computing system 20 can include a model trainer 47 and training data 48. In such embodiments, the models 26, 46 can be locally trained and used at the first computing system 20. As another example, in some embodiments, the computing system 20 may not be connected to other computing systems. Additionally, components illustrated or discussed as being included in one of the computing systems 20 or 40 can alternatively be included in the other of the computing systems 20 or 40.

[0237] The computing tasks discussed herein as being performed at a computing device remote from an autonomous platform (e.g., an autonomous vehicle) can alternatively be performed at the autonomous platform (e.g., via a vehicle computing system of the autonomous vehicle), and vice versa. Such a configuration can be achieved without departing from the scope of the present disclosure. The use of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functions among two or more components. Computer-implemented operations can be performed on a single component or across multiple components. Computer-implemented tasks or operations can be performed sequentially or in parallel. Data and instructions can be stored in a single memory device or across multiple memory devices.

[0238] Aspects of the present disclosure have been described in terms of its illustrative embodiments. Many other embodiments, modifications, or variations within the scope and spirit of the appended claims will occur to those of ordinary skill in the art upon reading the present disclosure. Any and all features in the following claims can be combined or rearranged in any possible way. Accordingly, the scope of the present disclosure is presented by way of example and not limitation, and the subject matter disclosed herein does not exclude including such modifications, variations, or additions to the subject matter that would be apparent to those of ordinary skill in the art. Additionally, terms are described herein using lists of example elements connected by conjunctions such as "and," "or," "but," etc. It should be understood that such conjunctions are provided for explanatory purposes only. A list connected by a particular conjunction such as "or" can, for example, refer to "at least one" or "any combination" of the example elements listed therein, where "or" is understood as "and / or" unless otherwise indicated. Further, terms such as "based on" should be understood as "at least partially based on."

[0239] Using the disclosure provided herein, those of ordinary skill in the art will understand that, without departing from the scope of the present disclosure, the elements of any of the claims, operations, or processes discussed herein can be modified, rearranged, extended, omitted, combined, or altered in various ways. For purposes of illustrative explanation, some claims are described with alphabetic reference numerals to the claim elements and are not meant to be limiting. The alphabetic reference numerals do not imply a particular order of operations. For example, alphabetic identifiers such as (a), (b), (c), ……, (i), (ii), (iii) ……, etc. can be used to illustrate operations. Such identifiers are provided for the convenience of the reader and do not denote a particular order of steps or operations. The operations illustrated by the list identifiers such as (a), (i), etc. can be performed before, after, or in parallel with another operation illustrated by the list identifier such as (b), (ii), etc.

Claims

1. A computer-implemented method, comprising: (a) obtaining sensor data, the sensor data including a plurality of image frames indicating actors in the environment of an autonomous vehicle; (b) for each corresponding image frame, (i) using a machine learning model to determine that the actor is an emergency vehicle; (ii) using the machine learning model to generate output data, the output data indicating that the emergency vehicle is in an activated state or a non-activated state in the corresponding image frame, and (iii) storing attribute data of the corresponding image frame, the attribute data including the output data of the machine learning model and the time associated with the corresponding image frame; (c) determining that the emergency vehicle is an activated emergency vehicle based on the output data associated with the corresponding image frame; and (d) performing an action for the autonomous vehicle based on the activated emergency vehicle being within the environment of the autonomous vehicle.

2. The computer-implemented method according to claim 1, wherein, (b)(ii) includes: using the machine learning model to determine the state of the lights of the emergency vehicle in the corresponding image frame, wherein the state of the lights includes an on state or an off state; and using the machine learning model to determine that the emergency vehicle is in the activated state or the non-activated state for the corresponding image frame based on the state of the lights.

3. The computer-implemented method according to any one of claims 1 and 2, wherein, (b)(i) includes: using the machine learning model to determine the category of the emergency vehicle from one of the following categories: (1) police vehicle; (2) ambulance; (3) fire truck; or (4) tow truck.

4. The computer-implemented method according to any one of claims 1-3, wherein, the output data indicates the category of the emergency vehicle.

5. The computer-implemented method according to any one of claims 1-4, wherein, (b)(i) includes: obtaining tracking data of the actor within the environment of the autonomous vehicle, the tracking data indicating the boundary shape of the actor; determining that the centroid of the boundary shape is within the projection field of view of the autonomous vehicle; and in response to determining that the centroid of the boundary shape is within the projection field of view, generating input data for the machine learning model based on the sensor data.

6. The computer-implemented method according to any one of claims 1-5, wherein, the output data further includes tracking data associated with the emergency vehicle.

7. The computer-implemented method according to any one of claims 1-6, wherein, (b)(iii) includes storing the attribute data in a buffer, and wherein the method further includes: determining that the buffer includes a threshold amount of attribute data for a plurality of image frames at a plurality of times; and using a second model to determine that the emergency vehicle is an activated emergency vehicle based on the attribute data for at least a subset of the plurality of image frames.

8. The computer-implemented method according to any one of claims 1-7, wherein (c) includes: using the second model to determine at least one of the following: (1) the light pattern of the emergency vehicle; (2) the color of the lights of the emergency vehicle; or (3) the light intensity of the lights of the emergency vehicle.

9. The computer-implemented method according to any one of claims 1-8, wherein, the machine learning model is trained based on labeled training data, wherein, The training data of the markings is based on point cloud data and image data, and wherein, the training data of the markings indicates a plurality of training actors, and the corresponding training actors are labeled with emergency vehicle tags.

10. The computer-implemented method according to claim 9, wherein, the emergency vehicle tag indicates the type of the emergency vehicle of the corresponding training actor or that the training actor is not an emergency vehicle.

11. The computer-implemented method according to any one of claims 9 or 10, wherein, the corresponding training actor includes an active tag, and the active tag indicates whether the corresponding training actor is in an active state or an inactive state based on the lights of the training actor.

12. The computer-implemented method according to any one of claims 1-11, wherein, the machine learning model is a convolutional neural network.

13. The computer-implemented method according to any one of claims 7 or 8, wherein, the second model is a rule-based smoothing model.

14. The computer-implemented method according to any one of claims 1-13, wherein, the actions of the autonomous vehicle include at least one of the following: (1) predicting the movement of the activated emergency vehicle; (2) generating a motion plan for the autonomous vehicle; or (3) controlling the movement of the autonomous vehicle.

15. The computer-implemented method according to any one of claims 1-14, wherein, (b) (iii) includes storing the attribute data for the corresponding image frame in a buffer, the attribute data including the output data of the machine learning model and the time associated with the corresponding image frame, and wherein, (c) includes determining that the buffer includes the threshold amount of attribute data for the plurality of image frames at the plurality of times.

16. The computer-implemented method according to any one of claims 1-15, wherein, the machine learning model is further configured to process one or more historical image frames to generate the output data, the one or more historical image frames being associated with one or more time steps before the time step associated with the corresponding time frame.

17. One or more non-transitory computer-readable media storing instructions that, when executed, cause one or more processors to perform operations, the operations including: (a) obtaining sensor data, the sensor data including a plurality of image frames indicating actors in the environment of the autonomous vehicle. (b) For a corresponding image frame, (i) using a machine learning model, determining that the actor is an emergency vehicle, (ii) using the machine learning model, generating output data indicating whether the emergency vehicle is in an active state or an inactive state in the corresponding image frame, and (iii) storing attribute data for the corresponding image frame, the attribute data including the output data of the machine learning model and the time associated with the corresponding image frame; (c) determining, based on the output data associated with the corresponding image frame, that the emergency vehicle is an activated emergency vehicle; and (d) Based on the activated emergency vehicle being within the environment of the autonomous vehicle, perform an action for the autonomous vehicle.

18. The one or more non-transitory computer-readable media according to claim 17, wherein, (b)(ii) includes: Using the machine learning model, determine the state of the lights of the emergency vehicle in the corresponding image frame, where the state of the lights includes an on state or an off state; and Using the machine learning model, based on the state of the lights, determine whether the emergency vehicle is in an activated state or a non-activated state for the corresponding image frame.

19. The one or more non-transitory computer-readable media according to any one of claims 17 or 18, wherein the output data indicates the category of the emergency vehicle.

20. One or more non-transitory computer-readable media according to any one of claims 17-19, wherein, (b)(i) includes: Obtain tracking data of the actor within the environment of the autonomous vehicle, the tracking data indicating the boundary shape of the actor; Determine that the centroid of the boundary shape is within the projection field of view of the autonomous vehicle; and In response to determining that the centroid of the boundary shape is within the projection field of view, generate input data for the machine learning model based on the sensor data.

21. An autonomous vehicle control system for controlling an autonomous vehicle, the autonomous vehicle control system comprising: One or more processors; And One or more non-transitory computer-readable media storing instructions that can be executed by the one or more processors to cause the autonomous vehicle control system to control the movement of the autonomous vehicle using an operating system; wherein the operating system detects an emergency vehicle by: (a) Obtain sensor data, the sensor data including a plurality of image frames indicating actors in the environment of the autonomous vehicle; (b) For a corresponding image frame, (i) Use a machine learning model to determine that the actor is the emergency vehicle, (ii) Use the machine learning model to generate output data, the output data indicating whether the emergency vehicle is in an activated state or a non-activated state in the corresponding image frame, and (iii) Store attribute data for the corresponding image frame, the attribute data including the output data of the machine learning model and the time associated with the corresponding image frame; And (c) Based on the output data associated with the corresponding image frame, determine that the emergency vehicle is an activated emergency vehicle.

22. The autonomous vehicle control system according to claim 21, wherein (b)(iii) includes storing the attribute data in a buffer, and wherein the operating system further detects the emergency vehicle by: Determine that the buffer includes a threshold amount of attribute data for a plurality of image frames at a plurality of times; and Use a second model to determine that the emergency vehicle is an activated emergency vehicle based on the attribute data for at least a subset of the plurality of image frames.