Method and system for a vehicle and storage medium

By analyzing ambient sound, camera images, and 3D data, and combining acoustic localization and machine learning techniques, the system accurately locates rescue vehicles, solving the problem of inaccurate detection of rescue vehicles in existing systems and improving the response capability of the vehicles.

CN115718484BActive Publication Date: 2025-10-28MOTIONAL AD LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111485949.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-08-26
Filing Date
2021-12-07
Publication Date
2025-10-28
Estimated Expiration
2041-12-07

AI Technical Summary

Technical Problem

Existing rescue vehicle detection systems may not be able to accurately detect the presence and location of rescue vehicles, especially when the flashlights are obstructed, making it impossible for vehicle operators to effectively respond to the needs of rescue vehicles.

Method used

By receiving ambient sound, camera images, and 3D data, and using a processor to analyze siren sounds, flashing lights, and the presence of objects, the location of rescue vehicles is determined by combining acoustic localization and machine learning techniques. A fusion estimation is then performed to confirm the presence and location of the rescue vehicles.

Benefits of technology

It enables accurate positioning and response to rescue vehicles, ensuring that vehicles can perform avoidance and yielding operations in accordance with national vehicle operator regulations, and improves the planning and operation accuracy of autonomous or semi-autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115718484B_ABST
    Figure CN115718484B_ABST
Patent Text Reader

Abstract

This invention provides a method and system for a rescue vehicle, as well as a storage medium. In an embodiment, the method includes: receiving ambient sound; determining whether the ambient sound includes a horn; determining a first location associated with the horn based on the determination that the ambient sound includes a horn; receiving a camera image; determining whether the camera image includes a flash; determining a second location associated with the flash based on the determination that the camera image includes a flash; receiving 3D data; determining whether the 3D data includes an object; determining a third location associated with the object based on the determination that the 3D data includes an object; determining the presence of a rescue vehicle based on the horn, the detected flash, and the detected object; determining an estimated location of the rescue vehicle based on the first, second, and third locations; and initiating actions related to the vehicle based on the determined presence and location of the rescue vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The following description pertains to autonomous vehicle systems. Background Technology

[0002] Autonomous vehicles include multiple sensors (e.g., cameras, radar, LiDAR) for collecting data related to the vehicle's operating environment. This data is used by the vehicle to predict the state of the agent in the operating environment and to plan and execute the vehicle's trajectory in the operating environment, taking into account various rules and constraints (such as map constraints (e.g., drivable areas), safety constraints (e.g., avoiding collisions with other objects), and occupant comfort constraints (e.g., minimizing sharp turns, hard braking, and rapid acceleration / deceleration)).

[0003] In typical vehicle operation environments, autonomous vehicles may encounter rescue vehicles (e.g., fire trucks, ambulances). In the United States, vehicle operators are required to detect and respond to the sirens and flashing lights of rescue vehicles. The operator's response depends on the location of the rescue vehicle. For example, if the rescue vehicle is behind the vehicle, the operator must give way to it. If the rescue vehicle is in front of the vehicle (potentially at the scene of a rescue situation), the operator is required in many situations to provide a safe buffer (e.g., minimum distance) between the vehicle and the rescue vehicle (e.g., "Move Over" laws).

[0004] Current rescue vehicle detection systems may only detect horns or flashing lights. While flashing lights are the primary method, their visibility can often be obstructed by trucks, signs, or buildings. Summary of the Invention

[0005] Technology for detection systems and methods for rescue vehicles is provided.

[0006] In one embodiment, a method includes: using at least one processor to receive ambient sound; using the at least one processor to determine whether the ambient sound includes a horn sound; based on the determination that the ambient sound includes a horn sound, using the at least one processor to determine a first location associated with the horn sound; using the at least one processor to receive a camera image; using the at least one processor to determine whether the camera image includes a flash; based on the determination that the camera image includes a flash, using the at least one processor to determine a second location associated with the flash; using the at least one processor to receive three-dimensional data, i.e., 3D data; using the at least one processor to determine whether the 3D data includes an object; based on the determination that the 3D data includes an object, using the at least one processor to determine a third location associated with the object; using the at least one processor to determine the presence of a rescue vehicle based on the horn sound, the detected flash, and the detected object; using the at least one processor to determine an estimated location of the rescue vehicle based on the first location, the second location, and the third location; and using the at least one processor to initiate actions related to the vehicle based on the determined presence and location of the rescue vehicle.

[0007] In an embodiment, determining whether the ambient sound includes a horn sound further includes: converting the ambient sound into a digital signal; analyzing the digital signal for error and blocking signatures; notifying a diagnostic system to determine the cause of the error or blocking signature if an error or blocking signature is detected; and analyzing the digital signal for the horn sound if no error or blocking signature is detected.

[0008] In an embodiment, determining whether the ambient sound includes a horn sound further includes: converting the ambient sound into a digital signal; filtering noise from the digital signal; comparing the digital signal with a reference signal; and determining horn sound detection based on the comparison result.

[0009] In an embodiment, determining whether the ambient sound includes a horn sound further includes: converting the ambient sound into a digital signal; using machine learning to analyze the digital signal; and determining horn sound detection based on the analysis results.

[0010] In one embodiment, the method further includes: determining the confidence level of the horn sound detection.

[0011] In one embodiment, the first location associated with the source of the horn sound is determined based on acoustic localization, wherein the acoustic localization is based on time-of-flight analysis, or TOF analysis, of signals output from multiple microphones located around the vehicle.

[0012] In one embodiment, the first location associated with the source of the horn sound is determined based on the location and orientation of a plurality of unidirectional microphones, and the first location is determined based on the strength of a key signature signal obtained from analysis of the output signals of the plurality of unidirectional microphones and the location and orientation of each of the plurality of unidirectional microphones.

[0013] In an embodiment, determining whether the camera image includes a flash further includes: analyzing the camera image for error and blocking signatures; if an error or blocking signature is detected, notifying a diagnostic system to determine the cause of the error or blocking signature; and if no error or blocking signature is detected, analyzing the camera image to determine whether the camera image includes a flash.

[0014] In one embodiment, determining whether the camera image includes a flash further includes: filtering the camera image to remove low-intensity light signals to obtain a high-intensity image; comparing the high-intensity image with a reference image; and determining flash detection based on the comparison result.

[0015] In one embodiment, determining whether the camera image includes a flash further includes: using machine learning to analyze the camera image; and predicting flash detection based on the results of the analysis.

[0016] In one embodiment, the method further includes: determining the confidence level of the flash detection.

[0017] In one embodiment, the analysis using machine learning includes using a neural network to analyze the camera images, the neural network being trained to distinguish between bright flashes and other bright lights.

[0018] In an embodiment, the second location associated with the flash is determined by: detecting flash signatures in multiple camera images captured by multiple cameras; and determining the radial location of the flash based on the position and orientation of each of the multiple cameras; or determining the second location based on triangulation of the position and orientation of the multiple cameras.

[0019] In an embodiment, determining whether the 3D data includes an object further includes: analyzing the 3D data for error and blocking signatures; notifying a diagnostic system to determine the cause of the error or blocking signature if an error or blocking signature is detected; and analyzing the 3D data to determine whether the 3D data includes an object if no error or blocking signature is detected.

[0020] In an embodiment, determining whether the 3D data includes an object further includes: filtering the 3D data to remove noise; dividing the 3D data into groups with similar characteristics; and using machine learning to analyze the groups to determine object detection.

[0021] In one embodiment, the method further includes: determining a confidence level for the object detection.

[0022] In one embodiment, determining the third location includes determining the radial location of the object based on an analysis of the sensor position and orientation of the sensor that detected the object.

[0023] In one embodiment, determining the third location includes determining the radial location of the object based on an analysis of the sensor position and orientation of the sensor that detected the object.

[0024] In one embodiment, determining the presence and location of the rescue vehicle further includes: comparing siren detection, flashing light detection, and object detection to determine whether the siren detection, flashing light detection, and object detection match; and generating a signal indicating the presence of the rescue vehicle based on the determination of a match.

[0025] In one embodiment, the method further includes: tracking the rescue vehicle based on its location.

[0026] In one embodiment, a system includes: at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform any of the methods described above.

[0027] In one embodiment, a computer-readable storage medium stores instructions that, when executed by at least one processor, cause the at least one processor to perform any of the methods described above.

[0028] One or more of the disclosed embodiments provide one or more of the following advantages. The vehicle can accurately determine the presence and location of a rescue vehicle and perform vehicle behavior (such as avoidance and yielding) in response to detection according to national vehicle operator regulations. In embodiments, the planner, perception module, vehicle controller, or any other application or subsystem of a semi-autonomous or fully autonomous vehicle can more accurately determine the agent state (e.g., position, speed, heading) in the vehicle's operating environment, including but not limited to determining that other vehicles will slow down or merge away from the rescue vehicle. In embodiments, cloud-based remote vehicle assistance can be enabled to monitor / manage rescue vehicle encounters and take actions as needed, such as taking temporary control of the vehicle.

[0029] These and other aspects, features, and implementations may be referred to as methods, apparatus, systems, components, program products, manner or steps for performing a function, and other means. These and other aspects, features, and implementations will become clear from the following description, including the claims. Attached Figure Description

[0030] Figure 1 Examples of autonomous vehicles (AVs) with autonomous capabilities according to one or more embodiments are shown.

[0031] Figure 2 An example "cloud" computing environment is illustrated according to one or more embodiments.

[0032] Figure 3 Examples of computer systems according to one or more embodiments.

[0033] Figure 4 An example architecture of an AV is shown according to one or more embodiments.

[0034] Figure 5 This is a block diagram of a rescue vehicle detection system according to one or more embodiments.

[0035] Figure 6 According to one or more embodiments Figure 5 The block diagram shown is of the horn sound detection and positioning system 501.

[0036] Figure 7 According to one or more embodiments Figure 5 The diagram shown is a block diagram of the flashing light detection and positioning system.

[0037] Figure 8 According to one or more embodiments Figure 5 The diagram shows a block diagram of an object detection and localization system.

[0038] Figure 9 This is a flowchart of the rescue vehicle detection process performed by a rescue vehicle detection system according to one or more embodiments. Detailed Implementation

[0039] In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of the invention. However, it will be apparent that the invention may be practiced without these specific details. In other instances, well-known constructions and apparatuses are shown in block diagram form to avoid unnecessarily obscuring the invention.

[0040] In the accompanying drawings, for ease of description, a specific arrangement or order of schematic elements (such as those representing devices, modules, instruction blocks, and data elements) is shown. However, those skilled in the art will understand that the specific order or arrangement of the schematic elements in the drawings is not intended to imply a requirement for a particular processing order or sequence, or a separation of processing procedures. Furthermore, the inclusion of schematic elements in the drawings is not intended to imply that such elements are required in all embodiments, nor is it intended to imply that features represented by such elements cannot be included in some embodiments or cannot be combined with other elements in some embodiments.

[0041] Furthermore, in the accompanying drawings, connecting elements (such as solid or dashed lines or arrows) are used to illustrate connections, relationships, or associations between two or more other schematic elements. The absence of any such connecting element does not imply that connections, relationships, or associations cannot exist. In other words, connections, relationships, or associations between some elements are not shown in the drawings so as not to obscure the content of this disclosure. Additionally, for ease of illustration, a single connecting element is used to represent multiple connections, relationships, or associations between elements. For example, if a connecting element represents communication of signals, data, or instructions, those skilled in the art will understand that such an element represents one or more signal paths (e.g., a bus) that may be necessary to influence the communication.

[0042] Reference will now be made in detail to embodiments, examples of which are illustrated in the accompanying drawings. Numerous specific details are set forth in the following detailed description in order to provide a thorough understanding of the various embodiments described. However, it will be apparent to those skilled in the art that the various embodiments described can be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0043] The features described below can each be used independently of each other or in any combination with other features. However, any individual feature may not solve any of the problems discussed above, or may only solve one of the problems discussed above. Some of the problems discussed above may not be adequately solved by any of the features described herein. Although headings are provided, information relating to specific headings but not found in the sections bearing those headings can be found elsewhere in this specification. Embodiments are described herein based on the following summary:

[0044] 1. General Overview

[0045] 2. System Overview

[0046] 3. Autonomous Vehicle Architecture

[0047] 4. Rescue vehicle detection system and methods

[0048] General Overview

[0049] Techniques for a rescue vehicle detection system and method are provided. Ambient sound, camera images, and 3D sensor data (e.g., LiDAR, RADAR, SONAR) are captured and analyzed to detect horn sounds, flashing lights, and objects in the vehicle's operating environment (e.g., autonomous vehicle operating environment). The types of horn sounds, flashing lights, and objects are also determined. After detection, the locations of the horn sounds, flashing lights, and objects are estimated. The estimated locations and determined types are fused to determine the presence and location of the rescue vehicle in the operating environment, and in some implementations, the type of rescue vehicle is determined.

[0050] The estimated location can be used by downstream systems of the autonomous vehicle stack (e.g., planners, perception modules, vehicle controllers) to determine the vehicle's route or trajectory in the operating environment. Using the rescue vehicle's location, the vehicle's planner, perception module, or vehicle controller can more accurately determine the agent state in the vehicle's operating environment, such as determining that other vehicles will slow down or merge away from the rescue vehicle. Furthermore, rescue vehicle information can enable remote vehicle assistance to monitor / manage rescue vehicle encounters and take actions as needed (e.g., take temporary control of the vehicle).

[0051] System Overview

[0052] Figure 1 An example of an autonomous vehicle 100 with autonomous capabilities is shown.

[0053] As used herein, the term “autonomy” refers to a function, feature, or facility that enables a vehicle to operate partially or fully without real-time human intervention, including but not limited to fully autonomous vehicles, highly autonomous vehicles, and conditionally autonomous vehicles.

[0054] As used in this article, an autonomous vehicle (AV) is a vehicle with autonomous capabilities.

[0055] As used in this article, "vehicle" includes any mode of transport for goods or people. Examples include cars, buses, trains, airplanes, drones, trucks, ships, vessels, submersibles, spacecraft, motorcycles, bicycles, etc. Driverless cars are an example of vehicles.

[0056] As used herein, a “trajectory” refers to a path or route that operates an AV from a first spatiotemporal location to a second spatiotemporal location. In embodiments, the first spatiotemporal location is referred to as the initial location or starting point, and the second spatiotemporal location is referred to as the destination, final location, target, target location, or target position. In some examples, a trajectory consists of one or more road segments (e.g., segments of a road), and each road segment consists of one or more blocks (e.g., a lane or part of an intersection). In embodiments, spatiotemporal locations correspond to real-world locations. For example, a spatiotemporal location is a pick-up or drop-off point for people or goods to board or alight.

[0057] As used in this paper, “manifestation” refers to the trajectory generated by the sample-based maneuvering manifestation device described in this paper.

[0058] A "maneuver" is a change in the position, speed, or heading of an AV. All maneuvers are trajectories, but not all trajectories are maneuvers. For example, the trajectory of an AV traveling at a constant speed on a straight path is not a maneuver.

[0059] As used herein, “(one or more) sensors” includes one or more hardware components for detecting information relating to the environment surrounding the sensor. Some hardware components may include sensing components (e.g., image sensors, biometric sensors), transmission and / or receiving components (e.g., laser or radio frequency wave transmitters and receivers), electronic components (such as analog-to-digital converters), data storage devices (such as RAM and / or non-volatile memory), software or firmware components, and data processing components (such as ASICs (Application-Specific Integrated Circuits)), microprocessors, and / or microcontrollers.

[0060] As used in this article, a "road" is a physical area that can be traversed by vehicles and can correspond to a named arterial road (e.g., a city street, an interstate highway, etc.) or an unnamed arterial road (e.g., a driveway within a house or office building, a section of a parking lot, a section of an open space, a dirt road in a rural area, etc.). Because some vehicles (e.g., four-wheel drive pickup trucks, SUVs, etc.) can traverse a variety of physical areas that are not particularly suitable for vehicle travel, a "road" can be any physical area that is not formally defined as an arterial road by any municipality or other government or administrative agency.

[0061] As used herein, a “lane” is the portion of a road that can be traversed by vehicles and may correspond to most or all of the space between lane markings, or only a portion of the space between lane markings (e.g., less than 50%). For example, a road with lane markings spaced far apart may accommodate two or more vehicles, allowing one vehicle to overtake another without crossing the lane markings; therefore, it can be interpreted as a lane being narrower than the space between lane markings, or as having two lanes. Lanes can also be interpreted in the absence of lane markings. For example, a lane may be defined based on the physical characteristics of the environment (e.g., rocks in a rural area and trees along a road).

[0062] As used herein, a “rulebook” is a data structure that implements a priority structure on a set of rules arranged according to their relative importance, wherein for any particular rule in the priority structure, one or more rules(s) with a lower priority than that particular rule in the priority structure have lower importance than that particular rule. Possible priority structures include, but are not limited to: hierarchical structures (e.g., the overall or partial order of rules), non-hierarchical structures (e.g., a weighted system of rules), or hybrid priority structures in which subsets of rules are hierarchical, but the rules within each subset are non-hierarchical. Rules may include traffic laws, safety rules, ethical rules, local cultural rules, occupant comfort rules, and any other rules that can be used to evaluate vehicle trajectories provided by any source (e.g., humans, texts, regulations, websites).

[0063] As used herein, “ego vehicle” or “ego” refers to a virtual vehicle or AV having virtual sensors for sensing a virtual environment, which is used, for example, by a planner to plan the route of a virtual AV within that virtual environment.

[0064] "One or more" includes functions performed by a single element, functions performed by multiple elements, such as in a distributed manner, several functions performed by a single element, several functions performed by several elements, or any combination of the foregoing.

[0065] It will also be understood that, although in some cases the terms “first,” “second,” etc., are used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, without departing from the scope of the various described embodiments, a first contact may be referred to as a second contact, and similarly, a second contact may be referred to as a first contact. Both the first contact and the second contact are contacts, but they are not the same contact.

[0066] The terminology used in the description of the various embodiments described herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in the description of the various embodiments described and the appended claims, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that “and / or” as used herein refers to and includes any and all possible combinations of one or more of the relevant list items. It will also be understood that when the terms “comprising,” “including,” “possessing,” and / or “having” are used in this specification, they specifically indicate the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0067] As used herein, depending on the context, the term "if" may optionally be understood as meaning "when" or "at that time" or "in response to being determined" or "in response to being detected." Similarly, depending on the context, the phrase "if determined" or "if [the stated condition or event] has been detected" may optionally be understood as meaning "when determined" or "in response to being determined" or "when [the stated condition or event] is detected" or "in response to being detected."

[0068] As used herein, an AV system refers to an AV and an array of hardware, software, stored data, and real-time generated data that support AV operation. In embodiments, the AV system is incorporated within an AV. In embodiments, the AV system is distributed across several locations. For example, some of the software of the AV system is similar to that described below. Figure 2 The cloud computing environment described is implemented on the cloud computing environment 200.

[0069] Generally, this document describes technologies applicable to any vehicle with one or more autonomous capabilities, including fully autonomous vehicles, highly autonomous vehicles, and conditionally autonomous vehicles, such as so-called Level 5, Level 4, and Level 3 vehicles, respectively (see SAE International's standard J3016: Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems, the entire contents of which are incorporated herein by reference for further details on vehicle autonomy levels). The technologies described in this document also apply to partially autonomous vehicles and driver-assisted vehicles, such as so-called Level 2 and Level 1 vehicles (see SAE International's standard J3016: Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems). In embodiments, one or more Level 1, Level 2, Level 3, Level 4, and Level 5 vehicle systems may automatically perform certain vehicle operations (e.g., steering, braking, and map usage) under certain operating conditions based on the processing of sensor inputs. The techniques described in this document can benefit vehicles of any level, ranging from fully autonomous vehicles to human-operated vehicles.

[0070] refer to Figure 1 The AV system 120 enables the AV 100 to operate along a trajectory 198, traversing the environment 190 to the destination 199 (sometimes referred to as the final location), while avoiding objects (e.g., natural obstacles 191, vehicles 193, pedestrians 192, cyclists and other obstacles) and complying with road rules (e.g., operating rules or driving preferences).

[0071] In one embodiment, the AV system 120 includes means 101 for receiving and operating commands from a computer processor 146. In another embodiment, the computer processor 146 is referenced below. Figure 3 The processor 304 described is similar. Examples of the device 101 include a steering controller 102, a brake 103, a gear, an accelerator pedal or other acceleration control mechanism, a windshield wiper, a side door lock, a window controller, and a turn indicator.

[0072] In an embodiment, the AV system 120 includes sensors 121 for measuring or inferring attributes of the state or condition of the AV 100, such as the AV's position, linear velocity and linear acceleration, angular velocity and angular acceleration, and heading (e.g., orientation of the front of the AV 100). Examples of sensors 121 include a Global Navigation Satellite System (GNSS) receiver, an inertial measurement unit (IMU) for measuring both linear acceleration and angular rate of the vehicle, a wheel rate sensor for measuring or estimating wheel slip ratio, a wheel braking pressure or braking torque sensor, an engine torque or wheel torque sensor, and steering angle and angular rate sensors.

[0073] In an embodiment, sensor 121 also includes sensors for sensing or measuring properties of the AV's environment. Examples include a monocular or stereo camera 122 with visible, infrared, or thermal (or both) spectra, a LiDAR 123, a RADAR, an ultrasonic sensor, a time-of-flight (TOF) depth sensor, a rate sensor, a temperature sensor, a humidity sensor, and a precipitation sensor.

[0074] In one embodiment, the AV system 120 includes a data storage unit 142 and a memory 144 for storing machine instructions associated with a computer processor 146 or data collected by the sensor 121. In another embodiment, the data storage unit 142 is associated with the following... Figure 3 The described ROM 308 or storage device 310 is similar. In this embodiment, memory 144 is similar to main memory 306 described below. In this embodiment, data storage unit 142 and memory 144 store historical, real-time, and / or predictive information about environment 190. In this embodiment, the stored information includes maps, driving performance, traffic congestion updates, or weather conditions. In this embodiment, data related to environment 190 is transmitted from remote database 134 to AV100 via a communication channel.

[0075] In an embodiment, AV system 120 includes communication devices 140 for transmitting measured or inferred attributes of the state and conditions of other vehicles (such as position, linear and angular velocities, linear and angular accelerations, and linear and angular headings) to AV 100. These devices include vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) communication devices, as well as devices for wireless communication via point-to-point or ad hoc networks, or both. In an embodiment, communication device 140 communicates across the electromagnetic spectrum (including radio and optical communications) or other media (e.g., air and acoustic media). Combinations of vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V2I) communication (and in some embodiments, one or more other types of communication) are sometimes referred to as vehicle-to-everything (V2X) communication. V2X communication typically conforms to one or more communication standards for communication with and between autonomous vehicles.

[0076] In an embodiment, the communication device 140 includes a communication interface. For example, this may be a wired, wireless, WiMAX, Wi-Fi, Bluetooth, satellite, cellular, optical, near-field, infrared, or radio interface. The communication interface transmits data from a remote database 134 to the AV system 120. In an embodiment, the remote database 134 is embedded in, for example... Figure 2 In the cloud computing environment 200 described herein, communication interface 140 transmits data collected from sensor 121 or other data related to the operation of AV 100 to remote database 134. In some embodiments, communication interface 140 transmits information related to teleoperation to AV 100. In some embodiments, AV 100 communicates with other remote (e.g., "cloud") servers 136.

[0077] In this embodiment, the remote database 134 also stores and transmits digital data (e.g., data such as road and street locations). This data is stored in memory 144 on the AV 100 or transmitted from the remote database 134 to the AV 100 via a communication channel.

[0078] In one embodiment, the remote database 134 stores and transmits historical information (e.g., rate and acceleration distribution) related to driving attributes of vehicles that previously traveled along trajectory 198 at similar times of day. In one implementation, such data may be stored in memory 144 on the AV 100 or transmitted from the remote database 134 to the AV 100 via a communication channel.

[0079] The computing device 146 located on AV 100 generates control actions in an algorithmic manner based on both real-time sensor data and prior information, allowing AV system 120 to perform its autonomous driving capabilities.

[0080] In one embodiment, the AV system 120 includes a computer peripheral device 132 coupled to a computing device 146 for providing information and alerts to a user of the AV 100 (e.g., an occupant or a remote user) and receiving input from that user. In another embodiment, the peripheral device 132 is similar to the one described in the following reference. Figure 3 The discussed display 312, input device 314, and cursor controller 316 are coupled wirelessly or wiredly. Any two or more interface devices can be integrated into a single device.

[0081] Example cloud computing environment

[0082] Figure 2 Example of a "cloud" computing environment. Cloud computing is a service delivery model that enables convenient, on-demand access over a network to a shared pool of configurable computing resources, such as networks, network bandwidth, servers, processing power, memory, storage, applications, virtual machines, and services. In a typical cloud computing system, one or more large cloud data centers house the machines used to deliver the services provided by the cloud. Now refer to... Figure 2 The cloud computing environment 200 includes cloud data centers 204a, 204b, and 204c interconnected via cloud 202. Data centers 204a, 204b, and 204c provide cloud computing services to computer systems 206a, 206b, 206c, 206d, 206e, and 206f connected to cloud 202.

[0083] A cloud computing environment 200 includes one or more cloud data centers. Generally, a cloud data center (e.g.) Figure 2 The cloud data center 204a shown refers to the cloud (e.g., Figure 2 The physical arrangement of servers in cloud 202 (or a specific portion of the cloud) is illustrated. For example, servers are physically arranged in rooms, groups, rows, and racks within a cloud data center. A cloud data center has one or more regions, each containing one or more server rooms. Each room has one or more rows of servers, and each row includes one or more racks. Each rack includes one or more individual server nodes. In some implementations, servers in regions, rooms, racks, and / or rows are arranged into groups based on the physical infrastructure requirements of the data center facility, including power, energy, heat, heat sources, and / or other requirements. In this embodiment, server nodes are similar to... Figure 3 The computer system described herein. Data center 204a has many computer systems distributed across multiple racks.

[0084] Cloud 202 includes cloud data centers 204a, 204b, and 204c, and networks and network resources (e.g., network devices, nodes, routers, switches, and network cables) for connecting cloud data centers 204a, 204b, and 204c and facilitating access to cloud computing services by computer systems 206a-206f. In embodiments, the network represents one or more local area networks, wide area networks, or any combination of wired or wireless networks coupled using terrestrial or satellite connections. Data exchanged over the network is transmitted using various network layer protocols, such as Internet Protocol (IP), Multiprotocol Label Switching (MPLS), Asynchronous Transfer Mode (ATM), Frame Relay, etc. Furthermore, in embodiments where the network represents a combination of multiple sub-networks, different network layer protocols are used on each underlying sub-network. In some embodiments, the network represents one or more interconnected internetworks (such as the public Internet).

[0085] Computer systems 206a-206f or cloud computing service consumers connect to cloud 202 via network links and network adapters. In embodiments, computer systems 206a-206f are implemented as various computing devices, such as servers, desktops, laptops, tablets, smartphones, Internet of Things (IoT) devices, autonomous vehicles (including cars, drones, space shuttles, trains, buses, etc.), and consumer electronics. In embodiments, computer systems 206a-206f are implemented in other systems or as part of other systems.

[0086] Computer System

[0087] Figure 3 Example: Computer system 300. In implementation, computer system 300 is a dedicated computing device. The dedicated computing device is hardwired to perform these techniques, or includes a digital electronic device persistently programmed to perform the aforementioned techniques, such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), or may include one or more general-purpose hardware processors programmed to perform these techniques according to program instructions in firmware, memory, other storage units, or combinations thereof. Such a dedicated computing device may also combine custom hardwired logic, ASICs, or FPGAs with custom programming to perform these techniques. In various embodiments, the dedicated computing device is a desktop computer system, a portable computer system, a handheld device, a network device, or any other device that includes hardwired and / or program logic to implement these techniques.

[0088] In one embodiment, the computer system 300 includes a bus 302 or other communication mechanism for conveying information, and a hardware processor 304 coupled to the bus 302 to process information. The hardware processor 304 is, for example, a general-purpose microprocessor. The computer system 300 also includes a main memory 306, such as random access memory (RAM) or other dynamic storage device, coupled to the bus 302 to store information and instructions executed by the processor 304. In one implementation, the main memory 306 is used to store temporary variables or other intermediate information during the execution of instructions to be executed by the processor 304. When these instructions are stored in a non-transitory storage medium accessible to the processor 304, the computer system 300 becomes a dedicated machine customized to perform the operations specified in the instructions.

[0089] In an embodiment, the computer system 300 further includes a read-only memory (ROM) 308 or other static storage device coupled to the bus 302 for storing static information and instructions of the processor 304. A storage device 310, such as a disk, optical disk, solid-state drive, or three-dimensional cross-point memory, is provided and coupled to the bus 302 to store information and instructions.

[0090] In this embodiment, the computer system 300 is coupled via a bus 302 to a display 312, such as a cathode ray tube (CRT), liquid crystal display (LCD), plasma display, light-emitting diode (LED) display, or an organic light-emitting diode (OLED) display for displaying information to a computer user. An input device 314, including alphanumeric keys and other keys, is coupled to the bus 302 for transmitting information and command selections to the processor 304. Another type of user input device is a cursor controller 316, such as a mouse, trackball, touchscreen, or cursor arrow keys, for transmitting directional information and command selections to the processor 304 and for controlling the movement of the cursor on the display 312. Such input devices typically have two degrees of freedom on two axes (a first axis (e.g., the x-axis) and a second axis (e.g., the y-axis)), which allow the device to specify a position on a plane.

[0091] According to one embodiment, the techniques described herein are executed by computer system 300 in response to processor 304 executing one or more sequences of one or more instructions contained in main memory 306. These instructions are read into main memory 306 from another storage medium, such as storage device 310. Executing the sequence of instructions contained in main memory 306 causes processor 304 to perform the process steps described herein. In alternative embodiments, hardwired circuitry is used instead of or in combination with software instructions.

[0092] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that enable a machine to operate in a particular manner. Such storage media include non-volatile media and / or volatile media. Non-volatile media include, for example, optical discs, magnetic disks, solid-state drives, or three-dimensional cross-point memory such as storage device 310. Volatile media include dynamic memory, such as main memory 306. Common forms of storage media include, for example, floppy disks, floppy disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with perforations, RAM, PROMs and EPROMs, FLASH-EPROMs, NV-RAMs, or any other memory chips or memory cartridges.

[0093] Storage media differ from transmission media, but can be used in conjunction with them. Transmission media participate in the information transmission between storage media. For example, transmission media include coaxial cables, copper wires, and optical fibers, which include wires with a bus 302. Transmission media can also take the form of sound waves or light waves, such as those generated during radio wave and infrared data communication.

[0094] In embodiments, various forms of media involve carrying one or more sequences of one or more instructions to processor 304 for execution. For example, these instructions may initially be executed on a disk or solid-state drive of a remote computer. The remote computer loads the instructions into its dynamic memory and transmits them over a telephone line using a modem. A local modem of computer system 300 receives data over the telephone line and converts the data into an infrared signal using an infrared transmitter. An infrared detector receives the data carried in the infrared signal, and appropriate circuitry places the data on bus 302. Bus 302 carries the data to main memory 306, from which processor 304 retrieves and executes the instructions. The instructions received by main memory 306 may optionally be stored on storage device 310 before or after execution by processor 304.

[0095] Computer system 300 also includes a communication interface 318 coupled to bus 302. Communication interface 318 provides bidirectional data communication coupled to network link 320 connected to local network 322. For example, communication interface 318 is an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem used to provide data communication connectivity with a corresponding type of telephone line. As another example, communication interface 318 is a Local Area Network (LAN) card used to provide data communication connectivity with a compatible LAN. In some implementations, a wireless link is also implemented. In any such implementation, communication interface 318 transmits and receives electrical, electromagnetic, or optical signals carrying digital data streams representing various types of information.

[0096] Network link 320 typically provides data communication to other data devices via one or more networks. For example, network link 320 provides connectivity to host computer 324 or to a cloud data center or device operated by Internet Service Provider (ISP) 326 via local network 322. ISP 326, in turn, provides data communication services via a worldwide packet data communication network now commonly referred to as the "Internet" 328. Both local network 322 and Internet 328 use electrical, electromagnetic, or optical signals that carry digital data streams. Signals through various networks and signals on network link 320 via communication interface 318 are example forms of transmission media carrying digital data entering and leaving computer system 300. In embodiments, network 320 includes the aforementioned cloud 202 or a portion of cloud 202.

[0097] Computer system 300 sends messages and receives data including program code through one or more networks, network links 320, and communication interfaces 318. In an embodiment, computer system 300 receives code for processing. The received code is executed by processor 304 upon receipt and / or stored in storage device 310, or in other non-volatile storage devices for later execution.

[0098] Autonomous Vehicle Architecture

[0099] Figure 4 This illustrates the use of autonomous vehicles (e.g., Figure 1 The example architecture 400 of the AV 100 shown is illustrated. Architecture 400 includes a sensing module 402 (sometimes called a sensing circuit), a planning module 404 (sometimes called a planning circuit), a control module 406 (sometimes called a control circuit), a positioning module 408 (sometimes called a positioning circuit), and a database module 410 (sometimes called a database circuit). Each module plays a role in the operation of the AV 100. Commonly, modules 402, 404, 406, 408, and 410 can be... Figure 1 This is part of the AV system 120 shown. In some embodiments, any of modules 402, 404, 406, 408, and 410 is a combination of computer software (e.g., executable code stored on a computer-readable medium) and computer hardware (e.g., one or more microprocessors, microcontrollers, application-specific integrated circuits (ASICs), hardware memory devices, other types of integrated circuits, other types of computer hardware, or any or all combinations of these hardware).

[0100] In use, the planning module 404 receives data representing the destination 412 and determines data representing the trajectory 414 (sometimes called the route) that the AV100 can travel to reach (e.g., arrive at) the destination 412. In order for the planning module 404 to determine the data representing the trajectory 414, the planning module 404 receives data from the sensing module 402, the positioning module 408, and the database module 410.

[0101] The sensing module 402 is used, for example, as follows Figure 1 One or more sensors 121 are shown to identify nearby physical objects. The objects are classified (e.g., grouped into types such as pedestrians, bicycles, cars, traffic signs, etc.), and a scene description including the classified objects 416 is provided to the planning module 404.

[0102] The planning module 404 also receives data representing the location 418 of the AV from the positioning module 408. The positioning module 408 determines the location of the AV by using data from the sensor 121 and data (e.g., geographic data) from the database module 410. For example, the positioning module 408 uses data from a GNSS receiver and geographic data to calculate the longitude and latitude of the AV. In embodiments, the data used by the positioning module 408 includes high-precision maps with lane geometry properties, maps describing road network connectivity properties, maps describing lane physical properties (such as traffic speed, traffic volume, number of vehicle and bicycle lanes, lane width, lane traffic direction, or lane marking type and location, or combinations thereof), and maps describing spatial locations of road features (such as intersections, traffic signs, or various types of other traffic signals).

[0103] The control module 406 receives data representing trajectory 414 and data representing AV position 418, and operates the AV's control functions 420a-420c (e.g., steering, throttle, braking, ignition) in a manner that will cause the AV 100 to travel along trajectory 414 to reach destination 412. For example, if trajectory 414 includes a left turn, the control module 406 will operate the control functions 420a-420c in such a way that the steering angle of the steering function will cause the AV 100 to turn left, and the throttle and brake will cause the AV 100 to pause before turning and wait for passing pedestrians or vehicles.

[0104] In this embodiment, any of the aforementioned modules 402, 404, 406, and 408 can send a request to the rule-based trajectory verification system 500 to verify the planned trajectory and receive the score of that trajectory, as shown in the reference. Figures 5 to 8 As described in further detail.

[0105] Rescue Vehicle Inspection System and Method

[0106] Figure 5 This is a block diagram of a rescue vehicle detection system 500 according to one or more embodiments. The rescue vehicle detection system 500 includes a siren detection and location pipeline 501, a flashing light detection and location pipeline 502, an object detection and location pipeline 503, and a fusion module 504.

[0107] The horn detection and localization pipeline 501 receives ambient audio captured by one or more microphones and analyzes the audio to detect the presence of a horn. If a horn is detected, the location of the horn in the vehicle's operating environment is estimated.

[0108] Flash detection and localization pipeline 502 receives camera images from one or more camera systems and analyzes the images to detect the presence of flashes. If a flash is detected, the location of the flash in the vehicle's operating environment is estimated. Object detection and localization pipeline 503 receives echoes from light, radio frequency (RF), or sound waves emitted into the environment and analyzes the echoes (waves reflected from objects in the environment) to detect one or more objects. If an object is detected, the location of the object in the vehicle's operating environment is estimated.

[0109] The estimated locations output from pipelines 501, 502, and 503 are input to fusion module 504, which determines whether the estimated location matches. If the degree of matching of the estimated location is within a specified matching criterion, the matching location of the rescue vehicle can be used by downstream systems or applications. For example, the perception module 402, planner 404, and / or controller module 406 of AV 100 can use the estimated location of the rescue vehicle to predict the behavior of other agents in the operating environment (such as avoiding or yielding to the side of the road), and then plan the route or trajectory of the vehicle in the operating environment based on other predictions of agent behavior. Reference will now be made to... Figures 6 to 8 Provide a more detailed description of production lines 501, 502, and 503.

[0110] Figure 6 According to one or more embodiments Figure 5The diagram shows a block diagram of a horn sound detection and localization system 501. An audio interface 601 receives ambient sounds and converts these sounds into digital signals. In an embodiment, the ambient sound receiver includes an analog front end (AFE) coupled to one or more microphones or microphone arrays (e.g., phased array microphone arrays, MEMS microphone arrays, linear microphone arrays, beamforming microphone arrays). Each microphone is coupled to a preamplifier (e.g., a two-stage differential amplifier) ​​including a bandpass filter to provide a flat response, for example, between 1.6 Hz and 24 kHz. The preamplifier increases the gain of the microphone output signal and differentially connects this signal to an analog-to-digital converter (ADC). In an embodiment, the ADC is a multi-bit (e.g., 24-bit) sigma-delta ADC. The ADC can be sourced from the same clock to guarantee low jitter, which is important for time-dependent signals and algorithms based on Time Differential Arrival (TDOA) (described in further detail below). An application processor or logic circuit (such as a field-programmable gate array (FPGA)) is coupled to the output of one or more ADCs via a serial or parallel bus interface (e.g., a Serial Peripheral Interface (SPI)).

[0111] In an embodiment, an application embedded in one or more processors or FPGAs receives digital signals output by one or more ADCs and adds metadata such as channel number, time, and error detection codes (e.g., cyclic redundancy check (CRC) codes). These digital signals are then input to a data verifier 602, which can be implemented in software or logic.

[0112] Data verifier 602 analyzes the digital signal for error and blocking signatures, including but not limited to: lost data, muffled signals, and error protocols. For example, lost or leaked data can be detected by error detection / correction codes or by dropped packets. Muffled signals can be detected by monitoring for drops in the signal level (e.g., unexpected drops that do not follow the expected pattern of known horns, as discussed below). If an error is detected, the diagnostic system is notified so that the error can be reported, analyzed, and / or corrected if possible. The output of data verifier 602 is input to horn detector 603.

[0113] The horn detector 602 analyzes a horn sound signal. This analysis includes, but is not limited to, filtering noise (e.g., wind noise, road noise, tire noise, "city" noise) and comparing the filtered signal with known horn sounds. This comparison can be facilitated by first transforming the signal to the frequency domain, for example, using a Fast Fourier Transform (FFT), to determine the signal's spectrum, and then searching for energy in a specific frequency band known to contain the horn sound. In an embodiment, a frequency matching algorithm can be used to match the signal's spectrum with a reference spectrum of the horn content. In an embodiment, the matching algorithm can store and use characteristics of the reference spectrum for different horn types (e.g., yelip, wail, hi-lo, power call, air horn, and howler horns).

[0114] In one embodiment, a machine learning model (e.g., a neural network) is used to detect horn sounds in the signal output by the data validator 602. In another embodiment, the machine learning model also estimates and outputs a confidence level for the horn detection (e.g., the probability of horn sound detection). In yet another embodiment, if a specific confidence level for a horn sound detection is below a defined confidence threshold, the horn sound detection is considered a false detection and excluded from the input of the fusion module 504.

[0115] If the horn detector 603 detects a horn sound with a confidence level higher than a confidence threshold, the sound source locator 604 estimates the location of the horn sound source. Multiple sensors (e.g., directional microphones, phased-array microphone arrays, MEMS microphones) placed at different locations around the vehicle can be used, and sound source localization methods can be applied to the microphone signals to determine the location of the horn sound source. Some examples of sound source localization methods include, but are not limited to, time of arrival (TOA) measurement, TDOA measurement, or direction of arrival (DOA) estimation, or by utilizing controlled response power (SRP) functionality. In an embodiment, three-dimensional (3D) sound source localization is achieved by the sound source locator 604 using, for example, microphone arrays, machine learning models, maximum likelihood, multiple signal classification (MUSIC), arrays of acoustic vector sensors (AVS), and controlled beamformers (e.g., delay-addition (DAS) beamformers). In an embodiment, each sensor is a unidirectional microphone with a beam pattern (main lobe) sensitive along a narrow audio "field of view." Positioning can be determined by comparing the strength of microphone signals and the pointing direction of the microphone that outputs the strongest signal among all microphone signals.

[0116] After detecting and locating the horn sound source, the estimated location of the horn sound source is input into the fusion module 504.

[0117] Figure 7According to one or more embodiments Figure 5 The block diagram shown illustrates a flash detection and positioning system 502. A camera interface 701 receives camera images and converts them into digital signals. In an embodiment, the camera interface 701 is a camera system including an image sensor for receiving incident light (photons) converged by a lens or other optics. One or more sensors (e.g., CMOS sensors) convert photons into electrons, electrons into analog voltages, and then one or more ADCs convert the analog voltages into digital values. An application processor or logic circuit (such as an FPGA) is coupled to the output of the one or more ADCs via a serial or parallel bus interface (e.g., an SPI interface).

[0118] In an embodiment, an application embedded in one or more processors or an FPGA receives digital signals output by the one or more ADCs. These digital signals are input to a data verifier 702, which can be implemented in software or logic. The data verifier 702 analyzes these digital signals for error and blocking signatures, including but not limited to: lost data, muffled signals, and error protocols. Lost or leaked data can be detected by error detection / correction codes or dropped packets, etc. Muffled signals can be detected by monitoring a drop in signal level, etc. If an error is detected, a diagnostic system is notified so that the error can be reported, analyzed, and / or repaired if possible. The output of the data verifier 702 is input to a flash detector 703.

[0119] Flash detector 703 analyzes flash signals. This analysis includes, but is not limited to: filtering a sequence of camera images of low-intensity light signals and comparing the resulting high-intensity image sequence with a reference image sequence of known flashes. If the degree of matching between one or more high-intensity camera images and one or more reference images is within a defined matching criterion, then flash detection is considered to have occurred.

[0120] In this embodiment, a machine learning model (e.g., a neural network) is used to detect flashlights. The output of the machine learning model may be a confidence level (e.g., a probability) of detecting a flashlight. In this embodiment, if a particular confidence level of a specific flashlight detection is lower than a defined confidence threshold, the flashlight detection is considered a false detection and is excluded from the input of the fusion module 504. Some examples of suitable deep convolutional neural networks include, but are not limited to: Bochkovskiy, A., Wang, CY., Hong-Yuan, MLYOLOv4: Optimal Speed ​​and Accuracy of Object Detection (retrieved from http: / / arxiv.org / pdf / 2004.10934.pdf), Shaoqing, R., Kaiming, H., Girshick R., Sun, J. Faster R-CNN: Towards Real-Time Object Detection with Regional Proposal Networks (retrieved from http: / / arxiv.org / pdf / 1506.01497.pdf), and Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, CY, Berg, A. SSD: Single Shot MultiBox Detector (retrieved from http: / / arxiv.org / pdf / 1512.02325.pdf).

[0121] If the flash detector 703 detects a flash with a confidence level (e.g., detection probability) higher than a confidence threshold, the flash source locator 704 estimates the location of the flash source. In one embodiment, the estimated location of the flash source can be achieved by detecting objects in an image and additionally estimating their radial distance. This can be transformed into a location in world coordinates when combined with known camera positions. In another embodiment, triangulation can be used to determine the estimated location of the flash, in which multiple camera systems are positioned around the vehicle, and the location is determined by observing key signatures within camera images produced by the multiple cameras. For example, the key signature could be a set of flashing, high-intensity pixels in the camera images.

[0122] In an embodiment, such as described in Yabuuchi K., Hirano M., Senoo T., Kishi N., Ishikawa M. Real-Time Traffic Light Detection with Frequency Patterns Using a High-speed Camera. Sensors (Basel). 2020; 20(14):4035 (published on July 20, 2020. doi:10.3390 / s20144035), the method for detecting flashing traffic lights can be used to detect the flashing lights of rescue vehicles in real time using one or more high-speed cameras. In an alternative embodiment, for all objects, a neural network is used in a single shot to simultaneously learn the position, size, and classification of the object. For those objects of interest (such as rescue vehicles) that may have flashing lights, the neural network learns additional conditional attributes based on whether a flashing light is detected. Since the neural network will only learn to associate with flashing lights that are already part of the object, doing this in a single shot allows for fast inference time while learning conditional attributes with less confusion with other flashing lights.

[0123] Using the method described by Yabuuchi K et al., the camera image is converted to grayscale, and a bandpass filter is applied to the grayscale camera image in the frequency domain to enhance the flash region in the camera image. A binarization module (e.g., using a Kalman filter) estimates the state of the flash dynamics, which includes flash amplitude, offset, and phase. The state estimation is used to determine a threshold for binarizing the filtered image. The estimated threshold is then used to convert the filtered image into a binary image. A buffer module forwards this binary image to a detection module, which extracts contours from the peak binary image and then uses the contours to exclude candidate pixels to prevent false detections. A classification module (e.g., a support vector machine) then uses the contours and the original RGB camera image to classify the light color into three categories labeled red, yellow, and green. The location pixels labeled "red" can then be used as key signatures for localization as described above. The flash detector 603 can also use this technique to perform flash detection as described above.

[0124] Figure 8 According to one or more embodiments Figure 5 The block diagram shown is of the object detection and localization system 503.

[0125] The 3D sensor interface 801 receives echo signals from one or more 3D object detection sensors (e.g., LiDAR, radar, ultrasound) and converts the echo signals into digital signals. In an embodiment, the 3D sensor interface 801 is a LiDAR system including one or more sensors for receiving incident light (photons) reflected / returned from objects in the vehicle's operating environment. The one or more sensors convert the photons into electrons, then from electrons into analog voltages, and subsequently convert the analog voltages into digital signals using one or more ADCs. These digital signals are also referred to as "point clouds." An application processor or logic circuit (such as an FPGA) is coupled to the output of one or more ADCs via a serial or parallel bus interface (e.g., an SPI interface). In an embodiment, an application embedded in one or more processors or FPGAs receives the digital signals output by the one or more ADCs.

[0126] These digital signals are input to a data verifier 802, which can be implemented in software or logic. The data verifier 802 analyzes these digital signals for error and blocking signatures, including but not limited to: lost data, attenuated signals, and error protocols. Lost or leaked data can be detected by error detection / correction codes or dropped packets, etc. Attenuated signals can be detected by monitoring for drops in signal levels, etc. If an error is detected, the vehicle's diagnostic system or a cloud-based diagnostic application is notified to enable reporting, analysis, and / or repair of the error, where possible.

[0127] Object detector 803 (deep convolutional neural network) receives sensor data from data validator 802 and analyzes the sensor data for objects. In an embodiment, object detector 803 filters noise from the sensor data (such as echoes from dust, leaves, snow, rain, humidity, etc.) and combines or groups echoes with similar characteristics (such as range, rate of change of range, height, local density, etc.). The combined or grouped echoes are then classified as object detections.

[0128] In this embodiment, machine learning is used for object detection. For example, a trained deep convolutional neural network (CNN) (e.g., a VGG-16 network) can be used to analyze the echo and output labeled bounding boxes around the detected objects. The CNN can also output a confidence level for the object detection (e.g., the probability of object detection). If the confidence level is low, the detected object is considered a false detection and excluded from the input of the fusion model 504. The object locator 804 then receives object data from the object detector 803 and determines the range and azimuth of the object by determining the radial location through analysis of the sensor position and orientation, or by determining the range location through analysis of the sensor echo, via TOF and / or echo intensity analysis.

[0129] The estimated locations output from pipelines 501, 502, and 503 are input to a fusion module 504 used to determine whether the estimated locations match. If the degree of matching of the estimated locations is within a specified matching criterion, the matching location of the rescue vehicle can be used by downstream systems or applications. For example, the perception module 402, planner 404, and / or controller module 406 of AV 100 can use the estimated location of the rescue vehicle to predict the behavior of other agents in the operating environment (such as avoiding or yielding to the side of the road), and then plan the route or trajectory of the vehicle in the operating environment based on other predictions of agent behavior.

[0130] In an embodiment, the fusion module 504 receives horn, flashing light, and / or object detection as input and matches the horn, flashing light, and / or object detection by: 1) comparing similar ranges, azimuths, range change rates, intensities, etc.; 2) comparing horn types with vehicle types; and / or 3) using a trained deep convolutional neural network to compare the detections.

[0131] In this embodiment, the fusion module 504 generates a rescue vehicle presence signal and sends it to one or more downstream modules (such as the planning module 402, the sensing module 404, or the controller module 406). For example, the presence signal may indicate "present", "present and near", "present and far", "multiple rescue vehicles detected", range, rate of change of range, azimuth, etc.

[0132] Example processing

[0133] Figure 9 This is a flowchart of a rescue vehicle detection process 900 according to one or more embodiments. For example, references may be used. Figure 3 The described computer system 300 is used to implement processing 900.

[0134] Processing 900 can begin by receiving input data (901) from various sensors of the vehicle. The input data may include, but is not limited to: ambient sound captured by one or more microphones, camera images captured by one or more cameras, and echoes from light waves, RF waves, or sound waves emitted into the environment.

[0135] Process 900 continues by verifying the input data (902) against error and obstruction signatures. For example, ambient sounds, flashing lights, and light echoes are converted into digital signals, which are analyzed against error and obstruction signatures. If an error or obstruction signature is detected as a result of this analysis, the cause of the error or obstruction signature is determined using a local vehicle diagnostic system and / or a cloud-based diagnostic system. However, if no error or obstruction signature is detected, the digital signals are analyzed against horn sounds, flashing lights, and objects.

[0136] Process 900 continues by detecting the presence of a rescue vehicle based on the input data (903). For example, as referenced... Figures 6 to 8 More comprehensively, the analysis includes analyzing ambient sound captured by one or more microphones to determine the presence of a siren, analyzing camera images captured by one or more cameras to detect the presence of a flash, and analyzing 3D data from one or more 3D sensors to determine the presence of static or dynamic objects in the environment. If a rescue vehicle is detected based on the detection of the siren and flash (904), the location of the rescue vehicle is estimated based on the location of the siren and flash (905). For example, the fusion module 504 receives siren, flash, and / or object detection as input and matches the siren, flash, and / or object detection by: 1) comparing similar ranges, azimuths, range change rates, and intensities of locations; 2) comparing the siren type with the rescue vehicle type using frequency signatures; and / or 3) using one or more deep CNNs trained on the siren, flash, and / or object detections as training data to compare the detections. Training data can be augmented to cover a variety of environments (e.g., dense urban areas, rural areas) and environmental conditions (e.g., nighttime, daytime, rain, fog, snow), thereby increasing the accuracy of the predictions output by the CNN. Detection of rescue vehicles provides a higher confidence level compared to relying solely on horn detection, flashing light detection, or object detection.

[0137] In an embodiment, object detection pipeline 503 predicts 2D or 3D bounding boxes for objects labeled with rescue vehicle type (e.g., police car, fire truck, ambulance) based on the physical appearance of the rescue vehicle (e.g., shape, size, color, and / or outline, etc.), and augments the labels with data based on the output of pipeline 502 indicating whether one or more lights of the rescue vehicle are flashing, and / or augments the labels with data based on the output of pipeline 501 indicating whether a siren has been activated. For example, the planning module 404 of AV 100 can use 2D or 3D bounding boxes to plan routes or trajectories in the operational environment. For example, if the presence of a rescue vehicle is detected, the planning module 404 can plan a maneuver to the side of the road based on the estimated location of the rescue vehicle in the operational environment and the physical and emotional states of other agents (e.g., other vehicles), which may respond in a predictable manner to the siren or flashing lights of the rescue vehicle in accordance with national regulations (e.g., lane change, stopping, leaving the road).

[0138] In the preceding description, embodiments of the invention have been described with reference to numerous specific details, which may vary from implementation to implementation. Therefore, the specification and drawings should be considered illustrative rather than restrictive. The sole and exclusive indication of the scope of the invention, and what the applicant expects to be the scope of the invention, is the literal and equivalent scope of the claims published from this application in the specific form of the claims, including any subsequent amendments. Any definitions of terms expressly set forth herein for inclusion in such claims should be taken as meaning as such terms are used in the claims. Furthermore, when the term “comprising” is used in the preceding specification or appended claims, what follows that phrase may be an additional step or entity, or a sub-step / sub-entity of a previously stated step or entity.

Claims

1. A method for a vehicle, comprising: Utilize at least one processor to receive ambient sound; The at least one processor is used to determine whether the ambient sound includes a horn sound; Based on the determination that the ambient sound includes a horn sound, the at least one processor is used to determine a first location associated with the horn sound; The at least one processor is used to receive camera images; The at least one processor is used to determine whether the camera image includes a flash; Based on the determination that the camera image includes a flash, the at least one processor is used to determine a second location associated with the flash; The at least one processor is used to receive three-dimensional data, i.e., 3D data. The at least one processor is used to determine whether the 3D data includes an object; Based on the determination that the 3D data includes an object, the at least one processor is used to determine a third location associated with the object; Using the at least one processor, the presence of a rescue vehicle is determined based on the siren sound, the detected flashing lights, and the detected objects; Using the at least one processor, the estimated location of the rescue vehicle is determined based on the first location, the second location, and the third location; as well as Using the at least one processor, actions related to the rescue vehicle are initiated based on the determined presence and location of the rescue vehicle.

2. The method according to claim 1, wherein, Determining whether the ambient sound includes horn sounds also includes: Convert the ambient sound into a digital signal; The digital signal is analyzed in response to errors and blocking signatures; Upon detecting an erroneous or blocked signature, the diagnostic system is notified to determine the cause of the erroneous or blocked signature; and The digital signal is analyzed based on the horn sound if no errors or blocked signatures are detected.

3. The method according to claim 1, wherein, Determining whether the ambient sound includes horn sounds also includes: Convert the ambient sound into a digital signal; Noise from the digital signal is filtered; Compare the digital signal with a reference signal; and The horn sound detection is determined based on the comparison results.

4. The method according to claim 1, wherein, Determining whether the ambient sound includes horn sounds also includes: Convert the ambient sound into a digital signal; Machine learning is used to analyze the digital signal; and The detection of horn sounds is determined based on the results of the analysis.

5. The method according to claim 4, further comprising: Determine the confidence level for the horn sound detection.

6. The method according to claim 1, wherein, The first location associated with the source of the horn sound is determined based on acoustic localization, wherein the acoustic localization is based on time-of-flight analysis, or TOF analysis, of signals output from multiple microphones located around the vehicle.

7. The method according to claim 1, wherein, The first location associated with the source of the horn sound is determined based on the location and orientation of multiple unidirectional microphones, and the first location is determined based on the strength of the key signature signal obtained from the analysis of the output signals of the multiple unidirectional microphones and the location and orientation of each of the multiple unidirectional microphones.

8. The method according to claim 1, wherein, Determining whether the camera image includes a flash also includes: The camera images are analyzed based on erroneous and blocking signatures. Upon detecting an erroneous or blocked signature, the diagnostic system is notified to determine the cause of the erroneous or blocked signature; and In the absence of detected errors or blocked signatures, the camera image is analyzed to determine whether the camera image includes a flash.

9. The method according to claim 1, wherein, Determining whether the camera image includes a flash also includes: The camera image is filtered to remove low-intensity light signals, thereby obtaining a high-intensity image; The high-intensity image is compared with a reference image; and The flash detection is determined based on the comparison results.

10. The method according to claim 1, wherein, Determining whether the camera image includes a flash also includes: Machine learning is used to analyze the camera images; and Predicting flash detection based on the analysis results.

11. The method of claim 10, further comprising: Determine the confidence level of the flash detection.

12. The method according to claim 11, wherein, The analysis using machine learning includes using a neural network to analyze the camera images, the neural network being trained to distinguish between bright flashes and other bright lights.

13. The method according to claim 1, wherein, The second location associated with the flashlight is determined by the following steps: Detect flash signatures in multiple camera images captured by multiple cameras; as well as The radial location of the flash is determined based on the position and orientation of each of the plurality of cameras; or The second location was determined based on triangulation of the positions and orientations of the multiple cameras.

14. The method according to claim 13, wherein, Determining whether the 3D data includes an object also includes: The 3D data is analyzed in response to errors and blocking signatures. Upon detecting an erroneous or blocked signature, the diagnostic system is notified to determine the cause of the erroneous or blocked signature; and In the absence of detected errors or blocked signatures, the 3D data is analyzed to determine whether the 3D data includes an object.

15. The method according to claim 13, wherein, Determining whether the 3D data includes an object also includes: The 3D data is filtered to remove noise; The 3D data is divided into groups with similar characteristics; and Machine learning is used to analyze the group to determine object detection.

16. The method of claim 15, further comprising: Determine the confidence level for the object detection.

17. The method according to claim 13, wherein, Determining the third location includes: The radial location of the object is determined based on the analysis of the sensor position and orientation of the sensor that detected the object.

18. The method according to claim 1, wherein, Determining the presence and location of the rescue vehicle also includes: The horn sound detection, flashing light detection, and object detection are compared to determine whether they match; and The presence of the rescue vehicle is generated based on the match determined.

19. The method of claim 18, further comprising: The rescue vehicle is tracked based on its location.

20. A system for a vehicle, comprising: At least one processor; A memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations, including: Receive ambient sound; Determine whether the ambient sound includes horn sounds; Based on the determination that the ambient sound includes a horn sound, a first location associated with the horn sound is determined; Receives images from the camera; Determine whether the camera image includes a flash; Based on the determination that the camera image includes a flash, a second location associated with the flash is determined; Receive three-dimensional data, i.e., 3D data; Determine whether the 3D data includes an object; Based on the determination that the 3D data includes objects, a third location associated with the objects is determined; The presence of the rescue vehicle is determined based on the siren sound, the detected flashing lights, and the detected objects; The estimated location of the rescue vehicle is determined based on the first location, the second location, and the third location; and Actions related to the rescue vehicle are initiated based on the presence and location of the identified rescue vehicle.

21. A computer-readable storage medium having instructions stored thereon, the instructions causing the at least one processor to perform an operation when executed by at least one processor, the operation comprising: Receive ambient sound; Determine whether the ambient sound includes horn sounds; Based on the determination that the ambient sound includes a horn sound, a first location associated with the horn sound is determined; Receives images from the camera; Determine whether the camera image includes a flash; Based on the determination that the camera image includes a flash, a second location associated with the flash is determined; Receive three-dimensional data, i.e., 3D data; Determine whether the 3D data includes an object; Based on the determination that the 3D data includes objects, a third location associated with the objects is determined; The presence of the rescue vehicle is determined based on the siren sound, the detected flashing lights, and the detected objects; The estimated location of the rescue vehicle is determined based on the first location, the second location, and the third location; as well as Actions related to the rescue vehicle are initiated based on the presence and location of the identified rescue vehicle.

Citation Information

Patent Citations

  • Vehicle brake control system and vehicle emergency brake avoiding method

    CN102582599A

  • Driverless car management and control system and method

    CN109116850A