COLLISION AVOIDANCE USING ACOUSTIC DATA

Acoustic sensors enhance obstacle detection in autonomous vehicles by processing audio data to identify hidden vehicles, addressing the limitations of visual sensors and improving collision avoidance.

DE102017103374B4Active Publication Date: 2025-07-10FORD GLOBAL TECH LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE102017103374
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2016-02-26
Filing Date
2017-02-20
Publication Date
2025-07-10
Estimated Expiration
2037-02-20

AI Technical Summary

Technical Problem

Autonomous vehicles struggle to detect obstacles that are not within the field of view of image sensors, particularly parked vehicles with running engines that may pose a hazard due to their potential movement.

Method used

Utilizing acoustic sensors to detect and classify potential obstacles by processing audio data through preprocessing, noise suppression, and machine learning models to enhance the detection of vehicles not visible to cameras, combined with image data for validation and obstacle avoidance strategies.

Benefits of technology

Enhances obstacle detection capabilities by identifying hidden vehicles with running engines, improving collision avoidance systems in autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for obstacle detection in an autonomous vehicle, the method comprising: Receiving, by a controller including one or more processing devices, two or more audio streams from two or more microphones mounted on the autonomous vehicle; Detecting audio features in the two or more audio streams by the controller; Detecting a direction to a noise source according to the audio characteristics by the controller; Identifying a class for the noise source according to the audio characteristics by the controller; and Determine, by the controller, that the class for the noise source is a parked vehicle, in response to determining that the class for the noise source is a vehicle, invoking obstacle avoidance with respect to a potential path of the parked vehicle, Receiving, by the controller, at least one image stream from at least one camera mounted on the autonomous vehicle; Determining, by the controller, that an image of a vehicle in the image stream of the at least one camera is within an angular tolerance from the direction of the noise source; in response to determining that the image of the vehicle in the image stream of the at least one camera is within the angular tolerance from the direction to the noise source, increasing a confidence value to indicate that the noise source is the parked vehicle, Retrieving map data for a current location of the autonomous vehicle by the controller; Determining, by the controller, that the map data indicates at least one parking area within the angular tolerance from the direction to the source of the audio features; and in response to determining that the map data indicates that at least one parking area is within the angular tolerance from the direction to the source of the audio features, increasing the confidence value to indicate that the source of the audio features is the parked vehicle, Classifying the audio features by inputting the audio features into a machine learning model by the controller, wherein the machine learning model is programmed to output a confidence score; Increase the trust value according to the trust rating by the controller.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUNDTECHNICAL FIELD OF THE INVENTIONThis invention relates to performing obstacle avoidance in autonomous vehicles.BACKGROUND OF THE INVENTIONAutonomous vehicles are equipped with sensors that detect their environment. An algorithm evaluates the output of the sensors and identifies obstacles. A navigation system may then steer, brake, and / or accelerate the vehicle to both avoid identified obstacles and reach a desired destination. Sensors may include image acquisition systems, e.g., video cameras, as well as RADAR or LIDAR sensors.DE 101 36 981 A1 discloses a method and a device for determining a stationary and / or moving object, in particular a vehicle, in which acoustic signals emitted by the object and / or reflected on another object are detected, on the basis of which the object in question is detected, evaluated and / or identified. Such a method enables, in the manner of self-localization on the basis of sound waves, both a stationary and a moving object, e.g. a vehicle, to be detected, evaluated and identified acoustically on the basis of self-noises and / or extraneous noises with respect to the own movement profile with respect to one or more coordinate axes (x-, y-axis).US 2010 / 0 228 482 A1 discloses a device and method for supporting collision avoidance between a vehicle and an object. The apparatus and method includes receiving an acoustic signal from the object using microphones worn by the vehicle and determining a position of the object using the acoustic signal and an acoustic model of the environment, the environment including a structure that blocks direct view of the object from the vehicle.The systems and methods disclosed herein provide an improved approach to detecting obstacles.BRIEF DESCRIPTION OF THE DRAWINGSIn order that the advantages of the invention may be readily understood, a more detailed description of the invention briefly described above will be given by reference to specific embodiments illustrated in the appended drawings. In understanding that these drawings depict only typical embodiments of the invention and are therefore not to be considered limiting of its scope, the invention will be described and explained with additional specificity and detail by use of the accompanying drawings, in which: FIG. 1 is a schematic block diagram of a system for implementing embodiments of the invention; FIG. 2 is a schematic block diagram of an example computing device suitable for implementing methods according to embodiments of the invention; FIG. 3 is a diagram illustrating obstacle detection using acoustic data; FIG. 4 is a schematic block diagram of components for performing obstacle detection using acoustic data; and FIG. 5 is a process flow diagram of a method for performing collision avoidance based on acoustic data, according to an embodiment of the present invention.DETAILED DESCRIPTIONIt will be readily understood that the components of the present invention as generally described herein and illustrated in the figures could be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the invention as represented in the figures is not intended to limit the scope of the invention as claimed, but is merely representative of certain examples of presently contemplated embodiments according to the invention. The presently described embodiments are best understood by reference to the drawings, wherein like parts are designated by like reference numerals throughout.Embodiments according to the present invention may be embodied as an apparatus, method or computer program product. Accordingly, the present invention may take the form of an embodiment entirely as hardware, an embodiment entirely as software (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, which may be referred to generally herein as a "module" or "system.". Furthermore, the present invention may take the form of a computer program product embodied as any tangible medium of expression, wherein computer usable program code is embodied in the medium.Any combination of one or more computer usable or computer readable media may be utilized. For example, a computer readable medium may include one or more of the following: a portable computer diskette, a hard disk, a random access memory (RAM) device, a read-only memory (ROM) device, an erasable programmable read-only memory (EPROM or Flash memory), a portable compact disc read-only memory (CDROM) device, an optical storage device, and a magnetic storage device. In selected embodiments, a computer-readable medium may comprise any non-transitory medium that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.Computer program code for carrying out operations of the present invention may be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code may execute entirely on a computer system as a stand-alone software package, on a stand-alone hardware unit, partly on a remote computer spaced some distance from the computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet Service Provider).The present invention will be described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, may be implemented by computer program instructions or code. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus for generating a machine such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram.These computer program instructions may also be stored in a non-transitory computer readable medium that can direct a computer or other programmable data processing device to function in a particular manner such that the instructions stored in the computer readable medium produce an article of manufacture including instruction means that implement the function / act specified in one or more blocks of the flowchart and / or block diagram.The computer program instructions may also be loaded onto a computer or other programmable data processing device to cause a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process such that the instructions executed on the computer or other programmable device provide processes for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram.Referring to FIG. 1, a controller 102 may be housed in a vehicle. The vehicle may include any vehicle known in the art. The vehicle may have all structures and features of any vehicle known in the art, including wheels, a powertrain coupled to the wheels, an engine coupled to the powertrain, a steering system, a braking system, and other systems known in the art to be included in a vehicle.As explained in more detail herein, the controller 102 may perform autonomous navigation and collision avoidance. In particular, image data and audio data may be analyzed to identify obstacles. In particular, audio data may be used to identify vehicles that are not within the field of view of one or more cameras or other imaging sensors, as described in detail below with reference to FIGS. 3 and 4.The controller 102 may receive one or more image streams from one or more image capture devices 104. For example, one or more cameras may be mounted on the vehicle and output image streams received from the controller 102. The controller 102 may receive one or more audio streams from one or more microphones 106. For example, one or more microphones or microphone arrays may be mounted on the vehicle and output audio streams received from the controller 102. The microphones 106 may include directional microphones having a sensitivity varying with the angle.The controller 102 may execute a collision avoidance module 108 that receives the image streams and audio streams, identifies possible obstacles, and takes action to avoid them. In the embodiments disclosed herein, only image and audio data are used to perform collision avoidance. However, other sensors may also be used to detect obstacles, such as radio detection and ranging (RADAR), light detection and ranging (LIDAR), sound navigation and ranging (SONAR), and the like. Accordingly, "image streams" received from the controller 102 may include optical images detected by a camera and / or objects and topology detected using one or more sensor devices. The controller 102 may then analyze both images and detected objects and topology to identify potential obstacles.The collision avoidance module 108 may include an audio detection module 110 a. The audio detection module 110 amay include an audio preprocessing module 112 aprogrammed to process the one or more audio streams to identify features that could correspond to a vehicle. The audio detection module 110 amay further include a machine learning module 112 bthat implements a model that evaluates features in processed audio streams from the preprocessing module 112 aand attempts to classify the audio features. The machine learning module 112 bmay output a confidence score indicating a likelihood that a classification is correct. The function of the modules 112 a, 112 bof the audio detection module 110 ais described in more detail below with reference to the method 500 of FIG. 5.The audio detection module 110 amay further include an image correlation module 112 c programmed to evaluate image outputs from the one or more image capture devices 104 and attempt to identify a vehicle in the image data within an angular tolerance from an estimated direction to the source of sound corresponding to a vehicle, such as a parked vehicle that is running but not moving. When a vehicle is displayed within the angular tolerance, confidence that the noise corresponds to a vehicle is increased.The audio detection module 110 amay further include a map correlation module 112 d. The map correlation module 112 dvaluates map data to determine whether a parking space, entry, or other parking area is within the angular tolerance from the direction to the source of noise corresponding to a running engine vehicle, particularly a parked vehicle. If so, confidence that the noise corresponds to a parked vehicle with the engine running is increased.The collision avoidance module 108 may further include an obstacle identification module 110 b, a collision prediction module 110 c, and a decision module 110 d. The obstacle identification module 110 banalysis the one or more image streams and identifies potential obstacles including people, animals, vehicles, buildings, curbstones and other objects and structures. Specifically, the obstacle identification module 110 bmay identify vehicle images in the image stream.The collision prediction module 110 cpredicts which obstacle images are likely to collide with the vehicle based on its current trajectory or path intended at present. The collision prediction module 110 cmay evaluate the likelihood of collision with objects identified by the obstacle identification module 110 bas well as objects detected by the audio detection module 110 a. In particular, running engine vehicles identified with threshold confidence from audio detection module 110 amay be added to a set of potential obstacles, particularly the potential movements of such vehicles. The decision module 110 dmay make a decision to stop, accelerate, turn, etc. to avoid obstacles. The manner in which the collision prediction module 110 cpredicts potential collisions and the manner in which the decision module 110 dtakes action to avoid potential collisions may be according to any method or system known in the autonomous vehicle art.The decision module 110 dmay control the trajectory of the vehicle by actuating one or more actuators 114 that control the direction and speed of the vehicle. For example, the actuators 114 may include a steering actuator 116 a, an accelerator actuator 116 b, and a brake actuator 116 c. The configuration of the actuators 116 a- 116 cmay correspond to any implementation of such actuators known in the art of autonomous vehicles.The controller 102 may be network enabled and retrieve information via a network 118. For example, map data 120 may be accessed from a server system 122 to identify potential parking spaces proximate to the autonomous vehicle in which the controller 102 is housed.FIG. 2 is a block diagram illustrating an example computing device 200. The computing device 200 may be used to perform various procedures such as those discussed herein. The controller 102 may have some or all of the attributes of the computing device 200.Computing device 200 includes one or more processors 202, one or more memory devices 204, one or more interfaces 206, one or more mass storage devices 208, one or more input / output (I / O) devices 210, and a display device 230, all of which are coupled to a bus 212. The processor or processors 202 include one or more processors or controllers that execute instructions stored in one or more memory devices 204 and / or mass storage devices 208. The processor or processors 202 may also include various types of computer readable media, such as cache memories.The one or more storage devices 204 include various computer readable media such as volatile memory (e.g., random access memory (RAM) 214) and / or nonvolatile memory (e.g., read only memory (ROM) 216). The one or more storage devices 204 may also include rewritable ROM, such as flash memory.The one or more mass storage devices 208 include various computer readable media such as magnetic tapes, magnetic disks, optical disks, semiconductor memory (e.g., flash memory), and so forth. As shown in Figure 2, a particular mass storage device is a hard disk drive 224. Various drives may also be included in mass storage device(s) 208 to enable reading from and / or writing to the various computer readable media. The one or more mass storage devices 208 include removable media 226 and / or non-removable media.The one or more I / O devices 210 include various devices that allow data and / or other information to be input to or retrieved from the computing device 200. The one or more example I / O devices 210 include cursor control devices, keyboards, keypads, microphones, monitors or other display devices, speakers, printers, network interface cards, modems, lenses, CCDs or other image capture devices, and the like.The display device 230 includes any type of device capable of displaying information to one or more users of the computing device 200. Examples of the display device 230 include a monitor, a display terminal, a video projection device, and the like.The one or more interfaces 206 include various interfaces that allow the computing device 200 to interact with other systems, devices, or computing environments. The one or more example interfaces 206 include any number of different network interfaces 220, such as interfaces to local area networks (LANs), wide area networks (WANs), wireless networks, and the Internet. One or more other interfaces include a user interface 218 and a peripheral device interface 222. The one or more interfaces 206 may also include one or more peripheral interfaces, such as interfaces for printers, pointing devices (mice, track pads, etc.), keyboards, and the like.The bus 212 allows the one or more processors 202, the one or more memory devices 204, the one or more interfaces 206, the one or more mass storage devices 208, the one or more I / O devices 210, and the display device 230 to both communicate with each other, and with other devices or components coupled to the bus 212. Bus 212 represents one or more of several types of bus structures, such as a system bus, PCI bus, IEEE 1394 bus, USB bus, and so forth.For purposes of illustration, programs and other executable program components are shown herein as discrete blocks, although it should be understood that such programs and components may reside in different memory components of computing device 200 at different times and are executed by the one or more processors 202. Alternatively, the systems and procedures described herein may be implemented in hardware or in a combination of hardware, software, and / or firmware. For example, one or more application specific integrated circuits (ASICs) may be programmed to execute the systems and / or procedures described herein.Referring now to FIG. 3, in many cases, a vehicle housing the controller 102 (hereinafter, the vehicle 300) may be prevented from optically detecting a potential obstacle 302 such as another vehicle, a cyclist, a pedestrian, or the like. For example, the obstacle 302 may be hidden by a concealing object, such as a parked vehicle, a building, a tree, a sign, etc., in a line of sight of the driver or image sensor 104. Accordingly, image capture devices 104 may not be effective in detecting such obstacles. Moreover, a parked vehicle is not moving and therefore may not be detected as a potential hazard by imaging sensors 104. However, when the engine of a parked vehicle is running, it is actually possible for it to move into the path of the vehicle 300 shortly.The vehicle 300 may be close enough to detect sounds generated by a hidden vehicle 304 or a parked vehicle 304. Although the methods disclosed herein are particularly useful when an occluding object is present, the identification of obstacles described herein may be performed when image data is available and, for example, confirm the location of an obstacle that is also visible to image capture devices 104. Similarly, the existence of a parked vehicle with the engine running may be confirmed using image capture devices 104, as described in more detail below.Referring now to FIG. 4, the microphone 106 may include multiple microphones 106 a- 106 d. The output signal of each microphone 106 a- 106 dmay be input to a corresponding pre-processing module 112 a- 1- 112 a- 4. The output of each preprocessing module 112 a- 1- 112 a- 4 may be further processed by a noise suppression filter 400 a- 400 d. The output of the noise suppression module 400 a- 400 dmay then be input to the machine learning module 112 b. In particular, the outputs of the noise suppression modules 400 a- 400 dmay be input to a machine learning model 402 that classifies features in the outputs as corresponding to a particular vehicle. The machine learning model 112 bmay further output confidence in the classification.The preprocessing modules 112 a- 1- 112 a- 4 may process the raw outputs from the microphones 106 a- 106 dand produce processed outputs that are input to the noise suppression modules 400 a- 400 dor directly to the machine learning module 112 b. The processed outputs may be a filtered version of the raw outputs, where the processed outputs have improved audio characteristics relative to the raw outputs. The enhanced audio features may be segments, frequency bands, or other components of the raw outputs that are likely to correspond to a vehicle. Accordingly, the preprocessing modules 112 a- 1- 112 a- 4 may include a band pass filter that passes a portion of the raw outputs in a frequency band corresponding to sounds generated by vehicles and vehicle engines while blocking portions of the raw outputs outside of this frequency band. The preprocessing modules 112 a- 1- 112 a- 4 may be digital filters in which coefficients are chosen to pass signals having spectral content and / or a temporal profile corresponding to a vehicle engine or other vehicle noise, such as an adaptive filter with experimentally selected coefficients that passes noise generated by a vehicle while attenuating other noise. The output of the preprocessing modules 112 a- 1- 112 a- 4 may be a time domain signal or a frequency domain signal or both. The output of the preprocessing modules 112 a- 1- 112 a- 4 may include multiple signals, including time domain and / or frequency domain signals. For example, signals that are the result of filtering using different band pass filters may be output in either the frequency or time domain.The noise cancellation modules 400 a- 400 dmay include any noise cancellation filter known in the art or implement any noise cancellation approach known in the art. In particular, the noise suppression modules 400 a- 400 dmay further use, as inputs, the speed of the vehicle 300, a speed of an engine of the vehicle 300, or other information describing a status of the engine, a speed of a ventilation fan of the vehicle 300, or other information. This information may be used by the noise cancellation modules 400 a- 400 dto remove noise caused by the engine and fan as well as vehicle wind noise.The machine learning model 402 may be a deep neural network, but other types of machine learning models may also be used, such as a decision tree, clustering, Bayesian network, genetic, or other type of machine learning model. The machine learning model 402 may be trained with different types of sounds in different types of situations. In particular, sounds recorded using the array of microphones 106 a- 106 d(or an array with similar specifications) may be recorded from a known source at various relative locations, at various relative speeds, as well as with and without background sounds.The machine learning model 402 may then be trained to detect the sounds from the known source. For example, the model may be trained using < audio input, sound source class> entries that respectively pair audio records using the microphones 106 a- 106 din the various situations mentioned above and the class of the sound source. The machine learning algorithm may then use these entries to train with a machine learning model 402 to output the class of sound source for a given audio input. The machine learning algorithm may train the machine learning model 402 for different classes of noise sources. Accordingly, a set of training entries may be generated for each class of noise sources and the model trained therewith, or separate models trained for each class of noise sources. The machine learning model 402 may output both a decision and a confidence score for that decision. Accordingly, the machine learning model 402 may produce an output indicating whether or not input signals correspond to a particular class, as well as a confidence score that that output is correct.The machine learning module 112 bmay further include a microphone array processing module 404. The microphone array processing module 404 may evaluate the time of arrival of an audio feature from various microphones 106 a- 106 dto estimate a direction to a source of the audio feature. For example, an audio feature may be the sound of a vehicle that begins in the outputs of the noise cancellation modules 400 a- 400 dat time T 1, T 2, T 3, and T 4. Accordingly, knowing the relative positions of the microphones 106 a- 106 dand the speed of sound S, the difference in distance from the microphones 106 a- 106 dto the source may be determined, e.g., D 2=S / (T 2-T 1), D 3=S / (T 3-T 1), D 4=S / (T 4-T 1), where D 2, D 3, D 4 is the estimated difference in distance travelled by the audio feature relative to a reference microphone, in this example microphone 106 a.For example, the angle A to the source of noise may be calculated as an average of Asin(D2 / R2), Asin(D3 / R3), and Asin(D4 / R4), where R2is the distance between microphone 106 aand microphone 106 b, R3is the distance between microphone 106 cand microphone 106 a, and R4is the distance between microphone 106 dand microphone 106 a. In this approach, it is assumed that the source of the sound is located at a great distance from the microphones 106 a- 106 d, so that the incident sound wave can be approximated as a planar wave. Other approaches to identifying the direction to a sound based on different arrival times, as known in the art, may also be used. Also, instead of simply determining a direction, a sector or range of angles may be estimated, i.e. a range of uncertainty with respect to each estimated direction, the range of uncertainty being a limitation on the accuracy of the direction estimation technique used.The direction estimated by the microphone array processing module 404 and the classification and confidence score generated by the machine learning model 402 may then be provided as output 406 from the machine learning model 112 b. For example, the obstacle identification module 110 bmay add a vehicle having the identified class and located in the estimated direction to a set of potential obstacles, where the set of potential obstacles includes all obstacles identified by other means such as using image capture devices 104. The collision prediction module 110 cmay then identify potential collisions with the set of potential obstacles and the decision module 110 dmay then determine actions to perform to avoid the potential collisions, such as turning the vehicle, applying the brakes, accelerating, or the like.FIG. 5 illustrates a method 500 that may be performed by the controller 102 by processing audio signals from the microphones 106 a- 106 d. The method 500 may include generating 502 audio signals representing detected sounds using the microphones 106 a- 106 d, as well as pre-processing 504 the audio signals to improve the audio characteristics. This may include performing each of the filter functions described above with respect to the preprocessing modules 112 a- 1- 112 a- 4. In particular, the pre-processing 504 may include generating one or more pre-processed signals in the time or frequency domain, wherein each output may be a band-pass filtered version of an audio signal from one of the microphones 106 a- 106 d, or may be filtered or otherwise processed using other techniques, such as using an adaptive filter or other audio processing techniques. The pre-processing 504 may further include performing the noise cancellation on either the input or the output of the pre-processing modules 112 a- 1- 112 a- 4, as described above with respect to the noise cancellation modules 400 a- 400 d.The method 500 may further include inputting 506 the pre-processed signals to the machine learning model 402. The machine learning model 402 then classifies 508 the origin of the sounds, i.e., the attributes of the audio features in the preprocessed signals are processed according to the machine learning model 402, which then outputs one or more classifications and confidence scores for the one or more classifications.The method 500 may further include estimating 510 a direction to the origin of the sounds. As described above, this may include invoking the functionality of the microphone array processing module 404 to evaluate the differences in the time of arrival of the audio features in the preprocessed outputs to determine a direction to author the audio features or a range of possible angles to author the audio features.The method 500 may further include attempting to validate the classification performed in step 508 using one or more other information sources. For example, method 500 may include attempting to correlate the direction of the sound origin with a vehicle image located in an output of imaging sensor 104 in a position corresponding to the direction of the sound origin. For example, each vehicle image that is in an angular region that includes the direction may be identified in step 512. If a vehicle image is found in the image stream of the imaging sensor 104 within the angular region, a confidence value may be increased.The method 500 may include attempting 514 to correlate the direction to the sound origin with map data. For example, if the map data is determined to indicate that a parking space or other legal parking area is in proximity (e.g., within a threshold radius) to the autonomous vehicle and within an angular region that includes the direction toward the sound origin, the confidence value may be increased, otherwise it is not increased based on the map data. The angular region used to determine whether a parking area is within tolerance of the direction to the sound origin may be the same as or different from that used in step 512.The confidence score may be further increased based on the confidence score from the classification step 508. In particular, the confidence score may be increased relative to or as a function of the magnitude of the confidence score. As described herein, parked cars with the engine running may be detected. Accordingly, if the classification step 508 indicates detection of the noise of a parked vehicle with the engine running, the confidence score is increased based on the confidence score of step 508 as well as based on steps 512 and 514.The method 500 may include evaluating 516 whether the confidence score of step 508 exceeds a threshold. For example, if no classifications at step 508 have a confidence score above a threshold, method 500 may include determining that the audio features that were the basis of the classification are likely not to correspond to a vehicle. Otherwise, if the confidence score exceeds a threshold, the method 500 may include adding 518 a potential obstacle to a set of obstacles identified by other means, such as using image capture devices 104. The potential obstacle may be defined as a potential obstacle that is in the direction or range of the angles determined in step 510.In a classification indicating a parked vehicle with the engine running, if the confidence value based on all steps 508, 512, and 514 exceeds a threshold corresponding to parked vehicles, step 516 may include determining that a parked vehicle is a potential obstacle and may move away from its current location. In particular, a potential path or range of potential paths of the parked vehicle may be added 518 to a set of potential obstacles. For example, since the estimated direction to the parked vehicle is known, the potential movement to a side of the estimated direction by the parked vehicle may be considered a potential obstacle.At each result of step 516, obstacles are detected using other detection systems, such as the image capture devices 104, and obstacles detected using these detection systems are added 520 to the obstacle set. The collision avoidance is performed 522 with respect to the obstacle set. As noted above, this may include detecting potential collisions and activating a steering actuator 116 aand / or an accelerator actuator 116 band / or a brake actuator 116 cto avoid the obstacles of the obstacle set and guide the vehicle to a designated destination.The present invention may be embodied in other specific forms without departing from the spirit or essential characteristics thereof. The described embodiments are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalence of the claims are to be included within their scope.

Claims

A method for obstacle detection in an autonomous vehicle, the method comprising: receiving, by a controller including one or more processing devices, two or more audio streams from two or more microphones mounted on the autonomous vehicle; detecting, by the controller, audio features in the two or more audio streams; detecting, by the controller, a direction toward a noise source according to the audio features; identifying, by the controller, a class for the noise source according to the audio features; determining, by the controller, that the class for the noise source is a parked vehicle, responsive to determining that the class for the noise source is a vehicle, invoking obstacle avoidance with respect to a potential path of the parked vehicle, receiving, by the controller, at least one image stream from at least one camera mounted on the autonomous vehicle; determining, by the controller, that an image of a vehicle in the image stream of the at least one camera is within an angular tolerance from the direction toward the noise source; in response to determining that the image of the vehicle in the image stream of the at least one camera is within the angular tolerance from the direction to the sound source, increasing a confidence value to indicate that the sound source is the parked vehicle, retrieving, by the controller, map data for a current location of the autonomous vehicle; determining, by the controller, that the map data indicates at least one parking area within the angular tolerance from the direction to the source of the audio features; In response to determining that the map data indicates that at least one parking area is within the angular tolerance from the direction to the source of the audio features, increasing the confidence value to indicate that the source of the audio features is the parked vehicle, classifying the audio features by inputting the audio features into a machine learning model by the controller, wherein the machine learning model is programmed to output a confidence score; increasing the confidence value according to the confidence score by the controller.The method of claim 1, further comprising invoking obstacle avoidance with respect to the direction to the noise source by actuating a steering actuator and / or accelerator actuator and / or brake actuator of the autonomous vehicle that cause a collision with the noise source to be avoided.The method of claim 1, wherein the machine learning model is a deep neural network.The method of claim 1, further comprising: receiving, by the controller, image outputs from one or more sensors mounted on the autonomous vehicle; identifying, by the controller, a set of potential obstacles in the image outputs; evaluating, by the controller, possible collisions between the autonomous vehicle and the set of potential obstacles and the source of audio features; and activating, by the controller, a steering actuator and / or an accelerator actuator and / or a brake actuator of the autonomous vehicle that cause collisions with the set of potential obstacles to be avoided.The method of claim 4, wherein the one or more sensors comprise cameras and / or LIDAR sensors and / or RADAR sensors.The method of claim 1, wherein identifying the audio features comprises filtering, by the controller, the two or more audio streams to obtain two or more filtered signals each comprising one or more of the audio features.The method of claim 6, wherein filtering the two or more audio streams further comprises removing ambient sounds from the two or more audio streams.

Citation Information

Patent Citations

  • Method and device for determining a stationary and / or moving object

    DE10136981A1

  • Collision avoidance system and method

    US20100228482A1