Method and system for acoustic machine perception for an aircraft
By combining microphone arrays with other sensors on the aircraft to acquire data and train machine learning models, the problem of inaccurate object detection by the vehicle perception system in severe weather conditions is solved, achieving higher detection confidence and autonomous training capabilities.
Patent Information
- Application Number
- CN202010343231.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-13
- Filing Date
- 2020-04-27
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2040-04-27
AI Technical Summary
Existing vehicle perception systems have low confidence in detecting objects in adverse weather conditions and rely on expensive LIDAR, radar, and camera modalities that may not be reliable enough.
A microphone array is deployed on the aircraft to collect data together with light detection and ranging sensors, radar sensors and cameras. The machine learning model is trained to predict the azimuth, range and type of the object, reducing dependence on existing models.
It improves the confidence of object detection in adverse weather conditions, reduces dependence on expensive sensors, and enhances the autonomous training capability and recognition accuracy of the perception system.
Smart Images

Figure CN112083378B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates generally to a machine perception system for a vehicle, and more particularly, to an acoustic-based machine perception system for an aircraft. BACKGROUND
[0002] Perception systems can be used in vehicles, such as cars and aircraft, to facilitate various vehicle operations. For example, a perception system can acquire an image or series of images of the environment surrounding the vehicle and attempt to identify whether an object (e.g., a car, an aircraft, a person, or other obstacle) is present in the image. Additionally, the perception system can determine a location of such an object and / or other information associated with such an object. From there, the perception system and / or related computing systems of the vehicle can cause the vehicle to perform an action related to the detected object, such as notifying a driver or passenger of the vehicle of the object’s presence or adjusting a heading of the vehicle to follow or avoid the object. In other examples, the perception system can also detect other information about the environment surrounding the vehicle, such as road signs and weather.
[0003] Existing perception systems, including those located on-board a vehicle or remote from a vehicle, often use modalities such as light detection and ranging (LIDAR), radar, and / or object recognition in camera images. However, these modalities can be expensive to implement and can not be reliable in all environmental conditions. For example, adverse weather conditions can obscure camera images or make it difficult to reconstruct LIDAR point cloud data, thereby reducing a confidence level with which the perception system detects objects and determines object information. Furthermore, if one or more radar sensors, one or more LIDAR sensors, and / or one or more cameras stop functioning, this confidence level can be further reduced.
[0004] What is needed is a perception system that reduces reliance on existing modalities and improves a confidence level with which the perception system can detect objects and object information. SUMMARY
[0005] In one example, a method is described. The method includes causing one or more sensors disposed on an aerial vehicle to acquire, within a time window, first data associated with a first object within an environment surrounding the aerial vehicle, wherein the one or more sensors include one or more of a light detection and ranging sensor, a radar sensor, and a camera. The method also includes causing a microphone array disposed on the aerial vehicle to acquire, within approximately the same time window as the acquiring of the first data, first acoustic data associated with the first object, and training, by a processor, a machine learning model using the first acoustic data as input values to the machine learning model and using an azimuth angle of the first object, a range of the first object, and a type of the first object as ground truth output labels to the machine learning model, wherein the machine learning model is configured to predict, based on target acoustic data subsequently acquired by the microphone array and associated with a target object within the environment surrounding the aerial vehicle, an azimuth angle of the target object, a range of the target object, and a type of the target object.
[0006] In another example, a method is described. The method includes causing a microphone array disposed on an aerial vehicle to acquire target acoustic data associated with a target object within an environment surrounding the aerial vehicle, and executing, by a processor, a machine learning model to predict, based on the target acoustic data, an azimuth angle of the target object, a range of the target object, and a type of the target object, wherein the machine learning model is trained by: (i) causing one or more sensors disposed on the aerial vehicle to acquire first data associated with a first object within the environment surrounding the aerial vehicle, wherein the one or more sensors include one or more of a light detection and ranging sensor, a radar sensor, and a camera; (ii) causing the microphone array disposed on the aerial vehicle to acquire, within approximately the same time window as the acquiring of the first data, first acoustic data associated with the first object; (iii) the processor using the first acoustic data as input values to the machine learning model; and (iv) the processor using an azimuth angle of the first object, a range of the first object, and a type of the first object as ground truth output labels to the machine learning model.
[0007] In another example, a system is described that includes an aircraft, one or more sensors disposed on the aircraft, a microphone array disposed on the aircraft, and a computing device. The one or more sensors include one or more of a light detection and ranging sensor, a radar sensor, and a camera. The computing device includes a processor and a memory storing instructions executable by the processor to perform a set of operations, the set of operations comprising: causing the one or more sensors to acquire first data associated with a first object in an environment surrounding the aircraft within a time window; causing the microphone array to acquire first acoustic data associated with the first object within a time window approximately the same as the time window in which the first data was acquired; and training a machine learning model by using the first acoustic data as input to the machine learning model and using an azimuth of the first object, a range of the first object, and a type of the first object identified from the first data as ground truth output labels for the machine learning model, wherein the machine learning model is configured to predict an azimuth of a target object, a range of the target object, and a type of the target object based on target acoustic data subsequently acquired by the microphone array and associated with the target object in the environment surrounding the aircraft.
[0008] The features, functions, and advantages that have been discussed can be achieved independently in various examples or may be combined in yet other examples. Further details of the examples can be seen with reference to the following description and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The novel features which are believed to be characteristic of the illustrative examples are set forth in the appended claims. The illustrative examples (and preferred modes of use), further objects and description thereof will be best understood by reference to the following detailed description of illustrative examples of the disclosure when read in conjunction with the accompanying drawings, in which:
[0010] Figure 1 is a block diagram of a system and objects in an aircraft surrounding the system according to an example implementation.
[0011] Figure 2 is a diagram depicting the example implementation for training Figure 1 The machine learning model of the system and use the machine learning model to predict Figure 1 A diagram of the data involved in an example process of associating information with a target object.
[0012] Figure 3 is a perspective view of an aircraft according to an example implementation.
[0013] Figure 4 A method for training according to an example implementation is shown. Figure 1 A flowchart of an example method for a machine learning model of a system.
[0014] Figure 5 a method for performing the causing function of the method of Figure 4 a flowchart of an example method of the method of
[0015] Figure 6 a flowchart of an example method of the method of Figure 4 together with a flowchart of an example method of the training function of the method of Figure 4
[0016] Figure 7 a flowchart of another example method of the method of Figure 4 together with a flowchart of another example method of the training function of the method of Figure 4
[0017] Figure 8 a flowchart of another example method of the method of Figure 4 together with a flowchart of another example method of the training function of the method of Figure 4
[0018] a flowchart of another example method of the training function of the method of Figure 9 Figure 4
[0019] Figure 10 a flowchart of an example method of the system of Figure 1 together with a flowchart of an example method of the system of Figure 1
[0020] a flowchart of an example method of the method of Figure 11 Figure 10 a flowchart of an example method of the method of
[0021] Figure 12 a flowchart of an example method of the causing function of the method of Figure 10 DETAILED DESCRIPTION
[0022] The disclosed examples will now be described more fully below with reference to the accompanying drawings, in which some, but not all, disclosed examples are shown. Indeed, these examples can be described in the general context of methods and processes, which can include software program modules executed by computers. Generally, a computer can include memory storage, a processor, an input device, and an output device. Generally, a computer can also include computer readable media, which can include computer readable storage media, which can be removable media devices, volatile media devices, non-volatile media devices, and / or a combination thereof. Computer readable storage media can include volatile media devices, non-volatile media devices, removable media devices, and / or a combination thereof. Computer readable storage media can include media implemented in a physical form, such as solid-state memory, optical media, magnetic media, and / or the like. Computer readable storage media can also include media implemented in a physical form, such as solid-state memory, optical media, magnetic media, and / or the like. Computer readable storage media can also include computer readable instructions encoded on such media for execution by a processor.
[0023] The terms “substantially,” “about,” “approximately,” and “near” as used herein refer to allowing for a margin of error or variation, for example, due to tolerances, measurement error, measurement precision limitations, and other factors that are known to those skilled in the art.
[0024] The elements depicted in the figures are not necessarily to scale unless otherwise specifically noted.
[0025] Within examples, methods and systems for training and using machine learning models, in particular, acoustic machine perception systems, are described herein. In the present disclosure, examples are primarily described with respect to aircraft. However, it should be appreciated that in other implementations, the disclosed methods and systems can be implemented with vehicles other than aircraft, such as automobiles.
[0026] The disclosed methods and systems relate to an acoustic machine perception system for an aircraft, referred to hereinafter for brevity as “the system.” The system supports both audio and visual modalities. In particular, the system includes one or more sensors (i.e., one or more LIDAR sensors, one or more radar sensors, and / or one or more cameras) arranged (e.g., mounted) on the aircraft and used to acquire “data” (referred to hereinafter as “first data”) associated with an object (referred to hereinafter as “first object”) within the environment surrounding the aircraft (e.g., another aircraft). Among other possible object information, the first data can identify an azimuth of the first object, a range of the first object, a height of the first object, and a type of the first object. The system also includes a microphone array arranged (e.g., mounted) on the aircraft and used to obtain acoustic data (referred to hereinafter as “first acoustic data”) associated with the first object.
[0027] At the same time, the processor of the system uses the first data and the first acoustic data as supervised training data for a machine learning model, so that the machine learning model can thereafter use acoustic data from a target object (hereinafter referred to as “target acoustic data”) to predict object information associated with the target object, such as the azimuth, range, altitude, and type of the target object, among other possibilities. More specifically, the first acoustic data is used as input values to train the machine learning model, and the azimuth, range, altitude, type, etc. of the first object are used as ground truth output labels to train the machine learning model. Thus, during operation of the aircraft (e.g., while the aircraft is on the ground at an airport or while in the air), the system can periodically or continuously use sensors and microphone arrays on the aircraft to acquire data about various objects encountered in the environment around the aircraft in real time, including the relative 3DOF (three degrees of freedom) positions of such objects. Moreover, the system can autonomously determine the supervised training data and use the supervised training data to train the machine learning model, i.e., to first train the machine learning model and / or to thereafter update / refine the machine learning model.
[0028] The disclosed system, arranged and trained in the manner described above, can advantageously replace or supplement existing perception systems rooted in modalities such as LIDAR, radar, and cameras. Moreover, the system’s ability to autonomously determine supervised training data (and in particular ground truth) can improve the efficiency of training a perception system. For example, some existing perception systems involving machine learning can involve or require human (i.e., by a person) or otherwise manual labor or time-consuming processes to train the model and acquire ground truth output labels. But the disclosed system can reduce or eliminate the need for human intervention in training and subsequently use the machine learning model to detect objects and learn information about such objects in the environment. Moreover, where some existing perception systems can use multiple machine learning models for object recognition, the disclosed system can use only a single machine learning model.
[0029] As will be described in greater detail elsewhere herein, in some examples, the disclosed system can use additional information as ground truth output labels, such as data indicating weather conditions of the surrounding environment and / or engine state of the aircraft. The additional information can be used to train the machine learning model to predict weather conditions and / or state of the engine of the aircraft. Moreover, in addition to or in lieu of the system using object information gathered from the first data acquired from sensors of the aircraft, the system can also use object information broadcast from other objects within the environment around the aircraft. For example, where the first object is another aircraft, the first object can be configured to broadcast its azimuth, range, altitude, type, heading, speed, and / or other information, which the system can use as additional ground truth output labels.
[0030] Unless otherwise noted, spatial information for a particular object (such as azimuth, range, and altitude) estimated or determined using one or more sensors of the aerial vehicle or using a machine learning model can take the form of the azimuth, range, and altitude of the particular object relative to the aerial vehicle whose LIDAR sensor, radar sensor, and / or camera was being used to acquire data for training the machine learning model. In the context of information broadcast by a particular object, the broadcast azimuth, range, and altitude of the particular object can be azimuth, range, and altitude relative to the Earth. Further, in some implementations, azimuth, range, altitude, and / or other spatial information predicted for a target object using a machine learning model can be spatial information relative to the aerial vehicle, or in other implementations can be spatial information relative to the Earth. Other frames of reference are also possible.
[0031] These and other improvements will be described in greater detail below. The implementations described below are for purposes of example. The implementations described below, and other implementations, can provide other improvements.
[0032] Referring now to the drawings, Figure 1 is an example of a system 100 for acoustic machine perception. The system 100 includes an aerial vehicle 102, one or more sensors 104 arranged on the aerial vehicle 102, a microphone array 106 arranged on the aerial vehicle 102, and a computing device 108 having a processor 110 and a memory 112 storing instructions 114 executable by the processor 110 to perform operations related to training and using a machine learning model 116. As shown, the computing device 108 includes the machine learning model 116. The one or more sensors 104 include one or more of a LIDAR sensor 118, one or more radar sensors 120, and one or more cameras 122.
[0033] As further shown, Figure 1 An example of a first object 124 is depicted, which represents an object within a surrounding environment of the aerial vehicle 102, and from which the system 100 can acquire information for training the machine learning model 116. Figure 1 An example of a target object 126 is also depicted, which represents another object within the surrounding environment of the aerial vehicle 102.
[0034] The aerial vehicle 102 can take the form of various types of manned or unmanned aerial vehicles, such as a commercial aircraft, a helicopter, or a drone, among other possibilities.
[0035] A LIDAR sensor of the one or more LIDAR sensors 118 from the one or more sensors 104 can take the form of an instrument configured to measure a distance to an object by emitting a laser pulse on a surface of the object and measuring the reflected pulse. The computing device 108 or another computing device can then use the difference in laser return times and wavelengths to image the object, e.g., by creating a three-dimensional (3D) representation (e.g., model) of the object.
[0036] Further, a radar sensor of the one or more radar sensors 120 can take the form of an instrument that includes (i) a transmitter (e.g., an electronic device with an antenna) for producing and emitting electromagnetic waves (e.g., radio waves) and (ii) a receiver (e.g., an electronic device with an antenna (possibly the same as the transmitter’s antenna)) for receiving electromagnetic waves reflected from an object. In some examples, a particular radar sensor or a group of radar sensors of the one or more radar sensors 120 can include its own processor, different from the processor 110 of the computing device 108, that is configured to determine properties of a detected object (e.g., the object’s position, velocity, range, angle, etc.) based on the reflected electromagnetic waves. In alternative examples, the processor 110 of the computing device 108 can be configured to determine the properties of the detected object.
[0037] Still further, a camera of the one or more cameras 122 can take the form of an optical instrument configured to capture still images and / or record videos of the surrounding environment of the aircraft 102, including objects in the surrounding environment.
[0038] The microphone array 106 can be configured to convert sound into electrical signals and can include one or more groups of microphones arranged on the aircraft 102. Each such group of microphones can include one or more microphones. The microphone array 106 can be located at one or more locations on the aircraft 102, such as a surface of a wing, a wingtip, a tail, a tail tip, a nose, and / or a body (e.g., fuselage) of the aircraft 102 (e.g., mounted above and / or below the belly of the fuselage). The microphone array 106 can be spatially distributed in a manner that can maximize the time differences of audio captured in parallel across the microphone array. Further, at least two microphones of the microphone array 106 can be mounted to the aircraft 102 such that they can have a direct line of sight to the first object 124 (or the target object 126) no matter where the object is located. In certain situations, it can be advantageous to place at least a portion of the microphone array 106 away from where the engines of the aircraft 102 are located, so that the sound of the engines does not drown out other important sounds in the surrounding environment.
[0039] The computing device 108 can take the form of a client device (e.g., a computing device actively operated by a user), a server, or some other type of computing platform. In some examples, the computing device 108 can be located on the aircraft 102, for example as part of a navigation system of the aircraft 102. In alternative examples, the computing device 108 can be located at a location remote from the aircraft 102, for example as part of a ground control station or a satellite. Other examples are possible.
[0040] The processor 110 can be a general purpose processor or a special purpose processor (e.g., a digital signal processor, an application specific integrated circuit, etc.). As described above, the processor 110 can be configured to execute instructions 114 (e.g., computer-readable program instructions including computer-executable code) stored in the memory 112 and executable to provide various operations described herein. In alternative examples, the computing device 108 can include additional processors configured in the same manner.
[0041] The memory 112 can take the form of one or more computer-readable storage media that the processor 110 can read and access. The computer-readable storage media can include volatile and / or non-volatile storage components, such as optical, magnetic, organic, or other memory or disc storage, which can be either removable and / or built-in. The memory 112 is considered a non-transitory computer-readable medium. In some examples, the memory 112 can be implemented using a single physical device (e.g., one optical, magnetic, organic, or other memory or disc storage unit), while in other examples, the memory 112 can be implemented using two or more physical devices.
[0042] The machine learning model 116 can take the form of an artificial intelligence-based computational model that can be executed by the processor 110 (or other processor), built using a machine learning algorithm, and trained to make predictions or decisions without being explicitly programmed to do so. For example, the machine learning model 116 can be a neural network in which many basic units (also referred to as “nodes”) can work independently and in parallel to make predictions or decisions, or otherwise solve complex problems, without central control. Each node in the neural network can represent a mathematical function that receives one or more inputs and produces an output.
[0043] As noted above, the machine learning model 116 can be trained using some form of supervised training data. Generally, the supervised training data can include, for example, a collection of input values and ground truth output values (referred to herein as “ground truth output labels”), where each collection of one or more input values has a corresponding collection of one or more desired ground truth output labels. That is, a ground truth output label represents an output value that is known to be an accurate classification of the corresponding input value, or an output value that is expected to be associated with the input value and thus expected to be produced by the machine learning model 116 when provided the input value.
[0044] In some examples, the machine learning model 116 can be trained by applying input values and producing output values, after which the produced output values can be compared to the ground truth output labels. In particular, a loss function (e.g., mean squared error or another metric) can be used to evaluate the error between the produced output values and the ground truth output labels, and the machine learning model 116 can be adjusted to reduce this error. This process can be performed iteratively until the error falls below a desired threshold. Other example processes for training the machine learning model 116 are possible and can additionally or alternatively be used to those described herein.
[0045] The first object 124 can take the form of another aircraft, a human, another type of vehicle (e.g., a ground support vehicle at an airport), or an air traffic control device (among other possibilities). The target object 126 can take one of the forms listed above, or can take another form.
[0046] In operation, a process for training the machine learning model 116 can involve causing the one or more sensors 104 to acquire first data 128 associated with the first object 124 over a time window (e.g., thirty seconds). The act of causing the one or more sensors 104 to acquire the first data 128 can involve causing one or more LIDAR sensors, one or more radar sensors, and / or one or more cameras to acquire the first data 128. As such, the first data 128 can include, for example, LIDAR data, radar data, camera images, and / or a series of camera images.
[0047] Additionally, the process can involve causing the microphone array 106 to acquire first acoustic data 130 associated with the first object 124 within approximately the same time window as the first data 128 is acquired. The first acoustic data 130 and the first data 128 can be acquired within approximately the same time window so that the first acoustic data 130 can more accurately represent the sounds present in the ambient environment and better correlate with the first data 128 when the first object 124 is in the ambient environment. The first acoustic data 130 can take the form of an audio signal representing the sounds caused by the first object 124 and / or audio present in the ambient environment during the time window.
[0048] The one or more sensors 104 can be caused to acquire the first data 128 in various ways, such as by receiving instructions from the computing device 108 or another computing device. The computing device 108 or other computing device can be configured to automatically send such instructions at a predetermined time and / or can send such instructions in response to receiving a manual input from a user. Other examples are possible. The first acoustic data 130 can be obtained in the same or different ways.
[0049] From the first data 128, the computing device 108 (e.g., the processor 110), another computing device, and / or an operator can identify various information associated with the first object 124, including but not limited to: an azimuth of the first object 124; a range of the first object 124; a height of the first object 124; and a type of the first object 124. The term “type” as used in this context can refer to information that identifies a detected object with varying degrees of granularity, such as (i) whether the object is a moving object or not moving, (ii) whether the object is an aircraft, a car, a human, a ground support vehicle, etc., (iii) what type of aircraft (e.g., a helicopter, a commercial passenger jet, etc.), car (e.g., a sedan, a tanker, a baggage car, etc.), human (e.g., a signalman, a pilot, etc.) is present, and / or (iv) a model number or other identifier associated with the aircraft, car, etc. In some cases, the first object 124 (or the target object 126) can be the ground, in which case the type can identify the first object 124 (or the target object 126) as the ground. Other examples are possible.
[0050] In some implementations, the process can also involve the processor 110 or other computing device performing preprocessing to convert the first data 128 and / or the first acoustic data 130 into a form that is best suited for training the machine learning model 116. For example, the audio signal of the first acoustic data 130 can be converted into a spectrogram that visually represents the frequencies of the audio signal over the time window.
[0051] The processor 110 can then train the machine learning model 116 by using the first acoustic data 130 as input value(s) to the machine learning model 116 and using the azimuth, range, height, and type of the first object 124 identified from the first data 128 as ground truth output labels for the machine learning model 116. That is, the processor 110 or other device can be configured to estimate values for the azimuth, range, height, and type of the first object 124 from the first data 128 and input these values as ground truth into the machine learning model 116.
[0052] The training process can also involve other operations. For example, waveforms of the first acoustic data 130 can be time-aligned within milliseconds of a particular timestamp (e.g., using Coordinated Universal Time). The processor 110 or other device can then determine (e.g., using Kalman filtering or other techniques) the azimuth, range, and height of the first object 124 corresponding to the particular timestamp. In addition, the first acoustic data 130 (e.g., the time-aligned waveforms) can be loaded into indexed data blocks and the ground truth output labels (i.e., the estimated azimuth, range, height, etc. of the first object 124) can be loaded into corresponding indexed data blocks, thereby relating the first acoustic data 130 to the ground truth output labels. These and / or other operations described herein can be performed for any number of audio, LIDAR, camera, radar, etc. samples (e.g., thousands of samples) as part of the training process.
[0053] By training the machine learning model 116 in this way, the machine learning model 116 can thus be configured to predict the azimuth of a target object 126, the range of the target object 126, the height of the target object 126, the type of the target object 126, and other possible information based on target acoustic data 132 subsequently acquired by the microphone array 106 (i.e., at a point in time after the machine learning model 116 has been trained using the first acoustic data 130 and the ground truth output labels) and associated with a target object 126 in the surrounding environment of the aircraft 102. For example, the microphone array 106 can acquire an audio signal associated with the target object 126, convert the audio signal into a spectrogram, and use the spectrogram as an input value into the machine learning model 116 to cause the machine learning model 116 to generate (as an output value) the azimuth, range, height, and type of the target object 126 (or an estimate believed to be the azimuth, range, height, and type of the target object 126).
[0054] Once trained to a desired degree (e.g., the machine learning model 116 can identify at least a majority of the objects with a confidence that exceeds a predefined threshold), the machine learning model 116 can be deployed for use in the aircraft 102, and perhaps equally in other aircraft as well. For example, the machine learning model 116 can help the aircraft 102 navigate with respect to the target objects. As a more specific example, the system 100 or other computing system of the aircraft 102 can use the azimuth, range, altitude, or type of the target objects 126 predicted using the machine learning model 116 to control the aircraft 102 to move and avoid the target objects 126, e.g., if the target objects 126 are another aircraft on a runway. As another example, such information can be used to control the aircraft 102 to adjust its heading to match the heading of the target objects 126. As yet another example, the machine learning model 116 can be used to help the aircraft 102 determine the presence and / or proximity of the ground (i.e., the ground is the target object 126 at this time) as the aircraft 102 approaches the ground, as the sound can change as the aircraft 102 approaches the ground. Other examples are possible as well.
[0055] Figure 2 is a diagram depicting data involved in an example process for training the machine learning model 116 and using the machine learning model 116 to predict information associated with the target objects 126. As shown, for example, the first acoustic data 130 is used, along with various information (including the azimuth, range, altitude, and type of the first objects 124 (labeled as first object azimuth 136, first object range 138, first object altitude 139, and first object type 140, respectively, in Figure 2 ) as ground truth output labels 134 to train the machine learning model 116. Once so trained, the machine learning model 116 can predict the azimuth, range, and type of the target objects 126 (labeled as target object azimuth 142, target object range 144, target object altitude 145, and target object type 146, respectively, in Figure 2 ).
[0056] These operations enable the machine learning model 116 to be automatically trained during operation of the aircraft 102 (e.g., while the aircraft 102 is moving in the air or on a runway, or while the engines and systems of the aircraft 102 are running but not moving), without having to manually label ground truth or create training scenarios for the system 100. In other words, the system 100 can desirably reduce or eliminate human intervention in the training process. Moreover, as noted above, once trained, the machine learning model 116 can advantageously be used to supplement existing LIDAR, radar, and camera perception systems, and to identify objects with increased confidence. Additionally or alternatively, a well-trained version of the machine learning model 116 can enable the aircraft 102 to identify objects based on sound alone, thereby enabling the system 100 to replace existing LIDAR, radar, and camera perception systems or to be used in events where such LIDAR, radar, and camera perception systems are offline or otherwise unavailable.
[0057] In some implementations, in addition to or instead of the azimuth, range, altitude, and type of the first object 124 identified from the first acoustic data 130, other information can be used as ground truth output labels 134. For example, the processor 110 can detect Automatic Dependent Surveillance-Broadcast (ADS-B) information 148 broadcast by the first object 124 within approximately the same time window as the first data 128 and the first acoustic data 130 are acquired. To this end, the first object 124 can include a transponder and an associated computing system configured to obtain various information about the first object 124, such as the azimuth, range, altitude, type, heading, speed, and / or global positioning system (GPS) location of the first object 124. One or more of these can be indicated by the ADS-B information 148. The transponder can then broadcast this ADS-B information 148 in real time, which can be received by an antenna or other device located locally or remotely from the system 100 and transmitted to the processor 110. The processor 110 can then use the azimuth, range, altitude, type, heading, speed, and / or GPS location of the first object 124 indicated by the ADS-B information 148 as at least part of the ground truth output labels 134 used to train the machine learning model 116. As noted above, the spatial information broadcast as part of the ADS-B information 148, such as the azimuth, range, and altitude, can be spatial information relative to the Earth, while the first object azimuth 136, the first object range 138, and the first object altitude 139 can be spatial information relative to the aircraft 102.
[0058] In some cases, ADS-B information 148 can be more accurate than information gathered from LIDAR, radar, and / or camera data, in which case it can be helpful to use ADS-B information 148 in addition to or as an alternative to first data 128 (e.g., to account for any differences between the two). ADS-B information 148 is another form of ground truth that system 100 can obtain in real-time during operation of aircraft 102, as various objects in the surrounding environment of aircraft 102 at an airport, in the air, etc. can be configured to broadcast such information. However, in some cases, such as when first object 124 is a type of object that is not configured to broadcast such information, ADS-B information 148 can not be available.
[0059] Another example of ground truth used to train machine learning model 116 is an engine state of aircraft 102. To this end, processor 110 can receive information indicative of an engine state of aircraft 102 (labeled as aircraft engine state information 150 in FIG. 1) within approximately the same time window as first data 128 and first acoustic data 130 are acquired. For example, an avionics system of aircraft 102 can be configured to provide various types of information, any one or more of which can constitute aircraft engine state information 150, such as values related to a revolutions per minute (RPM) gauge, engine oil pressure, and / or engine fuel quantity associated with the engine. Other examples are also possible. Figure 2
[0060] Upon receiving aircraft engine state information 150, processor 110 can then use this aircraft engine state information 150 as one of ground truth output labels 134 for training machine learning model 116. Thus, machine learning model 116 can be configured to use target acoustic data 132 to predict whether an engine of aircraft 102 is generating normal acoustic signatures (e.g., the engine is operating as expected) or abnormal acoustic signatures (e.g., the engine is making a rattling noise (louder than expected) and / or is shutting down slowly). For example, machine learning model 116 can output data indicative of an engine generating abnormal acoustic signatures (labeled as “target object abnormal acoustic signatures 152” in FIG. 1). Figure 2
[0061] So configured, system 100 can quickly and accurately determine an engine state, corroborate another indication of an engine state, and / or identify an anomaly in an engine state. For example, target acoustic data 132 can indicate that the engine sounds like it is not operating in an expected manner, such as when acoustic features of the engine identified from target acoustic data 132 do not match (or are not within a threshold confidence of) any known acoustic features of the engine identified from previously acquired acoustic data. In such a case, machine learning model 116 can output data indicating that there is a problem with the engine.
[0062] System 100 can also use engine states in other ways. For example, in addition to being trained to identify and predict about the engine of aircraft 102 itself, machine learning model 116 can be trained to identify acoustic features of other types of engines used in other aircraft or vehicles. Thus, machine learning model 116 can be used to help system 100 detect or verify the identity and presence of other aircraft or vehicles based on identified acoustic features of their respective engines (e.g., acoustic features indicating that a nearby vehicle has the same engine or a similar engine as aircraft 102). As another example, because aircraft engine state information 150 can be used to help system 100 identify sounds emitted by the engine of aircraft 102 during normal operation, system 100 or another computing system can perform pre-processing on acoustic data to filter out engine sounds before using the acoustic data to predict information about a target object. This can be useful in situations where the engine can be loud and can make it difficult to detect subtle differences in sounds produced by a target object. Other examples are possible as well.
[0063] In certain situations, weather conditions such as rain, snow, wind, and fog can negatively impact LIDAR, radar, and / or camera data, which can reduce the confidence with which a machine perception system uses such modalities to identify objects. Accordingly, one or more weather conditions (labeled as weather conditions 154 in Figure 2 may be used to train machine learning model 116 as one or more of ground truth output labels 134, in addition to or in lieu of other types of ground truth discussed above, thus using acoustic perception to increase the confidence with which system 100 identifies objects. This can serve as a practical and automated way to account for changing weather conditions that can be present in the surroundings of aircraft 102.
[0064] Weather conditions 154 can be determined in a variety of ways. As an example, weather conditions 154 can be estimated and identified from first data 128. For example, system 100 or another system of aircraft 102 can be configured to identify the presence of rain or snow in camera images or video.
[0065] Additionally or alternatively, the weather condition 154 can be determined by correlating a weather report with the GPS location of the aircraft 102. To this end, the processor 110 or other computing device of the aircraft 102 can determine the GPS location of the aircraft 102 within approximately the same time window as the first data 128 and the first acoustic data 130 are acquired. Additionally, the processor 110 or other computing device can receive a weather report indicating a weather condition 154 for at least one geographic location (e.g., a county, city, or region of a state or country). In this way, the processor 110 or other computing device can determine whether the weather condition 154 indicated in the weather report is present at or near the GPS location of the aircraft 102 (i.e., where the aircraft 102 is located) during approximately the same time window as the first data 128 and the first acoustic data 130 are acquired.
[0066] In cases where the weather condition 154 is used to train the machine learning model 116, the machine learning model 116 can be configured to predict whether the weather condition 154 is present in the surrounding environment of the aircraft 102 based on the target acoustic data 132. The weather condition 154 that is predicted to be present is labeled as the target weather condition 156 in Figure 2
[0067] In some implementations, the microphone array 106 that acquires the first acoustic data 130 (or similarly, the target acoustic data 132) can include one or more microphones arranged on each wing of the aircraft 102, on or near the wingtips. In further examples, one or more microphones can be arranged at the front of the aircraft 102, such as on the nose and / or tail of the aircraft 102. Such microphones can be located at the tail, wingtips, etc. to avoid the sound of the engines of the aircraft 102 interfering with the sound from the first object 124. Spacing the microphones in this way can also help to account for various positions of the first object 124 relative to the aircraft 102. For example, if the first object 124 is behind the aircraft 102, the system 100 can be able to obtain better first acoustic data 130 using one or more microphones on the tail of the aircraft 102 than if there were no microphones on the tail of the aircraft 102. Other examples are possible as well.
[0068] As an illustrative example of the locations of the microphone arrays, Figure 3 A perspective view of the aircraft 102 is shown. In particular, Figure 3 A respective set of microphones can be located on each wing of the aircraft 102 (e.g., near the wingtips) as well as on the tail of the aircraft 102, as shown. (Because each set of microphones is part of the microphone array 106, the microphone arrays 106 are not shown in Figure 3 , each set of microphones is indicated by reference numeral 106.) In alternative examples, the microphone array 106 may be located elsewhere.
[0069] In some implementations, the machine learning model 116 may be a single model that is specifically designed for use with (i.e., trained on) the acoustic data acquired by each microphone in the microphone array 106. In an alternative example, the machine learning model 116 may be one of a plurality of machine learning models that are each specifically designed for use with acoustic data acquired by a corresponding location on the aircraft 102 where one or more microphones are disposed. For example, one model may be specific to the microphones on the wingtips of the aircraft 102, while another model may be specific to the microphones on the tail of the aircraft 102. Other examples are also possible.
[0070] Figure 4 Shows that it can be Figure 1 A flow chart of an example of a method 200 for use with the system 100 is shown, particularly for training the machine learning model 116. The method 200 may include one or more operations, functions, or actions, as illustrated by one or more of blocks 202-206.
[0071] At block 202 , method 200 includes causing one or more sensors disposed on an aircraft to acquire first data associated with a first object within an environment surrounding the aircraft within a time window, wherein the one or more sensors include one or more of a LIDAR sensor, a radar sensor, and a camera.
[0072] At block 204 , method 200 includes causing a microphone array disposed on an aircraft to acquire first acoustic data associated with a first object within approximately the same time window as the first data was acquired.
[0073] At box 206, method 200 includes training a machine learning model by a processor in the following manner: using the first acoustic data as an input value of the machine learning model, and using the azimuth of the first object, the range of the first object, the altitude of the first object, and the type of the first object identified from the first data as ground truth output labels of the machine learning model, wherein the machine learning model is configured to predict the azimuth of the target object, the range of the target object, the altitude of the target object, and the type of the target object based on target acoustic data subsequently acquired by the microphone array and associated with the target object within the aircraft's surrounding environment.
[0074] Figure 5A flowchart showing an example method for performing the causing functionality as shown in block 204 is shown. At block 208, the functionality includes causing one or more microphones arranged on each wing of the aircraft to acquire first acoustic data associated with the first object.
[0075] In some implementations, the azimuth of the first object, the range of the first object, and the height of the first object identified from the first data can be relative to the aircraft. In such implementations, additional spatial information associated with the first object can be obtained and used as additional ground truth. Figure 6 A flowchart showing an example method for use with the method 200 is shown. At block 210, the functionality includes detecting, by the processor, ADS-B information broadcast by the first object within approximately the same time window as the first data and the first acoustic data are acquired. The ADS-B information can indicate an azimuth of the first object relative to the Earth, a range of the first object relative to the Earth, a height of the first object relative to the Earth, a type of the first object, a heading of the first object, and a speed of the first object, among other possible information. Figure 6 A flowchart showing an example method for performing the training as shown in block 206 is also shown, particularly in implementations where the processor detects ADS-B information as shown in block 210. At block 212, the functionality includes training the machine learning model by using the azimuth of the first object relative to the Earth, the range of the first object relative to the Earth, the height of the first object relative to the Earth, the type of the first object, the heading of the first object, and the speed of the first object indicated by the ADS-B information broadcast by the first object as at least a portion of ground truth output labels for the machine learning model.
[0076] Figure 7 A flowchart showing another example method for use with the method 200 is shown. At block 214, the functionality includes receiving, by the processor, information indicating an engine state of the aircraft within approximately the same time window as the first data and the first acoustic data are acquired. Figure 7 A flowchart showing an example method for performing the training as shown in block 206 is also shown, particularly in implementations where the processor receives information indicating an engine state as shown in block 214. At block 216, the functionality includes training the machine learning model by using the engine state as one of the ground truth output labels for the machine learning model, where the machine learning model is configured to predict that the engine is generating anomalous acoustic features based on the target acoustic data.
[0077] Figure 8A flow chart of another example method for use with method 200 is shown. At block 218, the functionality includes determining a GPS position of the aircraft within approximately the same time window as when the first data and the first acoustic data were acquired. At block 220, the functionality includes receiving a weather report. At block 222, the functionality includes determining that one or more weather conditions indicated by the weather report are present in the surrounding environment at the GPS position within approximately the same time window as when the first data and the first acoustic data were acquired. Figure 8 Also shown is a flow chart of an example method for performing training as shown in block 206, particularly in an implementation in which the functionality of blocks 218, 220, and 222 is performed. At block 224, the functionality includes training a machine learning model by using one or more weather conditions of the surrounding environment as one or more of the ground truth output labels of the machine learning model, wherein the machine learning model is configured to predict the presence of the one or more weather conditions in the surrounding environment based on the target acoustic data.
[0078] Figure 9 A flow chart is shown of another example method for performing training as shown in block 206. At block 226, the functionality includes training a machine learning model by using one or more weather conditions of the surrounding environment identified from the first data as one or more of the ground truth output labels of the machine learning model, wherein the machine learning model is configured to predict the presence of the one or more weather conditions in the surrounding environment based on the target acoustic data.
[0079] Figure 10 Shows that it can be Figure 1 1. A flowchart of an example of a method 300 for use with the system 100 is shown, particularly for using the machine learning model 116 after the machine learning model has been trained to predict information associated with a target object (e.g., the target object 126). The method 300 may include one or more operations, functions, or actions as illustrated in one or more of blocks 302-304.
[0080] At block 302 , the functionality includes causing a microphone array disposed on an aircraft to acquire target acoustic data associated with a target object within an environment surrounding the aircraft.
[0081] At block 304, the functionality includes executing, by the processor, a machine learning model to predict an azimuth of the target object, a range of the target object, a height of the target object, and a type of the target object based on the target acoustic data, wherein the machine learning model is trained by: (i) causing one or more sensors disposed on the aerial vehicle to acquire first data associated with a first object within an environment surrounding the aerial vehicle, wherein the one or more sensors include one or more of a LIDAR sensor, a radar sensor, and a camera; (ii) causing an array of microphones disposed on the aerial vehicle to acquire first acoustic data associated with the first object within substantially the same time window as the first data is acquired; (iii) the processor using the first acoustic data as input values for the machine learning model; and (iv) the processor using an azimuth of the first object, a range of the first object, a height of the first object, and a type of the first object identified from the first data as ground truth output labels for the machine learning model.
[0082] Figure 11 A flowchart of an example method for use with the method 300 is shown. At block 306, the functionality includes causing the aerial vehicle to move and avoid the target object based on one or more of the azimuth of the target object, the range of the target object, the height of the target object, or the type of the target object.
[0083] Figure 12 A flowchart of an example method for performing the causing functionality as shown in block 302 is shown. At block 308, the functionality includes causing one or more microphones disposed on each wing of the aerial vehicle to acquire target acoustic data associated with the target object. The microphones used to acquire the target acoustic data can be the same or different from the microphones used to acquire the first acoustic data.
[0084] Further, the present disclosure includes examples in accordance with the following clauses:
[0085] Clause 1. A method comprising: causing one or more sensors disposed on an aerial vehicle to acquire, within a time window, first data associated with a first object within a surrounding environment of the aerial vehicle, wherein the one or more sensors comprise one or more of a light detection and ranging sensor, a radar sensor, and a camera; causing a microphone array disposed on the aerial vehicle to acquire, within approximately the same time window as the acquisition of the first data, first acoustic data associated with the first object; and training, by a processor, a machine learning model using the first acoustic data as input values to the machine learning model and using an azimuth angle of the first object, a range of the first object, a height of the first object, and a type of the first object, as identified from the first data, as ground truth output labels to the machine learning model, wherein the machine learning model is configured to predict an azimuth angle of a target object, a range of the target object, a height of the target object, and a type of the target object based on target acoustic data subsequently acquired by the microphone array and associated with the target object within the surrounding environment of the aerial vehicle.
[0086] Clause 2. The method of clause 1, wherein causing the microphone array disposed on the aerial vehicle to acquire the first acoustic data associated with the first object comprises causing one or more microphones disposed on each wing of the aerial vehicle to acquire the first acoustic data associated with the first object.
[0087] Clause 3. The method of any one of clauses 1-2, wherein the azimuth angle of the first object, the range of the first object, and the height of the first object, as identified from the first data, are relative to the aerial vehicle, the method further comprising: detecting, by the processor within approximately the same time window as the acquisition of the first data and the first acoustic data, automatic dependent surveillance broadcast information broadcast by the first object, wherein the automatic dependent surveillance broadcast information indicates an azimuth angle of the first object relative to the Earth, a range of the first object relative to the Earth, a height of the first object relative to the Earth, a type of the first object, a heading of the first object, and a speed of the first object, and wherein training the machine learning model comprises training the machine learning model using the azimuth angle of the first object relative to the Earth, the range of the first object relative to the Earth, the height of the first object relative to the Earth, the type of the first object, the heading of the first object, and the speed of the first object, as indicated by the automatic dependent surveillance broadcast information broadcast by the first object, as at least a portion of the ground truth output labels to the machine learning model.
[0088] Clause 4. The method of any of clauses 1-3, further comprising: receiving, by the processor, information indicative of an engine state of the aircraft within substantially the same time window as the first data and the first acoustic data are acquired, wherein training the machine learning model comprises: training the machine learning model using the engine state as one of the ground truth output labels for the machine learning model, wherein the machine learning model is configured to predict that the engine is generating anomalous acoustic characteristics based on the target acoustic data.
[0089] Clause 5. The method of any of clauses 1-4, further comprising: determining a global positioning system location of the aircraft within substantially the same time window as the first data and the first acoustic data are acquired; and receiving a weather report; and determining that one or more weather conditions indicated by the weather report are present in a surrounding environment at the global positioning system location within substantially the same time window as the first data and the first acoustic data are acquired, wherein training the machine learning model comprises: training the machine learning model using the one or more weather conditions of the surrounding environment as one or more ground truth output labels for the machine learning model, wherein the machine learning model is configured to predict that the one or more weather conditions are present in the surrounding environment based on the target acoustic data.
[0090] Clause 6. The method of any of clauses 1-5, wherein training the machine learning model comprises: training the machine learning model using the one or more weather conditions of the surrounding environment identified from the first data as one or more ground truth output labels for the machine learning model, wherein the machine learning model is configured to predict that the one or more weather conditions are present in the surrounding environment based on the target acoustic data.
[0091] Clause 7. The method of any of clauses 1-6, wherein the method is performed while the aircraft is in motion.
[0092] Clause 8. A method comprising: causing a microphone array disposed on an aerial vehicle to acquire target acoustic data associated with a target object within a surrounding environment of the aerial vehicle; and executing, by a processor, a machine learning model to predict an azimuth of the target object, a range of the target object, an altitude of the target object, and a type of the target object based on the target acoustic data, wherein the machine learning model is trained by: (i) causing one or more sensors disposed on the aerial vehicle to acquire first data associated with a first object within the surrounding environment of the aerial vehicle, wherein the one or more sensors include one or more of a light detection and ranging sensor, a radar sensor, and a camera; (ii) causing the microphone array disposed on the aerial vehicle to acquire first acoustic data associated with the first object within approximately the same time window as the first data is acquired; (iii) the processor using the first acoustic data as input values to the machine learning model; and (iv) the processor using an azimuth of the first object, a range of the first object, an altitude of the first object, and a class of the first object identified from the first data as ground truth output labels to the machine learning model.
[0093] Clause 9. The method of clause 8, further comprising: causing the aerial vehicle to move and avoid the target object based on one or more of the azimuth of the target object, the range of the target object, the altitude of the target object, or the type of the target object.
[0094] Clause 10. The method of any one of clauses 8-9, wherein causing the microphone array disposed on the aerial vehicle to acquire the target acoustic data associated with the target object comprises: causing one or more microphones disposed on each wing of the aerial vehicle to acquire the target acoustic data associated with the target object.
[0095] Clause 11. The method of any one of clauses 8-10, wherein the azimuth of the first object, the range of the first object, and the altitude of the first object identified from the first data are relative to the aerial vehicle, and wherein the machine learning model is further trained by: detecting, by the processor within approximately the same time window as the first data and the first acoustic data are acquired, automatic dependent surveillance broadcast information broadcasted by the first object, wherein the automatic dependent surveillance broadcast information indicates an azimuth of the first object relative to the Earth, a range of the first object relative to the Earth, an altitude of the first object relative to the Earth, a type of the first object, a heading of the first object, and a speed of the first object, and using the azimuth of the first object relative to the Earth, the range of the first object relative to the Earth, the altitude of the first object relative to the Earth, the type of the first object, the heading of the first object, and the speed of the first object indicated by the automatic dependent surveillance broadcast information broadcasted by the first object as at least a portion of ground truth output labels to the machine learning model.
[0096] Clause 12. The method of any of clauses 8-11, wherein the machine learning model is further trained by receiving, by the processor, information indicative of an engine state of the aircraft within approximately the same time window as the first data and the first acoustic data are acquired and using the engine state as one of the ground truth output labels for the machine learning model, wherein the machine learning model is configured to predict that the engine is generating anomalous acoustic characteristics based on the target acoustic data.
[0097] Clause 13. The method of any of clauses 8-12, wherein the machine learning model is further trained by determining a global positioning system location of the aircraft within approximately the same time window as the first data and the first acoustic data are acquired; receiving a weather report; determining that one or more weather conditions indicated by the weather report are present in a surrounding environment at the global positioning system location within approximately the same time window as the first data and the first acoustic data are acquired and using the one or more weather conditions of the surrounding environment as one or more of the ground truth output labels for the machine learning model, wherein the machine learning model is configured to predict that the one or more weather conditions are present in the surrounding environment based on the target acoustic data.
[0098] Clause 14. The method of any of clauses 8-13, wherein the machine learning model is further trained by using one or more weather conditions of a surrounding environment identified from the first data as one or more of the ground truth output labels for the machine learning model, wherein the machine learning model is configured to predict that the one or more weather conditions are present in the surrounding environment based on the target acoustic data.
[0099] Clause 15. A system comprising: an aerial vehicle; one or more sensors disposed on the aerial vehicle, wherein the one or more sensors comprise one or more of a light detection and ranging sensor, a radar sensor, and a camera; a microphone array disposed on the aerial vehicle; and a computing device having a processor and a memory storing instructions executable by the processor to perform a set of operations comprising: causing the one or more sensors to acquire, within a time window, first data associated with a first object within a surrounding environment of the aerial vehicle; causing the microphone array to acquire, within approximately the same time window as the first data, first acoustic data associated with the first object; and training a machine learning device by using the first acoustic data as input values to the machine learning model and using an azimuth angle of the first object, a range of the first object, a height of the first object, and a type of the first object identified from the first data as ground truth output labels to the machine learning model, wherein the machine learning model is configured to predict an azimuth angle of a target object, a range of the target object, a height of the target object, and a type of the target object based on target acoustic data subsequently acquired by the microphone array and associated with the target object in the surrounding environment of the aerial vehicle.
[0100] Clause 16. The system of clause 15, wherein the computing device is on the aerial vehicle.
[0101] Clause 17. The system of any one of clauses 15-16, wherein the set of operations are performed while the aerial vehicle is in motion.
[0102] Clause 18. The system of any one of clauses 15-17, wherein the microphone array comprises one or more microphones disposed on each wing of the aerial vehicle.
[0103] Clause 19. The system of any one of Clauses 15-18, wherein the azimuth of the first object, the range of the first object, and the altitude of the first object identified from the first data are relative to the aircraft, the set of operations further comprising: detecting, within approximately the same time window as the first data and the first acoustic data are acquired, an automatic dependent surveillance broadcast information broadcast by the first object, wherein the automatic dependent surveillance broadcast information indicates an azimuth of the first object relative to the Earth, a range of the first object relative to the Earth, an altitude of the first object relative to the Earth, a type of the first object, a heading of the first object, and a speed of the first object, wherein training the machine learning model comprises training the machine learning model using the azimuth of the first object relative to the Earth, the range of the first object relative to the Earth, the altitude of the first object relative to the Earth, the type of the first object, the heading of the first object, and the speed of the first object indicated by the automatic dependent surveillance broadcast information broadcast by the first object as at least a portion of ground truth output labels for the machine learning model.
[0104] Clause 20. The system of any one of Clauses 15-19, the set of operations further comprising: receiving, within approximately the same time window as the first data and the first acoustic data are acquired, information indicating a state of an engine of the aircraft, wherein training the machine learning model comprises training the machine learning model using the state of the engine as one of the ground truth output labels for the machine learning model, wherein the machine learning model is configured to predict that the engine is generating anomalous acoustic features based on the target acoustic data.
[0105] The logic functions shown in Figure 4 , Figure 5 , Figure 6 , Figure 7 , Figure 8 , Figure 9 , Figure 10 , Figure 11 , and Figure 12 may be performed using a device or system, or configured to perform them. In some cases, components of the device and / or system can be configured to perform these functions, such that the components are actually configured and constructed (using hardware and / or software) to enable such performance. In other examples, components of the device and / or system can be arranged to be suitable, capable or adapted to perform these functions, e.g., when operated in a particular manner. Although shown in a sequential order, some of the Figure 4 , Figure 5 , Figure 6 , Figure 7 , Figure 8 , Figure 9 , Figure 10 , Figure 11 , and Figure 12The blocks in the flow diagrams herein represent code modules, segments, or portions of code which include one or more instructions for implementing specific logical functions (or steps) in the process. M odules can be stored on a computer-readable medium, which can include any type of disk including floppy disks, optical disks, CD-ROMs, and magnetic- or optical-storage disks, semiconductor memory such as random access memory (RAM), and tape. Alternatively, or additionally, the modules can be transmitted as generated code sequences over the Internet in the manner described, using any one of a number of well-known transfer protocols (e.g., HTTP). Each block or portion of a block in the flow diagrams can represent a module, segment, or portion of code which includes one or more instructions for implementing the specified logical function(s) or step(s) in the process. Additionally, in some alternative implementations, the functions or aspects can be executed out of order from that shown or discussed, including substantially concurrently or in reverse order, depending on the functionality involved. The flow diagrams depict processes that can be executed to implement the disclosed subject matter. In this regard, each block in the flow diagram can represent a module, segment, or portion of code, which includes one or more executable instructions for implementing the specified logical functions (or steps) in the process. It should also be noted that each block in the flow diagram can represent a portion of code that includes one or more instructions executable to implement the specified logical functions (or steps) in the process. The program code may
[0106] It should be understood that, for the herein-disclosed processes and methods, the flow diagrams represent one possible implementation of the present examples. In this regard, each block in the flow diagrams can represent a module, segment, or portion of code, which includes one or more instructions that can be executed by a processor to implement the specific logical functions or steps in the process. The program code can be stored on any type of computer readable medium or data storage device, including a storage device such as a disk or hard drive. Additionally, the program code can be encoded in a machine-readable format on a computer readable storage medium, or encoded on other non-transitory media or articles of manufacture. The computer readable medium can include non-transitory computer readable media or memory, such as a short term memory, including a computer readable medium such as a register memory, processor cache, and random access memory (RAM). The computer readable medium can also include a non-transitory medium, such as a secondary or persistent long term storage medium, such as a read only memory (ROM), an optical disc, or a magnetic disk. The computer readable medium can also be any other volatile or non-volatile storage system. For example, the computer readable medium can be considered tangible computer readable storage media.
[0107] Additionally, Figure 4 , Figure 5 , Figure 6 , Figure 7 , Figure 8 , Figure 9 , Figure 10 , Figure 11 and Figure 12 Each block or portion of a block in the flow diagrams can represent a circuit that is wired to perform the particular logical function in the process. Alternative implementations include within the scope of the examples of the disclosure, where the functions can be performed in an order other than that shown or discussed, including substantially simultaneously, or in the reverse order, depending on the functionality involved, as will be understood by those skilled in the art.
[0108] The various examples of the systems, devices, and methods disclosed herein include various components, features, and functions. It should be understood that the various examples of the systems, devices, and methods disclosed herein can include any combination or any sub-combination of any of the components, features, and functions disclosed herein, and all such possibilities are intended to fall within the scope of the present disclosure.
[0109] The description of the different advantageous arrangements has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the examples disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art. Moreover, different advantageous arrangements can provide different advantages as compared to other advantageous arrangements. The chosen examples were chosen and described in order to best explain the principles of the examples, the practical application, and to enable others skilled in the art to understand the various examples with various modifications in accord with various examples.
Claims
1. A method for acoustic machine perception of an aircraft, the method comprising the following steps: causing one or more sensors disposed on the aircraft to acquire first data associated with a first object in a surrounding environment of the aircraft within a time window, wherein the one or more sensors include one or more of a light detection and ranging sensor, a radar sensor, and a camera; causing a microphone array disposed on the aircraft to acquire first acoustic data associated with the first object within the same time window as that for acquiring the first data; and training, by a processor, a machine learning model by using the first acoustic data as input values of the machine learning model and using the azimuth of the first object, the range of the first object, the altitude of the first object, and the type of the first object identified from the first data as ground truth output labels of the machine learning model, wherein the machine learning model is configured to predict the azimuth of the target object, the range of the target object, the altitude of the target object, and the type of the target object based on target acoustic data subsequently acquired by the microphone array and associated with the target object in the surrounding environment of the aircraft, The method further comprises the following steps: receiving, by the processor, information indicating a status of an engine of the aircraft within the same time window as acquiring the first data and the first acoustic data, wherein the step of training the machine learning model comprises training the machine learning model by using the state of the engine as one of the ground truth output labels of the machine learning model, Wherein, the machine learning model is configured to predict that the engine is generating abnormal acoustic characteristics based on the target acoustic data.
2. The method according to claim 1, wherein The step of causing the microphone array arranged on the aircraft to acquire the first acoustic data associated with the first object includes causing one or more microphones arranged on each wing of the aircraft to acquire the first acoustic data associated with the first object.
3. The method according to claim 1 or 2, wherein: The azimuth of the first object, the range of the first object, and the altitude of the first object identified from the first data are relative to the aircraft, and the method further includes the following steps: detecting, by the processor, automatic dependent surveillance broadcast information broadcast by the first object within the same time window as acquiring the first data and the first acoustic data, wherein the automatic dependent surveillance broadcast information indicates an azimuth of the first object relative to the earth, a range of the first object relative to the earth, an altitude of the first object relative to the earth, a type of the first object, a heading direction of the first object, and a speed of the first object, and Wherein, the step of training the machine learning model includes training the machine learning model in the following manner: using the azimuth of the first object relative to the earth, the range of the first object relative to the earth, the altitude of the first object relative to the earth, the type of the first object, the heading of the first object, and the speed of the first object indicated by the automatic dependent surveillance broadcast information broadcast by the first object as at least part of the ground truth output label of the machine learning model.
4. The method according to claim 1 or 2, further comprising the steps of: determining a global positioning system position of the aircraft within the same time window as acquiring the first data and the first acoustic data; Receive weather reports; as well as determining that one or more weather conditions of the surrounding environment indicated by the weather report exist at the global positioning system location within the same time window as acquiring the first data and the first acoustic data, The step of training the machine learning model comprises training the machine learning model by using the one or more weather conditions of the surrounding environment as one or more of the ground truth output labels of the machine learning model, Wherein, the machine learning model is configured to predict the presence of the one or more weather conditions in the surrounding environment based on the target acoustic data.
5. The method according to claim 1 or 2, wherein: The step of training the machine learning model includes training the machine learning model by using one or more weather conditions of the surrounding environment identified from the first data as one or more of the ground truth output labels of the machine learning model, Wherein, the machine learning model is configured to predict the presence of the one or more weather conditions in the surrounding environment based on the target acoustic data.
6. The method according to claim 1 or 2, wherein: The method is performed while the aircraft is in motion.
7. A system for acoustic machine perception of an aircraft, the system comprising: the aircraft; one or more sensors disposed on the aircraft, wherein the one or more sensors include one or more of a light detection and ranging sensor, a radar sensor, and a camera; a microphone array disposed on the aircraft; and A computing device having a processor and a memory storing instructions, the instructions executable by the processor to perform a set of operations, the set of operations comprising: causing the one or more sensors to acquire first data associated with a first object within a surrounding environment of the aircraft within a time window; causing the microphone array to acquire first acoustic data associated with the first object within the same time window as that for acquiring the first data; and training a machine learning model by using the first acoustic data as input values of the machine learning model and using the azimuth of the first object, the range of the first object, the altitude of the first object, and the type of the first object identified from the first data as ground truth output labels of the machine learning model, wherein the machine learning model is configured to predict the azimuth of the target object, the range of the target object, the altitude of the target object, and the type of the target object based on target acoustic data subsequently acquired by the microphone array and associated with the target object in the surrounding environment of the aircraft, The set of operations also includes: receiving information indicative of a status of an engine of the aircraft within the same time window as acquiring the first data and the first acoustic data, wherein training the machine learning model comprises training the machine learning model by using the state of the engine as one of the ground truth output labels of the machine learning model, Wherein, the machine learning model is configured to predict that the engine is generating abnormal acoustic characteristics based on the target acoustic data.
8. The system according to claim 7, wherein: The computing device is onboard the aircraft.
9. The system according to claim 7 or 8, wherein: The set of operations is performed while the aircraft is in motion.
10. The system according to claim 7 or 8, wherein: The microphone array includes one or more microphones arranged on each wing of the aircraft.
11. The system according to claim 7 or 8, wherein: The azimuth of the first object, the range of the first object, and the altitude of the first object identified from the first data are relative to the aircraft, the set of operations further comprising: detecting automatic dependent surveillance broadcast information broadcast by the first object within the same time window as acquiring the first data and the first acoustic data, wherein the automatic dependent surveillance broadcast information indicates an azimuth of the first object relative to the earth, a range of the first object relative to the earth, an altitude of the first object relative to the earth, a type of the first object, a heading direction of the first object, and a speed of the first object, and Wherein, training the machine learning model includes training the machine learning model in the following manner: using the azimuth of the first object relative to the earth, the range of the first object relative to the earth, the altitude of the first object relative to the earth, the type of the first object, the heading of the first object, and the speed of the first object indicated by the automatic dependent surveillance broadcast information broadcast by the first object as at least part of the ground truth output label of the machine learning model.
Citation Information
Patent Citations
Collision Avoidance System and Method for Unmanned Aircraft
US20180196435A1
Deep neural network of multiple audio streams for location determination and environment monitoring
US20190057715A1