Extracting audio signal from audio mixed signal using neural network

By using a deep learning-based multi-channel sound separation system and target phase correlation spectrogram, the challenge of audio source separation in complex industrial environments is solved, achieving accurate separation and state estimation of multiple audio sources, and improving the accuracy of anomaly detection and health monitoring.

CN121794751APending Publication Date: 2026-04-03MITSUBISHI ELECTRIC CORP
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing audio source separation technologies struggle to effectively separate multiple audio sources in complex industrial environments, particularly failing to accurately identify and separate audio signals generated by similar components, leading to inaccurate anomaly detection and health monitoring.

Method used

A deep learning-based multi-channel sound separation system is adopted. The neural network is trained by using the target phase correlation spectrogram and the complex U-net architecture, combined with the position information of the microphone array, to separate different audio sources in the audio mixture signal.

Benefits of technology

It enables accurate separation of multiple audio sources in complex industrial environments, supports anomaly detection and health monitoring, and improves the accuracy of machine component condition estimation and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121794751A_ABST
    Figure CN121794751A_ABST
Patent Text Reader

Abstract

The present disclosure provides an audio system, method, and system for facilitating machine operation. The machine includes an actuator that assists the tool in performing the task. In an example, an audio system is configured to receive an audio mixed signal of a signal generated by an audio source that includes at least one of a tool or an actuator that is performing a task. An audio source forming an audio mixed signal is identified by a relative position to each microphone of a microphone array that measures the audio mixed signal. The audio system is configured to extract an audio signal from an audio mixed signal generated by the identified audio source based on a correlation of spectral features in a multi-channel spectrogram of the audio mixed signal and directional information indicative of a relative position of the identified audio source. The audio system outputs the extracted audio signal to facilitate operation of the machine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to sound separation using machine learning, and more specifically to extracting isolated audio signals from acoustically mixed signals using machine learning. Background Technology

[0002] Monitoring and controlling safety and quality are crucial in machinery where fast and powerful devices or equipment can execute complex sequences of operations at high speeds. Deviations from the intended sequence or timing can reduce quality, waste raw materials, cause downtime and equipment damage, and reduce output. Furthermore, in some cases, any deviation can also pose a danger to workers. Therefore, it is important to design processes to minimize unforeseen events, and safety measures must be implemented using various sensors and emergency switches.

[0003] One practical approach to increasing safety and minimizing material and output losses is to detect when a machine malfunctions and, if necessary, stop it. One way to achieve this is by defining a permissible operating zone based on the range of measurable variables (e.g., temperature, pressure, etc.) and detecting operating points outside that zone. This approach is common in process manufacturing industries (e.g., oil refining), where the permissible ranges of physically measurable variables and quality metrics for product quality, defined directly from these variables, are generally well understood.

[0004] However, in some cases, deviations from normal operating procedures can have significantly different characteristics. For example, anomalies may include the incorrect execution of one or more tasks or the incorrect order of tasks. Even in anomalous situations, physical variables (such as temperature or pressure) are typically not out of range, so directly monitoring such variables cannot reliably detect such anomalies.

[0005] In some cases, complex systems can comprise combinations of processes and operations. When processes and operations are intertwined on signal production line machines, anomaly detection methods designed for different types of manufacturing may be inaccurate. Therefore, it is natural to design different anomaly detection methods for different categories of operations. One such anomaly detection method could be based on audio processing of sounds generated by the machine during its operation.

[0006] In the field of audio processing, isolating different audio sources from combined audio mixtures has always been a complex and challenging problem. This problem arises in a variety of scenarios, such as isolating human voices from musical instruments, separating multiple speakers in a recorded dialogue, extracting specific sound events from a noisy background, extracting specific sounds from tools or actuators in a machine, and extracting the sounds of specific vehicles from a convoy.

[0007] In the example, when monitoring machine performance and health, highly skilled human operators can use their ears to listen to the sounds generated during machine operation. With increasing automation, algorithms that use microphones to monitor machine sounds are becoming increasingly important. Therefore, the use of anomalous sound detection techniques, where automated algorithms can determine whether sounds generated during machine operation are normal or abnormal, has increased.

[0008] Recently, there has been increasing research addressing the challenging problem setting where only normal data is available for training. In this case, domain offset leads to variations in the sound signal regardless of the presence or absence of anomalous sounds. However, much of the existing literature on machine sound analysis treats recorded sounds as being generated by a single machine component. In practice, however, most industrial machinery consists of multiple sound-generating components or parts, and it may be desirable to monitor the health of each component individually.

[0009] In these cases, audio source separation can be a useful preprocessing step to isolate sound from each machine component, and the separated sound signals from each component can then be used as input for downstream processing, such as anomalous sound detection or other types of audio monitoring. Audio source separation techniques have been used in the fields of speech enhancement, speech separation, music source separation, and general sound separation. Traditional methods of audio source separation rely primarily on basic filtering techniques, such as bandpass filters or spectral subtraction. While these techniques provide a degree of separation, they typically lack the ability to effectively distinguish closely overlapping audio sources or manage complex sound environments with different spectral and temporal characteristics.

[0010] In recent years, advancements in digital signal processing have led to the development of more complex algorithms for audio source separation. This is evidenced by the use of techniques such as Isolated Component Analysis (ICA), Nonnegative Matrix Factorization (NMF), and deep learning-based methods for audio signal extraction or separation. However, these methods still face limitations in separating sources with high similarity in spectral content and achieving near real-time processing. Furthermore, existing deep learning-based solutions utilize a fully supervised framework, where a database of isolated audio signals is used to create artificially mixed signals, and the source separation model is trained to separate using ground-based isolated signals as targets. Therefore, collecting a database of isolated signals from individual machine components may be impractical for separating machine parts, as all components may need to operate simultaneously for the machine to function, necessitating the consideration of methods with limited supervision.

[0011] Existing deep learning solutions often focus on specific types of audio sources or require isolated or separate audio signals from different audio sources for training, which may not be readily available in many practical situations. For example, for machines that include multiple tools and / or actuators that can operate separately or in combination to perform one or more tasks, separate audio signals from different tools and / or actuators may be unavailable or practically impossible to obtain.

[0012] Therefore, there is a need for an improved audio separation technique that addresses the limitations of current methods, provides enhanced separation performance even in complex and challenging acoustic environments, and offers the flexibility to manage a wide range of audio source types without heavily relying on predefined models or assumptions. Summary of the Invention

[0013] One implementation aims to provide a system and method adapted to acoustically separate mixed audio signals from a complex industrial system having multiple actuators that actuate one or more tools to perform one or more tasks. Additionally or alternatively, one implementation aims to use machine learning to estimate the execution state of these tasks, detect anomalies in the operation of the industrial system, and control the system accordingly.

[0014] Some implementations are based on the understanding that deep learning-based techniques can be used for audio source separation and for abnormal sound detection. In these systems, automated algorithms determine whether sounds generated during machine operation are normal or abnormal.

[0015] Some implementations are based on the understanding that it is natural to equate the state of a system with the execution state of a task, since it is usually possible to simply measure or observe the state of the system, and if enough observations are collected, the state of the system can indeed represent the execution state. However, in some cases, the state of the system is difficult to measure, and the execution state of a task is difficult to define.

[0016] For example, consider Computer Numerical Control (CNC) for machining workpieces with cutting tools. The state of the system includes the state of the actuators that move the cutting tool along the toolpath. The machine's execution state is the actual cutting state. Industrial CNC systems can have multiple different, and sometimes redundant, actuators with multiple state variables of complex nonlinear relationships, each of which is difficult to observe in terms of the CNC system's state. However, it may also be difficult to measure the machining state of the workpiece.

[0017] Some implementations are based on the understanding that the execution state of a task can be represented by acoustic signals generated through such execution. For example, the execution state of CNC machining of a workpiece can be represented by acoustic signals caused by the deformation of the workpiece during machining. Therefore, if the acoustic signals can be measured, various classification techniques, including machine learning methods, can be used to analyze the acoustic signals to estimate the execution state of the task and select appropriate control actions to control the execution.

[0018] However, this approach faces the problem of a lack of isolation in the acoustic signals. For example, in a system comprising multiple actuators that actuate one or more tools to perform one or more tasks, the acoustic signals generated by the tools performing the tasks are always mixed with signals generated by the actuators actuating the tools. For instance, the acoustic signals generated by workpiece deformation are always mixed with signals from the motor of a moving cutting tool. If the acoustic signals could be synthesized or captured in an isolated manner in some way, machine learning systems such as neural networks could be trained to extract such signals. However, in many cases, including CNC machining, generating isolated signals representing task execution is impractical. Similarly, recording different signals separately using multiple microphones may also be impractical.

[0019] However, some implementations are based on the understanding that, for certain machines or complex industrial systems, only normal data can be used for training, and domain offsets can cause changes in the sound signal regardless of the presence or absence of anomalous sounds. In the example, for complex industrial systems or factory automation setups, isolated sounds for each component may not be available. Furthermore, some implementations are based on the understanding that existing systems for machine sound analysis or sound separation treat recorded sounds as generated by individual machine components.

[0020] Some implementations are based on the understanding that most industrial machinery consists of multiple sound-generating components, and the health of each component may need to be monitored individually. For example, when a machine includes multiple actuators for operating one or more tools and / or performing one or more tasks, the audio sounds generated by the machine include the sounds of all these actuators and / or tools. In this case, audio source separation can be a useful preprocessing step to isolate the sound from each machine component. Furthermore, the separated sound signals from each component can then be used, for example, for abnormal sound detection or other types of audio monitoring.

[0021] Some implementations are based on the understanding that deep learning-based audio source separation techniques utilize a fully supervised framework. In such techniques, a database of isolated sound signals is used to create an artificially mixed signal, and the source separation model is trained to separate this artificially mixed signal using a ground-based isolated signal as the target. However, some implementations are based on the understanding that collecting a database of isolated sound signals from individual machine components (i.e., audio sources) may be impossible for separating the audio signals from the machine itself, as all components may need to operate simultaneously to keep the machine running.

[0022] Some implementations are based on the understanding that unsupervised audio separation algorithms are limited to applications where all audio sources to be separated are isolated (i.e., have different spectral characteristics). However, unsupervised audio separation algorithms cannot separate related sources, such as music signals.

[0023] Some implementations are based on the understanding that beamforming techniques can be used for sound separation to separate multi-channel speech signals in a fully supervised setting with prior information conditioned on the source location. For example, beamforming can be used to map and measure the direction of sound arrival. In examples, beamforming can be used to obtain a form of array-based measurement and for source location mapping, particularly from medium to long distances. For example, the source location can be determined by estimating the amplitude of a plane wave at the angle of incidence in a specific direction.

[0024] However, some implementations are based on the understanding that beamforming-based sound separation is suboptimal due to the high levels of noise present in industrial or factory settings and the desire to isolate close-proximity machine components. For example, beamforming-based sound separation only achieves a slight reduction in background noise without separating the acoustic signals of individual machine components. Therefore, beamforming-based sound separation cannot reliably and accurately output isolated sound signals corresponding to each component of a machine in a complex industrial system.

[0025] To overcome the aforementioned drawbacks and achieve accurate and reliable sound separation, embodiments of this disclosure provide a multi-channel sound separation system. Some embodiments are based on the recognition that a multi-channel spectrogram of an audio mixture carries information about the relative positions (i.e., distance and orientation) between the different audio sources forming the audio mixture and the microphone array measuring the audio mixture. For example, this information can indicate the inter-channel phase difference between channels in the multi-channel spectrogram of the audio mixture.

[0026] Some implementations are based on the recognition that directional information can be used to separate signals that form mixed audio signals that can be generated by a machine. In examples, a machine may include one or more categories of components, such as motors, tools, chains, bearings, and other machine parts. Conventional sound separation techniques rely on the audio characteristics of specific categories of components to identify and extract audio signals that can be generated by those specific categories. However, in industrial environments, heavy machinery may include a large number of components and multiple components belonging to the same category. For example, a machine may include multiple motors, multiple tools, etc., making it impossible for conventional sound separation techniques to separate audio signals that may be generated by different components belonging to the same category (such as different motors of a machine).

[0027] Some implementations are based on the understanding that separating and extracting audio signals from various components (or audio sources) of a machine is useful for anomaly detection. In the example, the extracted audio signals belonging to different motors of the machine can be used to detect anomalies in those motors. It is worth noting that conventional audio separation techniques cannot identify or separate audio signals belonging to the same category of components, such as audio signals from different motors of the machine. Therefore, conventional anomaly detection or health monitoring techniques may not reliably monitor the health of individual machine components or determine their anomalies.

[0028] Some implementations are based on the understanding that the spatial location of the machine's audio source can be used to identify and extract audio signals from different audio sources belonging to the machine, regardless of the type of component.

[0029] Furthermore, some implementations are based on the understanding that spatial location information of different audio sources (such as components belonging to the same or different categories) is useful for anomaly detection of various components of the machine.

[0030] This is advantageous because the audio sources (i.e., components) that form the audio mixture of machine operation signals are often not possible to record individually. For example, separating the sound of the tool performing the task from the sound of the motor moving the tool during machine operation can be challenging and / or impractical to record. Instead, the relative or spatial positions of the machine's tools and / or actuators are usually known, so these relative or spatial positions can be used to facilitate audio signal separation even without recording the individual sounds that form the mixture signal.

[0031] Some implementations are based on the recognition of the advantages of target phase correlation spectrograms in training neural networks for subsequent audio signal separation. Based on this understanding, some implementations use target phase correlation spectrograms as training targets when individual source signals cannot be collected to train deep neural networks.

[0032] Some implementations are based on the understanding that since the structure of a machine is generally known, and the audio source or component that generates the audio signal is also generally known, weak labels corresponding to the spatial locations of the audio source or component that may generate audio mixture signals during machine operation can be provided. To this end, some implementations have developed methods capable of learning to separate sounds in an audio mixture signal when training data of only weakly labeled audio mixture signals is available.

[0033] Therefore, some implementations train neural networks to separate one or more audio signals generated by one or more audio sources (such as tools performing corresponding tasks and one or more actuators of actuating tools) from an audio mixture. For example, the neural network is trained to separate different audio signals from the audio mixture such that, during machine operation, each separated audio signal belongs to an audio source having a corresponding relative distance from the microphone array and / or may belong to a different category of signal, and the sum of all isolated signals constitutes the audio mixture. Weak labels identify the relative position and / or category of signals present during machine operation.

[0034] Therefore, to simplify the automation of industrial systems, such as monitoring the health of each component of a machine in an industrial system, it is necessary to separate the audio signals from the audio mixture signals of each audio source. Thus, some implementations aim to train neural networks to perform sound separation of the audio signals from audio sources that form mixed signals, even in the absence of isolated audio signals from such audio sources. As used herein, audio sources that generate mixed signals or audio mixture signals can occupy different relative positions or spaces in the corresponding environment.

[0035] Therefore, one embodiment discloses that the machine includes one or more actuators that assist one or more tools in performing one or more tasks. The audio system includes an audio input interface configured to receive an audio mixture signal generated by a plurality of audio sources, the plurality of audio sources including at least one of the following: the one or more tools performing one or more tasks, or the one or more actuators operating the one or more tools. At least one of the audio sources forming the audio mixture signal is identified by its relative position to the position of each microphone in a microphone array measuring the audio mixture signal. The audio system includes a processor configured to extract an audio signal generated by the identified audio sources from the audio mixture signal based on the correlation between spectral characteristics in a multi-channel spectrogram of the audio mixture signal and directional information indicating the relative position of the identified audio sources among the plurality of audio sources. The audio system includes an output audio interface configured to output the extracted audio signal to facilitate the operation of the machine.

[0036] According to an additional system implementation, the processor is configured to use a neural network to extract the audio signal generated by the identified audio source.

[0037] According to an additional system implementation, the spectral features may include inter-channel phase differences between channels in the multi-channel spectrogram of the audio mixture. In an example, the directional information includes the target phase difference (TPD) of sound propagating from the relative position of the identified audio source to different microphones in the microphone array. Furthermore, the correlation between the spectral features and the directional information is represented by a target phase correlation spectrogram. For example, the values ​​of different time-frequency intervals of the target phase correlation spectrogram quantify the alignment of the inter-channel phase differences with the target phase difference in the corresponding time-frequency interval, and the target phase difference is the expected phase difference in the time-frequency interval that indicates the characteristics of sound propagation.

[0038] According to an additional system implementation, the processor is further configured to: determine the target phase correlation spectrogram; and process the target phase correlation spectrogram using the neural network to extract the audio signal.

[0039] According to the additional system implementation, the target phase correlation spectrum includes complex numbers, and the neural network is a complex neural network for processing the complex numbers of the target phase correlation spectrum.

[0040] According to the additional system implementation, the complex neural network has a complex U-net architecture.

[0041] According to an additional system implementation, the complex U-net architecture includes: a complex convolutional encoder; a complex bidirectional long short-term memory (BLSTM) network module, the complex BLSTM module being arranged to process the output of the complex convolutional encoder; and a complex convolutional decoder, the complex convolutional decoder being arranged to process the output of the complex convolutional encoder and the output of the complex BLSTM module.

[0042] According to an additional system implementation, the neural network is trained to extract signals from multiple identified audio sources. In an example, the complex U-net architecture includes at least one complex convolutional decoder for each identified audio source.

[0043] According to an additional system implementation, the processor is further configured to: use the neural network to determine a target phase correlation spectrum of the audio mixture signal.

[0044] According to an additional system implementation, in order to train the neural network, the processor is further configured to receive a training audio mixture signal generated by one or more audio sources, said one or more training audio sources including at least one of the following: one or more tools performing one or more tasks, or one or more actuators operating said one or more tools. In an example, at least one of the one or more training audio sources forming the training audio mixture signal is identified by positional data relative to the position of each microphone in the microphone array measuring the training audio mixture signal. The processor is also configured to generate one or more training target phase correlation spectrograms associated with corresponding training audio sources, said one or more training target phase correlation spectrograms being generated based on the correlation between the spectral characteristics of the training audio mixture signal and the directional characteristics indicating the positional data of the one or more training audio sources forming the training audio mixture signal. For example, each time-frequency (TF) interval of the one or more training target phase correlation spectrograms defines a feature that quantifies the match between the observed inter-channel phase difference and the corresponding expected phase difference in the spectral features of the measured training audio mixture signal, the corresponding expected phase difference indicating the sound propagation characteristics of corresponding location data of the one or more training audio sources relative to the position of each microphone in the microphone array. The processor is also configured to train the neural network to extract training audio signals corresponding to the one or more training audio sources based on the corresponding one or more training target phase correlation spectrograms. For example, for components with known locations, maximizing the target phase correlation is emphasized as the loss function for neural network training.

[0045] According to an additional system implementation, in order to train the neural network, the processor is further configured to train the neural network based on a set of loss functions. In an example, the set of loss functions includes at least one of the following: a positional loss function corresponding to each of the separated training audio signals of the training audio source, or a reconstruction loss function associated with the sum of the extracted training audio signals used to reconstruct the training audio mixture signal.

[0046] According to an additional system implementation, in order to calculate the position loss function, the processor is configured to calculate an ideal target phase correlation spectrogram using the physical characteristics of sound propagation of the one or more training audio sources with associated position data. The processor is configured to calculate an estimated training target phase correlation spectrogram associated with a corresponding training audio source. The one or more estimated training target phase correlation spectrograms are generated based on the correlation between spectral features associated with corresponding separated training audio signals and directional features indicating the position data of the one or more training audio sources forming the training audio mixture signal. Furthermore, each time-frequency (TF) interval of the one or more estimated training target phase correlation spectrograms defines a feature that quantifies the match between the inter-channel phase difference observed in the spectral features of the corresponding separated training audio signal and a corresponding expected phase difference, which indicates the sound propagation characteristics of the corresponding position data of the one or more training audio sources relative to the position of each microphone in the microphone array. The processor is configured to determine the difference between the estimated training target phase correlation spectrogram and the corresponding ideal target phase correlation spectrogram for each of the one or more training audio sources, wherein the difference indicates the position loss function.

[0047] According to an additional system implementation, in order to train the neural network, the processor is further configured to collect the training audio mixture signal generated by the one or more training audio sources by moving the microphone array at different locations near the machine.

[0048] According to an additional system implementation, the processor is further configured to: transform the received audio mixture signal using a Fourier transform to generate a multi-channel short-time Fourier transform (STFT) of the received audio mixture signal; determine the inter-channel phase difference (IPD) between different channels of the multi-channel STFT. The processor is further configured to: determine the target phase difference (TPD) of sound propagating from the relative position of an identified audio source to different microphones in the microphone array; correlate the IPD with the TPD to generate a target phase correlation spectrogram; combine the target phase correlation spectrogram with the multi-channel STFT and frequency position encoding to generate a channel concatenation of the received audio mixture signal; and process the channel concatenation of the received audio mixture signal using a neural network to extract the audio signal. For example, the values ​​of the target phase correlation spectrograms in different time-frequency intervals quantify the alignment of the IPD with the TPD in the corresponding time-frequency interval, and wherein the TPD is the expected phase difference of the time-frequency interval indicating the sound propagation characteristics.

[0049] According to an additional system implementation, in order to determine the IPD, the processor is further configured to compare complex values ​​of different channels with a reference channel in the multi-channel STFT to generate an inter-channel phase difference (IPD) with respect to a reference microphone in the microphone array. Furthermore, the processor is configured to represent the IPD as a complex number. For example, each of the complex numbers has a real part indicating the cosine of the corresponding phase difference and an imaginary part indicating the sine of the corresponding phase difference, to generate the complex conjugate of each of the represented complex IPDs.

[0050] According to an additional system implementation, in order to determine the TPD, the processor is configured to: calculate the target phase difference (TPD) between different channels relative to the reference channel when sound propagating from an identified source in the audio mix arrives at the different channels, based on the position values ​​of the machine and the microphone in the microphone array. Furthermore, the processor is configured to represent the TPD as a complex number. For example, each of the complex numbers has a real part indicating the cosine of the corresponding expected target phase difference from the generated target phase difference and an imaginary part indicating the sine of the corresponding target phase difference, to produce the complex conjugate of each of the represented complex TPDs.

[0051] According to the additional system implementation, in order to determine the target phase correlation spectrum, the processor is configured to: calculate the product of each of the complex conjugates of the complex IPD and the corresponding complex conjugate of the complex TPD for each time-frequency interval, and determine the sum of each product on all non-reference channels.

[0052] According to an additional system implementation, the processor is further configured to: generate control commands for operating the machine based on the extracted signals; and transmit the control commands to the machine via a communication channel.

[0053] According to an additional system implementation, the processor is further configured to: analyze an extracted audio signal generated from the audio mixture by the identified sound sources to generate an execution state of a task; select a control command from a set of control commands based on the execution state of the task; and cause the machine to execute the control command. In the example, the set of control commands corresponds to different execution states of the one or more tasks.

[0054] According to an additional system implementation, the processor is further configured to: determine an anomaly score for the identified sound source based on extracted audio signals corresponding to the identified audio source. In an example, the anomaly score indicates the correlation between the anomaly type and the state of the identified sound source. The processor is also configured to: compare the anomaly score with an anomaly threshold; when the anomaly score is greater than the anomaly threshold, select a control command from a set of control commands to be executed by the machine; and send the selected control command to the machine to overcome the anomaly at the identified sound source.

[0055] Another embodiment discloses a system for facilitating the operation of a machine, the machine including one or more actuators that assist one or more tools in performing one or more tasks. The system includes: a processor; and a memory storing instructions that cause the processor to: receive an audio mixture signal generated by a plurality of audio sources, the plurality of audio sources including at least one of: the one or more tools performing the one or more tasks, or the one or more actuators operating the one or more tools; extract audio signals generated by the identified audio sources from the audio mixture signal based on a correlation between spectral features in a multi-channel spectrogram of the audio mixture signal and directional information indicating the relative positions of identified audio sources among the plurality of audio sources; and output the extracted audio signals to facilitate the operation of the machine. For example, at least one of the audio sources forming the audio mixture signal is identified by its relative position to the position of each microphone in a microphone array measuring the audio mixture signal.

[0056] Another embodiment discloses a method for facilitating the operation of a machine, the machine including one or more actuators that assist one or more tools in performing one or more tasks. The method includes: receiving, using an audio input interface, an audio mixture signal generated by a plurality of audio sources, the plurality of audio sources including at least one of the following: one or more tools performing one or more tasks, or one or more actuators operating the one or more tools. For example, at least one of the audio sources forming the audio mixture signal is identified by its relative position to the position of each microphone in a microphone array measuring the audio mixture signal. The method includes: using a processor, extracting an audio signal generated by an identified audio source from the audio mixture signal based on the correlation between spectral characteristics in a multi-channel spectrogram of the audio mixture signal and directional information indicating the relative position of the identified audio source among the plurality of audio sources. The method includes outputting the extracted audio signal using an output audio interface to facilitate the operation of the machine.

[0057] The embodiments disclosed herein will be further explained with reference to the accompanying drawings. The drawings are not necessarily drawn to scale, but generally focus on illustrating the principles of the embodiments disclosed herein. Attached Figure Description

[0058] [ Figure 1 ]

[0059] Figure 1 A block diagram is shown of an environment in which an audio system for extracting audio signals from a mixed audio signal is implemented, according to an example embodiment.

[0060] [ Figure 2A ]

[0061] Figure 2A A flowchart illustrating the input of a neural network for audio signal extraction in an audio system according to an example embodiment is shown.

[0062] [ Figure 2B ]

[0063] Figure 2B A block diagram of an audio system according to some embodiments is shown, which is used to extract an audio signal from an audio mix during machine operation and to use the extracted signal for subsequent monitoring tasks.

[0064] [ Figure 3A ]

[0065] Figure 3A A machine capable of generating mixed audio signals according to an example implementation is shown.

[0066] [ Figure 3B ]

[0067] Figure 3B A spectrogram of an audio mixture signal generated by a machine performing a task, according to some implementations, is shown.

[0068] [ Figure 4A ]

[0069] Figure 4A Example illustrations are shown of phase angle differences in microphone arrays used to calculate target phase angle difference (TPD) and corresponding IPD, according to some implementation methods.

[0070] [ Figure 4B ]

[0071] Figure 4B Example illustrations are shown of phase angle differences in microphone arrays used to calculate target phase angle difference (TPD) and corresponding IPD, according to some implementation methods.

[0072] [ Figure 5A ]

[0073] Figure 5A Example illustrations of the ideal target phase correlation spectrum and the calculated target phase correlation spectrum according to some implementation methods are shown respectively.

[0074] [ Figure 5B ]

[0075] Figure 5B Example illustrations of the ideal target phase correlation spectrum and the calculated target phase correlation spectrum according to some implementation methods are shown respectively.

[0076] [ Figure 6A ]

[0077] Figure 6A An exemplary flowchart is shown, according to some implementations, for training a neural network to extract audio signals corresponding to an audio source.

[0078] [ Figure 6B ]

[0079] Figure 6B An example of collecting training audio mixture signals according to some implementation methods is shown.

[0080] [ Figure 6C ]

[0081] Figure 6C An example of training a neural network based on a set of loss functions is shown according to some implementation methods.

[0082] [ Figure 6D ]

[0083] Figure 6D An example flowchart illustrating the calculation of a location loss function for training a neural network, according to some implementations, is shown.

[0084] [ Figure 7A ]

[0085] Figure 7A The network architecture of a neural network according to some implementations is shown.

[0086] [ Figure 7B ]

[0087] Figure 7B A flowchart illustrating the calculation of input features of a neural network according to some embodiments is shown.

[0088] [ Figure 7C ]

[0089] Figure 7C A flowchart illustrating the calculation of a target phase correlation spectrum for extracting an audio signal is shown according to some embodiments.

[0090] [ Figure 8 ]

[0091] Figure 8 A schematic diagram illustrating anomaly detection-based control of processing operations is shown according to some embodiments. Detailed Implementation

[0092] In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of this disclosure. However, it will be apparent to those skilled in the art that this disclosure may be practiced without these specific details. In other instances, systems and methods are shown only in block diagram form to avoid obscuring this disclosure.

[0093] Figure 1 A block diagram of an environment 100 according to some embodiments is shown, in which an audio system 106 is implemented for extracting audio signals from an audio mix 104. For example, the audio system 106 is configured to separate audio signals corresponding to different audio sources that co-generate the audio mix 104 during machine operation. According to some embodiments, the separated audio signals can be used for performance analysis and machine operation control.

[0094] In an implementation, microphone array 102 may be an array of multiple microphones configured to measure audio mixture signal 104. Specifically, microphone array 102 may include multiple microphones (e.g., two or more microphones) to record sound. The microphones in microphone array 102 may work together to record sound simultaneously. For example, the microphones of microphone array 102 may be similarly matched to ensure uniform and consistent sound recording. In this example, microphone array 102 may be configured to measure audio mixture signal 104 generated by the machine during its operation.

[0095] In the example, microphone array 102 captures a mixed audio signal 104 generated by the machine's various actuators and / or tools during machine operation. In some cases, some actuators or tools may not operate independently. For example, actuators assisting tools in performing a task may only operate together when working in conjunction with the tool. Subsequently, the audio signals generated by the execution of the actuators, tools, and / or tasks may only be captured simultaneously by microphone array 102. Therefore, the mixed audio signal 104 captured by microphone array 102 is an acoustic mixture signal, which consists of the sum of audio signals generated by different components of the machine performing one or more tasks. In some embodiments, at least some audio signals in the spectrogram of the mixed audio signal 104 may occupy the same time and spectrum.

[0096] In the example, some conventional sound separation techniques may be able to separate sounds based on audio characteristics such as frequency. For instance, different categories of audio sources can produce audio signals with different audio characteristics (e.g., different musical instruments or different people may produce audio signals with different audio characteristics). Therefore, based on the identification of the audio characteristics of different categories of audio sources when they operate individually (such as the audio characteristics of different instruments played individually) and knowledge of the categories of components that produce the mixed audio signal, audio signals belonging to different audio sources can be separated.

[0097] However, in some cases, the audio characteristics of an audio source may not be known individually. For example, isolated audio signals from different machine components that must operate simultaneously for the machine to function properly cannot be recorded separately. Therefore, the individual signals of each component may be unavailable and may hinder the training of neural networks used for audio signal separation or extraction.

[0098] Furthermore, the problem of audio source separation is exacerbated by the unavailability of isolated signals from different components of the machine and the presence of multiple components of the same type or category (such as two or more motors or two or more actuators). In this situation, extracting the audio signal generated by a specific audio source from the audio mixture becomes challenging because it is impossible to record the separated audio sources individually, and multiple audio signals may have the same or similar audio characteristics.

[0099] In practice, most industrial machinery may consist of multiple sound-generating components, and the health of each component may need to be monitored individually. In these cases, separate or isolated audio signals from each component may be required for downstream processing, such as abnormal sound detection or other types of audio monitoring. Once the audio signals from different audio sources of the machine are isolated from the audio mix 104, such audio signals can be used to estimate the execution status, health, and task execution of the corresponding audio source.

[0100] Some embodiments of this disclosure are based on the understanding that the presence of the microphone array 102 is known and the relative position of the machine's audio source (i.e., component) to the microphone array 102. In examples, the relative position of the audio source can be determined based on a schematic diagram of the machine and / or the environment in which the machine is located, or based on sensor data from an image sensor. For this purpose, the audio source forming the audio mixture signal 104 can be identified by the corresponding position (referred to as the audio source position) relative to the position of the microphone array 102 measuring the audio mixture signal 104 (hereinafter referred to as the microphone array position).

[0101] Some embodiments of this disclosure are based on the understanding that the location of an audio source can be used as a weak label for learning to separate audio signals from different audio sources.

[0102] Some implementations are based on the understanding that the location of the audio source relative to the microphone array location is useful for anomaly detection within the audio source.

[0103] In the example, audio system 106 may include an audio input interface 108 configured to receive a mixed audio signal 104 generated by multiple audio sources of the machine. For example, various components of the machine can generate sound, such as during control or task execution. In some cases, combinations of machine components can generate sound; for example, actuators assisting a tool in performing a task can combine to generate sound. Therefore, the various components that generate sound individually or in combination may correspond to multiple audio sources of the machine.

[0104] The audio system 106 may include a processor 110 configured to extract audio signals generated by identified audio sources from an audio mix 104. For example, the identified audio sources may be pre-selected from a plurality of audio sources in the machine, or they may be randomly selected. An audio input interface 108 may provide the audio mix 104 to the processor 110 for extracting the audio signals from the identified audio sources. Furthermore, the processor 110 may be configured to extract spectral features 112 from the audio mix 104. The processor 110 is also configured to determine directional information 114 related to the audio sources that generated the audio mix 104. For example, spectral features 112 may indicate a multi-channel spectrogram of the audio mix 104; and directional information 114 may indicate the relative position of the identified audio sources. Furthermore, the audio signals of the identified audio sources are separated based on the correlation between the spectral features 112 of the audio mix 104 and the directional information 114 of the identified audio sources.

[0105] In the example, directional information 114 can be used to identify different audio sources, regardless of the category of the audio source. Such directional information 114 can provide spatial location information of the audio source relative to the microphone array 102. This enables audio separation when isolated audio signals of the audio source are unavailable, and when multiple audio sources belonging to the same sound category (such as motors, actuators, etc.) may exist.

[0106] In the example, the audio system 106 includes a neural network configured to isolate or extract audio signals generated by identified audio sources from the audio mix 104. Once the audio signals from the identified audio sources and / or all audio sources are isolated and extracted from the audio mix 104, they can be used to monitor and / or control the operation of the machine.

[0107] Furthermore, the audio system 106 may include an output audio interface 116 configured to output the extracted audio signal to facilitate machine operation. For example, the extracted audio signal may be rendered or played back for further analysis, such as anomaly detection, determining the condition of identified audio sources, determining the execution status of tasks performed by the identified audio sources, etc. In one example, the extracted audio signal may be rendered as playback audio, for example, to analyze the health of the identified audio source. In another example, the extracted audio signal may be provided to a downstream processing system that can be used to analyze the extracted audio signal.

[0108] Figure 2A A flowchart 200 is shown illustrating the inputs of a neural network 202 for audio signal extraction according to some embodiments. In the example, the neural network 202 may be configured to extract audio signals that can be generated by the identified audio sources from different audio sources of the machine.

[0109] In the example, neural network 202 is configured to extract audio signals from audio mixture 104 based on the correlation between spectral features 112 and directional information 114 regarding the relative position of the identified audio sources. Notably, spectral features 112 can be frequency-based features in a multi-channel spectrogram of the audio mixture 104. Since the audio mixture 104 is recorded using multiple microphones of microphone array 102, a multi-channel spectrogram of the audio mixture 104 is extracted from the same signal recorded using different microphones. For example, spectral features 112 can be obtained by converting the time-based signal of the audio mixture 104 to the frequency domain using, for example, a Fourier transform. In the example, spectral features 112 can indicate, for example, fundamental frequency, frequency components, spectral centroid, spectral flux, spectral density, spectral roll-off, etc.

[0110] In one example, spectral feature 112 includes the inter-channel phase difference (IPD) between channels in a multi-channel spectrogram of the audio mixture signal 104. Inter-channel phase difference (IPD) refers to the phase shift or phase offset between audio signals from different channels (i.e., multiple channels of microphones in microphone array 102) in a multi-channel audio system. In this example, the IPD of the multi-channel spectrogram can indicate the time difference or alignment difference of waveforms between two or more audio channels or microphones. The time difference between microphones in microphone array 102 indicated by the IPD can provide information about the physical location of the audio source, as sound waves will arrive at different microphones at different times depending on the physical location of the identified audio source.

[0111] Furthermore, the directional information 114 can indicate the relative distance and relative direction between each microphone of the microphone array 102 and different audio sources. In an example, the directional information 114 may include the target phase difference (TPD) of the sound propagating from the relative position of the identified audio source to the different microphones in the microphone array 102.

[0112] In one example, TPD corresponds to the expected time delay of sound propagating from the identified audio source location to multiple channels of microphone array 102. TPD refers to the desired, ideal, or anticipated phase relationship between different audio channels or elements within the multiple channels of microphone array 102. For an identified audio source at a known physical location relative to microphone array 102, TPD corresponds to the IPD that can be expected from the identified audio source at the known physical location given the characteristics of sound propagation. Combined with, for example... Figure 4A Describe the details of TPD.

[0113] Furthermore, TPD can be correlated with IPD to generate a target phase correlation spectrum 204. In this example, the correlation between spectral features 112 and directional information 114 is represented by the target phase correlation spectrum 204. When IPD (Inter-channel Phase Difference) and TPD (Target Phase Difference) are correlated in audio, the actual phase relationship of the audio mixture 104 across the audio channels can be compared with the expected or desired phase relationship of different channels in the multiple channels of the microphone array 102.

[0114] For example, IPD can indicate the measured or actual phase difference between different audio channels, such as the phase difference between multiple microphones in microphone array 102 due to the spatial difference in the positions of multiple microphones and the machine's audio sources. Furthermore, TPD can indicate the target or ideal phase relationship of the machine's audio sources, such as a specific phase difference that might ideally exist due to the relative positions of the audio sources. In this example, IPD and TPD can be plural.

[0115] In the example, a target phase correlation spectrum 204 can be generated by comparing or correlated with the expected TPD. For example, the target phase correlation spectrum 204 can indicate whether the actual or measured phase relationship aligns with the target phase relationship corresponding to the TF interval. If the IPD closely matches the TPD, it indicates the presence of a single audio source observed by the microphone array 102, and its location is the physical location used to calculate the TPD. If there is a significant discrepancy, it may mean that the IPD is calculated based on an audio mixture rather than a single audio source. Combined with, for example... Figure 4B Further details on the association between IPD and TPD are described.

[0116] Furthermore, the values ​​of different time-frequency (TF) intervals in the target phase correlation spectrum 204 can quantify the alignment of IPD with the target phase difference in the corresponding TF interval; that is, the TF intervals of the target phase correlation spectrum 204 can quantify the alignment of IPD with the TPD of the associated TF interval. Moreover, given a specific source location, TPD can be the expected phase difference or expected phase relationship of the corresponding TF interval and the expected value of the corresponding IPD. Therefore, the target phase correlation spectrum 204 can refer to a visual representation or analysis of the phase relationship over time between different audio channels based on the expected phase relationship.

[0117] In the example, a target phase correlation spectrogram 204 can be determined based on the correlation or comparison of the measured phase difference of the IPD across multiple channels of the microphone array 102 or the expected or target phase difference of the identified audio source. In the example, the TF interval of the target phase correlation spectrogram 204 defines a feature that quantifies the degree of matching between the phase difference (i.e., inter-channel phase difference or IPD) observed in the measured spectrogram or spectral feature 112 and the directional information 114 or the expected phase difference (i.e., target phase difference or TPD). For example, when the measured IPD is correlated with the TPD corresponding to a given source location of the identified audio source (such as the source location of a component of a machine), the TF interval of the IPD that matches the TPD well is emphasized by the neural network 202 and / or used to separate audio signals that can be generated by the identified audio source.

[0118] Furthermore, to extract audio signals that can be generated from the identified audio sources, the target phase correlation spectrum 204, along with the measured spectral features 112, can be provided to the neural network 202. The target phase correlation spectrum 204 can be processed by the neural network 202 to extract the audio signals corresponding to the identified audio sources. The presence of the target phase correlation spectrum 204 as input to the neural network 202 is crucial for achieving any separation from the audio mixture signal 104. Otherwise, the neural network 202 may lack a conditional mechanism for indicating the location of the source to be separated. In conjunction with, for example... Figures 6A to 6D Describe the training details of neural network 202; and combine, for example... Figure 7A , Figure 7B and Figure 7C The method by which neural network 202 extracts the audio signal of the identified audio source using the target phase correlation spectrum 204 is described in detail.

[0119] Figure 2B A block diagram of an audio system 106 for extracting audio signals from an audio mix 104 during operation of machine 234 is shown. According to some embodiments, the audio system 106 can be configured to use the extracted audio signals for subsequent monitoring tasks.

[0120] In this example, audio system 106 may receive audio mix 104 from microphone array 102 via network 230. For example, processor 110 may be configured to process audio mix 104.

[0121] For this purpose, the audio system 106 may include various modules executed by the processor 110 to process the audio mixture signal 104 and control the operation of the machine 234. The processor 110 may use a neural network 202 to process the audio mixture signal 104 to extract audio signals generated by identified audio sources (e.g., tools performing tasks) from the audio mixture signal 104. For example, the processor 110 may be configured to use an anomaly detector 206 to analyze the extracted audio signals to identify anomalies in the identified audio sources and / or the execution status of the task performed by the identified audio sources. In some cases, the processor 110 may be configured to send control commands selected by the controller 208 based on anomalies and / or execution status of the task to the machine 234. For example, the selected control commands may be transmitted to the machine 234 via a control interface 224 for applications such as fault avoidance and maintaining smooth operation within the machine 234.

[0122] Audio system 106 may have multiple input interfaces 108 and output interfaces 116 for connecting audio system 106 to other systems and devices. For example, network interface controller (NIC) 220 is adapted to connect audio system 106 to network 230 via bus 218. Through network 230 (wirelessly or via cable), audio system 106 can receive audio mix signal 104 as an input signal. In some embodiments, human-machine interface (HMI) 216 within audio system 106 connects audio system 106 to keyboard 212 and pointing device 214, wherein pointing device 214 may include a mouse, trackball, touchpad, joystick, pointing stick, stylus, or touchscreen, etc. Through HMI 216 or NIC 220, audio system 106 can receive data, such as audio mix signal 104 generated during operation of machine 234.

[0123] The audio system 106 includes an output interface 116 configured to output separate acoustic signals corresponding to different audio sources of the audio mixture signal 104 generated during operation of the machine 234. In this example, the output interface 116 may also output the outputs of the anomaly detector 206 and / or the controller 208, i.e., anomaly scores of the audio sources of the machine 234 and / or control commands for the audio sources of the machine 234. For example, the output interface 116 may include a memory to store and / or output the separate audio signals or anomaly detector results. For example, the audio system 106 may be linked to a display interface 226 via a bus 218, which is adapted to connect the audio system 106 to a display device 228, such as a speaker, headphones, a computer monitor, a camera, a television, a projector, or a mobile device. The audio system 106 may also be connected to an application interface 222 adapted to connect the audio system 106 to a device 232 to perform various operations.

[0124] The audio system 106 includes a processor 110 configured to execute stored instructions and a memory 210 storing instructions executable by the processor 110. The processor 110 may be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. The memory 210 may include random access memory (RAM), read-only memory (ROM), flash memory, or any other suitable memory system. The processor 110 is connected to one or more input and output devices via a bus 218. These instructions implement methods for extracting audio signals generated during operation of the machine 234 for purposes such as anomaly detection, performance estimation, and future control.

[0125] Figure 3A A machine 234 capable of generating an audio mixed signal 104 is shown. In this example, machine 234 may include one or more actuators 302 (depicted as actuators) connected to and / or assisting one or more tools (depicted as tool 304). , and For example, tool 304 may be configured to perform one or more tasks associated with the operation of machine 234. Examples of tasks performed by actuator 302 and / or tool 304 in combination may include, for example, machining, cutting, welding, assembly, motion control, etc. In some embodiments, actuator 302 and tool 304 may operate simultaneously and / or in a specific combination. For example, actuator 302 may assist and / or control tool 304 individually or in combination. For example, in some embodiments, audio system 106 controls the operation of machine 234, including actuator 302 assisting tool 304 in performing one or more tasks, based on extracted audio signals corresponding to the respective audio sources (i.e., actuator 302 and tool 304).

[0126] In the example, the multiple audio sources of machine 234 may include, for example, one or a combination of a tool 304 performing one or more tasks and an actuator 302 operating the tool 304. In particular, either the actuator 302 and / or the tool 304 may produce sound together or in some combination. Thus, such actuators 302 and tools 304 together or in some combination can form the audio sources of machine 234. For example, the audio sources (i.e., actuators 302 and / or tools 304) may generate an audio mixture signal 104.

[0127] For example, during task execution, actuator 302 and / or tool 304 can generate corresponding audio signals. Therefore, actuator 302 and tool 304 can correspond to audio sources that generate the corresponding audio signals. For this purpose, the sum of the audio signals generated by each audio source (i.e., actuator 302 and tool 304) can form an audio mixed signal 104.

[0128] In this example, microphone array 102 can measure different audio signals generated by machine 234 (i.e., different components of machine 234) during its operation. The sum of these audio signals can form an audio mix signal 104. For example, microphone array 102 can provide audio mix signal 104 to audio system 106, specifically audio input interface 108. In some cases, audio system 106 may include memory 210, such that audio input interface 108 allows the received audio mix signal 104 to be stored in memory 210. Furthermore, audio input interface 108 can provide audio mix signal 104 to processor 110 for processing and extracting or separating audio signals from identified audio sources, such as actuators. The extracted audio signal is used to generate an audio signal. In some cases, the processor 110 can store the extracted audio signal in the memory 210. Furthermore, the output audio source 116 can be configured to render, output, or play the extracted audio signal. In the example, the output audio source 116 can retrieve the extracted audio signal from the memory 210 for output. In some other cases, the output audio source 116 can provide the extracted audio signal from the identified audio source for downstream processing, such as anomaly detection, health checks, etc. For this purpose, the neural network 202 utilizes the extracted audio signal to enable accurate association between the separated audio signal and its corresponding audio source. Furthermore, in this respect, using anomaly detection with an audio signal accurately mapped to the corresponding audio source can improve the accuracy of anomaly detection.

[0129] According to embodiments of this disclosure, spatial location information corresponding to the audio sources of machine 234 can be used to extract audio signals associated with each audio source. In this regard, in an example, processor 110 can be configured to generate a location map 306 indicating spatial information of several audio sources (i.e., actuator 302 and tool 304). In an example, location map 306 can be generated with the aid of a known schematic diagram of machine 234, a known schematic diagram of the environment in which machine 234 is placed, and / or based on sensor data. In one example, a sensor such as an image sensor or a camera can be used to scan the environment in which machine 234 is placed. Furthermore, based on the environmental scan, different components can be identified. In some cases, user input can be requested, for example, from a user operating machine 234, to accurately identify components of machine 234, i.e., actuator 302 and tool 304.

[0130] For this purpose, each known component of machine 234 can be labeled or marked with corresponding position data. Such position data could be the relative position of the component with respect to the position of each microphone in the microphone array 102 in the environment.

[0131] Furthermore, the orientation information 114 of the audio source of machine 234 can be extracted based on the position map 306 of machine 234. It is worth noting that this orientation information 114 of the audio source is crucial for separating audio signals belonging to different audio sources from the audio mixture signal 104.

[0132] refer to Figure 3B The spectrum 308 shows an audio mixture signal 104 generated by the actuator 302 of the machine 234 performing the task, according to some embodiments. The spectrum 308 can depict a time-frequency representation of the sound signal. In other words, the spectrum 308 can show the audio source forming the audio mixture signal 104. and The acoustic characteristics of the audio signal. It is worth noting that the spectrogram 308 of the audio mixture signal 104 is shown only for a single microphone in the microphone array 102. The spectrograms of the other microphones may look similar, but may have phase shifts caused by delays between microphones. In the example, when multiple actuators 302 can belong to the same category of audio sources, they can produce sound signals with similar characteristics, which can completely overlap on the spectrogram 308, such as... Figure 3B In and As shown. In this case, the signal from the source corresponding to the sound source will be used. The directional or spatial information of the microphone array 102 allows the neural network 202 to distinguish audio sources belonging to the same category. (i.e., different audio signals of actuator 302).

[0133] For example, the audio mixture signal 104 can be generated by actuator 302 actuating tool 304 and the tool performing a task. However, for ease of depiction and interpretation, the acoustic characteristics of actuator 302 are shown in the spectrogram. This should not be considered a limitation. The audio mixture signal 104 may also include audio signals generated by tool 304.

[0134] Some implementations are based on the understanding that audio signals from different audio sources can be isolated by individually identifying the characteristics of the sound produced by each audio source. For example, machine 234 includes... The actuator 302 is indicated, and when the characteristics of the sound produced by each actuator 302 in the isolated state are known, the audio signal produced by each audio source can be isolated. However, since some components of the machine 234 always operate in coordination, for example, the actuators... and They can work together to assist the operation of tool 304 and can operate together, thus the sound isolation characteristics of each component (i.e.) (This may be unknown.)

[0135] Furthermore, in the audio mixture generated during the operation of machine 234, different audio signals generated by different audio sources of machine 234 may overlap, for example, in time and frequency, as shown in spectrum diagram 308. According to this example, the audio characteristics of the audio signals generated by actuator 302 may be similar. For example, the frequencies of the audio signals generated by actuator 302 may be similar. Therefore, the audio signals may not be separable by conventional techniques.

[0136] Therefore, in order to simplify the operation of the machine 234, which includes the actuator 302, which includes the actuation tool 304 to perform one or more tasks, it is necessary to separate the audio signal from the audio mixture signal 104 generated by the combination of signals from different audio sources when performing the task. Therefore, the aim of some embodiments is to train a neural network for sound separation or audio signal extraction for the different audio sources that generate the audio mixture signal 104 in the absence of isolated sound for each audio source.

[0137] In some cases, at least some audio sources occupy the same time and / or frequency spectrum in the spectrogram 308 of the audio mixture signal 104. For example, in one embodiment, the audio sources forming the audio mixture signal 104... It can occupy different areas of an environment such as a room. Therefore, as shown in spectrum diagram 308, when the time-frequency characteristics are insufficient to isolate sound, spatial or positional information corresponding to actuator 302 can be identified. For example, identifying information corresponding to the actuator... Location Identify the actuator Location And identify the actuator Location Additionally, the acoustic mixture signal 104 can be measured solely from the output of multiple microphones (i.e., multiple channels) of the microphone array 102. In this respect, the multiple channels of the microphone array 102 can receive the audio mixture signal differently. Subsequently, position information... , and It can be used for audio signal separation.

[0138] Some implementations are based on the recognition that the multichannel spectrogram of the audio mix 104 carries information about the relative positions of the different audio sources forming the audio mix 104 and the microphone array 102 measuring the audio mix 104. For example, this information may indicate the inter-channel phase difference (IPD) between the channels in the multichannel spectrogram of the audio mix 104.

[0139] Some implementations are based on the understanding that such relative positional information can be used to separate the signals forming the audio mix 104 to facilitate the operation of the machine 234. This is advantageous because the audio sources of the audio mix 104 that forms the operation of the machine 234 are typically not recorded separately. For example, isolating the sound of the tool 304 performing the task from the sound of the actuator 302 or motor operating or assisting the tool 304 for separate recording may be challenging and / or impractical. Instead, the relative positions of the tool 304 and / or actuator 302 of the machine 234 may be known, and therefore can be used to facilitate audio signal separation even without recording of the isolated sounds forming the audio mix 104.

[0140] Some implementations are based on the understanding that because the relative positions of the tool 304 and / or actuator 302 of machine 234 are known, even when the sound signals completely overlap in time and frequency, such as Figure 3B C1 and C3 in the diagram can also use location information as labels for the audio sources of the audio mixture representing the operation of the machine. To this end, some implementations provide techniques for separating sound from the audio mixture, training the acoustic mixture data using only weak labels based on the location of the audio sources.

[0141] Therefore, some implementations train neural network 202 to separate audio signals generated by identified audio sources from audio mixture 104. For example, neural network 202 can be trained to separate different audio signals from audio mixture 104, such that each separated audio signal belongs to an audio source present in the operation of machine 234, and the separated audio signals constitute audio mixture 104. Weak tag identification determines the relative positions of the audio sources that generate the audio signals present in the operation of machine 234.

[0142] refer to Figure 4A An example illustration exists of the phase angle difference in the microphone array 102 used to calculate the target phase angle difference (TPD) 404 according to an embodiment. It is worth noting that directional information (i.e., target phase correlation (TPD)) can be calculated for a single microphone pair (i.e., 102a and 102b) of the microphone array 102 at a single time-frequency interval. However, this depiction of microphone pairs 102a and 102b is merely exemplary and should not be construed as limiting.

[0143] Figure 4A An exemplary machine 402 is shown having a component 402a (such as an actuator or tool) for generating sound. Such a component 402a can be an audio source. It is worth noting that the machine 402 depicted as having a single component 402a is merely exemplary and should not be construed as limiting. The machine 402 may include other components.

[0144] Regarding component 402a, the TPD 404 of the sound generated by component 402a can be calculated based on microphone pairs 102a and 102b. Specifically, the TPD 404 of component 402a corresponds to the time difference between when the sound reaches a reference microphone (such as microphone 102a in microphone array 102) and when the sound reaches a non-reference microphone (such as microphone 102b in microphone array 102). For this purpose, the TPD 404 is calculated using the characteristics of sound propagation between the known location of each audio source to be separated (i.e., component 402a) and the known location of each microphone in microphone array 102 (i.e., microphones 102a and 102b). Specifically, the time difference depends on the frequency value and speed of sound in the time-frequency interval from which the TPD is calculated.

[0145] In one example, TPD can be calculated using different microphone pairs at the same or different TF intervals. For example, different microphone pairs may include reference microphone 102a and any other non-reference microphones of microphone array 102. Such TPD can indicate the target, ideal, or expected phase difference in the propagation of sound generated by component 402a across different microphones (i.e., channels) of microphone array 102.

[0146] also, Figure 4BThe calculation of a single time-frequency interval in the target phase correlation spectrum 408 of TPD 404 and the corresponding IPD 406 according to some embodiments is shown. It is noteworthy that both IPD 406 and TPD 404 can be modeled as complex numbers. Furthermore, the target phase correlation spectrum 408 can indicate directional characteristics. In particular, the target phase correlation spectrum 408 can indicate the angular difference between the complex numbers of IPD 406 and TPD 404.

[0147] Figure 5A and Figure 5B Example illustrations of an ideal target phase correlation spectrum 502 and a calculated target phase correlation spectrum 504, according to some embodiments, are shown respectively. It is worth noting that the values ​​of the target phase correlation spectra 502 and 504 are complex numbers; however, for illustrative purposes and simplicity, they are not shown as such. Figure 5A and Figure 5B Only the real part is shown.

[0148] In the example, target phase correlation spectrograms 502 and 504 may include multiple time-frequency intervals (described as TF intervals 506a, 506b, 506c, 506d, 506e, 508a, 508b, 508c, and 508d; collectively referred to below as TF intervals 506 and 508). For example, a TF interval from TF intervals 506 and 508 refers to a specific unit or element representing a local segment of time and frequency within the corresponding visualization of spectrogram 502 or 504. For example, target phase correlation spectrograms 502 and 504 may be a 2D representation of a signal, such as the frequency content of an audio mix signal 104 over time, where the x-axis represents time and the y-axis represents frequency. Furthermore, the color or intensity of each of TF intervals 506 and 508 represents the magnitude of the signal energy at that time and frequency.

[0149] In this regard, for example, the TF intervals 506a, 506b, 506c, 506d, and 506e (depicted as black squares) in the calculated target phase correlation spectrum 502 and the ideal target phase correlation spectrum 504 can indicate that the values ​​of the target phase correlation spectrum at these intervals reach a maximum value based on the total number p of microphones present in the microphone array 102. In the example, the amplitude measurement of the target phase correlation spectrum calculated from the audio mixed signal measures the degree of match between the measured IPD and the expected TPD (e.g., component 402a) corresponding to the source location of the identified audio source. When the IPD and TPD of a given TF interval match perfectly, it will take the maximum value corresponding to TF intervals 506a, 506b, 506c, 506d, and 506e, depicted as a black square.

[0150] refer to Figure 5AThe ideal target phase correlation spectrum 502 can correspond to, for example, an isolated signal from an audio source. Specifically, each TF interval (such as TF intervals 506a, 506b, and 506c in the ideal target phase correlation spectrum 502) is represented as being at its maximum value, i.e., the corresponding IPD based on the total number of microphones p. The maximum value is equal to the total number of microphones p because the IPD and corresponding TPD can be calculated for each pair of microphones p, and the maximum value of the target phase correlation spectrum for a single pair of microphones is equal to 1. Furthermore, after summing over p pairs of microphones, the maximum value of the target phase correlation spectrum over all p pairs of microphones can be equal to p.

[0151] refer to Figure 5B The TF intervals 508a, 508b, 508c, and 508d in the calculated target phase correlation spectrum 504 can indicate that the corresponding IPD does not match well with the expected TPD corresponding to the source location of the identified audio source. In the example, the TF intervals 508a, 508b, 508c, and 508d, depicted as gray squares, can indicate that the corresponding IPD is less than p. In the example, the gray TF intervals 508a, 508b, 508c, and 508d can indicate that the amplitude of the target phase correlation spectrum calculated from the audio mixture signal contains measured IPDs that do not match well with the expected TPD (e.g., component 402a) corresponding to the source location of the identified audio source, i.e., they are less than the maximum values ​​at the time and frequency points corresponding to the TF intervals 508a, 508b, 508c, and 508d. In particular, only some TF intervals 506c and 506d, which are depicted as black squares, are at their maximum values ​​in the target phase correlation spectrum, while other TF intervals 508a, 508b, 508c, and 508d, which are depicted as gray squares, are not at their maximum values ​​because other audio sources corresponding to different TPDs may dominate at the corresponding time and frequency, and noise and reverberation may also affect the amplitude.

[0152] Some embodiments of this disclosure are based on the understanding that when the isolation signal corresponding to the identified audio source is unavailable, the calculated target phase correlation spectrum 504 of the identified source or component 402a should be presented as close as possible to the ideal target phase correlation spectrum 502, so that the calculated target phase correlation spectrum 504 can be used as a metric for evaluating the quality of the source separation model.

[0153] In particular, TF intervals 506d and 506e corresponding to the black squares can be emphasized for extracting the audio signal of the identified audio source (e.g., component 402a) from the audio mixture signal, while TF intervals 508a, 508b, 508c, and 508d corresponding to the gray squares can be left unemphasized. The calculated target phase correlation spectrum 504 can be provided as input to neural network 202, which separates the audio sources corresponding to known locations and extracts the audio signals that can be generated by the audio sources.

[0154] Figure 6A An exemplary flowchart for training a neural network 202 to extract audio signals corresponding to an audio source, according to some embodiments, is shown. In the example, the neural network 202 is trained using a training dataset based on a target phase correlation spectrogram. For example, it is advantageous to train the neural network 202 using a target phase correlation spectrogram for subsequent audio signal separation because the target phase correlation spectrogram provides useful runtime features as network input due to its strong cues for separating the audio source at known locations.

[0155] At position 602, a training audio mixture signal can be received. For example, such a training audio mixture signal can be generated by one or more training audio sources. Training audio sources can include one or more tools performing one or more tasks and / or one or more actuators operating one or more tools. For example, such a training audio source can be part of a machine. Combined with, for example... Figure 6B It describes a method for collecting mixed training audio signals from training audio sources.

[0156] In this example, the training audio mix can be a combination of multiple sound sources or an acoustic mix of audio signals from multiple sound sources in a corresponding acoustic environment. For example, each acoustic signal in the training audio mix can have its own unique characteristics and spatial location. According to this example, one or more training audio sources forming the training audio mix can be identified by positional data relative to the position and orientation of a microphone array (such as microphone array 102) measuring the training audio mix.

[0157] Therefore, separating or isolating individual sound sources or audio signals from a training audio mix is ​​a challenging task, especially when the individual sound signals of one or more training audio sources are unavailable. For example, separating or isolating corresponding audio sources within an audio mix can be used for purposes such as noise reduction, anomaly detection, and audio source maintenance.

[0158] At position 604, a phase correlation spectrogram of one or more training targets associated with the corresponding training audio sources is generated. For example, one or more training target phase correlation spectrograms can be generated based on the correlation between the spectral characteristics of the training audio mixture and the directional characteristics of the location data indicating one or more training audio sources that form the training audio mixture.

[0159] In the example, the spectral characteristics of the training audio mixture signal may include the inter-channel phase difference of the training audio mixture signal. For example, the inter-channel phase difference of the training audio mixture signal may indicate the phase difference, such as delay, in the measured training audio mixture signal across different channels (i.e., microphones) of the microphone array 102. Furthermore, the directional characteristics of the training audio mixture signal may include the target or expected phase difference of a given audio source from which an audio signal is to be extracted. For example, the target phase difference of a given audio source may indicate the target or expected phase difference between the time when sound from the given audio source should arrive at a reference channel of the microphone array 102 and the time when sound should arrive at a non-reference channel of the microphone array 102. Notably, the target phase difference of a given audio source from one or more training audio sources can be calculated based on sound propagation characteristics. For this purpose, the target phase difference of a given audio source may indicate, for example, the sound propagation characteristics of known location data of the given audio source relative to the known location of each microphone in the microphone array 102. In the example, the target phase difference can be calculated for each of the one or more training audio sources that form the audio mixture signal and from which an audio signal is to be extracted or separated.

[0160] Furthermore, the inter-channel phase differences of the training audio mixture signal can be correlated with the target phase differences of a given audio source to calculate the target phase correlation spectrum of the given audio source. For example, based on the correlation between the inter-channel phase differences of the audio mixture signal and the different target phase differences of one or more audio sources, one or more corresponding training target phase correlation spectra can be calculated.

[0161] To this end, each time-frequency (TF) interval of the training target phase correlation spectrogram defines a feature that quantifies the match between the inter-channel phase difference observed in the measured spectrogram of the training audio mixture at a specific time and frequency and the corresponding expected phase difference of a given audio source at that specific time and frequency. Notably, the TF intervals of one or more training target phase correlation spectrograms can indicate the degree to which the phase difference observed in the measured spectrogram of the inter-channel phase difference matches or aligns with the corresponding expected or target phase difference of each of one or more training audio sources.

[0162] At 606, a neural network can be trained to extract training audio signals corresponding to a given audio source from one or more training audio sources forming a training audio mixture signal, based on one or more corresponding training target phase correlation spectrograms. For example, when the measured inter-channel phase difference is correlated with the target phase difference at a given source location of the given audio source—that is, in the target phase correlation spectrogram of the given audio source—TF intervals with inter-channel phase differences that match the target phase difference are emphasized, and TF intervals that do not match well are not emphasized, for training neural network 202. In the example, if the correlation magnitude between the inter-channel phase difference and the target phase difference in a TF interval is high, the neural network can utilize this information to emphasize the characteristics of that TF interval during training. Additionally or alternatively, the neural network is trained to train and execute using the correlation magnitude.

[0163] According to embodiments of this disclosure, when isolated source signals cannot be collected to train the deep neural network 202, the one or more training target phase correlation spectrograms can be used as training targets. In particular, for isolated audio sources at known locations, such as a given audio source at a given source location, the ideal target phase correlation spectrogram can be calculated using the physical properties of sound propagation and the measured inter-channel phase difference.

[0164] Furthermore, in some implementations, the difference between the calculated phase correlation spectrograms of one or more training targets and the ideal values ​​of the target phase correlation spectrograms can be used to generate a set of loss functions to train the neural network. Combined with, for example... Figure 6C and Figure 6D The details of this training of neural network 202 based on this set of loss functions are described.

[0165] Figure 6B An example illustration is shown of collecting training audio mixture signals according to some implementations. In this regard, a scenario is considered in which a machine 608 with two sound generation training audio sources (depicted as training audio sources 608a and 608b) exists in an acoustic environment.

[0166] Furthermore, in order to collect the training audio mixture signal, the microphone array 102 can be arbitrarily positioned, but with some constraints. For example, the microphone array 102 can be located at any position, described as the measurement location. and Such measurement locations and This indicates situations where, for example, due to obstacles in the acoustic environment, it may not always be possible to place the microphone array 102 in an ideal location (such as near machine 608 or between two training audio sources 608a and 608b). This ensures that the training audio mixture signal is similar to that measured under realistic conditions.

[0167] In this example, microphone array 102 may include 11 microphones. For example, the microphones in the microphone array may be harmonic-spaced. The spacing (in cm) between the microphones of microphone array 102 may be, for example, 16.8, 8.4, 4.2, 2.1, 2.1, 2.1, 2.1, 4.2, 8.4, or 16.8. Therefore, the total span or length of microphone array 102 may be, for example, 67.2 cm to 68 cm. It is important to note that these dimensions and harmonic spacing between microphones are merely exemplary and should not be construed as limiting. In other embodiments, the harmonic spacing may have different values, or the microphones may be uniformly spaced, or their placement may not be linear.

[0168] In addition, to collect the mixed training audio signals, the microphone array 102 can be located at the measurement position. and At the measurement location. and At this location, the audio mixed signal generated by the training audio sources 608a and 608b can be measured.

[0169] In addition, at the measurement location At this point, the relative position between the training audio source 608a and the reference microphone 102d of the microphone array 102 can be determined as follows: Furthermore, the relative positions between the training audio source 608b and the reference microphone 102d of the microphone array 102 can be determined as follows: Furthermore, at the measurement location At this point, the relative position between the training audio source 608a and the reference microphone 102d can be determined as follows: Furthermore, the relative position between the training audio source 608b and the reference microphone 102d can be determined as follows: Similarly, at the measurement location At this point, the relative position between the training audio source 608a and the reference microphone 102d can be determined as follows: Furthermore, the relative position between the training audio source 608b and the reference microphone 102d can be determined as follows: It is worth noting that the reference microphone 102d can be arbitrarily selected or predefined by the user, for example. In the example, the reference microphone 102d can be a microphone that can be located at the center of the microphone array 102. Furthermore, when the array geometry and array orientation of the microphone array 102 are known, the relative position between the training audio source 608b in the microphone array 102 and other non-reference microphones (such as microphones 102a and 102b) can be determined based on the relative position between the training audio source 608b and the reference microphone 102d. For this purpose, such relative positions can indicate the distance and direction information between the microphones of the microphone array 102 and the training audio sources 608a and 608b at the corresponding measurement positions.

[0170] For example, such relative positions (such as the relative positions between training audio source 608b and reference microphone 102d at different measurement locations, and the relative positions between training audio source 608b and non-reference microphone at different measurement locations) can be used to create training datasets for identifying the relative positions of training audio sources 608a and 608b from microphone array 102 and separating the training audio signals of training audio sources 608a and 608b based on spatial information. For this purpose, the mixed training audio signals generated by training audio sources 608a and 608b can be collected or measured by moving microphone array 102 at different locations near machine 608.

[0171] Once from different measurement locations in a real acoustic environment or a virtual acoustic environment (such as measurement location) and The training audio mixture signal is measured and can then be used to train the neural network 202. For example, the target phase correlation spectrogram corresponding to audio sources 608a and 608b can be processed by the neural network 202 to learn and extract the training audio signals from the training audio sources 608a and 608b.

[0172] Figure 6C An example illustration is shown of training a neural network 202 based on a set of loss functions according to some implementations. Notably, the microphone array 102 can measure the training audio mixture 610. The training audio mixture 610 can be the sum of audio signals generated by training audio sources 608a and 608b of the machine 608. The training audio mixture 610 (i.e., the spectral characteristics of the training audio mixture 610, including a subset of the complex short-time Fourier transform spectrograms, inter-channel phase difference spectrograms, and target phase correlation spectrograms of the known locations of the training audio sources 608a and 608b) can be processed by the neural network 202 to extract isolated training audio signals from the training audio sources 608a and 608b.

[0173] In the example, assume the training audio mix is ​​a P-channel audio mix, where It has a length of N samples, which are from the training audio source. 608a and audio source The sum of the reverberation isolation source signals of the 608b, i.e. The neural network 202 can process time-frequency (TF) mixed signals. As input, neural network 202 is configured to estimate for i=0,1 (i.e., corresponding to training audio sources 608a and 608b). .

[0174] In the example, neural network 202 can have a complex U-net architecture. It is worth noting that the U-Net architecture of neural network 202 can include an encoder ( Figure 6C (not shown in the image) and decoder ( Figure 6C (Not shown in the image), where corresponding layers between the encoder and decoder are connected via skip connections. Specifically, the encoder may include alternating complex convolutional layers and complex dense blocks, where each dense block of the encoder may have skip connections to a corresponding block in the decoder, and each dense block may be batch normalized. In the example, the encoder portion of U-Net resembles a convolutional neural network (CNN) architecture, such as including multiple convolutional and pooling layers that progressively reduce the spatial dimension of the input training target phase correlation spectrogram for the spatial location identified in each of the training audio sources 608a and 608b, while extracting high-level features from the training input features consisting of a subset of the complex short-time Fourier transform spectrogram, the inter-channel phase difference spectrogram, and the target phase correlation spectrogram.

[0175] Furthermore, for example, the decoder can be a multi-head architecture that can be configured to upsample the feature maps from the encoder and expand them back to the number of time-frequency intervals present in the spectrogram of the input training audio mixture. For example, the decoder can include a series of upsampling layers, such as upsampling layers using transposed convolution or interpolation, followed by convolutional layers, to generate dense prediction maps of the extracted training audio signals from the identified training audio sources 608a and 608b, which match the shape of the spectrogram of the input training audio mixture. Additionally, cross-layer connections can directly connect corresponding layers between the encoder and decoder, which can help preserve detailed spatial information from the encoder and improve the accuracy of segmenting the training audio signal from the audio mixture.

[0176] Furthermore, during training, the neural network 202 can be trained on, for example, blocks of STFT frames from the audio mixture, a phase correlation spectrogram of the training target, and / or other input features corresponding to a specific time frame. Once the neural network has been trained on the training subset or the training audio mixture, training results can be generated. The training results may include separate training audio signals 612a and 612b corresponding to training audio sources 608a and 608b, respectively.

[0177] If the isolated source signal is available, then the real source signal The complex TF representation can be used as a training target. However, for machine 608, the isolated source signals from training audio sources 608a and 608b are unavailable. When isolated sources or isolated signals are unavailable for training, some embodiments of this disclosure aim to ensure that the separated sources (i.e., training audio signals 612a and 612b) output by neural network 202 can reconstruct the training audio mixture signal 610. In this regard, verification of training audio signals 612a and 612b to reconstruct the training audio mixture signal 610 can be performed based on this set of loss functions.

[0178] In this example, the set of loss functions may include positional loss functions 614a and 614b, which correspond to the degree of matching between the target phase correlation spectrograms of the separated audio signal estimates 612a and 612b and the known positions of the audio sources 608a and 608b, respectively. For example, to ensure the separation of the training audio mixture signal 610, losses based on the target phase correlation spectrograms, i.e., positional loss functions 614a and 614b, are used. According to this example, separated source estimates are used... (That is, training audio signal estimates 612a and 612b) to calculate the position loss. In the example, the position loss functions 614a and 614b can be defined as:

[0179] in It uses the i-th estimated separation training audio signal 612a or 612b and the true position of the corresponding training audio source 608a or 608b. (Right now, and The target phase correlation spectrum is calculated. In one example, the position losses 614a and 614b that separate the training audio signals 612a or 612b ensure that the separated training audio signals 612a or 612b match the true positions, where When the inter-channel phase difference of the separated training audio signals 612a or 612b matches the target or expected phase delay of the known source positions of the training audio sources 608a and 608b. For example, the estimated output order of the separated training audio signals 612a or 612b can be explicitly determined by the order of the source positions input to the neural network 202. For example, position loss functions 614a and 614b are determined for each TF interval of the separated training audio signals 612a and 612b corresponding to the training audio sources 608a and 608b.

[0180] refer to Figure 6D An example flowchart illustrating the calculation of position loss functions 614a and 614b (collectively referred to as position loss function 614) for training neural network 202 according to some embodiments is shown. Notably, the multi-channel separated source spectrogram 618 includes the estimated separated training audio signals 612a and 612b. Furthermore, the source position and microphone position 620 can respectively indicate the known positions of the training audio sources 608a and 608b. and The known positions of the microphones in the microphone array 102, and the known positions of the microphones, can be used, for example, to calculate the target phase difference 622 using the sound propagation characteristics and the known positions of the training audio sources 608a and 608b and the microphones.

[0181] Furthermore, for example, the target phase difference 622 can be compared or correlated with the inter-channel phase difference 624 of the separated training audio signals 612a and 612b to obtain the directional characteristics or target phase correlation spectrum 626 of the separated training audio signals 612a and 612b.

[0182] In one example, the physical characteristics of sound propagation from training audio sources 608a and 608b can be used to calculate the ideal training target phase correlation spectrum. The distance and location of the training audio sources 608a and 608b are known. Furthermore, the difference between the target phase correlation spectrum calculated using the separated audio signals 612a and 612b and the corresponding ideal target phase correlation spectrum of each of the training audio sources 608a and 608b can be determined. For example, the difference in the target phase correlation spectrum can indicate the location loss functions 614a and 614b.

[0183] In the example, when the inter-channel phase difference 624 of the separated training audio signals 612a and 612b matches the expected phase delay or target phase difference 622 of the known source positions of the training audio sources 608a and 608b, the target phase correlation spectrum calculated for the position loss can be... Subsequently, position losses 614a and 614b (collectively referred to as position loss 614) can be calculated based on the 628 terms of squared error and the difference between the target phase correlation spectrum and the ideal target phase correlation spectrum.

[0184] return Figure 6C This set of loss functions may include reconstruction loss function 616. In the example, reconstruction loss function 616 may include two components: spectral loss and temporal spatial covariance loss. For example, the complex spectrum of the input audio mixture signal. and associated estimates The spectral loss can be defined as:

[0185] Furthermore, the temporal spatial covariance loss can be defined as:

[0186] Going further, This can be the estimated mixed signal reconstructed from the separated training audio signals 612a and 612b. To this end, the reconstruction loss function only ensures that the sum of the combined outputs of the neural network 202 (i.e., the training audio signals 612a and 612b) is approximately equal to the input training audio mixed signal 610. In the example, the reconstruction loss function... This can be used to verify whether the training audio mix 610 is completely reconstructed from the training audio signals 612a and 612b without any loss of audio content. For example, the reconstruction loss function 616 is associated with the sum of the extracted training audio signals 612a and 612b used to reconstruct the training audio mix 610.

[0187] The other total loss or loss function can be defined as:

[0188] in , and These are hyperparameters that weight each term of the loss function. For example, based on the loss function... The learning of neural network 202 can be verified. In some cases, neural network 202 can be retrained based on the same training dataset or a different training dataset. In this way, neural network 202 can be trained to extract or separate audio signals from audio mixtures. Once training is complete, neural network 202 can be deployed, for example, in industrial settings or environments with complex machinery, to separate audio signals from different audio sources operating in concert. Combined with, for example... Figure 7A , Figure 7B as well as Figure 7C and Figure 8To describe the details of how neural network 202 operates in different acoustic environments.

[0189] Figure 7A The network architecture of a neural network 202 according to some embodiments is shown. In this regard, a multi-channel audio mixture signal 702 (such as audio mixture signal 104) can be provided as input to the neural network 202. Furthermore, the known relative positions of the microphones of the microphone array 102 and the known relative positions of audio sources (such as actuators 302 and / or tools 304 of the machine 234 that can form the audio mixture signal 702) can be provided as input to the neural network 202. Such known positions are depicted as source and microphone positions 704.

[0190] In one implementation, a preprocessing step of feature extraction 706 can be performed on the multi-channel audio mixture signal 702 and the source and microphone locations 704. These extracted features at 706 can be provided as input to the neural network 202. Combined with, for example... Figure 7B Describe the details of feature extraction.

[0191] Some implementations are based on the understanding that a target phase correlation spectrogram can be used to train a neural network 202, which indicates a comparison between the spectral characteristics or inter-channel phase differences of the measured multi-channel audio mixture 702 and the target phase differences based on the source locations of the identified audio sources that form the multi-channel audio mixture 702. In an example, the neural network 202 may have a complex U-net architecture.

[0192] In this example, the U-Net architecture may include a complex convolutional encoder 708 (hereinafter referred to as encoder 708), a complex bidirectional long short-term memory (BLSTM) module 710, and complex convolutional decoders 1 712a and 2 712b (hereinafter collectively referred to as decoders 712). Although this example describes two complex convolutional decoders 712a and 712b, this should not be construed as limiting. In other examples, the U-Net architecture may include a scalable number of decoders, for example, based on the locations of multiple known audio sources forming the audio mixture to be separated.

[0193] For example, encoder 708 may include alternating complex convolutional layers and complex dense blocks, wherein each dense block of encoder 708 may have cross-layer connections to corresponding blocks in decoder 712, and each dense block may have batch normalization. During operation, encoder 708 may be configured to reduce the spatial dimension of input features, which may include multi-channel STFTs, IPDs, frequency position encoding, and targeted phase correlation spectrograms of the identified spatial locations of the identified audio sources for forming audio mixture signal 702, to identify and extract separated or isolated audio signals from the audio sources. Furthermore, a complex BLSTM 710 operates between encoder 708 and decoder 712. In the example, the BLSTM may be arranged to process the output of complex convolutional encoder 708.

[0194] Furthermore, for example, all convolutional layers of encoder 708 can use a stride of 1 in the time dimension and a stride of 2 in the frequency dimension, such that the output of encoder 708 or the input of BLSTM 710 transforms the input features from the input of BLSTM 710 into a shape of 256×T×1. Transformation. Furthermore, the number of frequency ranges can be... And the number of channels This can depend on the number of microphones and input features; for example, 11 microphones, 10 IPDs, 2 target phase correlation spectrograms, and 10 frequency position coding channels, totaling [number missing]. It is worth noting that only a subset of the input features, namely the frequency location code or the target phase correlation spectrogram, need not be included as input to the network.

[0195] In the example, each dense block of encoder 708 may contain cross-layer connections and batch normalization. Furthermore, decoder 712 may be a multi-head decoder because neural network 202 outputs separate audio signals for two audio sources rather than a single source from the audio mixture signal 702. For example, decoder 712 may be used to process the output of complex convolutional encoder 708 and the output of complex BLSTM module 710. In the example, neural network 202 may output the separate audio signals as a multi-channel complex spectrogram. For this purpose, neural network 202 may be configured to extract audio signals from multiple identified audio sources forming the audio mixture signal 702. For example, a complex U-net architecture includes at least one complex convolutional decoder for each identified audio source, such as two decoders for reconstructing the separate audio signals from the two corresponding audio sources.

[0196] It is worth noting that a key advantage of the U-Net architecture for processing the target phase correlation spectrogram is its ability to learn features of both local and global properties at multiple scales while still maintaining an output with the same shape as the input, such as a multi-channel complex spectrogram of an observed audio mixture signal. Each layer in the convolutional encoder 708 progressively learns essentially a global feature representation. The convolutional decoder 712 then reconstructs the corresponding separated audio signal of the audio source, starting from the global representation at the output of the encoder 708, with each decoder layer progressively learning a more local feature representation. Furthermore, cross-layer connections between each corresponding encoder 708 and decoder 712 at a common scale are crucial, allowing information from the encoder 708 to be re-injected into the decoder 712 for local and global details that should be present in the processed spectrogram (i.e., the U-Net output) of the separated audio signal of the identified audio source. These characteristics make the U-Net architecture a robust deep network architecture for audio source separation.

[0197] Using the multi-channel complex spectrogram of the audio mixture signal trained and observed based on the U-net architecture as input, the neural network 202 can be configured to output estimated multi-channel complex spectrogram part 1 and estimated multi-channel complex spectrogram part 2. For example, estimated multi-channel complex spectrogram part 1 can be a spectrogram corresponding to audio source 1 or a separated audio signal; and estimated multi-channel complex spectrogram part 2 can be a spectrogram corresponding to audio source 2 or a separated audio signal. It is worth noting that separating the multi-channel audio mixture signal 702 into estimated multi-channel complex spectrogram part 1 and estimated multi-channel complex spectrogram part 2 is merely exemplary.

[0198] refer to Figure 7B The diagram illustrates a flowchart of the calculation of input features for a neural network 202 according to some embodiments. The calculation of input features corresponds to... Figure 7A The feature extraction process involves 706 steps.

[0199] In this regard, the multi-channel audio mixture signal 702 measured by the microphone array 102, along with the known source and microphone locations 704, can be used for feature extraction 706. In an example, the multi-channel audio mixture signal 702 can be processed to compute a multi-channel short-time Fourier transform (STFT) 722. For instance, the measured multi-channel audio mixture signal 702 can be transformed to generate a multi-channel STFT 722 of the received audio mixture signal 702. In particular, the Fourier transform can be applied, for example, using a windowing function to small overlapping portions of the audio mixture signal 702. The resulting multi-channel STFT 722 can be used to extract local frequency information of the audio mixture signal 702. Based on the frequency information of the measured audio mixture signal 702, for example, a 2D representation or spectrogram showing the frequency variation over time can be generated. Therefore, the multi-channel STFT 722 corresponds to representing the audio mixture signal 702 as a signal for processing by the neural network 202.

[0200] Furthermore, based on the complex spectrum of the measured audio mixture signal 702, the inter-channel phase difference (IPD) 718 can be calculated. This IPD 718 may be caused by the phase difference between different channels of the microphone array 102 or within the microphones measuring the audio mixture signal 702. Therefore, the IPD 718 between different channels in the multi-channel STFT is determined. In an example, IPD 718 can represent the relative phase difference between each non-reference channel in the microphone array 102 and the reference channel in the microphone array 102. Combined with, for example... Figure 7C Details of IPD are described below. While the calculation of IPD and the target phase correlation spectrum is shown here only for the measured audio mixture signal 702, the same mechanism can be used to calculate the IPD and the estimated training target phase correlation spectrum of the estimated separated training audio signals 612a and 612b. For example, the estimated IPD and the estimated training target phase correlation spectrum of the separated training audio signals 612a and 612b can be used to determine losses, such as positional losses, which can be used to retrain the neural network and / or update the weights of the neural network.

[0201] Subsequently, using directional information indicating the position or location of the audio source and microphone array 102, the target phase difference (TPD) of the sound propagating from the relative position of the identified audio source to the different microphones in the microphone array can be determined. In the example, reference and non-reference microphones in microphone array 102 can be used to determine the TPD. Combined with, for example... Figure 7C The details of TPD are described.

[0202] In the example, a frequency location code 724 is determined to ensure that the neural network 202 can perform frequency-related processing in the early convolutional layers of the U-Net. To generate the frequency location code 724, the frequency interval index of the measured spectrogram of the audio mixture signal 702 can be represented as a 10-dimensional vector, for example, using a sine function with interval indices. In this way, the relative location information of the audio sources can be embedded in the frequency information of the measured spectrogram. In particular, the frequency location code 724 adds an additional layer of control and manipulation to the spectral characteristics of the audio mixture signal 702, thereby enabling more efficient audio processing for audio signal separation. Techniques used to generate the frequency location code 724 can include, but are not limited to, logarithmic scaling, perceptual weighting, or psychoacoustic modeling.

[0203] Continuing further, IPD 718 is associated with TPD 716 to generate a target phase correlation spectrum 720. For example, the values ​​of the target phase correlation spectrum for different time-frequency (TF) intervals quantize the alignment of the IPDs in the time-frequency intervals with the TPDs in the same time-frequency interval, which are the expected phase differences in that time-frequency interval. For example, the target phase correlation spectrum 720 encodes location-related information of the identified audio source to be separated from the audio signal. (Referring to Figure 3...) Figure 4A , Figure 4B , Figure 5A and Figure 5B Describe the details of the target phase correlation spectrum.

[0204] In the example, the target phase correlation spectrogram 720 can be combined with a multi-channel STFT 722 and frequency position encoding 724 to generate a channel cascade 726 of the received audio mixture signal 702. In one example, the channel cascade 726 of multiple sets of features is used as input to a neural network 202. For example, each set of features is specifically encoded with different information, such as the array shape in the IPD 718, the target position of the identified audio source in the target phase correlation spectrogram 720, the audio mixture signal in the multi-channel STFT 722, and frequency correlation processing is ensured by the frequency position encoding 724. In the example, the channel cascade 726 combines multiple sets of features by stacking features along a new dimension that can be called the channel dimension. It is worth noting that in practice, if some features are unavailable due to limited computation, data availability, or other constraints, the channel cascade block 726 may only use a subset of one or more sets of features from the multiple sets of features.

[0205] Once multiple sets of feature channel cascades 726 are generated, the neural network 202 can be used to process the received audio mixture signal 702 channel cascades 726 to extract the audio signal from the identified audio source. Combined with... Figure 7AThe method by which neural network 202 separates the audio signal from the identified audio source is described.

[0206] Figure 7C A flowchart illustrating the calculation of a target phase correlation spectrogram 720 according to some embodiments is shown. This target phase correlation spectrogram can be used as an input feature and to compute a loss function for extracting an audio signal. In the example, the target phase correlation spectrogram 720 may include complex numbers, i.e., having real and imaginary or complex numerical values. For example, the complex numbers of the target phase correlation spectrogram 720 may correspond to values ​​in the TF interval of the target phase correlation spectrogram 720. Furthermore, the neural network 202 with a U-Net architecture is a complex neural network for processing the complex numbers of the target phase correlation spectrogram 720.

[0207] It is understood that the target phase correlation spectrum 720 is calculated based on IPD 718 and TPD 716. In the example, the phase difference values ​​in IPD 718 and TPD 716 can be complex numbers.

[0208] According to this example, audio mix 702 is considered to be a P-channel audio mix, where It has a length of N samples. Based on the Fourier transform of the audio mixture signal, an input time-frequency (TF) mixture signal can be generated. In addition to using the complex multichannel spectrogram Y as input to the neural network 202, other input features are also considered. In multichannel scenarios, inter-ear or inter-channel phase difference (IPD) can be used to indicate spatial features. For example, IPD can be defined as:

[0209] in, Let 102d be the reference microphone, and IPD be calculated for each non-reference microphone, i.e., p=1, …, P-1. To mitigate the discontinuities caused by phase entanglement, the IPD characteristic is typically mapped to a complex number, defined as:

[0210] For example, to determine the IPD 718, a Fourier transform of the measured audio mixture signal across multiple channels of microphone array 102 can be used. In this respect, the STFT of each channel in microphone array 102 can be calculated by performing a Fourier transform on a small window segment of the measured audio mixture signal. In one example, the windowed segment's time period can range from 20 milliseconds (ms) to 50 ms. The STFT output (i.e., a complex spectrogram) can be represented as a complex number in amplitude-phase format.

[0211] Subsequently, the phase angle difference between each non-reference microphone and the reference microphone can be calculated. For example, the phase angle of different non-reference channels (e.g., non-reference microphone STFT 732) can be calculated as the difference between the phase angle of the STFT of the reference channel at the reference microphone 102d (i.e., reference microphone STFT 730) in the multi-channel STFT. Based on the comparison of the non-reference microphone STFTs of different microphones, inter-channel phase angle difference (IPD) can be generated for various non-reference channels (e.g., non-reference microphone 102b relative to reference microphone 102d or the reference channel in microphone array 102). For example, reference microphone 102d may be a microphone located at the center of microphone array 102. In this way, each channel or microphone can be analyzed independently to examine the frequency content of each channel over time, and further identify the phase difference of each non-reference channel STFT 730 relative to the reference channel STFT 732.

[0212] In the example, to avoid phase entanglement discontinuities, the phase angle difference, or IPD 718, can be converted into a complex number. As can be understood, each complex number can have a real part and an imaginary part. According to IPD 718, the real part of the complex number can indicate the cosine 734 of the corresponding phase angle difference, while the imaginary part indicates the sine 736 of the corresponding phase angle difference. The real part (i.e., cosine 734) and imaginary part (i.e., sine 736) of the represented complex number IPD can be used to generate the complex conjugate 738 of the phase angle difference of the represented complex number IPD 718 by summing the real and imaginary components and changing the sign of the imaginary component of the complex number.

[0213] Similarly, to determine TPD 716, the time delay can be calculated using the known location information or positioning of the microphones of microphone array 102 and the identified audio source. In this regard, firstly, the distance between the identified or known source location 744 of the audio source and the location value 740 of the reference microphone 102d can be calculated. Then, this distance can be divided by the speed of sound c to obtain the time to reach the reference microphone. The same process can then be repeated using the location value 742 of the non-reference microphone and the relative source location 744 of the audio source. In one example, the time to reach the reference microphone is subtracted from the time to reach the non-reference microphone to obtain the time difference to reach the non-reference microphone.

[0214] For example, when relative to a reference microphone 102d The defined source location of the identified audio source When known, pure time delay can be defined as: from the location located The source location propagates to the non-reference microphone. The audio signal or sound is transmitted to the reference microphone. The signals between them are measured in seconds. In the example, the source The target phase difference (TPD) is defined as:

[0215] The arrival time difference can then be converted into TPD for each time-frequency interval of the spectrogram using equation (7). It is noteworthy that TPD can be calculated based on the time delay between when sound (e.g., an audio mixture 702 propagating from the source location of the identified audio source) arrives at the reference channel or reference microphone 102d and when sound arrives at the non-reference channel or non-reference microphone 102b. For example, TPD 716 can be calculated based on the time delay of sound propagating from the source location 744 of the identified audio source to the reference microphone 102d and non-reference microphone 102b, as well as the position values ​​of the microphones in the machine and microphone array 102.

[0216] Furthermore, TPD 716 is also represented as a complex number. For example, each of the complex numbers in TPD 716 can have a real part and an imaginary part. According to TPD 716, the real part of the complex number can indicate the cosine 746 of the corresponding target phase angle difference, while the imaginary part of the complex number indicates the sine 748 of the corresponding target phase angle difference. The real part (i.e., cosine 746) and the imaginary part (i.e., sine 748) of the represented complex number TPD 716 can be used to generate a complex number 750 by summing the real and imaginary components of the target phase angle difference of the represented complex number TPD 716.

[0217] Furthermore, the product of each complex conjugate 738 of IPD 718 with the complex number 750 of TPD for each time-frequency interval can be determined. For example, the complex conjugate of IPD can be multiplied with the complex number of TPD corresponding to the same TF interval. In this way, the complex coordinate 738 of each in IPD 718 can be multiplied with the corresponding complex number 750 of TSD. Such products can then be summed or aggregated over all non-reference channels to generate the target phase correlation spectrum 720.

[0218] In the example, target phase correlation is used as a positional adjustment for the input feature, which indicates the spectrogram. Does the TF interval in the source location originate from the source location? The audio source at that location is dominant. The target phase correlation spectrum diagram 720 is defined as:

[0219] in yes The complex conjugate of the target phase correlation spectrum 720 can be used to process the target phase correlation spectrum 720 to extract isolated audio signals from identified audio sources such as component 402a. Based on the isolated audio signals that can be generated by the identified audio source 402a, certain audio analyses can be performed. In the example, equation (8) can also be applied to the estimated isolated training audio signals, such as training audio signals 612a and 612b used to calculate the training loss. In this case, (Y) can be replaced by the spectrum of the isolated training audio signals 612a and 612b to calculate the estimated training target phase correlation spectrum of the isolated training audio signals 612a and 612b. Furthermore, the estimated training target phase correlation spectrum of the isolated training audio signals 612a and 612b can be used to calculate the loss during training.

[0220] A target phase correlation spectrum 720 is calculated for each time-frequency interval, thus ensuring complete synchronization with other input features (such as complex spectrum) and features derived from the complex spectrum (such as IPD). Furthermore, any time-frequency intervals where IPD does not match well with the expected TPD will not be emphasized, while well-matched time-frequency intervals will be emphasized to ensure accurate separation of the audio signal from the audio mixture signal 702.

[0221] According to an exemplary embodiment of this disclosure, neural network 202 is configured to take a target phase correlation spectrum 720 as input. In one example, the neural network module of neural network 202 may be configured to use a complex BLSTM module 710 between a convolutional encoder 708 and convolutional decoders 712a and 712b. The complex BLSTM module 710 can output a multi-channel complex target phase correlation spectrum of the audio mixture signal.

[0222] In the example, based on the analysis of extracted audio signals from an identified audio source, control commands for the operation of the machine using the identified audio source can be identified. For example, such control commands can be associated with the current operating conditions of the machine and / or the identified audio source, the required operating conditions of the machine and / or the identified audio source, the health of the machine and / or the identified audio source, etc., and can be transmitted to the machine via wired or wireless communication channels to enable the machine to perform certain tasks. For example, when the analysis of extracted audio signals from an identified audio source indicates low fuel or raw material levels in a manufacturing machine, control commands for refilling fuel tanks or raw material containers can be transmitted to the machine.

[0223] It is worth noting that this analysis of the extracted audio signal from the identified audio source used to determine fuel level is merely exemplary. The extracted audio signal can be used, for example, for any kind of spatial audio analysis, surround sound processing, etc. Combined with, for example... Figure 8 This provides an example implementation for defining the analysis of the extracted audio signal.

[0224] refer to Figure 8 The diagram illustrates a schematic 800 illustrating control of machining operations based on anomaly detection according to some embodiments. Schematic 800 includes a machine 802 comprising components 802a and 802b. For example, components 802a and 802b of machine 802 can generate audio signals during operation. For example, component 802a can be a tool, and component 802b can be an actuator for actuating the tool. Therefore, components 802a and 802b can operate in unison to perform tasks such as machining tasks, assembly tasks, manufacturing tasks, etc.

[0225] In the example, machine 802 can use sensors to collect data. These sensors can be digital sensors, analog sensors, or combinations thereof. The collected data can be used for two purposes: some data is stored in a training data pool and used as training data for neural network 202, and some data can be used by the anomaly detection model as operational time data to detect anomalies. Both neural network 202 and the anomaly detection model can use the same data.

[0226] To detect abnormal operation of machine 802, specifically, training data can be collected first during the operation of any individual component 802a and / or 802b. The training data for anomaly detection can be used to train an anomaly detection model. The training data can include labeled or unlabeled data. Labeled data may already be labeled, for example, abnormal or normal. Unlabeled data has no label but is generally assumed to be non-abnormal data collected during normal machine operation. Different training methods are applied to the machine learning-based anomaly detection model based on the type of training data. Supervised learning is typically used for labeled training data, and unsupervised learning is typically applied for unlabeled training data. In this way, different implementations can manage different types of data.

[0227] In the example, the machine learning-based anomaly detection model learns the features and patterns of the training data, including normal and anomalous data patterns. The anomaly detection model uses trained parameters and collected operational time data (such as the separated audio signals from parts 802a and / or 802b) to perform anomaly detection. The operational time data (i.e., the separated audio signals) can be identified as normal or anomalous. For example, using normal data patterns, the trained machine learning-based anomaly detection model can classify operational time data into normal and anomalous data. Once an anomaly is detected, the necessary actions are taken.

[0228] In operation, components 802a and 802b can individually generate sound or audio signals. The sum of the audio signals can be measured by microphone array 102 as an audio mixture signal 104. For example, microphone array 102 can provide audio mixture signal 104 to audio system 106, specifically, audio input interface 108 of audio system 106. Furthermore, audio input interface 108 can receive audio mixture signal 104 and provide it to processor 110. For example, processor 110 can be configured to generate spectral characteristics, initially composed of multi-channel STFT 722, which can be used to calculate inter-channel phase angle differences in the multiple channels of microphone array 102 when measuring audio mixture signal 104. Furthermore, processor 110 can be configured to identify the audio source, i.e., components 802a and 802b, based on the known position of components 802a and 802b relative to the position of each microphone in microphone array 102. For example, processor 110 can be configured to identify components 802a and 802b based on relative position using an image sensor or a schematic diagram of the environment in which machine 802 is located. Processor 110 can represent the directional information of the microphones and audio sources 802a and 802b within the target phase angle difference. The TPD can encode the spatial information corresponding to the microphones of microphone array 102 and audio sources 802a and 802b in the order of time delay during sound propagation. Furthermore, processor 110 is configured to correlate spectral features (i.e., the IPD of the measured multi-channel spectrogram) with the target or expected phase angle to generate a target phase correlation spectrogram 720.

[0229] The target phase correlation spectrum 720 can indicate the directional characteristics of the audio mixture signal 104. For example, the processor 110 can be configured to process the target phase correlation spectrum 720 using the neural network 202 to extract audio signals, such as the estimated multi-channel complex spectrum portion 1 corresponding to component 802a and the estimated multi-channel complex spectrum portion 2 corresponding to component 802b. The generated estimated multi-channel complex spectrum portion 1 and estimated multi-channel complex spectrum portion 2 can produce an audio mixture signal when added together. For example, the neural network 202 can extract several features of the audio mixture signal 104, components 802a and 802b, and the source locations of components 802a and 802b, such as the array shape based on IPD, the target locations of components 802a and 802b based on the target phase correlation spectrum 720, the audio mixture signal based on the multi-channel STFT of the audio mixture signal 104, and frequency-related information based on frequency location encoding.

[0230] Once the separate audio signals constituting the audio mixture signal 104 are separated, they can be used to monitor the execution status of individual components or as control signals to isolate and control different components 802a and 802b. For example, the extracted audio signals can be provided to anomaly detectors 206a and 206b (collectively referred to as anomaly detector 206). For example, audio signal 1 corresponding to component 802a can be fed to anomaly detector 206a, while audio signal 2 corresponding to component 802b can be fed to anomaly detector 206b. Anomaly detectors 206a and 206b can include machine learning-based anomaly detection models. The anomaly detection model of anomaly detector 206a can be configured to identify whether isolated audio signal 1 meets the normal behavior or normal operation of component 802a. Similarly, the anomaly detection model of anomaly detector 206b can be configured to identify whether isolated audio signal 2 meets the normal behavior or normal operation of component 802b.

[0231] In one example, anomaly detector 206 can be configured to determine an anomaly score for each of the identified sound sources 802a and 802b. For example, isolated audio signal 1 fed to anomaly detector 206a can determine an anomaly score 804a for component 802a based on isolated audio signal 1 corresponding to component 802a. Similarly, isolated audio signal 2 fed to anomaly detector 206b can determine an anomaly score 804b for component 802b based on isolated audio signal 2 corresponding to component 802b. In this example, isolated audio signals 1 and 2 can be analyzed based on predictive patterns associated with possible anomalies in components 802a and 802b to generate anomaly scores 804a and 804b. For example, if components 802a and 802b operate normally, the anomaly score may be low, such as less than 1. However, if components 802a and 802b operate abnormally, the anomaly score may be high, such as greater than 7. For example, the anomaly score can be defined on a 0-10, 0-1 scale based on alphabetic characters or levels, etc.

[0232] Therefore, anomaly scores 804a and 804b can indicate the correlation between the type of anomaly in components 802a and 802b and the state of the identified sound source (i.e., components 802a and 802b). For example, processor 110 or anomaly detector 206 can be configured to compare the generated anomaly scores 804a and 804b with anomaly thresholds. In this example, the anomaly thresholds can be predefined, for example, by the manufacturers of components 802a and 802b, the manufacturer of machine 802, the user operating machine 802, etc. In some other cases, the anomaly thresholds can be determined dynamically.

[0233] In one example, the accuracy of anomaly detection can be improved by using embodiments of this disclosure for sound separation and anomaly detection. For example, using neural network 202, the extracted audio signals (such as audio signals 1 and 2) are accurately mapped to corresponding audio sources, such as components 802a and 802b. This is particularly advantageous when there may be a large number of components in a machine. In these cases, if all components except one are operating normally and the sound signals are not separated, the sound of the abnormal component may be masked by the sound of the normally operating component (i.e., undetectable) because the normally operating component may generate a larger sound signal. However, by separating the sound signal of each component, anomalies in specific components can be detected. Furthermore, it is beneficial to use the extracted audio signals for anomaly detection because when an anomaly is detected, it will be associated with a specific component, which will minimize the effort required to repair the abnormal component, as it has already been identified, and the necessary tests required to identify the abnormal component will be eliminated. This ensures cost-effective repairs, reduced downtime, and increased machine operational reliability.

[0234] Furthermore, the processor 110 or the anomaly detector 206 can be configured to select a control command 806 from a set of control commands to be executed by the machine 802 when either anomaly score 804a or 804b exceeds an anomaly threshold. For example, the control command 806 could be an instruction for the machine 802 to restore its operation to a normal operating range. The selected control command 806 can then be transmitted to the machine 802 to overcome the anomaly at either the audio source or either component 802a or 802b. In this way, the operating mode of the machine 802 can be changed, for example, by altering the operating conditions of components 802a and / or 802b. In some cases, the control command can shut down the machine 802 to prevent further abnormal or malfunctioning operation. This can be crucial in the event of a sudden failure, such as in heavy electrical equipment.

[0235] In one example, control command 806 could be an instruction that causes machine 802 to transfer control to a downstream controller. For example, such a downstream controller could be configured to perform operations to bring machine 802 back to normal operating range to resolve abnormal behavior of components 802a and / or 802b. In some cases, processor 110 or anomaly detector 206 could be configured to transmit anomaly score 804 to the downstream controller for, for example, handling faults in machine 802 and its components 802a and 802b, monitoring heat, monitoring operation, etc.

[0236] In another example, an audio signal extracted from the audio mixture signal 104 generated by the identified sound sources 802a and 802b can be analyzed to generate the state of the task performed by components 802a and 802b. In this example, a state estimator or anomaly detector 206 is trained on signals generated by the tools and / or actuators performing the task to estimate the execution state of the task. For example, in some implementations, the anomaly detector 206 can be configured to detect predictive patterns indicating the state of the task performed by components 802a and 802b. For example, real-valued time series of isolated signals 1 and 2 collected over a time period may include normal regions and, in some cases, anomalous regions leading to a failure point in either or both of components 802a and 802b. The anomaly detector 206 can be configured to detect any such anomalous regions to prevent failure of either component 802a or 802b. For example, in some implementations, the anomaly detector 206 may use a Shapelet discovery method to search for predictive patterns until an optimal predictive pattern is found. Therefore, based on the identification of any abnormal prediction patterns that may lead to malfunctions in machine 802, control commands can be selected from a set of control commands. In the example, this set of control commands may be based on different execution states of different tasks, different prediction patterns, different abnormal behaviors, etc. The selected control commands can then be transmitted to machine 802 to cause machine 802 to execute the control commands. For example, such control commands, when executed, can stop machine 802, stop abnormal components of machine 802, issue a request for maintenance of machine 802 and / or any component 802a or 802b, or utilize technical experts to schedule diagnostic or maintenance activities for the machine, etc.

[0237] In practice, state estimation based on the extracted signals allows control to be adapted to various types of complex manufacturing. However, some implementations are not limited to factory automation. For example, in one implementation, the controlled machine is a gearbox to be monitored for potential anomalies, and the gearbox can only be recorded in the presence of vibrations from the motor, coupling, or other vibrations from moving parts.

[0238] Benefiting from the teachings presented in the foregoing description and related drawings, those skilled in the art to which this disclosure pertains will conceive of numerous modifications and other embodiments of the disclosure set forth herein. Therefore, it should be understood that this disclosure is not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. Furthermore, although the foregoing description and related drawings describe exemplary embodiments in the context of certain example combinations of elements and / or functions, it should be understood that different combinations of elements and / or functions may be provided by alternative embodiments without departing from the scope of the appended claims. In this regard, for example, combinations of elements and / or functions different from those explicitly described above are also contemplated as being set forth in some of the appended claims. Although specific terminology is used herein, it is used only in a general and descriptive sense and not for limiting purposes.

Claims

1. An audio system for facilitating the operation of a machine, the machine including one or more actuators assisting one or more tools in performing one or more tasks, the audio system comprising: An audio input interface configured to receive an audio mixture signal generated by a plurality of audio sources, the plurality of audio sources including at least one of: the one or more tools performing one or more tasks, or the one or more actuators operating the one or more tools, wherein at least one of the audio sources forming the audio mixture signal is identified by its relative position to the position of each microphone in a microphone array that measures the audio mixture signal; A processor configured to extract audio signals generated by the identified audio sources from the audio mixture based on the correlation between spectral features in a multi-channel spectrogram of the audio mixture and directional information indicating the relative positions of identified audio sources among the plurality of audio sources; and An output audio interface is configured to output extracted audio signals to facilitate the operation of the machine.

2. The audio system according to claim 1, wherein, The processor is configured to use a neural network to extract the audio signal generated from the identified audio source.

3. The audio system according to claim 2, wherein, The spectral features include the inter-channel phase difference between channels in the multi-channel spectrum of the audio mixture signal. The directional information includes the target phase difference (TPD) of the sound propagating from the relative position of the identified audio source to the different microphones in the microphone array. The correlation between the spectral features and the directional information is represented by a target phase correlation spectrum, wherein the values ​​of the target phase correlation spectrum in different time-frequency intervals quantify the alignment of the inter-channel phase difference with the target phase difference in the corresponding time-frequency interval, and wherein the target phase difference is the expected phase difference of the time-frequency interval that indicates the characteristics of sound propagation. The processor is further configured as follows: Determine the target phase correlation spectrum; and The target phase correlation spectrum is processed using the neural network to extract the audio signal.

4. The audio system according to claim 3, wherein, The target phase correlation spectrum includes complex numbers, and the neural network is a complex neural network for processing the complex numbers of the target phase correlation spectrum.

5. The audio system according to claim 4, wherein, The complex neural network has a complex U-net architecture.

6. The audio system according to claim 5, wherein, The complex U-net architecture includes: Complex convolution encoder; Complex bidirectional long short-term memory (BLSTM) modules, the complex BLSTM modules being arranged to process the output of the complex convolutional encoder; and A complex convolutional decoder is arranged to process the output of the complex convolutional encoder and the output of the complex BLSTM module.

7. The audio system according to claim 6, wherein, The neural network is trained to extract signals from multiple identified audio sources, and the complex U-net architecture includes at least one complex convolutional decoder for each identified audio source.

8. The audio system according to claim 2, wherein, The processor is also configured to: The neural network is used to determine the target phase correlation spectrum of the audio mixture signal.

9. The audio system according to claim 2, wherein, In order to train the neural network, the processor is further configured to: A training audio mix signal is received from signals generated by one or more training audio sources, said one or more training audio sources including at least one of the following: one or more tools performing one or more tasks, or one or more actuators operating said one or more tools, wherein at least one of said one or more training audio sources forming the training audio mix signal is identified by position data relative to the position of each microphone in the microphone array that measures the training audio mix signal; Generate one or more training target phase correlation spectrograms associated with corresponding training audio sources, said one or more training target phase correlation spectrograms being generated based on the correlation between the spectral characteristics of the training audio mixture signal and the directional characteristics of the position data of said one or more training audio sources forming said training audio mixture signal, wherein each time-frequency (TF) interval of said one or more training target phase correlation spectrograms defines a feature that quantifies the match between the observed inter-channel phase difference and the corresponding expected phase difference in the spectral characteristics of the measured training audio mixture signal, said corresponding expected phase difference indicating the sound propagation characteristics of the corresponding position data of said one or more training audio sources relative to the position of each microphone of said microphone array; and The neural network is trained to extract training audio signals corresponding to the one or more training audio sources based on phase correlation spectrograms of one or more training targets.

10. The audio system according to claim 9, wherein, In order to train the neural network, the processor is further configured to: The neural network is trained based on a set of loss functions, which includes at least one of the following: a position loss function corresponding to each of the separated training audio signals of the training audio source, or a reconstruction loss function associated with the sum of the extracted training audio signals used to reconstruct the training audio mixture signal.

11. The audio system according to claim 9, wherein, In order to calculate the location loss function, the processor is configured as follows: Based on the position data of each of the one or more training audio sources, the physical characteristics of sound propagation of the one or more training audio sources are used to calculate the phase correlation spectrum of the ideal target. as well as Calculate estimated training target phase correlation spectrograms associated with corresponding training audio sources. One or more estimated training target phase correlation spectrograms are generated based on the correlation between spectral features associated with corresponding separated training audio signals and directional features indicating the positional data of the one or more training audio sources forming the training audio mixture signal. Each time-frequency (TF) interval of the one or more estimated training target phase correlation spectrograms defines a feature that quantifies the match between the observed inter-channel phase difference and a corresponding expected phase difference in the spectral features of the corresponding separated training audio signal, the expected phase difference indicating the acoustic propagation characteristics of the corresponding positional data of the one or more training audio sources relative to the position of each microphone in the microphone array. Determine the difference between the estimated training target phase correlation spectrogram and the corresponding ideal target phase correlation spectrogram for each of the one or more training audio sources, wherein the difference indicates the position loss function.

12. The audio system according to claim 9, wherein, In order to train the neural network, the processor is further configured to: The training audio mix signal generated by the one or more training audio sources is collected by moving the microphone array at different locations near the machine.

13. The audio system according to claim 1, wherein, The processor is also configured to: The received audio mixture signal is transformed using Fourier transform to generate a multi-channel short-time Fourier transform (STFT) of the received audio mixture signal. Determine the inter-channel phase difference (IPD) between different channels of the multi-channel STFT; Determine the target phase difference (TPD) of the sound propagating from the relative position of the identified audio source to the different microphones in the microphone array; The IPD is correlated with the TPD to generate a target phase correlation spectrum, wherein the values ​​of the target phase correlation spectrum in different time-frequency intervals quantify the alignment of the IPD with the TPD in the corresponding time-frequency interval, and wherein the TPD is the expected phase difference of the time-frequency interval indicating the characteristics of sound propagation. The target phase correlation spectrogram is combined with the multi-channel STFT and frequency position encoding to generate a channel cascade of the received audio mixed signal; and The received audio mixed signal is processed by channel cascading using a neural network to extract the audio signal.

14. The audio system according to claim 13, in, To determine the IPD, the processor is configured to: The complex values ​​of different channels are compared with a reference channel in the multi-channel STFT to generate the inter-channel phase difference (IPD) with respect to a reference microphone in the microphone array. The IPD is represented as a complex number, each of which has a real part indicating the cosine of the corresponding phase angle difference and an imaginary part indicating the sine of the corresponding phase angle difference, to produce the complex conjugate of each of the represented complex numbers IPD. In order to determine the TPD, the processor is configured as follows: Based on the position values ​​of the machine and microphone in the microphone array, calculate the target phase angle difference (TPD) between the sound propagating from the identified source in the audio mixture and the reference channel as the sound reaches different channels. The TPD is represented as a complex number, each of which has a real part indicating the cosine of the corresponding target phase angle difference from the generated target phase angle difference and an imaginary part indicating the sine of the corresponding target phase angle difference, to produce the complex conjugate of each of the represented complex numbers TPD, and In order to determine the target phase correlation spectrum, the processor is configured to: For each time-frequency interval, calculate the product of each complex conjugate of the complex number IPD and the corresponding complex conjugate of the complex number TPD, and Determine the sum of each product on all non-reference channels.

15. The audio system according to claim 1, wherein, The processor is also configured to: Based on the extracted signals, control commands for the operation of the machine are generated; and The control commands are transmitted to the machine via a communication channel.

16. The audio system according to claim 15, wherein, The processor is also configured to: The extracted audio signal generated from the audio mixture by the identified sound sources is analyzed to generate the execution state of the task; The control command is selected from a set of control commands based on the execution state of the task, wherein the set of control commands corresponds to different execution states of the one or more tasks; and The machine is made to execute the control command.

17. The audio system according to claim 15, wherein, The processor is also configured to: An anomaly score for the identified sound source is determined based on the extracted audio signal corresponding to the identified sound source, wherein the anomaly score indicates the correlation between the anomaly type and the state of the identified sound source. The anomaly score is compared with the anomaly threshold. When the anomaly score is greater than the anomaly threshold, the control command to be executed by the machine is selected from a set of control commands; and The selected control commands are sent to the machine to overcome the anomalies at the identified sound source.

18. The audio system according to claim 1, wherein, The plurality of audio sources that generate the audio mixture signal belong to the same category, and the processor is further configured to: Based on the correlation between the spectral features in the multi-channel spectrogram of the audio mixture signal and the directional information indicating the relative positions of the plurality of audio sources, the audio signal generated by each of the plurality of audio sources is extracted from the audio mixture signal.

19. A system for facilitating the operation of a machine, the machine including one or more actuators that assist one or more tools in performing one or more tasks, the system comprising: processor; as well as The memory stores instructions that cause the processor to: Receive an audio mixture signal generated by a plurality of audio sources, the plurality of audio sources including at least one of: the one or more tools performing the one or more tasks, or the one or more actuators operating the one or more tools, wherein at least one of the audio sources forming the audio mixture signal is identified by its relative position to the position of each microphone in the microphone array measuring the audio mixture signal; Based on the correlation between the spectral features in the multi-channel spectrogram of the audio mixture and the directional information indicating the relative positions of the identified audio sources among the plurality of audio sources, the audio signals generated by the identified audio sources are extracted from the audio mixture; and The extracted audio signal is output to facilitate the operation of the machine.

20. A method for facilitating the operation of a machine, the machine comprising one or more actuators that assist one or more tools in performing one or more tasks, the method comprising: An audio mixture signal is received using an audio input interface, which is generated by a plurality of audio sources, including at least one of the following: one or more tools performing one or more tasks, or one or more actuators operating the one or more tools, wherein at least one of the audio sources forming the audio mixture signal is identified by its relative position to the position of each microphone in the microphone array that measures the audio mixture signal. Using a processor, based on the correlation between spectral features in the multi-channel spectrogram of the audio mixture and directional information indicating the relative positions of the identified audio sources among the plurality of audio sources, the processor extracts the audio signals generated by the identified audio sources from the audio mixture; and The extracted audio signal is output using the output audio interface to facilitate the operation of the machine.

Citation Information

Cited By

  • Method for determining risk score, storage medium, electronic device and product

    CN122221039A