Audio processing method, audio processing model acquisition method and related equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-12
- Publication Date
- 2026-03-13
AI Technical Summary
In audio acquisition scenarios, specific noise signals are difficult to distinguish from ambient sound signals, causing traditional noise filtering techniques to lose some useful signals when removing noise, thus affecting audio quality.
Specific noise signals are identified and processed using a target noise processing model to ensure the integrity of the ambient sound signal, and a non-periodic, non-stationary signal processing method is used for separation.
It achieves the goal of preserving the integrity of ambient sound signals while removing specific noise signals, thereby improving the quality of the output audio signal.
Smart Images

Figure CN121666618A_ABST
Abstract
Description
Audio processing and its model acquisition methods and related equipment Technical Field
[0001] This disclosure relates to the field of audio processing technology, and in particular to an audio processing method and related equipment for obtaining audio models. Background Technology
[0002] In some audio acquisition scenarios, the input audio signal captured by recording equipment is often subject to noise due to environmental influences. The presence of noise affects the quality of the input audio signal and may even render it unusable. Noise filtering techniques can reduce the impact of noise to some extent, but certain types of noise can overlap with the useful signal, causing some of the useful signal to be filtered out during noise filtering, thus affecting the quality of the acquired audio.
[0003] Summary of the Invention
[0004] In a first aspect, embodiments of this disclosure provide an audio processing method, the method comprising: acquiring an input audio signal of a target scene, the input audio signal including an ambient sound signal and a specific noise signal of the target scene, wherein the ambient sound signal is a useful signal and the specific noise signal is a non-periodic, non-stationary signal; inputting the input audio signal into a target noise processing model to identify the specific noise signal that is distinguishable from the ambient sound signal; and processing the identified specific noise signal to obtain an output audio signal including the ambient sound signal after separating the specific noise signal.
[0005] Secondly, embodiments of this disclosure provide an audio processing apparatus, comprising: at least one processor; and at least one memory containing a computer program, wherein the at least one memory cooperates with the at least one processor to cause the apparatus to perform at least the following operations: acquiring an input audio signal of a target scene, the input audio signal including an ambient sound signal and a specific noise signal of the target scene; inputting the input audio signal into a target noise processing model to identify a specific noise signal that is distinguishable from the ambient sound signal; and processing the identified specific noise signal to obtain an output audio signal including the ambient sound signal after separating the specific noise signal, wherein the ambient sound signal is a useful signal and the specific noise signal is a non-periodic, non-stationary signal.
[0006] Thirdly, embodiments of this disclosure provide an audio acquisition device, comprising: a microphone for acquiring input audio signals of a target scene, the input audio signals including ambient sound signals and specific noise signals of the target scene; the ambient sound signals being useful signals, and the specific noise signals being non-periodic and non-stationary signals; and a processor communicatively connected to the microphone for receiving the input audio signals acquired by the microphone and processing the input audio signals, wherein processing the input audio signals includes: inputting the input audio signals into a target noise processing model to identify specific noise signals distinguishable from the ambient sound signals; and processing the identified specific noise signals to obtain an output audio signal including the ambient sound signals after separating the specific noise signals.
[0007] Fourthly, embodiments of this disclosure provide a multimedia acquisition system, comprising: an audio acquisition device for acquiring input audio signals of a target scene, the input audio signals including ambient sound signals and specific noise signals of the target scene; the ambient sound signals being useful signals, and the specific noise signals being non-periodic and non-stationary signals; a video acquisition device communicatively connected to the audio acquisition device, or the audio acquisition device being disposed on the video acquisition device, for acquiring video of the target scene, and for setting the audio acquisition parameters and various functions of the audio acquisition device on / off and / or setting the level; and at least one processor disposed in the audio acquisition device and / or the video acquisition device, for processing the input audio signals, wherein processing the input audio signals includes: inputting the input audio signals into a target noise processing model to identify specific noise signals distinguishable from the ambient sound signals; and processing the identified specific noise signals to obtain an output audio signal including the ambient sound signals after separating the specific noise signals.
[0008] Fifthly, embodiments of this disclosure provide a multimedia acquisition system, comprising: an audio acquisition device for acquiring input audio signals of a target scene, the input audio signals including ambient sound signals and specific noise signals of the target scene; the ambient sound signals being useful signals, and the specific noise signals being non-periodic and non-stationary signals; a video acquisition device for acquiring video of the target scene and for setting the audio acquisition parameters and various functions of the audio acquisition device by switching them on and / or setting their levels; and a terminal communicatively connected to the audio acquisition device and / or the video acquisition device, for controlling the audio acquisition parameters of the audio acquisition device. The system includes a switch and / or level setting for various functions of the audio acquisition device, and a switch and / or level setting for various functions of the video acquisition device; at least one processor, configured on at least one of the audio acquisition device, the video acquisition device, and the device terminal, is used to process the input audio signal, wherein processing the input audio signal includes: inputting the input audio signal into a target noise processing model to identify a specific noise signal that is distinguishable from the ambient sound signal; and processing the identified specific noise signal to obtain an output audio signal that includes the ambient sound signal after separating the specific noise signal.
[0009] In this embodiment, after acquiring the input audio signal, a specific noise signal that is distinguishable from the ambient sound signal is identified using a target noise processing model. The identified specific noise signal is then processed to obtain an output audio signal that includes the ambient sound signal after the specific noise signal has been separated. Since the specific noise signal is a non-periodic and non-stationary signal, it is indistinguishable from the ambient sound. However, the ambient sound, as a useful signal, should be preserved rather than processed along with the specific noise signal. Therefore, by employing the target noise processing model of this embodiment, the indistinguishable ambient sound signal can be separated from the specific noise signal. This allows processing only the specific noise signal without affecting the ambient sound signal in the output audio signal, preserving the useful signal for the user and thus ensuring the audio quality of the output audio signal.
[0010] In a sixth aspect, embodiments of this disclosure provide an audio processing method, the method comprising: acquiring an input audio signal of a target scene, the input audio signal including an ambient sound signal and a wind noise signal of the target scene; wherein the ambient sound signal is a useful signal, and the wind noise signal overlaps with the ambient sound signal in the frequency spectrum; inputting the input audio signal into a pre-trained target noise processing model to identify the wind noise signal that is distinguishable from the ambient sound signal; and suppressing the identified wind noise signal to obtain an output audio signal that includes the ambient sound signal but does not include the wind noise signal.
[0011] In this embodiment, after acquiring the input audio signal, a target noise processing model is used to identify wind noise signals that are distinct from ambient sound signals. The identified wind noise is then processed to obtain an output audio signal that includes ambient sound signals but excludes wind noise signals. Since wind noise is a non-periodic, non-stationary signal, it overlaps with ambient sound in the frequency spectrum. Therefore, it cannot be completely distinguished from ambient sound using relevant methods. Ambient sound, as a useful signal, should ideally be fully preserved, rather than being suppressed along with specific noise signals, resulting in a loss of ambient sound signal. Therefore, by employing the target noise processing model of this embodiment, the indistinguishable ambient sound signals and wind noise signals can be completely separated, allowing processing only the wind noise signal to obtain an output audio signal that includes ambient sound signals but excludes wind noise signals, thereby ensuring the audio quality of the output audio signal.
[0012] In a seventh aspect, embodiments of this disclosure provide a method for obtaining an audio processing model, the method comprising: acquiring an input sample signal, the input sample signal including an ambient sound sample signal and a specific noise sample signal; acquiring an output signal of an initial noise processing model based on features of the input sample signal; determining a loss function based on the output signal; and training the initial noise processing model based on the loss function to obtain a target noise processing model, wherein the ambient sound sample signal is a useful signal, the specific noise sample signal is a non-periodic non-stationary signal, and the target noise processing model is capable of distinguishing between the ambient sound sample signal and the specific noise sample signal.
[0013] Eighthly, embodiments of this disclosure provide an audio processing model acquisition apparatus, comprising: at least one processor; and at least one memory containing a computer program, wherein the at least one memory cooperates with the at least one processor to cause the apparatus to perform at least the following operations: acquiring an input sample signal, the input sample signal including an ambient sound sample signal and a specific noise sample signal; acquiring an output signal of an initial noise processing model based on features of the input sample signal; determining a loss function based on the output signal; and training the initial noise processing model based on the loss function to obtain a target noise processing model, wherein the ambient sound sample signal is a useful signal, the specific noise sample signal is a non-periodic non-stationary signal, and the target noise processing model is capable of distinguishing between the ambient sound sample signal and the specific noise sample signal.
[0014] In this embodiment of the disclosure, the specific noise sample signal is a non-periodic, non-stationary signal, and the ambient sound sample signal is a useful signal. It is difficult to distinguish between the ambient sound sample signal and the specific noise sample signal in the input sample signal that combines the two. By training the initial noise processing model with the purpose of distinguishing between the ambient sound sample signal and the specific noise sample signal and based on the loss function, a target noise processing model that can distinguish between any input audio signal including the ambient sound signal and the specific noise signal can be trained. Therefore, the target noise processing model can be used to distinguish between the difficult-to-distinguish ambient sound signal and the specific noise signal, so that the quality of the ambient sound signal is not affected when the specific noise signal is processed subsequently.
[0015] Ninthly, embodiments of this disclosure provide a method for obtaining an audio processing model, the method comprising: obtaining an input sample signal, the input sample signal including an ambient sound sample signal and a wind noise sample signal, wherein the ambient sound sample signal is a useful signal and the wind noise sample signal is a non-periodic, non-stationary signal; obtaining the difference between the ambient sound sample signal and a predicted ambient sound signal output by an initial noise processing model based on the ambient sound features of the ambient sound sample signal; obtaining the predicted probability that the wind noise sample signal exists in the input sample signal, as output by the initial noise processing model based on the wind noise features of the wind noise sample signal; determining a loss function based on the difference and the predicted probability; and training the initial noise processing model based on the loss function to obtain a target noise processing model, wherein the target noise processing model is capable of distinguishing between ambient sound signals and wind noise signals that overlap in the spectrum.
[0016] In this embodiment, the wind noise sample signal is a non-periodic, non-stationary signal, and the ambient sound sample signal is a useful signal. The ambient sound signal and the wind noise signal overlap in their spectra, making them difficult to distinguish in the combined input sample signal. By training the noise processing model with the aim of distinguishing the ambient sound sample signal and the wind noise sample signal, and based on the loss function determined by the difference and prediction probability, a target noise processing model capable of distinguishing any input audio signal including both ambient sound and wind noise signals can be trained. That is, using this target noise processing model, the difficult-to-distinguish ambient sound signal and wind noise signal can be differentiated, so that subsequent suppression of the wind noise signal does not affect the quality of the ambient sound signal.
[0017] In a tenth aspect, embodiments of this disclosure provide an audio processing method, the method comprising: acquiring an output audio signal including an ambient sound signal of a target scene, the output audio signal being obtained by processing an input audio signal based on a target noise processing model, the input audio signal including the ambient sound signal and a specific noise signal that overlap in the spectrum; and performing sound effect matching between the target scene and the output audio signal, wherein the ambient sound signal is a useful signal and the specific noise signal is a non-periodic, non-stationary signal.
[0018] Eleventhly, embodiments of this disclosure provide an audio processing apparatus, comprising: at least one processor; and at least one memory containing a computer program, wherein the at least one memory cooperates with the at least one processor to cause the apparatus to perform at least the following operations: acquiring an output audio signal including an ambient sound signal of a target scene, the output audio signal being obtained by processing an input audio signal based on a target noise processing model, the input audio signal including the ambient sound signal and the specific noise signal overlapping in the spectrum; and performing sound effect matching between the target scene and the output audio signal, wherein the ambient sound signal is a useful signal and the specific noise signal is an aperiodic non-stationary signal.
[0019] In a twelfth aspect, embodiments of this disclosure provide an audio acquisition device, comprising: a microphone for acquiring an input audio signal of a target scene, the input audio signal including an ambient sound signal and a specific noise signal of the target scene; the ambient sound signal being a useful signal, the specific noise signal being a non-periodic, non-stationary signal, and the ambient sound signal and the specific noise signal overlapping in the frequency spectrum of the input audio signal; and a processor communicatively connected to the microphone for receiving the input audio signal acquired by the microphone and processing the input audio signal, wherein processing the input audio signal includes: processing the input audio signal based on a target noise processing model to obtain an output audio signal including the ambient sound signal; and performing sound effect matching between the target scene and the output audio signal.
[0020] In a thirteenth aspect, embodiments of this disclosure provide a multimedia acquisition system, comprising: an audio acquisition device for acquiring input audio signals of a target scene, the input audio signals including ambient sound signals and specific noise signals of the target scene; the ambient sound signals being useful signals, the specific noise signals being non-periodic and non-stationary signals, and the ambient sound signals and specific noise signals in the input audio signals overlapping in the frequency spectrum; a video acquisition device communicatively connected to the audio acquisition device, or the audio acquisition device being disposed on the video acquisition device, for acquiring video of the target scene, and for setting and / or adjusting the audio acquisition parameters and various functions of the audio acquisition device; and at least one processor disposed in the audio acquisition device and / or the video acquisition device, for processing the input audio signals, wherein the processing of the input audio signals includes: processing the input audio signals based on a target noise processing model to obtain an output audio signal including the ambient sound signals; and performing sound effect matching between the target scene and the output audio signal.
[0021] In this embodiment, the ambient sound signal is a useful signal, and the specific noise signal is a non-periodic, non-stationary signal. Since the output audio signal is obtained by processing the input audio signal based on the target noise processing model, it can distinguish between the ambient sound signal and the specific noise signal that overlap in the spectrum. This allows for subsequent sound effect matching of the distinguished specific noise signal and the distinguished ambient sound signal for different target scenarios, achieving the effect that related solutions cannot match for sound effects of specific noise signals and ambient sound signals that overlap in the spectrum, thus enriching the user experience.
[0022] In a fourteenth aspect, embodiments of this disclosure provide an audio processing method, the method comprising: inputting an input audio signal into a target noise processing model, the input audio signal including an ambient sound signal and a wind noise signal of a target scene, the ambient sound signal and the wind noise signal overlapping in the frequency spectrum; identifying a wind noise signal that is distinct from the ambient sound signal through the target noise processing model; processing the identified wind noise signal to obtain an output audio signal including the ambient sound signal after separating the wind noise signal; and performing sound effect matching between the target scene and the output audio signal.
[0023] In this embodiment of the disclosure, the ambient sound signal is a useful signal. Since the output audio signal is obtained by processing the input audio signal based on the target noise processing model, it can distinguish the ambient sound signal and the wind noise signal that overlap in the spectrum. Therefore, the output audio signal of the ambient sound signal after separating the wind noise signal can be matched with sound effects for different target scenarios. That is, sound effect matching can be performed on the ambient sound signal with only the wind noise removed, thus enriching the user experience.
[0024] In a fifteenth aspect, embodiments of the present disclosure provide a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the methods of the first, sixth, seventh, ninth, tenth, or fourteenth aspects of the present disclosure.
[0025] In a sixteenth aspect, embodiments of the present disclosure provide a computer program product having a computer program stored thereon, which, when executed by a processor, implements the methods of the first, sixth, seventh, ninth, tenth, or fourteenth aspects of the present disclosure.
[0026] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 is a schematic diagram of the input audio signal according to an embodiment of the present disclosure.
[0029] Figure 2 is a flowchart of an audio processing method according to an embodiment of the present disclosure.
[0030] Figure 3 is a schematic diagram of the input audio signal in the related technology.
[0031] Figure 4 is a schematic diagram of the input and output of the target noise processing model according to an embodiment of the present disclosure.
[0032] Figure 5 is a general flowchart of an embodiment of this disclosure.
[0033] Figure 6 is a flowchart of the model training process according to an embodiment of this disclosure.
[0034] Figure 7 is a schematic diagram of the input and output of the model training process according to an embodiment of the present disclosure.
[0035] Figure 8 is a schematic diagram of the adaptive control process of noise reduction intensity and the sound effect matching process of an embodiment of the present disclosure.
[0036] Figure 9 is a schematic diagram of sound effect matching methods under different target scenarios in embodiments of this disclosure.
[0037] Figure 10 is a flowchart of an audio processing method according to another embodiment of this disclosure.
[0038] Figure 11 is a flowchart of the method for obtaining the audio processing model according to an embodiment of the present disclosure.
[0039] Figure 12 is a flowchart of a method for obtaining an audio processing model according to another embodiment of this disclosure.
[0040] Figure 13 is a flowchart of an audio processing method according to another embodiment of the present disclosure.
[0041] Figure 14 is a flowchart of an audio processing method according to yet another embodiment of this disclosure.
[0042] Figure 15 is a schematic diagram of an audio processing apparatus according to an embodiment of the present disclosure.
[0043] Figure 16 is a schematic diagram of an audio acquisition device according to an embodiment of the present disclosure.
[0044] Figure 17 is a schematic diagram of a multimedia acquisition system according to an embodiment of the present disclosure.
[0045] Figure 18 is a schematic diagram of a multimedia acquisition system according to another embodiment of the present disclosure. Detailed Implementation
[0046] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0047] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0048] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0049] In some audio acquisition scenarios, the input audio signal obtained from picking up sound from a target scene (i.e., the sound collection process) typically includes not only useful signals (signals the user wants to retain) but also specific noise signals (signals the user wants to remove). As shown in Figure 1, useful signals can include ambient sound signals from the target scene. Furthermore, if the target scene includes objects capable of emitting speech signals (e.g., people, animals, robots, smart speakers, artificial intelligence, etc.), useful signals can include both ambient sound signals and speech signals. To improve the quality of the input audio signal, specific noise signals need to be processed (e.g., suppressed). These specific noise signals are those that are desired to be removed, but are difficult to distinguish easily from ambient sound signals.
[0050] For example, when using action cameras for audio and video recording in outdoor scenarios such as cycling, skiing, and the beach, wind noise is a frequent problem. Users only care about the ambient sound and human voices, but the superimposed wind noise significantly affects sound quality; sometimes, a single instance of wind noise can render an entire clip unusable. Therefore, wind noise needs to be processed, for example, suppressed. Similarly, in cycling scenarios, users want to preserve the sound of the bicycle chain and human voices. If there is strong wind, these effective sound information will be masked or affected by wind noise. Therefore, users need to suppress wind noise during recording without affecting ambient sound. Furthermore, during drone flight, propeller noise is constant and can interfere with human voices. If propeller noise can be selectively eliminated while preserving other ambient and human voices, it could create possibilities for customized audio pickup scenarios for remote controls.
[0051] However, specific noise signals may overlap with useful signals, making it difficult to effectively separate the specific noise signals. Therefore, when suppressing specific noise signals, a portion of the useful signal is also filtered out, thus affecting the quality of the audio signal.
[0052] Based on this, embodiments of this disclosure propose an audio processing scheme that first identifies specific noise signals that are distinct from ambient sound signals using a target noise processing model, and then processes the identified specific noise signals to obtain an output audio signal that includes the ambient sound signal after separating the specific noise signals. This avoids affecting the ambient sound signal in the output audio signal when processing (e.g., suppressing) the specific noise signals, thus ensuring the audio quality of the output audio signal. The scheme of embodiments of this disclosure will be described in detail below.
[0053] As shown in Figure 2, this embodiment of the present disclosure provides an audio processing method, the method comprising:
[0054] Step S11: Acquire the input audio signal of the target scene. The input audio signal includes the ambient sound signal and the specific noise signal of the target scene. The ambient sound signal is a useful signal, and the specific noise signal is a non-periodic and non-stationary signal.
[0055] Step S12: Input the input audio signal into the target noise processing model to identify specific noise signals that are distinguishable from the ambient sound signal; and
[0056] Step S13: Process the identified specific noise signal to obtain an output audio signal that includes the ambient sound signal after separating the specific noise signal.
[0057] In step S11, the target scene can be various scenarios. For example, based on whether the target scene is located in an enclosed area, the target scene can be an outdoor scene or an indoor scene; based on the weather of the target scene, the target scene can be a windy scene, a light breeze scene, a windless scene, etc.; based on the function and purpose of the target scene, the target scene can be a stage scene, a meeting scene, a call scene, a dining scene, etc.; based on the geographical location of the target scene, the target scene can be a city scene, a seaside scene, a mountain scene, a forest scene, etc.; based on the location of the target scene, the target scene can be a transportation scene, a scenic spot scene, a shopping mall scene, a residential area scene, etc. In addition to the scenarios listed above, the target scene of this embodiment can also be other types of scenarios, which will not be listed here.
[0058] In some embodiments, a recording element can be used to pick up sound from the target scene to obtain the input audio signal of the target scene. The recording element may include one or more microphones. Each microphone can pick up sound from the target scene, obtaining a set of audio signals. At any given time, the microphones picking up sound from the target scene may include some or all of the aforementioned microphones. The input audio signal of the target scene can be determined based on the audio signals picked up by each microphone. It should be noted that the input audio signal described in this disclosure refers to the sound signal obtained after collecting all sounds in the environment or scene; the term "input" refers to the audio acquisition device or noise processing model, not the sound input into the target scene.
[0059] In some embodiments, the recording element is disposed on the audio acquisition device. In examples where the recording element includes multiple microphones, the multiple microphones can be disposed at different locations on the audio acquisition device. For example, the multiple microphones may consist of two microphones, which are respectively disposed on the left and right sides of the audio acquisition device. Another example is that the multiple microphones may consist of four microphones, which are respectively disposed on the front, rear, left, and right sides of the audio acquisition device. The number and arrangement of the multiple microphones can also be configured in other ways, and are not limited here. By distributing multiple microphones at different locations on the audio acquisition device, a more realistic and stereo sound effect can be captured, and the location and direction of the sound source can be easily determined.
[0060] When the recording element may include a microphone, the audio signal picked up by the microphone can be determined as the input audio signal of the target scene, or the audio signal picked up by the microphone can be post-processed (e.g., audio editing, audio synthesis, etc.) to obtain the input audio signal of the target scene.
[0061] When the recording element includes multiple microphones, the input audio signal for the target scene can be determined jointly based on the audio signals collected by each microphone. For example, the audio signal with the lowest specific noise among the audio signals collected by each microphone can be obtained and used as the input audio signal for the target scene. The audio signal with the lowest specific noise can be obtained by performing correlation operations between multiple sets of audio signals picked up by multiple microphones. Specifically, the correlation between the audio signals of two adjacent microphones is closely related to a specific noise signal (e.g., wind noise). For two adjacent microphone signals, the higher the correlation between their audio signals, the lower the probability of the presence of a specific noise signal; conversely, the lower the correlation, the higher the probability of the presence of a specific noise signal. In other words, the higher the correlation between the audio signals of two adjacent microphones, the less affected they are by a specific noise signal; conversely, the lower the correlation, the more affected they are by a specific noise signal. Specifically, the pair of adjacent microphones with the highest correlation is identified, and the audio signals collected by these two adjacent microphones are jointly determined as the audio signal with the lowest specific noise signal. By selecting the audio signal with the lowest specific noise level as the input audio signal for the target scene, the impact of specific noise can be effectively reduced, improving the quality of the input audio signal. Alternatively, in some implementations, multiple sets of audio signals from the target scene captured by multiple microphones can be used as the input audio signal for the target scene. Since the sound captured by each microphone may have slight time and phase differences, using multiple sets of audio signals from multiple microphones as the input audio signal for the target scene allows for the acquisition of more information about the location and direction of the sound source, thereby enhancing the stereo effect and facilitating sound source localization and sound field reconstruction.
[0062] In some implementations, the input audio signal includes ambient sound signals of the target scene and specific noise signals. The ambient sound signals of the target scene are audio signals from the target scene, typically used to describe and reproduce the sound characteristics and atmosphere of the target environment. These ambient sound signals are useful signals, signals that the user wants to retain, unlike in relevant application scenarios (e.g., phone calls) where they are considered noise and removed. Ambient sound signals can serve as background noise for the target scene. By retaining these signals, users can experience a sense of immersion in the target scene, enhancing their realism and sense of presence. The specific type of ambient sound signal can vary depending on the target scene. For example, in a stage scene, the ambient sound signal may include music; in a seaside scene, it may include the sound of waves; and in an outdoor or traffic scene, it may include sounds generated by vehicles (such as car horns, the sound of a bicycle chain turning, or the sound of a car engine starting).
[0063] In some implementations, the specific noise signal is an interference signal, the signal that the user wishes to remove. This specific noise signal can be additive noise rather than multiplicative noise (e.g., echoes). It exists regardless of the presence of the useful signal, unlike multiplicative noise, which ceases to exist when the useful signal is absent. In some implementations, the specific noise signal is an aperiodic, non-stationary signal. An aperiodic, non-stationary signal is one that is neither periodic nor stationary in time. The specific noise signal is not periodic, meaning it does not exhibit a repeating periodic pattern on the time axis; it is not stationary, meaning its statistical characteristics (such as mean, variance, power spectrum, etc.) change over time. The frequency components of aperiodic, non-stationary noise can be very complex and extensive, and its statistical characteristics change over time, often exhibiting high randomness. Different specific noise signals are uncorrelated, making mathematical modeling and estimation difficult. Traditional noise reduction algorithms estimate noise based on the signal's spectral energy, and then perform Wiener filtering based on the estimated noise to achieve noise reduction. Because noise estimation requires time window tracking, traditional noise reduction methods are suitable for suppressing stationary noise but cannot estimate specific noise signals that are periodically non-stationary, let alone further distinguish between specific noise signals and ambient sound signals. Currently, there are also solutions that improve the hardware structure to suppress specific noise signals. For example, when the specific noise signal is wind noise, a multi-microphone structure is used for wind noise suppression. However, on the one hand, these methods require improvements to the hardware structure, increasing hardware costs; on the other hand, because specific noise signals have non-periodic and non-stationary characteristics, traditional noise suppression methods cannot completely suppress specific noise signals, resulting in poor suppression effects.
[0064] The above embodiments illustrate cases where specific noise signals include wind noise signals. Wind noise signals refer to environmental noise signals caused by wind. They typically manifest as sound produced by airflow vibrations caused by wind passing over or through the interior of an object. In other examples, specific noise signals may also include at least one of the following: propeller noise signals, road noise signals, collision noise signals, gimbal noise signals, and mechanical button noise signals. Propeller noise signals refer to the noise signals emitted when the propellers of devices such as drones, helicopters, ships, or underwater robots rotate; road noise signals refer to the noise signals generated during the activities of vehicles and pedestrians on the road; collision noise signals refer to the noise signals generated when multiple objects collide; gimbal noise signals refer to the noise generated during the rotation of a gimbal; and mechanical button noise signals refer to the noise signals generated when a mechanical button is pressed. Besides the noise types listed above, the specific noise signals in the embodiments of this disclosure may also be other signals with non-periodic and non-stationary characteristics.
[0065] It should be noted that the ambient sound signal and the specific noise signal may be different in different target scenarios, and the ambient sound signal in one type of target scenario may be the specific noise signal in another type of target scenario. A mapping relationship between each type of target scenario and the ambient sound signal and specific noise signal in that type of target scenario can be established in advance in order to determine the ambient sound signal and specific noise signal in each type of target scenario.
[0066] In some embodiments, the input audio signal further includes a speech signal. Including a speech signal in the input audio signal can mean that at least a portion of the audio segments in the input audio signal include a speech signal; or, it can mean that all audio segments in the input audio signal include a speech signal.
[0067] Speech signals can be sound signals (such as speaking, singing, or calling) emitted by humans, animals, robots, smart speakers, or artificial intelligence, and typically contain specific semantic information. The object emitting the speech signal can be an object within the target scene. In this case, the speech signal can be the speech emitted by the sound-collecting entity within the target scene, while the ambient sound signal is the background sound within the target scene, emitted by objects other than the sound-collecting entity. Taking human speech as an example, the target scene can be captured by a recording element while a person is speaking, resulting in an input audio signal that includes the speech signal, the scene sound signal of the target scene, and specific noise signals. Optionally, the speech signal and the scene sound signal of the target scene can be audio signals emitted by different types of objects within the target scene. For example, the speech signal could be an audio signal emitted by a person in the target scene, while the ambient sound signal could be an audio signal emitted by a bicycle chain, motorcycle engine, or other objects within the target scene.
[0068] Alternatively, the object emitting the speech signal may not be an object in the target scene. For example, the target scene, which does not include the object emitting the speech signal, can be captured by a recording element to obtain an audio signal that includes only the scene sound signal of the target scene and a specific noise signal, but does not include the speech signal. The audio signal obtained by the recording is then combined with the speech signal to obtain an input audio signal that includes the speech signal, the scene sound signal of the target scene, and the specific noise signal.
[0069] When the input audio signal includes both speech signals and ambient sound signals of the target scene, both are useful signals that the user wants to retain, while the specific noise signal is the interference signal that the user wants to remove. It should be noted that traditional audio noise reduction techniques are often used in call or conference scenarios. Referring to Figure 3, in traditional audio noise reduction, only the speech signal is typically considered useful, and all other signals are considered noise. When processing the target, only the speech signal in the input audio signal is retained, and all other signals in the input audio signal are suppressed. While this processing method can reduce the impact of noise on audio quality, it fails to retain the ambient sound signals that reflect the audio characteristics of the target scene, thus preventing the user from perceiving the audio characteristics of the target scene. In other words, this traditional audio noise reduction technique cannot be applied to scenarios such as outdoor sound pickup (using an action camera with a microphone to collect or acquire immersive ambient sounds of an outdoor scene) as described in this disclosure. For example, when applied to a cycling scenario, the sound of the bicycle chain, as an ambient sound signal, is suppressed along with wind noise, leaving only human voices. This makes it impossible to discern that the user is speaking or recording life while cycling. Unlike related technologies, in this embodiment, considering the user's need to preserve ambient sound, both ambient sound and speech signals are retained as useful signals. Only specific noise signals are processed (e.g., suppressed), allowing the user to perceive the audio characteristics of the target scene and thus creating an immersive auditory experience. For example, when applied to a cycling scenario, the sound of the bicycle chain, as an ambient sound signal, is preserved along with human speech, thus demonstrating that the user is speaking while cycling.
[0070] In some embodiments, before being input into the target noise processing model, the ambient sound signal of the target scene in the input audio signal overlaps with the specific noise signal in the spectrum. Overlap refers to the presence of aliasing, interleaving, or other signal mixing, making it difficult to separate into two independent signals. Since the noise pattern of the specific noise signal can be a non-periodic, non-stationary signal, it overlaps with both the ambient sound signal and the speech signal, making it indistinguishable from both. To separate the specific noise signal or reduce only the specific noise signal, the target noise processing model of this disclosure is needed to process the input audio signal. Specifically, the phase spectrum characteristics of the input audio signal can be used as a reference. These phase spectrum characteristics are input into the target noise processing model, allowing the noise processing model to process the phase information to distinguish the specific noise signal from the other two types of signals, and to suppress only the specific noise signal. In other embodiments, the spectra of the ambient sound signal of the target scene in the input audio signal may partially overlap with the speech signal; in other embodiments, the spectra of the specific noise signal in the input signal may also partially overlap with the speech signal.
[0071] In some embodiments, the amplitude spectra of the ambient sound signal of the target scene in the input audio signal may partially overlap with the amplitude spectra of the speech signal; in other embodiments, the amplitude spectra of the specific noise signal in the input signal may also partially overlap with the amplitude spectra of the ambient sound signal; in still other embodiments, the amplitude spectra of the specific noise signal in the input signal may also partially overlap with the amplitude spectra of the speech signal. To address this, the amplitude spectrum characteristics of the input audio signal can be used to input these characteristics into a target noise processing model, allowing the noise processing model to process the amplitude information to distinguish between the speech signal and the ambient sound signal. However, this still cannot distinguish between the specific noise signal and the ambient sound signal.
[0072] Optionally, before inputting the target noise processing model, the proportion of the amplitude spectrum where the ambient sound signal and the speech signal of the target scene overlap in the input audio signal is less than a preset proportion. This preset proportion can be set according to actual needs, for example, it can be set to 0 or a value close to 0. Optionally, the proportion of the amplitude spectrum where the ambient sound signal and the speech signal of the target scene overlap in the input audio signal is less than the proportion of the amplitude spectrum where the ambient sound signal and the specific noise signal of the target scene overlap in the input audio signal. Because the proportion of the amplitude spectrum where the ambient sound signal and the speech signal overlap in the input audio signal is small, the ambient sound signal and the speech signal are easier to distinguish than the ambient sound signal and the specific noise signal. In some embodiments, in a windless environment, the amplitude spectrum characteristics of the input audio signal can be used alone to input the amplitude spectrum characteristics of the input audio signal into the target noise processing model, allowing the noise processing model to process this amplitude information to distinguish between the speech signal and the ambient sound signal.
[0073] In step S12, the input audio signal can be input into the target noise processing model. The target noise processing model processes the identified specific noise signal to obtain an output audio signal that includes the environmental sound signal after separating the specific noise signal. Further, if the input audio signal includes both a speech signal and a scene sound signal of the target scene, the input audio signal can be input into the target noise processing model. The target noise processing model processes the identified specific noise signal to obtain an output audio signal that includes the speech signal after separating the specific noise signal and the environmental sound signal. By separating specific noise signals from ambient sound and speech signals, it is possible to process (e.g., suppress) only the specific noise signals, preserving the useful ambient sound and speech signals while avoiding any impact on their quality. Compared to related solutions that use methods such as finding the minimum value in the complex domain for low-frequency frequencies or replacing frequency bands, similar to algorithmic solutions that replace high-wind-noise frequency bands with low-wind-noise bands or adaptive high-pass filtering for noise reduction based on wind speed sensors, the solution disclosed here, because it uses a target noise processing model to distinguish between ambient sound and specific noise signals, does not affect the useful signals. Furthermore, the above method does not require any hardware structure improvements, reducing hardware costs.
[0074] Figure 4 shows the input audio signal and output speech signal of some embodiments. It can be seen that before the input audio signal is input to the target noise processing model, the ambient sound signal and the specific noise signal in the input audio signal are superimposed, and the speech signal and the specific noise signal are also superimposed. In the output audio signal processed by the target noise processing model, the ambient sound signal and the specific noise signal are separated, and the speech signal and the specific noise signal are also separated. Compared with traditional signal processing methods, the target noise processing model has a strong nonlinear fitting capability. Based on data-driven principles, the target noise processing model infers results from real data and can handle various complex non-periodic and non-stationary signals. Therefore, it can effectively separate the ambient sound signal from the specific noise signal, and the speech signal from the specific noise signal, thereby processing different signals more accurately and specifically in the audio processing stage. Compared with traditional noise suppression methods, the method of this disclosure can reduce the impact on the ambient sound signal and the speech signal when processing the specific noise signal in audio, and effectively improve the noise suppression effect, such as achieving complete suppression of the specific noise signal. Furthermore, the above method does not require any modifications to the hardware structure, thus reducing hardware costs.
[0075] In this embodiment, after acquiring the input audio signal, a specific noise signal that is distinguishable from the ambient sound signal is identified using a target noise processing model. The identified specific noise signal can then be processed to obtain an output audio signal that includes the ambient sound signal after separation of the specific noise signal. Since the specific noise signal is a non-periodic and non-stationary signal, it is indistinguishable from the ambient sound. However, the ambient sound, as a useful signal, should be preserved rather than processed along with the specific noise signal. Therefore, by employing the target noise processing model of this embodiment, the indistinguishable ambient sound signal can be separated from the specific noise signal. This allows processing only the specific noise signal without affecting the ambient sound signal in the output audio signal, preserving the useful signal for the user and thus ensuring the audio quality of the output audio signal.
[0076] Furthermore, the target noise processing model can be pre-trained to distinguish between ambient sound signals and specific noise signals, thereby obtaining specific noise signals separated from the ambient sound signals. Prior to this, no related technology had discovered the need to distinguish between overlapping noise signals in a target scene to reduce only specific noise, nor had it disclosed the ability to train a noise processing model to distinguish between ambient sound signals and specific noise signals. In other words, the discovery of the problem addressed by this disclosure and the solution itself are the result of creative effort. The following is an example illustrating the training process of the target noise processing model.
[0077] In some embodiments, referring to Figures 5, 6, and 7, the target noise processing model is trained based on the features of the input sample signal. First, in the data preparation stage, different types of input sample signals can be prepared, and the features of the input sample signals can be obtained. Specifically, the input sample signals include ambient sound sample signals and specific noise sample signals. The features of the input sample signals include the ambient sound features corresponding to the ambient sound sample signals and the specific noise features corresponding to the specific noise sample signals. The ambient sound sample signals can be pure ambient sound signals without any other sounds, and the specific noise sample signals can be pure specific noise signals without any ambient sound signals. The features of the input sample signals may only include the phase spectrum, with 90% of the feature information of the input sample signals recorded in its phase spectrum. In some embodiments, the phase spectrum of the input audio signal describes the rate of change; for example, the phase changes rapidly when the wind speed is high, while it changes slowly when the wind speed is low and relatively stable. The more complex the scene, the faster the phase spectrum changes. If the scene is simpler, the phase spectrum changes more slowly. Therefore, inputting only the phase spectrum of the input sample signal into the target noise processing model can be understood as inputting most of the features of the input sample signal into the target noise processing model. Then, the noise processing model can be trained to learn the feature differences between ambient sound signals and specific noise signals to obtain the target noise processing model. Thus, by using only the phase spectrum as a feature of the input sample signal, the trained target noise processing model can effectively identify ambient sound signals and specific noise signals.
[0078] Alternatively, referring to Figure 7, the features of the input sample signal may include the phase spectrum and at least one of the following: amplitude spectrum and acoustic perception features. By employing multiple features, the trained target noise processing model can more accurately and comprehensively learn the characteristic differences between ambient sound signals and specific noise signals, thereby further improving the accuracy and reliability of the target noise processing model in separating specific noise signals and ambient sound signals. The amplitude spectrum of a signal refers to the amplitude or magnitude information of each frequency component when the signal is represented in the frequency domain. Acoustic perception features refer to the subjective perception characteristics of sound during auditory processing, such as which frequency spectrums the user is more sensitive to. Acoustic perception features may include, but are not limited to, at least one of the following: frequency band features, equivalent rectangular bandwidth, Barker scale, and Mel scale. Frequency band features are used to represent the characteristics of a signal in the frequency domain. Equivalent rectangular bandwidth refers to the width of a bandwidth of a signal in the frequency domain, which is the bandwidth of a rectangular frequency response with the same power or energy. Barker scale and Mel scale are used to describe how the auditory system perceives sound frequencies. By using phase spectrum and acoustic sensing features, such as frequency band information (information obtained by frequency domain transformation methods other than fast Fourier transform, such as high-frequency information, empirical values, overlapping frequency bands, etc., which can make the judgment of wind noise signals more accurate), this disclosure can achieve complete preservation of ambient sound (or background sound) in windless scenarios, while avoiding microphone distortion in extreme scenarios, such as extremely windy scenarios.
[0079] Feature extraction can be performed on ambient sound sample signals and specific noise sample signals using a target noise processing model or a feature extraction model other than the target noise processing model, respectively, to obtain ambient sound features and specific noise features. For example, the CONV operator can be used for high-level feature extraction within the network, while the GRU operator is used for time-series information transfer. Other operators or combinations of operators can also be flexibly used, without restrictions. For example, a Transformer model or a Linear model can be used for feature extraction, and a Long Short-Term Memory (LSTM) network or a Recurrent Neural Network (RNN) model can be used for time-series information transfer. The feature extraction process can be implemented based on Fast Fourier Transform (FFT) or Discrete Cosine Transform (DCT), etc.
[0080] The output signal of the initial noise processing model based on the features of the input sample signal can be obtained. Based on this output signal, a loss function is determined, and the initial noise processing model is trained using this loss function to obtain the target noise processing model. The initial noise processing model can be a noise processing model that has not yet been fully trained. Alternatively, the initial noise processing model can be used to directly extract features from the input sample signal, obtaining its features, and then the initial noise processing model can predict the output signal based on these features. Or, a feature extraction model can be used to extract features from the input sample signal, obtaining its features, and then the extracted features can be input into the initial noise processing model so that the initial noise processing model can predict the output signal based on the input features.
[0081] In some embodiments, referring to Figure 6, the loss function includes a first loss function (i.e., target 1 or can also be characterized by target output 1), which is determined based on the difference between the ambient sound sample signal and the predicted ambient sound signal output by the initial noise processing model based on ambient sound features. By employing the first loss function, the predicted ambient sound signal output by the trained target noise processing model can be made as close as possible to the ambient sound sample signal, reducing the distortion of the ambient sound signal after processing by the target noise processing model.
[0082] The features of the input sample signal can be fed into an initial noise processing model to obtain the first network output of the initial noise processing model. This first network output is an audio signal that retains the ambient sound sample signal while suppressing specific noise sample signals. Then, a first loss function can be calculated based on the first network output and the ambient sound sample signal. Specifically, the similarity between the first network output and the ambient sound sample signal can be calculated, and the first loss function can be determined based on this similarity. Optionally, the first loss function may include a scale-invariant signal-to-noise ratio (SI-SDR) loss function and a mean square error (MSE) loss function.
[0083] In some embodiments, referring to Figure 6, the loss function includes a second loss function (i.e., target 2, or can also be characterized by target output 2), which is determined based on the predicted probability of a specific noise sample signal existing in the input sample signal output by the initial noise processing model based on specific noise features. By employing the second loss function, the trained target noise processing model can accurately predict the probability of a specific noise signal existing in the input audio signal. Generally, the probability of a specific noise signal existing in the input audio signal reflects the signal strength of the specific noise signal. The stronger the signal strength of the specific noise signal, the higher the probability of the specific noise signal existing in the input audio signal predicted by the target noise processing model tends to be; the weaker the signal strength of the specific noise signal, the lower the probability of the specific noise signal existing in the input audio signal predicted by the target noise processing model tends to be. By predicting the probability of a specific noise signal existing in the input audio signal, it is helpful to estimate the signal strength of the specific noise signal, thereby adaptively employing an appropriate method to process the specific noise signal according to its signal strength. For example, when the probability of the specific noise signal existing is low, the noise suppression intensity can be reduced, thereby ensuring that the ambient sound signal and speech signal are preserved to a greater extent. Conversely, when the probability of a specific noise signal being present is relatively high, the noise suppression strength can be enhanced, thereby ensuring that the specific noise signal can be better suppressed.
[0084] In examples where the loss function includes a second loss function, the input sample signal may further include label data indicating the presence of specific noise sample signals. The second loss function can be determined based on the predicted probabilities and the label data. An initial noise processing model can be trained to obtain a second network output, which is the predicted probability of the presence of specific noise sample signals in the input sample signal. Based on the second network output and the label data, the second loss function is calculated. For example, the input sample signal can be classified based on the predicted probabilities to obtain its category, including a first category with specific noise signals and a second category without specific noise signals. Specifically, if the predicted probability is greater than or equal to a preset probability threshold, the input sample signal is determined to be in the first category; otherwise, it is determined to be in the second category. The label data can also indicate the category of the input sample signal. Then, the difference between the category determined based on the predicted probabilities and the category indicated by the label data is used to determine the second loss function. Optionally, the second loss function can be a cross-entropy (CE) loss function.
[0085] Referring again to Figure 6, during training, the initial noise processing model can be trained to obtain the target noise processing model in response to the satisfaction of a preset training termination condition. Specifically, the training termination condition can be that the number of training iterations reaches a preset threshold, or that the loss function meets a preset condition. For example, the first and second loss functions mentioned above can be weighted to obtain the total loss function. The loss function meets the preset condition if the total loss function is less than a preset value, or if the value of the total loss function has not significantly improved in several consecutive training iterations.
[0086] Furthermore, when the input audio signal includes a speech signal, the target noise processing model can be pre-trained to distinguish between ambient sound signals, speech signals, and specific noise signals. Further, in the example where the input audio signal includes a speech signal, and the target noise processing model is used to distinguish between ambient sound signals, speech signals, and specific noise signals, the target noise processing model is trained based on the features of the input sample signals. The input sample signals include ambient sound sample signals, speech sample signals, and specific noise sample signals. The features of the input sample signals include ambient sound features corresponding to the ambient sound sample signals, speech features corresponding to the speech sample signals, and specific noise features corresponding to the specific noise sample signals. The ambient sound sample signals can be pure ambient sound signals without any other sounds, the speech sample signals can be pure speech signals without any other sounds, and the specific noise sample signals can be pure specific noise signals without containing either ambient sound signals or speech signals.
[0087] In this embodiment, the output signal of the initial noise processing model based on the features of the input sample signal can still be obtained. Based on this output signal, a loss function is determined, and the initial noise processing model is trained based on the aforementioned loss function to obtain the target noise processing model. The loss function includes a first loss function and a second loss function. Unlike the previous embodiments, when determining the first loss function, the first network output is an audio signal that retains the ambient sound sample signal and the speech sample signal, and suppresses specific noise sample signals. Otherwise, the training method of the target noise processing model can be found in the previous embodiments and will not be repeated here.
[0088] It should be noted that in related technologies, noise reduction is commonly used in conference call scenarios, where only human voice is retained and all other sounds are suppressed. Under the above objective, the training input of the target noise processing model is a speech signal and noise signals (all signals except the speech signal), and the target output is a speech signal. Its training objective is solely to ensure that the speech signal output by the target noise processing model has sufficient similarity to the speech signals used as training samples, thus avoiding signal distortion. Referring to Figure 6, the training input of the target noise processing model in this embodiment includes a speech signal, an ambient sound signal, and a specific noise signal. Its target output includes two components: one is the speech signal and the ambient sound signal (referred to as Target 1), and the other is the probability of the existence of the specific noise signal (referred to as Target 2).
[0089] In some embodiments, since the types of ambient sound signals and / or specific noise signals may be different in different target scenarios, the types of ambient sound signals and specific noise signals that a target noise processing model can handle are generally limited. Therefore, in order to adapt to the needs of different target scenarios, multiple noise processing models (hereinafter referred to as candidate noise processing models) can be pre-trained, and the target noise processing model can be selected from multiple candidate noise processing models according to actual needs.
[0090] For example, in response to a user's selection, a target noise processing model can be selected from multiple candidate noise processing models. A user interface can be provided, displaying multiple options corresponding to each candidate noise processing model. The user selects the target option from these options, and the candidate noise processing model corresponding to the target option is determined as the target noise processing model for the current target scenario. Alternatively, a correspondence between various target scenarios and noise processing models can be pre-established. After automatically identifying the target scenario (e.g., based on sensor data or public platform data), the candidate noise processing model corresponding to the identified target scenario is determined as the target noise processing model for that target scenario based on the aforementioned correspondence.
[0091] Data from each candidate noise processing model can be stored in the cloud or in an acquisition device or terminal connected to the cloud. After determining the target noise processing model from multiple candidate models, the data of the target noise processing model can be obtained from the cloud or from an acquisition device or terminal connected to the cloud, so that the specific noise signal can be processed using the target noise processing model. Simultaneously, since the data can be stored in the cloud, it does not occupy the memory of the audio acquisition device, acquisition terminal, or terminal, thus improving data processing efficiency. In some implementations, the terminal can be a mobile phone, watch, motion-sensing remote control, VR control device, wearable device, etc., as long as it can communicate with the acquisition device and transmit information to or control the acquisition device.
[0092] In step S13, the specific noise signal identified by the target noise processing model can be processed to obtain an output audio signal including the ambient sound signal after separating the specific noise signal.
[0093] Furthermore, in examples where the input audio signal also includes a speech signal, the specific noise signal identified by the target noise processing model can be processed to obtain an output audio signal that includes an ambient sound signal and a speech signal after separating the specific noise signal.
[0094] Furthermore, if the target noise processing model is also used to determine the probability of a specific noise signal existing in the input audio signal, the specific noise signal can be processed based on this probability. The suppression strength of the specific noise signal can be determined based on this probability; for example, the suppression strength can be determined as an intensity value positively correlated with the probability. It can also be determined based on this probability whether the function to suppress the specific noise signal needs to be enabled; for example, the function to suppress the specific noise signal is enabled when the probability is greater than a preset probability threshold, and disabled otherwise. Alternatively, the suppression method for the specific noise signal can be determined based on this probability, such as partial suppression or complete suppression; for example, partial suppression is used when the probability is greater than a preset threshold, and complete suppression is used otherwise.
[0095] In some embodiments, only the identified specific noise signals may be processed, including complete suppression or partial suppression. Complete suppression means completely eliminating the specific noise signal, while partial suppression may involve reducing the signal strength of the specific noise signal or suppressing only the specific noise signal within a certain frequency range. Users can choose to completely or partially suppress the identified specific noise signals according to their actual needs. For example, if high audio quality is required for the audio signal, the identified specific noise signal can be completely suppressed; if lower audio quality is required, the identified specific noise signal can be partially suppressed.
[0096] After distinguishing between ambient sound signals and specific noise signals, the target noise processing model identifies the specific noise signals. These specific noise signals can then be processed by a noise reduction module to achieve complete or partial suppression. The noise reduction module can be integrated within the target noise processing model, suppressing only the identified specific noise signals (e.g., completely suppressing them) to obtain an output audio signal that includes the ambient sound signals. Alternatively, the noise reduction module can be located outside the target noise processing model. After distinguishing between ambient sound signals and specific noise signals, the target noise processing model identifies the specific noise signals, which can then be input into the noise reduction module for suppression (e.g., complete suppression), while the ambient sound signals can be output as the audio signal.
[0097] In some embodiments, referring again to Figure 5, the processing parameters of the identified specific noise signal can be adaptively adjusted, and the identified specific noise signal can be processed based on the adjusted processing parameters to obtain a more suitable noise reduction effect. The processing parameters include, but are not limited to, the suppression intensity of the specific noise signal, the frequency range of the specific noise signal to be suppressed, and / or the algorithm parameters of the noise suppression algorithm. By adaptively adjusting the processing parameters, appropriate processing parameters can be selected to process the specific noise signal according to the actual characteristics of the target scene, thereby making the processing result more closely match the actual characteristics of the target scene and improving the processing effect.
[0098] Specifically, the processing parameters for the identified specific noise signal can be adaptively adjusted based on at least one of the following:
[0099] (1) The probability that the input audio signal contains a specific noise signal. This probability can be predicted by the target noise processing model. This probability reflects the signal strength of the specific noise signal in the input audio signal. Generally, the signal strength of the specific noise signal is positively correlated with the probability. Based on the probability, the processing parameters of the identified specific noise signal are adaptively adjusted so that the processing parameters are adapted to the signal strength of the specific noise signal. For example, the processing parameters may include the suppression strength of the specific noise signal, and this suppression strength is positively correlated with the probability. When the probability is low, it indicates that the signal strength of the specific noise signal in the input audio signal is weak, or that there is no specific noise signal in the input audio signal. Therefore, a smaller noise suppression strength can be used to suppress the specific noise signal. Conversely, a larger noise suppression strength can be used to suppress the specific noise signal, thereby reducing signal processing power consumption while ensuring noise suppression effect.
[0100] (2) Scene type of the target scene. The characteristics of specific noise signals differ under different scene types; therefore, the noise reduction requirements for target scenes of different scene types also differ. Adaptively adjusting the processing parameters of the identified specific noise signals based on the scene type of the target scene allows the processing effect of the specific noise signals to match the characteristics of the specific noise signals in the target scene and the noise reduction requirements of the target scene, thereby improving the processing effect. Scene types may include, but are not limited to, outdoor scenes, indoor scenes, windy scenes, light wind scenes, stage scenes, conference scenes, urban scenes, seaside scenes, and traffic scenes.
[0101] Users can either select the scene type of the target scene or perform scene recognition to determine the scene type. In the former case, to facilitate user operation, a user interface can be provided displaying several options corresponding to the scene type. Users select the target option from these options, and the scene type corresponding to the target option is determined as the scene type of the target scene. In the latter case, sensor data (such as anemometer data, GPS data, speed data, and illumination data) and / or public platform data (such as meteorological data: wind force, wind speed, etc.) can be acquired, and scene recognition of the target scene can be performed based on this sensor data and / or public platform data.
[0102] (3) Obtaining parameter adjustment instructions input by the user. In this embodiment, the user can directly issue parameter adjustment instructions to adjust the processing parameters of a specific noise signal, including the suppression strength of the specific noise signal. For example, a user interface can be provided, displaying at least one control, which the user can operate to issue parameter adjustment instructions. For example, the at least one control includes a first control for increasing the processing parameter and a second control for decreasing the processing parameter. The user can operate the first control to issue a parameter adjustment instruction to increase the processing parameter, and can also operate the second control to issue a parameter adjustment instruction to decrease the processing parameter. In this way, the user can precisely control the processing parameters, thereby facilitating personalized settings of processing parameters according to the needs of different users. In addition, by allowing users to participate in parameter adjustment, the user's sense of participation and control over the processing process can be enhanced, which helps to improve user satisfaction and overall experience.
[0103] In addition to the various ways of adaptively adjusting processing parameters listed above, other conditions can also be used to adaptively adjust processing parameters, which will not be listed here.
[0104] In some embodiments, the ambient sound signal after separating the specific noise signal can be used as at least a part of the output audio signal for playback, transmission, amplification, enhancement, or speech recognition. For example, the audio processing method of this disclosure embodiment can be executed by a device with audio acquisition and audio playback functions. After acquiring the ambient sound signal after separating the specific noise signal, the device can further play the ambient sound signal, or amplify or enhance the ambient sound signal before playback. As another example, the audio processing method of this disclosure embodiment can be executed by a device with audio acquisition functions. After acquiring the ambient sound signal after separating the specific noise signal, the device can further transmit the ambient sound signal to a device with audio playback functions for playback. Yet another example, the audio processing method of this disclosure embodiment can be executed by a device with audio acquisition functions. After acquiring the ambient sound signal after separating the specific noise signal, the device can further perform speech recognition on the ambient sound signal and output the speech recognition result.
[0105] Further, referring to Figure 5 and in conjunction with Figure 8, before the ambient sound signal after separating the specific noise signal is used as at least a part of the output audio signal for playback, transmission, amplification, enhancement, or speech recognition, the output audio signal can be matched with the target scene for sound effects. Sound effects matching refers to the process of setting the sound effects of the output audio signal so that the set sound effects match the target scene. By matching the output audio signal with the target scene for sound effects, the user's immersion and realism in a specific environment or situation can be enhanced. In related technologies, since the specific noise signal and the ambient sound signal are mixed together, it is difficult to achieve sound effects matching for the ambient sound signal with only the specific noise signal removed. The embodiments of this disclosure separate the specific noise signal and the ambient sound signal through a target noise processing model to obtain an output audio signal that includes only the ambient sound signal and does not include the specific noise signal. Then, sound effects matching is performed on the output audio signal with the target scene, thus achieving sound effects matching for the ambient sound signal with only the specific noise signal removed. It is understood that the sound effects matching step can also be performed after at least a part of the output audio signal has been played, transmitted, amplified, enhanced, or used for speech recognition, and this disclosure does not limit this.
[0106] In some embodiments, matching sound effects adapted to the target scene can be selected from a sound effects library, and the ambient sound signal in the output audio signal can be adjusted accordingly based on the matching sound effects. Specifically, a correspondence between the sound effect parameters of various sound effects in the sound effects library and the scene can be established in advance, and matching sound effects adapted to the target scene can be selected from the sound effects library based on the above correspondence. The sound effect parameters may include, but are not limited to, at least one of the following: volume, stereo position, reverberation effect, dynamic range of the sound effect (i.e., the difference between the loudest and softest sound in the sound effect), frequency response of the sound effect, and various special effects of the sound effect (such as echo, chorus, phase, distortion), etc.
[0107] If no matching sound effect for the target scene exists in the sound effect library, the ambient sound signal in the output audio signal can be adjusted based on default sound effect parameters. Default sound effect parameters can be pre-stored in a preset storage path, such as locally or in the cloud. When a matching sound effect for the target scene is found in the sound effect library, the default sound effect parameters can be read from the preset storage path and used to adjust the ambient sound signal. In some embodiments, after adjusting the ambient sound signal in the output audio signal based on the default sound effect parameters, the user can be prompted with the type of the target scene and the sound effect matching result. The type of the target scene can be a user-selected type or an automatically identified type. The sound effect matching result can include the type and / or value of the sound effect parameters after the ambient sound signal adjustment.
[0108] In some embodiments, sound effect matching between the output audio signal and the target scene can be performed based on at least one of the following information:
[0109] (1) At least one scene image of the target scene. Scene images can provide rich visual information. Based on scene images, the objects included in the target scene can be identified, thereby more accurately determining the scene type of the target scene and thus more accurately achieving sound effect matching. For example, a scene image may include lightning, indicating that the target scene is a thunderstorm scene. Therefore, the lightning effect in the output audio signal can be enhanced to obtain the sound effect of thunder (ambient sound signal, which may be regarded as noise signal and removed in related schemes).
[0110] (2) Information about the audio acquisition device used to acquire the input audio signal. This information includes, but is not limited to, the model, quantity, location of the audio acquisition device, and / or whether each microphone on the audio acquisition device is enabled. Since different audio acquisition devices acquire audio signals with different characteristics, sound effect matching based on the information of the audio acquisition device can be targeted to achieve sound effect matching based on the characteristics of the audio signal acquired by the audio acquisition device, thereby improving the sound effect matching effect. For example, if the audio acquisition device may have slight loss in the high-frequency range, compensation can be made to the high-frequency portion of the output audio signal to improve the audio quality of the high-frequency portion. As another example, if the audio acquisition device includes multiple microphones located in different directions, sound effect matching can enhance the stereo effect in the output audio signal, allowing users to more clearly perceive the direction and spatial location of the sound.
[0111] (3) Sensor data collected during input audio signal acquisition. The input audio signal can be acquired by an audio acquisition device, and the sensor data can be data collected by sensors on the audio acquisition device. Sensor data includes, but is not limited to, anemometer data, GPS data, speed data, and / or illumination data. Sensor data can reflect the scene type of the target scene or the characteristics of the audio signal in the target scene. For example, GPS data can reflect the area where the audio acquisition device is located or the type of area where the audio acquisition device is located, such as high-altitude areas, seaside, cities, etc. Input audio signals collected in different areas have different characteristics. For example, specific noise signals (such as wind noise) are often larger in high-altitude areas and seaside, and input audio signals collected in high-altitude areas travel farther. Input audio signals collected at the seaside are often accompanied by loud wave noise, while wind noise is relatively smaller in cities, and input audio signals collected in cities are often accompanied by the sound of motor vehicles. Speed data can reflect the speed at which the input audio signal is acquired, and acceleration data can reflect the acceleration at which the input audio signal is acquired. Input audio signals acquired at different speeds / accelerations may also have different characteristics. For example, at higher speeds, wind noise may also be higher; at lower speeds, wind noise may also be lower. Illumination data can reflect the weather conditions at the time the input audio signal was collected. In sunny weather, certain noise signals (such as wind noise) tend to be lower, while in stormy weather, certain noise signals tend to be higher. Sound effect matching based on sensor data can ensure that the matching results are compatible with the scene type or characteristics of the audio signals in the target scene as reflected by the sensor data.
[0112] (4) Collect public platform data when the input audio signal is acquired. This public platform data can include meteorological data, traffic flow monitoring data, etc. Meteorological data reflects the weather and climate characteristics at the time the input audio signal is acquired, such as sunny days, rainy days, strong winds, and blizzards. Different weather and climate conditions may result in different sound effects; for example, sunny days usually have less wind noise, while strong winds and blizzards may result in more wind noise. Traffic flow monitoring data reflects the traffic flow at the time the input audio signal is acquired, such as idle or congested conditions. Different traffic flows may result in different sound effects; for example, in low traffic flow, only a few vehicles pass by, the engine sound is relatively soft, and the overall sound is relatively calm; in medium traffic flow, the vehicles are closer together, and the sounds of acceleration and deceleration become more frequent; in high traffic flow, a large number of vehicles are traveling, the engine sounds intertwine, the overall sound is more noisy, and may be accompanied by frequent horn and brake sounds. Matching sound effects based on public platform data ensures that the sound effect matching results match the scene information reflected by the public platform data.
[0113] (5) User-inputted scene information. Scene information may include information indicating the scene type, such as a city scene, a stage scene, a dining scene, etc., and may also include descriptive information describing the target scene. For example, a descriptive message might be, "On a sunny, windless day, there were few vehicles on the road, and a flock of birds flew over the city." This descriptive information describes the weather (sunny), scene type (city type), and objects (vehicles, birds) of the target scene. Based on the scene information, the target scene can be characterized, thus facilitating the use of sound effect matching methods that match the target scene for sound effect matching.
[0114] In some embodiments, different sound effect adjustment methods are suitable for the output audio signals of different target scenarios. Therefore, a target adjustment method that matches the target scenario can be determined from multiple candidate adjustment methods, and the output audio signal can be adjusted based on the target adjustment method. These multiple candidate adjustment methods may include, but are not limited to, at least one of the following: no sound effect processing, adjustment of the signal strength of the speech signal in the output audio signal, and adjustment of the signal strength of the ambient sound signal in the output audio signal. Adjusting the signal strength can involve increasing or decreasing the signal strength, and the adjustment range can be selected from multiple preset strengths according to actual needs. For example, referring to Figure 9, when the target scenario is an outdoor interview scenario, a strong suppression level can be used to reduce specific noise signals (e.g., wind noise), enhance the speech signal (e.g., human voice), and slightly enhance the ambient sound signal; when the target scenario is an outdoor stage performance, the music rhythm can be enhanced, and the wind noise signal can be reduced; when the target scenario is a sunny day, a medium suppression level can be used to suppress wind noise; and when the target scenario is a windy day, a strong suppression level can be used to suppress wind noise.
[0115] The embodiments disclosed herein can be used to suppress specific noise signals encountered by recording devices, such as action cameras, handheld cameras, wireless microphones, pocket cameras, voice recorders, and drones, which are directly or indirectly equipped with recording functions, when recording audio and video scenes outdoors. These noises include wind noise, road noise, collision noise, gimbal noise generated by the pocket camera during gimbal rotation, propeller noise generated by the drone during flight, and button noise generated by the pocket camera's control joystick during use. The suppression of these noises will not affect the sound pickup effect of ambient sounds, thus greatly improving the user experience.
[0116] The present disclosure will now be described in more detail and with a more complete embodiment:
[0117] First, complete the training of the neural network (i.e., the target noise processing model) that distinguishes between background noise and specific noise.
[0118] Compared to related technical solutions, this disclosure has significantly different design goals in the early stages of neural network training:
[0119] The relevant technology center notes that noise reduction is commonly used in conference call scenarios, preserving only human voice (i.e., speech signal) and suppressing all other sounds. Its training methods include:
[0120] Training input: Clean human voice + all noise
[0121] Target output: Pure vocals
[0122] This disclosure can suppress specific noise signals, selectively reducing only specific noise signals while preserving human voice (i.e., speech signals) and background noise (i.e., ambient sound signals) other than the specific noise signals. It also outputs the probability of the presence of the characteristic noise signal. Its training methods include:
[0123] Training input: Pure human voice + background noise (ambient sound signal) other than a specific noise signal + specific noise signal
[0124] Target Output 1: Pure human voice + background sound (ambient sound signal) excluding specific noise signals.
[0125] Target Output 2: The probability of a specific noise signal being present
[0126] Specifically, in the data preparation phase, according to the design, it is first necessary to collect different types of data, classify them, and then use them as input signals. These different types of data include:
[0127] Pure human voice: A clean human voice that is free of any noise.
[0128] Specific noise signal: Pure noise of a specific type, without any other sounds mixed in.
[0129] c. Ambient sound signals other than specific noise: all kinds of ambient sound signals, but they cannot include specific noise signals and human voices.
[0130] d. Probability of the presence of a specific noise signal: It is necessary to prepare label data on the presence or absence of a specific noise signal in the input signal.
[0131] After data preparation is complete, input signal feature extraction is required. The input signal features include:
[0132] a. Phase spectrum.
[0133] b. Amplitude spectrum.
[0134] c. Acoustic sensing features, such as equivalent rectangular bandwidth (ERB).
[0135] After fusing the aforementioned input signal features, the feature data is fed into the neural network for training. The neural network training will have two outputs: one is the audio after preserving the background noise and suppressing specific noise signals; the other is the probability that the current signal contains specific noise signals. The neural network structure can use vector convolution (CONV) operators for high-level feature extraction within the neural network, while employing gated recurrent units (GRU) operators to pass temporal information.
[0136] Next, the scale-invariant signal-to-noise ratio (SI-SDR) and mean square error (MSE) are used to calculate the losses for network output 1 and target output 1, enabling gradient backpropagation to update the network weight parameters. The cross-entropy loss function (CE) is then used to calculate the losses for network output 2 and target output 2. Finally, once the losses for both target output 1 and target output 2 are less than a threshold, network training is complete.
[0137] Then, using the pre-trained neural network (target noise processing model), specific noise signals are suppressed based on features such as the phase spectrum. Specifically:
[0138] After training a neural network model that can distinguish between ambient sound signals and specific noise signals, signal processing can be performed on the input signal to preserve the ambient sound signal and suppress the specific noise signal. First, feature extraction is performed on the input signal to obtain the amplitude spectrum, phase spectrum, and acoustic perception features. At least the phase spectrum of the input signal is used as input to the trained model for inference, thereby obtaining the audio data after preserving the ambient sound signal and suppressing the specific noise signal, as well as the probability of the existence of the specific noise signal.
[0139] The network target output 1 has completed the task of reducing only specific noise signals, that is, preserving ambient sound signals and human voices, and suppressing specific noise signals.
[0140] Meanwhile, the probability of the presence of a specific noise in target output 2 can be used for adaptive adjustment of the suppression intensity of that specific noise signal. For example, when the probability of the presence of a specific noise signal is very low, the suppression intensity can be weakened to ensure that ambient background noise and human voices are preserved to a greater extent, thus achieving complete preservation of useful signals in windless environments. Conversely, when the probability of the presence of a specific noise signal is high, the suppression intensity should be strengthened to ensure that the specific noise signal can be better suppressed and to avoid distortion.
[0141] Finally, different sound effects are matched based on the probability of the presence of characteristic noise signals, adaptive intensity suppression, and different application scenarios (target scenarios).
[0142] After the neural network inference is completed, we can obtain the audio after suppressing only specific noise signals and the probability of the existence of those specific noise signals. First, by using the probability of the existence of specific noise signals and combining it with the audio after suppressing only specific noise signals, we can dynamically adjust the suppression intensity of those specific noise signals to achieve the most suitable noise reduction effect.
[0143] Building upon the above, sound effect matching is then performed based on different application scenarios, allowing for more subtle adjustments to the sound effects for different actual situations, thereby achieving better sound quality. This can be achieved by using at least one or a combination of image scene recognition, public platform data (such as meteorological data), and device sensor data (such as GPS information) to determine the current user's target scene. After obtaining the target scene, a corresponding matching sound effect is selected from the sound effect library to enhance the currently denoised audio effect. For example, if image recognition indicates that the subject is a person speaking, a sound effect highlighting the human voice frequency range will be selected and matched accordingly.
[0144] One or more embodiments of this disclosure have at least the following advantages:
[0145] (1) It can distinguish between specific noise signals and ambient sound signals. Through a novel target noise processing model training target design scheme, it is possible to distinguish between specific noise signals and ambient sound signals in the pre-trained target noise processing model. This breaks the original noise reduction training paradigm, innovatively introduces specific noise training targets, and achieves specific noise suppression that preserves ambient sound signals.
[0146] (2) Based on the phase spectrum input to the target noise processing model, specific noise signals and ambient sound signals are distinguished, so as to suppress only specific noise signals. Furthermore, amplitude spectrum and phase spectrum information are simultaneously input into the target noise processing model to better distinguish specific noise signals and ambient sound signals.
[0147] (3) Adaptive noise reduction intensity design based on multiple objectives. While training the target noise processing model, a specific noise existence probability is designed as the output objective, so as to more efficiently calculate the existence probability of a specific noise signal and complete the adaptive noise reduction intensity control.
[0148] (4) Sound effect matching based on different application scenarios. Target scenarios are identified based on phase spectrum information, image recognition, sensor data, or public platform data to match different sound effects, so that the output audio signal better matches the target scenario. Referring to Figure 10, this disclosure also provides an audio processing method, the method including:
[0149] Step S21: Acquire the input audio signal of the target scene, which includes the ambient sound signal and wind noise signal of the target scene; wherein, the ambient sound signal is the useful signal, and the wind noise signal overlaps with the ambient sound signal in the spectrum;
[0150] Step S22: Input the input audio signal into a pre-trained target noise processing model to identify wind noise signals that are distinct from ambient sound signals; and
[0151] Step S23: Suppress the identified wind noise signal to obtain an output audio signal that includes ambient sound signals but excludes wind noise signals.
[0152] In this embodiment, after acquiring the input audio signal, a target noise processing model is used to identify the wind noise signal, which is distinct from the ambient sound signal. The identified wind noise signal is then processed to obtain an output audio signal that includes the ambient sound signal after the wind noise signal has been separated. Since wind noise is a non-periodic and non-stationary signal, it overlaps with the ambient sound in the frequency spectrum. Therefore, it cannot be completely distinguished from the ambient sound using relevant methods. Ambient sound, as a useful signal, should be fully preserved, rather than being suppressed along with specific noise signals, resulting in a loss of ambient sound signal. Therefore, by employing the target noise processing model of this embodiment, the ambient sound signal and wind noise signal, which cannot be completely distinguished, can be completely separated. This allows processing only the wind noise signal to obtain an output audio signal that includes the ambient sound signal but not the wind noise signal. Thus, suppressing the wind noise signal does not affect the ambient sound signal in the output audio signal, thereby ensuring the audio quality of the output audio signal. Better still, compared to related solutions that either remove wind noise but greatly affect useful signals (such as human voice signals and / or ambient sound signals), or where useful signals are not affected but wind noise is too loud, the embodiments of this disclosure can completely remove specific noise signals and obtain ambient sound signals that are completely de-noised and whose quality is not affected, greatly improving the user experience.
[0153] The wind noise signal in this embodiment may be one of the specific noise signals in the foregoing embodiments. For details of the audio processing method, please refer to the foregoing embodiments, which will not be repeated here.
[0154] Referring to Figure 11, this embodiment of the present disclosure also provides a method for obtaining an audio processing model, the method comprising:
[0155] Step S31: Acquire the input sample signal, which includes the ambient sound sample signal and the specific noise sample signal;
[0156] Step S32: Obtain the output signal of the initial noise processing model based on the features of the input sample signal;
[0157] Step S33: Determine the loss function based on the output signal; and
[0158] Step S34: Train the initial noise processing model based on the loss function to obtain the target noise processing model, where the ambient sound sample signal is a useful signal and the specific noise sample signal is a non-periodic and non-stationary signal. The target noise processing model can be used to distinguish between the ambient sound sample signal and the specific noise sample signal.
[0159] In this embodiment of the disclosure, the specific noise sample signal is a non-periodic, non-stationary signal, and the ambient sound sample signal is a useful signal. It is difficult to distinguish between the ambient sound sample signal and the specific noise sample signal in the input sample signal that combines the two. By training the initial noise processing model with the purpose of distinguishing between the ambient sound sample signal and the specific noise sample signal and based on the loss function, a target noise processing model that can distinguish between any input audio signal including the ambient sound signal and the specific noise signal can be trained. Therefore, the target noise processing model can be used to distinguish between the difficult-to-distinguish ambient sound signal and the specific noise signal, so that the quality of the ambient sound signal is not affected when the specific noise signal is processed subsequently.
[0160] In some embodiments, a specific noise sample signal includes a wind noise signal.
[0161] In some embodiments, the specific noise sample signal further includes at least one of the following: blade noise signal, road noise signal, collision noise signal, gimbal noise signal, and mechanical button noise signal.
[0162] In some embodiments, the input sample signal further includes a speech sample signal, and the features of the input sample signal further include the speech features corresponding to the speech sample signal. The speech sample signal is a useful signal, and the target noise processing model can be used to distinguish a specific noise sample signal from the ambient sound sample signal and the speech sample signal.
[0163] In some embodiments, the features of the input sample signal include ambient sound features corresponding to the ambient sound sample signal; determining the loss function based on the output signal includes: obtaining the difference between the ambient sound sample signal and the predicted ambient sound signal output by the initial noise processing model based on the ambient sound features; and determining a first loss function based on the difference; wherein the loss function includes the first loss function.
[0164] In some embodiments, the features of the input sample signal include specific noise features corresponding to specific noise sample signals; determining the loss function based on the output signal includes: obtaining the predicted probability that a specific noise sample signal exists in the input sample signal output by the initial noise processing model based on the specific noise features; and determining a second loss function based on the predicted probability; wherein the loss function includes the second loss function.
[0165] In some embodiments, the input sample signal further includes label data containing sample signals with specific noise, and the second loss function is determined based on the specific noise features and the label data.
[0166] In some embodiments, before obtaining the output signal of the initial noise processing model based on the features of the input sample signal, the method further includes: performing feature extraction on the input sample signal to obtain the features of the input sample signal, wherein the ambient sound sample signal is a pure ambient sound signal without other sounds, and the specific noise sample signal is a pure specific noise signal without containing ambient sound signals; determining the loss function based on the output signal includes: inputting the features of the input sample signal into the initial noise processing model, obtaining a first network output of the initial noise processing model, wherein the first network output is an audio signal after retaining the ambient sound sample signal and suppressing the specific noise sample signal; calculating a first loss function based on the first network output and the ambient sound sample signal; wherein the loss function includes the first loss function.
[0167] In some embodiments, the initial noise processing model is further trained to output a second network output, the second network output being a predicted probability of the presence of the specific noise sample signal, wherein the input sample signal further includes label data indicating the presence of the specific noise sample signal; determining the loss function based on the output signal includes: calculating a second loss function based on the second network output and the label data; wherein the loss function includes the second loss function.
[0168] In some embodiments, training the initial noise processing model based on the loss function to obtain the target noise processing model includes: in response to the loss function satisfying a preset threshold, completing the training of the initial noise processing model to obtain the target noise processing model.
[0169] In some embodiments, the features of the input sample signal include only the phase spectrum; or, the features of the input sample signal include the phase spectrum and at least one of the following: amplitude spectrum and acoustic sensing features.
[0170] In some embodiments, the acoustic sensing features include at least one of the following: frequency band features, equivalent rectangular bandwidth, Buck scale, and Mel scale.
[0171] In some embodiments, the target noise processing model is trained in the cloud.
[0172] In some embodiments, the data of the target noise processing model is stored in the cloud or in a data acquisition device or terminal that is connected to the cloud.
[0173] For details of the training process in this embodiment, please refer to the aforementioned embodiments of the audio processing method, which will not be repeated here.
[0174] Referring to Figure 12, this disclosure also provides a method for obtaining an audio processing model, the method comprising:
[0175] Step S41: Obtain the input sample signal, which includes the ambient sound sample signal and the wind noise sample signal, wherein the ambient sound sample signal is a useful signal and the wind noise sample signal is a non-periodic non-stationary signal;
[0176] Step S42: Obtain the difference between the ambient sound sample signal and the predicted ambient sound signal output by the initial noise processing model based on the ambient sound features of the ambient sound sample signal;
[0177] Step S43: Obtain the predicted probability that the wind noise sample signal exists in the input sample signal, which is output by the initial noise processing model based on the wind noise features of the wind noise sample signal;
[0178] Step S44: Determine the loss function based on the difference and the predicted probability;
[0179] Step S45: Train the initial noise processing model based on the loss function to obtain the target noise processing model, wherein the target noise processing model can be used to distinguish between ambient sound signals and wind noise signals that overlap in the spectrum.
[0180] In this embodiment, the wind noise sample signal is a non-periodic, non-stationary signal, and the ambient sound sample signal is a useful signal. The ambient sound signal and the wind noise signal overlap in their spectra, making them difficult to distinguish in the combined input sample signal. By training the noise processing model with the aim of distinguishing the ambient sound sample signal and the wind noise sample signal, and based on the loss function determined by the difference and prediction probability, a target noise processing model capable of distinguishing any input audio signal including both ambient sound and wind noise signals can be trained. That is, using this target noise processing model, the difficult-to-distinguish ambient sound signal and wind noise signal can be differentiated, so that subsequent suppression of the wind noise signal does not affect the quality of the ambient sound signal.
[0181] The wind noise signal in this embodiment can be one of the specific noise signals in the foregoing embodiments. For details of the training process in this embodiment, please refer to the embodiments of the foregoing audio processing method, which will not be repeated here.
[0182] Referring to Figure 13, this disclosure also provides an audio processing method, the method comprising:
[0183] Step S51: Obtain an output audio signal including the ambient sound signal of the target scene. The output audio signal is obtained by processing the input audio signal based on the target noise processing model. The input audio signal includes the ambient sound signal and a specific noise signal that overlap in the frequency spectrum.
[0184] Step S52: Perform sound effect matching between the target scene and the output audio signal, wherein the ambient sound signal is a useful signal and the specific noise signal is a non-periodic, non-stationary signal.
[0185] In this embodiment, the ambient sound signal is a useful signal, and the specific noise signal is a non-periodic, non-stationary signal. Since the output audio signal is obtained by processing the input audio signal based on the target noise processing model, it can distinguish between the ambient sound signal and the specific noise signal that overlap in the spectrum. This allows for subsequent sound effect matching of the distinguished specific noise signal and the distinguished ambient sound signal for different target scenarios, achieving the effect that related solutions cannot match for sound effects of specific noise signals and ambient sound signals that overlap in the spectrum, thus enriching the user experience.
[0186] In some embodiments, the sound effect matching between the target scene and the output audio signal includes: identifying a sound source in the target scene based on the output audio signal; selecting a matching sound effect from a sound effect library that corresponds to the sound source; and adjusting the audio signal emitted by the sound source in real time using the matching sound effect.
[0187] In some embodiments, the output audio signal is obtained by: inputting the input audio signal into a target noise processing model to identify a specific noise signal that is distinct from the ambient sound signal; and processing the identified specific noise signal to obtain an output audio signal that includes the ambient sound signal after separating the specific noise signal.
[0188] In some embodiments, sound effect matching is performed between the target scene and the output audio signal based on at least one of the following: a scene image of the target scene; information of an audio acquisition device used to acquire the input audio signal; sensor data when the input audio signal is acquired; public platform data when the input audio signal is acquired; and scene information input by the user.
[0189] In some embodiments, the sound effect matching between the target scene and the output audio signal includes: selecting a matching sound effect from a sound effect library that is suitable for the target scene; and performing sound effect matching between the target scene and the output audio signal based on the matching sound effect.
[0190] In some embodiments, in response to the absence of a matching sound effect in the sound effect library corresponding to the sound source, the ambient sound signal in the output audio signal is adjusted based on default parameters, and the user is prompted with the type of the current target scene and the sound effect matching result.
[0191] In some embodiments, the sound effect matching between the target scene and the output audio signal includes: determining a target adjustment method that is compatible with the target scene from a plurality of candidate adjustment methods; and adjusting the output audio signal based on the target adjustment method.
[0192] In some embodiments, the plurality of candidate adjustment methods include at least one of the following: no sound effect adjustment, adjustment of the signal strength of the speech signal in the output audio signal, and adjustment of the signal strength of the ambient sound signal in the output audio signal.
[0193] In some embodiments, the output audio signal is obtained by processing the input audio signal based on a target noise processing model, including: identifying the specific noise signal in the input audio signal through the target noise processing model; adaptively adjusting the processing parameters of the identified specific noise signal; and processing the identified specific noise signal based on the adjusted processing parameters to obtain an output audio signal including the ambient sound signal after separating the specific noise signal.
[0194] In some embodiments, the processing parameters of the identified specific noise signal are adaptively adjusted based on at least one of the following: the probability that the input audio signal includes the specific noise signal, the probability being predicted by the target noise processing model; the scene type of the target scene; and the parameter adjustment instructions obtained by the user.
[0195] In some embodiments, the processing parameters are adaptively adjusted based on the probability that the input audio signal includes the specific noise signal, and the processing parameters include the suppression strength of the identified specific noise signal, the suppression strength being positively correlated with the probability that the input audio signal includes the specific noise signal.
[0196] In some embodiments, the scene type of the target scene is obtained based on at least one of the following: a scene image of the target scene; information of an audio acquisition device for acquiring the input audio signal; sensor data when acquiring the input audio signal; public platform data when acquiring the input audio signal; and scene information input by the user.
[0197] In some embodiments, the specific noise signal includes wind noise.
[0198] In some embodiments, the specific noise signal further includes at least one of the following: blade noise signal, road noise signal, collision noise signal, gimbal noise signal, and mechanical button noise signal.
[0199] In some embodiments, the input audio signal further includes a speech signal that overlaps with the specific noise signal in the spectrum; the step of obtaining an output audio signal including an ambient sound signal of the target scene includes: inputting the input audio signal into the target noise processing model to identify a specific noise signal that is distinguishable from the speech signal and the ambient sound signal; and processing the identified specific noise signal to obtain an output audio signal including the speech signal and the ambient sound signal after separating the specific noise signal.
[0200] In some embodiments, processing the identified specific noise signal includes processing only the identified specific noise signal, wherein the processing includes complete suppression or partial suppression.
[0201] In some embodiments, before inputting the target noise processing model, the proportion of the amplitude spectrum of the target scene's ambient sound signal and the speech signal overlapping in the input audio signal is less than a preset proportion.
[0202] In some embodiments, the target noise processing model is pre-trained to distinguish between the ambient sound signal and the specific noise signal, thereby obtaining a specific noise signal that is separated from the ambient sound signal.
[0203] In some embodiments, after the target noise processing model distinguishes between the ambient sound signal and the specific noise signal, the identified specific noise signal is input into the noise reduction module, which is used to process the identified specific noise signal.
[0204] In some embodiments, the noise reduction module is configured in the target noise processing model to obtain the output audio signal including the ambient sound signal after completely suppressing only the identified specific noise signal.
[0205] In some embodiments, the noise reduction module is located outside the target noise processing model. After the target noise processing model distinguishes between the ambient sound signal and the specific noise signal, the identified specific noise signal is input to the noise reduction module to completely suppress the specific noise signal, and the ambient sound signal is output as the output audio signal.
[0206] In some embodiments, the target noise processing model is trained based on the features of the input sample signal, wherein the input sample signal includes an ambient sound sample signal and a specific noise sample signal, and the features of the input sample signal include the ambient sound features corresponding to the ambient sound sample signal and the specific noise features corresponding to the specific noise sample signal.
[0207] In some embodiments, the target noise processing model is trained based on the features of the input sample signal, including: obtaining the output signal of the initial noise processing model based on the features of the input sample signal; determining a loss function based on the output signal; and training the initial noise processing model based on the loss function to obtain the target noise processing model.
[0208] In some embodiments, determining the loss function based on the output signal includes: obtaining the difference between the ambient sound sample signal and the predicted ambient sound signal output by the initial noise processing model based on the ambient sound features; and determining a first loss function based on the difference; wherein the loss function includes the first loss function.
[0209] In some embodiments, determining the loss function based on the output signal includes: obtaining the predicted probability that the specific noise sample signal exists in the input sample signal output by the initial noise processing model based on the specific noise feature; and determining a second loss function based on the predicted probability; wherein the loss function includes the second loss function.
[0210] In some embodiments, the input sample signal further includes label data containing the specific noisy sample signal, and the second loss function is determined based on the predicted probability and the label data.
[0211] In some embodiments, before obtaining the output signal of the initial noise processing model based on the features of the input sample signal, the method further includes: performing feature extraction on the input sample signal to obtain the features of the input sample signal, wherein the ambient sound sample signal is a pure ambient sound signal without other sounds, and the specific noise sample signal is a pure specific noise signal without containing ambient sound signals; determining the loss function based on the output signal includes: inputting the features of the input sample signal into the initial noise processing model, obtaining a first network output of the initial noise processing model, wherein the first network output is an audio signal after retaining the ambient sound sample signal and suppressing the specific noise sample signal; calculating a first loss function based on the first network output and the ambient sound sample signal; wherein the loss function includes the first loss function.
[0212] In some embodiments, the initial noise processing model is further trained to output a second network output, the second network output being a predicted probability of the presence of the specific noise sample signal, wherein the input sample signal further includes label data indicating the presence of the specific noise sample signal; determining the loss function based on the output signal includes: calculating a second loss function based on the second network output and the label data; wherein the loss function includes the second loss function.
[0213] In some embodiments, training the initial noise processing model based on the loss function to obtain the target noise processing model includes: in response to the loss function satisfying a preset threshold, completing the training of the initial noise processing model to obtain the target noise processing model.
[0214] In some embodiments, the features of the input sample signal include only the phase spectrum; or, the features of the input sample signal include the phase spectrum and at least one of the following: amplitude spectrum and acoustic sensing features.
[0215] In some embodiments, the acoustic sensing features include at least one of the following: frequency band features, equivalent rectangular bandwidth, Buck scale, and Mel scale.
[0216] For details of the training process in this embodiment, please refer to the aforementioned embodiments of the audio processing method, which will not be repeated here.
[0217] Referring to Figure 14, this disclosure also provides an audio processing method, the method comprising:
[0218] Step S61: Input the input audio signal into the target noise processing model. The input audio signal includes the ambient sound signal and the wind noise signal of the target scene. The ambient sound signal and the wind noise signal overlap in the spectrum.
[0219] Step S62: Identify the wind noise signal that is distinct from the ambient sound signal using the target noise processing model;
[0220] Step S63: Process the identified wind noise signal to obtain an output audio signal including the ambient sound signal after separating the wind noise signal; and
[0221] Step S64: Perform sound effect matching between the target scene and the output audio signal.
[0222] In this embodiment of the disclosure, the ambient sound signal is a useful signal. Since the output audio signal is obtained by processing the input audio signal based on the target noise processing model, it can distinguish the ambient sound signal and the wind noise signal that overlap in the spectrum. Therefore, the output audio signal of the ambient sound signal after separating the wind noise signal can be matched with sound effects for different target scenarios. That is, sound effect matching can be performed on the ambient sound signal with only the wind noise removed, thus enriching the user experience.
[0223] The wind noise signal in this embodiment can be one of the specific noise signals in the foregoing embodiments. For details of the training process in this embodiment, please refer to the embodiments of the foregoing audio processing method, which will not be repeated here.
[0224] Referring to Figure 15, this disclosure also provides an audio processing apparatus, including:
[0225] At least one processor 100; and
[0226] At least one memory 200 containing a computer program,
[0227] In this embodiment, at least one of the memory 200 cooperates with at least one of the processors 100 to enable the device to perform at least the method described in any of the foregoing embodiments.
[0228] Referring to Figure 16, this embodiment of the present disclosure also provides an audio acquisition device, the audio acquisition device comprising:
[0229] Microphone 300 is used to acquire input audio signals from a target scene, the input audio signals including ambient sound signals and specific noise signals from the target scene; the ambient sound signals are useful signals, and the specific noise signals are non-periodic and non-stationary signals.
[0230] The processor 100 is communicatively connected to the microphone and is used to receive the input audio signal collected by the microphone and process the input audio signal using the audio processing method described in any of the above embodiments.
[0231] The audio acquisition device provided in this disclosure, in addition to the beneficial effects of the audio processing method described in any embodiment, can also process the acquired input audio signal in real time based on its own built-in processor. It can integrate the recording end, processing end and output end into one unit, making it easy to carry and operate, and greatly improving the user experience.
[0232] Referring to Figure 17, this disclosure also provides a multimedia acquisition system, which includes:
[0233] The audio acquisition device 400 is used to acquire input audio signals of a target scene, wherein the input audio signals include ambient sound signals and specific noise signals of the target scene; the ambient sound signals are useful signals, and the specific noise signals are non-periodic and non-stationary signals.
[0234] The video acquisition device 500 is communicatively connected to the audio acquisition device 400, or the audio acquisition device is installed on the video acquisition device, for acquiring video of the target scene, and for setting the audio acquisition parameters and various functions of the audio acquisition device by switching them on and off and / or setting their levels.
[0235] At least one processor 100 is disposed in the audio acquisition device and / or the video acquisition device for processing the input audio signal using the audio processing method described in any of the above embodiments.
[0236] In addition to the beneficial effects of the audio processing method described in any embodiment, the multimedia acquisition system provided by this disclosure can also switch on / off (wind noise reduction switch) and / or set the level (wind noise reduction level) of the audio acquisition parameters (e.g., acquisition duration) and various functions of the audio acquisition device based on the video acquisition device that is communicatively connected to the audio acquisition device. Since the audio acquisition device generally does not have a screen or is small in size and difficult to operate, by backing up the control function on the video acquisition device and synthesizing audio and image data on the video acquisition device, the audio and video combined data can be directly acquired and output, resulting in rich data, fast output, and greatly improved user experience.
[0237] Referring to Figure 18, this disclosure also provides a multimedia acquisition system, which includes:
[0238] The audio acquisition device 400 is used to acquire input audio signals of a target scene, wherein the input audio signals include ambient sound signals and specific noise signals of the target scene; the ambient sound signals are useful signals, and the specific noise signals are non-periodic and non-stationary signals.
[0239] The video acquisition device 500 is used to acquire video of the target scene and to switch and / or set the audio acquisition parameters and various functions of the audio acquisition device.
[0240] Terminal 600 is communicatively connected to the audio acquisition device and / or the video acquisition device, and is used to switch and / or set the audio acquisition parameters and various functions of the audio acquisition device, and to switch and / or set the various functions of the video acquisition device.
[0241] At least one processor 100 is disposed on at least one of the audio acquisition device 400, the video acquisition device 500, and the device terminal 600, for processing the input audio signal using the audio processing method described in any of the above embodiments.
[0242] The multimedia acquisition system provided in this disclosure, in addition to the beneficial effects of the audio processing method described in any embodiment, can also enable on / off settings (wind noise reduction switch) and / or level settings (wind noise reduction level) for the audio acquisition parameters (e.g., acquisition duration) and various functions of the audio acquisition device based on a video acquisition device communicatively connected to the audio acquisition device; or, through a terminal communicatively connected to the audio acquisition device and / or the video acquisition device, enable on / off settings and / or level settings for the audio acquisition parameters and various functions of the audio acquisition device, as well as for enabling on / off settings and / or level settings for various functions of the video acquisition device. Since audio acquisition devices generally do not have a screen or are small in size and difficult to operate, by backing up this control function on the video acquisition device or terminal, and synthesizing audio and image data on the video acquisition device or terminal, audio and video combined data can be directly acquired and output, resulting in rich data, fast output, and a significantly improved user experience.
[0243] The various technical features in the above embodiments can be combined arbitrarily, as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they have not been described one by one. Therefore, the arbitrary combination of various technical features in the above embodiments also falls within the scope of this disclosure.
[0244] This disclosure can take the form of a computer program product implemented on one or more storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing program code. Computer-usable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to: phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0245] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the description disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0246] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
[0247] The above description is merely a preferred embodiment of this disclosure and is not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An audio processing method, characterized in that, The method includes: The input audio signal of the target scene is acquired. The input audio signal includes the ambient sound signal and the specific noise signal of the target scene. The ambient sound signal is a useful signal and the specific noise signal is a non-periodic and non-stationary signal. The input audio signal is input into the target noise processing model to identify specific noise signals that are distinguishable from the ambient sound signal; and The identified specific noise signal is processed to obtain an output audio signal that includes the ambient sound signal after separating the specific noise signal.
2. The method according to claim 1, characterized in that, The acquisition of the input audio signal of the target scene includes: The target scene is picked up by a recording element to obtain the input audio signal of the target scene.
3. The method according to claim 2, characterized in that, The recording element includes multiple microphones, which are arranged at different locations on the audio acquisition device; The step of picking up sound from the target scene using a recording element includes: The target scene is picked up simultaneously using multiple microphones.
4. The method according to claim 3, characterized in that, The step of picking up sound from the target scene using a recording element to obtain the input audio signal of the target scene includes: Correlation operations are performed on the multiple sets of audio signals picked up by the multiple microphones to obtain the audio signal with the minimum specific noise signal; The audio signal with the lowest noise level is used as the input audio signal for the target scene.
5. The method according to claim 3, characterized in that, The step of picking up sound from the target scene using a recording element to obtain the input audio signal of the target scene includes: Multiple sets of audio signals of the target scene obtained by the multiple microphones are used as the input audio signals of the target scene.
6. The method according to claim 1, characterized in that, Before the target noise processing model is input, the ambient sound signal of the target scene in the input audio signal overlaps with the specific noise signal in the spectrum.
7. The method according to claim 1, characterized in that, The specific noise signal includes wind noise.
8. The method according to claim 7, characterized in that, The specific noise signal also includes at least one of the following: blade noise signal, road noise signal, collision noise signal, gimbal noise signal, and mechanical button noise signal.
9. The method according to claim 1, characterized in that, The input audio signal also includes a speech signal; the step of inputting the input audio signal into the target noise processing model to identify specific noise signals that are distinguishable from the ambient sound signal includes: The input audio signal is fed into the target noise processing model to identify specific noise signals that are distinguishable from the speech signal and the ambient sound signal; The process of processing the identified specific noise signal to obtain an output audio signal including the ambient sound signal after separating the specific noise signal includes: The identified specific noise signal is processed to obtain an output audio signal that includes the speech signal after separating the specific noise signal and the ambient sound signal.
10. The method according to claim 9, characterized in that, The processing of the identified specific noise signal includes: Only the identified specific noise signal is processed, wherein the processing includes complete suppression or partial suppression.
11. The method according to claim 9, characterized in that, Before inputting the target noise processing model, the proportion of the amplitude spectrum of the target scene's ambient sound signal and the speech signal overlapping in the input audio signal is less than a preset proportion.
12. The method according to claim 1, characterized in that, The target noise processing model is pre-trained to distinguish between the ambient sound signal and the specific noise signal, thereby obtaining the specific noise signal that is separated from the ambient sound signal.
13. The method according to claim 12, characterized in that, After distinguishing between the ambient sound signal and the specific noise signal, the target noise processing model identifies the specific noise signal, which is then input into the noise reduction module. The noise reduction module is used to process the identified specific noise signal.
14. The method according to claim 13, characterized in that, The noise reduction module is set in the target noise processing model to completely suppress only the identified specific noise signals to obtain the output audio signal including the ambient sound signal.
15. The method according to claim 13, characterized in that, The noise reduction module is located outside the target noise processing model. After the target noise processing model distinguishes between the ambient sound signal and the specific noise signal, the identified specific noise signal is input into the noise reduction module to completely suppress the specific noise signal, and the ambient sound signal is output as the output audio signal.
16. The method according to claim 12, characterized in that, The target noise processing model is trained based on the features of the input sample signal. The input sample signal includes an ambient sound sample signal and a specific noise sample signal. The features of the input sample signal include the ambient sound features corresponding to the ambient sound sample signal and the specific noise features corresponding to the specific noise sample signal.
17. The method according to claim 16, characterized in that, The target noise processing model is trained based on the features of the input sample signal, including: Obtain the output signal of the initial noise processing model based on the features of the input sample signal; Based on the output signal, determine the loss function; and The initial noise processing model is trained based on the loss function to obtain the target noise processing model.
18. The method according to claim 17, characterized in that, The determination of the loss function based on the output signal includes: The difference between the ambient sound sample signal and the predicted ambient sound signal output by the initial noise processing model based on the ambient sound features is obtained; Based on the difference, a first loss function is determined; wherein the loss function includes the first loss function.
19. The method according to claim 17, characterized in that, The determination of the loss function based on the output signal includes: Obtain the predicted probability that the specific noise sample signal exists in the input sample signal, based on the specific noise feature, output by the initial noise processing model; Based on the predicted probability, a second loss function is determined; wherein the loss function includes the second loss function.
20. The method according to claim 19, characterized in that, The input sample signal also includes label data containing the specific noisy sample signal, and the second loss function is determined based on the predicted probability and the label data.
21. The method according to claim 17, characterized in that, Before obtaining the output signal of the initial noise processing model based on the features of the input sample signal, the method further includes: Feature extraction is performed on the input sample signal to obtain the features of the input sample signal, wherein the ambient sound sample signal is a pure ambient sound signal without any other sounds, and the specific noise sample signal is a pure specific noise signal without containing ambient sound signals; Determining the loss function based on the output signal includes: The features of the input sample signal are input into the initial noise processing model to obtain the first network output of the initial noise processing model. The first network output is an audio signal that retains the ambient sound sample signal and suppresses the specific noise sample signal. Based on the output of the first network and the ambient sound sample signal, a first loss function is calculated; wherein, the loss function includes the first loss function.
22. The method according to claim 17, characterized in that, The initial noise processing model is also trained to output a second network output, which is a predicted probability of the existence of the specific noise sample signal, wherein the input sample signal also includes label data of the existence of the specific noise sample signal; Determining the loss function based on the output signal includes: Based on the second network output and the label data, a second loss function is calculated; wherein the loss function includes the second loss function.
23. The method according to any one of claims 17-22, characterized in that, The step of training the initial noise processing model based on the loss function to obtain the target noise processing model includes: In response to the loss function satisfying the preset training termination condition, the training of the initial noise processing model is completed to obtain the target noise processing model.
24. The method according to claim 16, characterized in that, The characteristics of the input sample signal include only the phase spectrum; or... The features of the input sample signal include the phase spectrum and at least one of the following: amplitude spectrum and acoustic sensing features.
25. The method according to claim 24, characterized in that, The acoustic sensing features include at least one of the following: frequency band features, equivalent rectangular bandwidth, Buck scale, and Mel scale.
26. The method according to claim 12, characterized in that, The target noise processing model is pre-trained in the cloud to distinguish between the ambient sound signal and the specific noise signal in real time.
27. The method according to claim 1, characterized in that, The processing of the identified specific noise signal includes: The processing parameters for the identified specific noise signal are adaptively adjusted. The identified specific noise signal is processed based on the adjusted processing parameters.
28. The method according to claim 27, characterized in that, The processing parameters of the identified specific noise signal are adaptively adjusted based on at least one of the following: The probability that the input audio signal includes the specific noise signal is predicted by the target noise processing model. The scene type of the target scene; and The parameter adjustment instructions obtained from user input.
29. The method according to claim 28, characterized in that, The processing parameters are adaptively adjusted based on the probability that the input audio signal includes the specific noise signal. The processing parameters include the suppression strength of the specific noise signal, which is positively correlated with the probability that the input audio signal includes the specific noise signal.
30. The method according to claim 1, characterized in that, The ambient sound signal after separating the specific noise signal is used as at least a part of the output audio signal for playback, transmission, amplification, enhancement, or speech recognition.
31. The method according to claim 30, characterized in that, Before the ambient sound signal, after separating the specific noise signal, is used as at least a portion of the output audio signal for playback, transmission, amplification, enhancement, or speech recognition, the method further includes: The output audio signal is matched with the target scene for sound effects.
32. The method according to claim 31, characterized in that, The step of matching the output audio signal with the target scene in terms of sound effects includes: Select matching sound effects from the sound effects library that are suitable for the target scene; and The ambient sound signal in the output audio signal is adjusted accordingly based on the matching sound effect.
33. The method according to claim 32, characterized in that, In response to the absence of a matching sound effect in the sound effect library that matches the target scene, the ambient sound signal in the output audio signal is adjusted based on the default parameters, and the user is prompted with the type of the target scene and the sound effect matching result.
34. The method according to claim 31, characterized in that, The output audio signal is matched with the target scene based on at least one of the following information: The scene image of the target scene; Information about the audio acquisition device used to acquire the input audio signal; Sensor data is collected when the input audio signal is received; Public platform data collected when the input audio signal is acquired; as well as User-input scene information.
35. The method according to claim 31, characterized in that, The step of matching the output audio signal with the target scene in terms of sound effects includes: From multiple candidate adjustment methods, a target adjustment method suitable for the target scenario is determined; and The output audio signal is adjusted based on the target adjustment method.
36. The method according to claim 35, characterized in that, The plurality of candidate adjustment methods include at least one of the following: no sound effect processing, adjusting the signal strength of the speech signal in the output audio signal, and adjusting the signal strength of the ambient sound signal in the output audio signal.
37. The method according to claim 1, characterized in that, In response to the user's selection operation, the target noise processing model is selected from multiple candidate noise processing models.
38. The method according to claim 37, characterized in that, The data of the candidate noise processing model is stored in the cloud or in a data acquisition device or terminal that is connected to the cloud.
39. The method according to claim 1, characterized in that, The processing of the identified specific noise signal includes: The probability of the presence of the specific noise signal in the input audio signal is obtained through the target noise processing model; and The specific noise signal is processed based on the probability.
40. An audio processing method, characterized in that, The method includes: The input audio signal of the target scene is acquired, which includes the ambient sound signal and wind noise signal of the target scene; wherein the ambient sound signal is a useful signal, and the wind noise signal overlaps with the ambient sound signal in the spectrum; The input audio signal is fed into a pre-trained target noise processing model to identify wind noise signals that are distinct from the ambient sound signal; and The identified wind noise signal is suppressed to obtain an output audio signal that includes the ambient sound signal but does not include the wind noise signal.
41. An audio processing apparatus, characterized in that, include: At least one processor; as well as At least one memory containing a computer program, Wherein, at least one of the memory cooperates with at least one of the processor to cause the device to perform at least the following operations: Acquire the input audio signal of the target scene, wherein the input audio signal includes the ambient sound signal and a specific noise signal of the target scene; The input audio signal is input into the target noise processing model to identify specific noise signals that are distinguishable from the ambient sound signal; and The identified specific noise signal is processed to obtain a result including the separation of the specific noise signal. The output audio signal of the ambient sound signal is a useful signal, and the specific noise signal is a non-periodic, non-stationary signal.
42. An audio acquisition device, characterized in that, The audio acquisition device includes: A microphone is used to collect input audio signals from a target scene. The input audio signals include ambient sound signals and specific noise signals from the target scene. The ambient sound signals are useful signals, and the specific noise signals are non-periodic and non-stationary signals. The processor, communicatively connected to the microphone, is used to receive the input audio signal collected by the microphone and process the input audio signal. The processing of the input audio signal includes: The input audio signal is fed into a target noise processing model to identify specific noise signals that are distinguishable from the ambient sound signal; and The identified specific noise signal is processed to obtain an output audio signal that includes the ambient sound signal after separating the specific noise signal.
43. A multimedia acquisition system, characterized in that, The multimedia acquisition system includes: An audio acquisition device is used to acquire input audio signals from a target scene. The input audio signals include ambient sound signals and specific noise signals from the target scene. The ambient sound signals are useful signals, and the specific noise signals are non-periodic and non-stationary signals. A video acquisition device is communicatively connected to the audio acquisition device, or the audio acquisition device is installed on the video acquisition device, for acquiring video of the target scene, and for setting the audio acquisition parameters and various functions of the audio acquisition device to switch on and / or to adjust the level. At least one processor, disposed in the audio acquisition device and / or the video acquisition device, is used to process the input audio signal. The processing of the input audio signal includes: The input audio signal is fed into a target noise processing model to identify specific noise signals that are distinguishable from the ambient sound signal; and The identified specific noise signal is processed to obtain an output audio signal that includes the ambient sound signal after separating the specific noise signal.
44. A multimedia acquisition system, characterized in that, The multimedia acquisition system includes: An audio acquisition device is used to acquire input audio signals from a target scene. The input audio signals include ambient sound signals and specific noise signals from the target scene. The ambient sound signals are useful signals, and the specific noise signals are non-periodic and non-stationary signals. A video acquisition device is used to acquire video of the target scene and to switch and / or set the audio acquisition parameters and various functions of the audio acquisition device. The terminal is communicatively connected to the audio acquisition device and / or the video acquisition device, and is used to switch and / or set the audio acquisition parameters and various functions of the audio acquisition device, and to switch and / or set the various functions of the video acquisition device. At least one processor is disposed on at least one of the audio acquisition device, the video acquisition device, and the device terminal, for processing the input audio signal. The processing of the input audio signal includes: The input audio signal is fed into a target noise processing model to identify specific noise signals that are distinguishable from the ambient sound signal; and The identified specific noise signal is processed to obtain an output audio signal that includes the ambient sound signal after separating the specific noise signal.
45. A method for obtaining an audio processing model, characterized in that, The method includes: Acquire input sample signals, which include ambient sound sample signals and specific noise sample signals; Obtain the output signal of the initial noise processing model based on the features of the input sample signal; Based on the output signal, determine the loss function; and The initial noise processing model is trained based on the loss function to obtain the target noise processing model, wherein the ambient sound sample signal is a useful signal, the specific noise sample signal is a non-periodic and non-stationary signal, and the target noise processing model can be used to distinguish between the ambient sound sample signal and the specific noise sample signal.
46. The method according to claim 45, characterized in that, The specific noise sample signal includes wind noise signal.
47. The method according to claim 46, characterized in that, The specific noise sample signal also includes at least one of the following: blade noise signal, road noise signal, collision noise signal, gimbal noise signal, and mechanical button noise signal.
48. The method according to claim 45, characterized in that, The input sample signal also includes a speech sample signal, and the features of the input sample signal also include the speech features corresponding to the speech sample signal. The speech sample signal is a useful signal, and the target noise processing model can be used to distinguish the specific noise sample signal from the ambient sound sample signal and the speech sample signal.
49. The method according to claim 45, characterized in that, The features of the input sample signal include the ambient sound features corresponding to the ambient sound sample signal; determining the loss function based on the output signal includes: The difference between the ambient sound sample signal and the predicted ambient sound signal output by the initial noise processing model based on the ambient sound features is obtained; Based on the difference, a first loss function is determined; wherein the loss function includes the first loss function.
50. The method according to claim 45, characterized in that, The features of the input sample signal include specific noise features corresponding to the specific noise sample signal; determining the loss function based on the output signal includes: Obtain the predicted probability that the specific noise sample signal exists in the input sample signal, based on the specific noise feature, output by the initial noise processing model; Based on the predicted probability, a second loss function is determined; wherein the loss function includes the second loss function.
51. The method according to claim 50, characterized in that, The input sample signal also includes label data containing the specific noise sample signal, and the second loss function is determined based on the specific noise feature and the label data.
52. The method according to claim 45, characterized in that, Before obtaining the output signal of the initial noise processing model based on the features of the input sample signal, the method further includes: Feature extraction is performed on the input sample signal to obtain the features of the input sample signal, wherein the ambient sound sample signal is a pure ambient sound signal without any other sounds, and the specific noise sample signal is a pure specific noise signal without containing ambient sound signals; Determining the loss function based on the output signal includes: The features of the input sample signal are input into the initial noise processing model to obtain the first network output of the initial noise processing model. The first network output is an audio signal that retains the ambient sound sample signal and suppresses the specific noise sample signal. Based on the output of the first network and the ambient sound sample signal, a first loss function is calculated; wherein, the loss function includes the first loss function.
53. The method according to claim 45, characterized in that, The initial noise processing model is also trained to output a second network output, which is a predicted probability of the existence of the specific noise sample signal, wherein the input sample signal also includes label data of the existence of the specific noise sample signal; The determination of the loss function based on the output signal includes: Based on the second network output and the label data, a second loss function is calculated; wherein the loss function includes the second loss function.
54. The method according to any one of claims 45-53, characterized in that, The step of training the initial noise processing model based on the loss function to obtain the target noise processing model includes: In response to the loss function satisfying the preset conditions, the initial noise processing model is trained to obtain the target noise processing model.
55. The method according to claim 45, characterized in that, The characteristics of the input sample signal include only the phase spectrum; or... The features of the input sample signal include the phase spectrum and at least one of the following: amplitude spectrum and acoustic sensing features.
56. The method according to claim 55, characterized in that, The acoustic sensing features include at least one of the following: frequency band features, equivalent rectangular bandwidth, Buck scale, and Mel scale.
57. The method according to claim 45, characterized in that, The target noise processing model is trained in the cloud.
58. The method according to claim 45, characterized in that, The data of the target noise processing model is stored in the cloud or in a data acquisition device or terminal that is connected to the cloud.
59. A method for obtaining an audio processing model, characterized in that, The method includes: Acquire input sample signals, which include ambient sound sample signals and wind noise sample signals, wherein the ambient sound sample signals are useful signals and the wind noise sample signals are non-periodic and non-stationary signals; The difference between the ambient sound sample signal and the predicted ambient sound signal output by the initial noise processing model based on the ambient sound features of the ambient sound sample signal is obtained; The predicted probability that the wind noise sample signal exists in the input sample signal is obtained from the output of the initial noise processing model based on the wind noise characteristics of the wind noise sample signal; Based on the difference and the predicted probability, determine the loss function; The initial noise processing model is trained based on the loss function to obtain the target noise processing model, wherein the target noise processing model can be used to distinguish between ambient sound signals and wind noise signals that overlap in the spectrum.
60. An audio processing model acquisition device, characterized in that, include: At least one processor; as well as At least one memory containing a computer program, Wherein, at least one of the memories cooperates with at least one of the processors to enable the device to perform at least... The following steps are required: Acquire input sample signals, which include ambient sound sample signals and specific noise sample signals; Obtain the output signal of the initial noise processing model based on the features of the input sample signal; Based on the output signal, determine the loss function; and The initial noise processing model is trained based on the loss function to obtain the target noise processing model, wherein the ambient sound sample signal is a useful signal, the specific noise sample signal is a non-periodic and non-stationary signal, and the target noise processing model can be used to distinguish between the ambient sound sample signal and the specific noise sample signal.
61. An audio processing method, characterized in that, The method includes: An output audio signal is acquired, comprising ambient sound signals of the target scene. The output audio signal is obtained by processing the input audio signal based on a target noise processing model. The input audio signal includes the ambient sound signals and a specific noise signal that overlap in their spectra. The target scene is matched with the output audio signal for sound effects, wherein the ambient sound signal is a useful signal and the specific noise signal is a non-periodic, non-stationary signal.
62. The method according to claim 61, characterized in that, The sound effect matching between the target scene and the output audio signal includes: Based on the output audio signal, identify the sound source in the target scene; Select a matching sound effect from the sound effect library that corresponds to the sound source; and The audio signal emitted by the sound source is adjusted in real time using the matching sound effect.
63. The method according to claim 61, characterized in that, The output audio signal is obtained in the following way: The input audio signal is input into the target noise processing model to identify specific noise signals that are distinguishable from the ambient sound signal; and The identified specific noise signal is processed to obtain an output audio signal that includes the ambient sound signal after separating the specific noise signal.
64. The method according to claim 61, characterized in that, Based on at least one of the following pieces of information, sound effect matching is performed between the target scene and the output audio signal: The scene image of the target scene; Information about the audio acquisition device used to acquire the input audio signal; Sensor data is collected when the input audio signal is received; Public platform data collected when the input audio signal is acquired; User-input scene information.
65. The method according to any one of claims 61 to 64, characterized in that, The sound effect matching between the target scene and the output audio signal includes: Select matching sound effects from the sound effects library that are suitable for the target scene; and Based on the matching sound effects, sound effect matching is performed between the target scene and the output audio signal.
66. The method according to claim 65, characterized in that, In response to the absence of a matching sound effect in the sound effect library corresponding to the sound source, the ambient sound signal in the output audio signal is adjusted based on the default parameters, and the user is prompted with the type of the current target scene and the sound effect matching result.
67. The method according to claim 61, characterized in that, The sound effect matching between the target scene and the output audio signal includes: From multiple candidate adjustment methods, a target adjustment method suitable for the target scenario is determined; and The output audio signal is adjusted based on the target adjustment method.
68. The method according to claim 67, characterized in that, The plurality of candidate adjustment methods include at least one of the following: no sound effect adjustment, adjustment of the signal strength of the speech signal in the output audio signal, and adjustment of the signal strength of the ambient sound signal in the output audio signal.
69. The method according to claim 61, characterized in that, The output audio signal is obtained by processing the input audio signal based on the target noise processing model, including: The specific noise signal in the input audio signal is identified using the target noise processing model. The processing parameters for the identified specific noise signal are adaptively adjusted; and Based on the adjusted processing parameters, the identified specific noise signal is processed to obtain an output audio signal including the ambient sound signal after separating the specific noise signal.
70. The method according to claim 69, characterized in that, The processing parameters of the identified specific noise signal are adaptively adjusted based on at least one of the following: The probability that the input audio signal includes the specific noise signal is predicted by the target noise processing model. The scene type of the target scene; and The parameter adjustment instructions obtained from user input.
71. The method according to claim 70, characterized in that, The processing parameters are adaptively adjusted based on the probability that the input audio signal includes the specific noise signal. The processing parameters include the suppression strength of the identified specific noise signal, and the suppression strength is positively correlated with the probability that the input audio signal includes the specific noise signal.
72. The method according to claim 70, characterized in that, The scene type of the target scene is obtained based on at least one of the following: The scene image of the target scene; Information about the audio acquisition device used to acquire the input audio signal; Sensor data is collected when the input audio signal is received; Public platform data collected when the input audio signal is acquired; as well as User-input scene information.
73. The method according to claim 61, characterized in that, The specific noise signal includes wind noise.
74. The method according to claim 73, characterized in that, The specific noise signal also includes at least one of the following: blade noise signal, road noise signal, collision noise signal, gimbal noise signal, and mechanical button noise signal.
75. The method according to claim 61, characterized in that, The input audio signal also includes a speech signal that overlaps with the specific noise signal in the frequency spectrum; the acquisition of the output audio signal including the ambient sound signal of the target scene includes: The input audio signal is fed into the target noise processing model to identify specific noise signals that are distinguishable from the speech signal and the ambient sound signal; The identified specific noise signal is processed to obtain an output audio signal that includes the speech signal after separating the specific noise signal and the ambient sound signal.
76. The method according to claim 75, characterized in that, The processing of the identified specific noise signal includes: Only the identified specific noise signal is processed, wherein the processing includes complete suppression or partial suppression.
77. The method according to claim 75, characterized in that, Before inputting the target noise processing model, the proportion of the amplitude spectrum of the target scene's ambient sound signal and the speech signal overlapping in the input audio signal is less than a preset proportion.
78. The method according to claim 61, characterized in that, The target noise processing model is pre-trained to distinguish between the ambient sound signal and the specific noise signal, thereby obtaining the specific noise signal that is separated from the ambient sound signal.
79. The method according to claim 78, characterized in that, After distinguishing between the ambient sound signal and the specific noise signal, the target noise processing model identifies the specific noise signal, which is then input into the noise reduction module. The noise reduction module is used to process the identified specific noise signal.
80. The method according to claim 79, characterized in that, The noise reduction module is set in the target noise processing model to completely suppress only the identified specific noise signals to obtain the output audio signal including the ambient sound signal.
81. The method according to claim 79, characterized in that, The noise reduction module is located outside the target noise processing model. After the target noise processing model distinguishes between the ambient sound signal and the specific noise signal, the identified specific noise signal is input into the noise reduction module to completely suppress the specific noise signal, and the ambient sound signal is output as the output audio signal.
82. The method according to claim 78, characterized in that, The target noise processing model is trained based on the features of the input sample signal. The input sample signal includes an ambient sound sample signal and a specific noise sample signal. The features of the input sample signal include the ambient sound features corresponding to the ambient sound sample signal and the specific noise features corresponding to the specific noise sample signal.
83. The method according to claim 82, characterized in that, The target noise processing model is trained based on the features of the input sample signal, including: Obtain the output signal of the initial noise processing model based on the features of the input sample signal; Based on the output signal, determine the loss function; and The initial noise processing model is trained based on the loss function to obtain the target noise processing model.
84. The method according to claim 83, characterized in that, The determination of the loss function based on the output signal includes: The difference between the ambient sound sample signal and the predicted ambient sound signal output by the initial noise processing model based on the ambient sound features is obtained; Based on the difference, a first loss function is determined; wherein the loss function includes the first loss function.
85. The method according to claim 83, characterized in that, The determination of the loss function based on the output signal includes: Obtain the predicted probability that the specific noise sample signal exists in the input sample signal, based on the specific noise feature, output by the initial noise processing model; Based on the predicted probability, a second loss function is determined; wherein the loss function includes the second loss function.
86. The method according to claim 85, characterized in that, The input sample signal also includes label data containing the specific noisy sample signal, and the second loss function is determined based on the predicted probability and the label data.
87. The method according to claim 83, characterized in that, Before obtaining the output signal of the initial noise processing model based on the features of the input sample signal, the method further includes: Feature extraction is performed on the input sample signal to obtain the features of the input sample signal, wherein the ambient sound sample signal is a pure ambient sound signal without any other sounds, and the specific noise sample signal is a pure specific noise signal without containing ambient sound signals; Determining the loss function based on the output signal includes: The features of the input sample signal are input into the initial noise processing model to obtain the first network output of the initial noise processing model. The first network output is an audio signal that retains the ambient sound sample signal and suppresses the specific noise sample signal. Based on the output of the first network and the ambient sound sample signal, a first loss function is calculated; wherein, the loss function includes the first loss function.
88. The method according to claim 83, characterized in that, The initial noise processing model is also trained to output a second network output, which is a predicted probability of the existence of the specific noise sample signal, wherein the input sample signal also includes label data of the existence of the specific noise sample signal; Determining the loss function based on the output signal includes: Based on the second network output and the label data, a second loss function is calculated; wherein the loss function includes the second loss function.
89. The method according to any one of claims 83-88, characterized in that, The step of training the initial noise processing model based on the loss function to obtain the target noise processing model includes: In response to the loss function satisfying a preset threshold, the initial noise processing model is trained to obtain the target noise processing model.
90. The method according to claim 82, characterized in that, The characteristics of the input sample signal include only the phase spectrum; or... The features of the input sample signal include the phase spectrum and at least one of the following: amplitude spectrum and acoustic sensing features.
91. The method according to claim 90, characterized in that, The acoustic sensing features include at least one of the following: frequency band features, equivalent rectangular bandwidth, Buck scale, and Mel scale.
92. An audio processing method, characterized in that, The method includes: The input audio signal is fed into the target noise processing model, and the input audio signal includes the environment of the target scene. The ambient sound signal and the wind noise signal overlap in their frequency spectrum; The target noise processing model is used to identify wind noise signals that are distinct from the ambient sound signals. The identified wind noise signal is processed to obtain an output audio signal including the ambient sound signal after separating the wind noise signal; and The target scene is matched with the output audio signal for sound effects.
93. An audio processing apparatus, characterized in that, include: At least one processor; as well as At least one memory containing a computer program, Wherein, at least one of the memory cooperates with at least one of the processor to cause the device to perform at least the following operations: An output audio signal is acquired, comprising ambient sound signals of the target scene. The output audio signal is obtained by processing the input audio signal based on a target noise processing model. The input audio signal includes the ambient sound signals and a specific noise signal that overlap in their spectra. The target scene is matched with the output audio signal for sound effects, wherein the ambient sound signal is a useful signal and the specific noise signal is a non-periodic, non-stationary signal.
94. An audio acquisition device, characterized in that, The audio acquisition device includes: A microphone is used to collect input audio signals from a target scene. The input audio signals include ambient sound signals and specific noise signals from the target scene. The ambient sound signals are useful signals, and the specific noise signals are non-periodic and non-stationary signals. The ambient sound signals and the specific noise signals in the input audio signals overlap in their frequency spectra. The processor, communicatively connected to the microphone, is used to receive the input audio signal collected by the microphone and process the input audio signal. The processing of the input audio signal includes: The input audio signal is processed based on the target noise processing model to obtain an output audio signal including the ambient sound signal; and The target scene is matched with the output audio signal for sound effects.
95. A multimedia acquisition system, characterized in that, The multimedia acquisition system includes: An audio acquisition device is used to acquire input audio signals from a target scene, wherein the input audio signals include... The target scene includes ambient sound signals and specific noise signals; the ambient sound signals are useful signals, and the specific noise signals are non-periodic and non-stationary signals, and the ambient sound signals and specific noise signals in the input audio signal overlap in the frequency spectrum. A video acquisition device is communicatively connected to the audio acquisition device, or the audio acquisition device is installed on the video acquisition device, for acquiring video of the target scene, and for setting the audio acquisition parameters and various functions of the audio acquisition device to switch on and / or to adjust the level. At least one processor, disposed in the audio acquisition device and / or the video acquisition device, is used to process the input audio signal. The processing of the input audio signal includes: The input audio signal is processed based on the target noise processing model to obtain an output audio signal including the ambient sound signal; and The target scene is matched with the output audio signal for sound effects.
96. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the method described in any one of claims 1-40, 45-59, and 61-92.
97. A computer program product having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the method described in any one of claims 1-40, 45-59, and 61-92.