Audio processing method and apparatus, readable medium, and electronic device
By combining bone conduction and gas conduction microphone signal processing models, the directional beam and noise audio characteristics are extracted, and the problems of insufficient microphone number and high-frequency signal loss on wearable devices are solved, achieving high-quality enhancement and simplified operation of audio signals.
Patent Information
- Application Number
- PCT/CN2025/079290
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-29
- Filing Date
- 2025-02-26
- Publication Date
- 2025-09-04
AI Technical Summary
The number of air-conducting microphones on wearable devices is small, resulting in poor directional voice enhancement effect, serious high-frequency signal loss of bone-conducting microphones and large calculation amount, and the calculation amount of voiceprint speaker separation scheme is large and the process is troublesome.
Combining bone conduction and gas conduction microphone signals, the directional beam audio characteristics and noisy audio characteristics are extracted through the signal processing model, the audio signals in the specified direction are enhanced and audio signals in the non-specified direction are suppressed, and the miniaturization model is used to deploy on wearable devices.
It improves the quality of audio signals on wearable devices, improves the comprehensibility and recognizability of audio signals, simplifies the operation process, and reduces the calculation amount and power consumption.
Smart Images

Figure CN2025079290_04092025_PF_FP_ABST
Abstract
Description
Audio processing method, device, readable medium and electronic device
[0001] This application claims priority to Chinese Patent Application No. 202410232546.3 filed on February 29, 2024, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field
[0002] The present disclosure relates to an audio processing method, device, readable medium and electronic device. Background Art
[0003] Speech enhancement aims to improve the quality of speech signals and improve their intelligibility and recognizability during transmission and processing. Speech enhancement technology can be applied to voice communication, speech recognition and other fields.
[0004] In related technologies, directional speech enhancement refers to enhancing sounds in a specific direction in a signal while suppressing sounds in non-specific directions. This solution requires multiple air-conduction microphones to form an array to receive speech signals. However, the number of air-conduction microphones on wearable devices is relatively small, and the effect is not good in actual applications. Summary of the Invention
[0005] This section is provided to briefly introduce the concepts that will be described in detail in the detailed description below. This section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0006] In a first aspect, the present disclosure provides an audio processing method, the audio processing method comprising:
[0007] Acquire a bone conduction signal collected by a bone conduction microphone and an air conduction signal collected by an air conduction microphone;
[0008] obtaining a directional beam audio feature and a noise audio feature based on the air conduction signal and a first signal processing model, and obtaining a bone conduction audio feature based on the bone conduction signal and a second signal processing model, wherein the first signal processing model is used to extract the directional beam audio feature in a specified direction and the omnidirectional noise audio feature from the air conduction signal, and the second signal processing model is used to extract the audio feature from the bone conduction signal;
[0009] A target enhanced audio signal in the specified direction is obtained based on the air conduction signal, the bone conduction audio feature, the noise audio feature, and the directional beam audio feature.
[0010] In a second aspect, the present disclosure provides an audio processing device, the audio processing device comprising:
[0011] An acquisition module, used to acquire bone conduction signals collected by a bone conduction microphone and air conduction signals collected by an air conduction microphone;
[0012] a model processing module, configured to obtain directional beam audio features and noise audio features based on the air conduction signal and a first signal processing model, and to obtain bone conduction audio features based on the bone conduction signal and a second signal processing model, wherein the first signal processing model is configured to extract directional beam audio features in a specified direction and omnidirectional noise audio features from the air conduction signal, and the second signal processing model is configured to extract audio features from the bone conduction signal;
[0013] The enhancement module is used to obtain the target enhanced audio signal in the specified direction based on the air conduction signal, the bone conduction audio feature, the noise audio feature and the directional beam audio feature.
[0014] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of any one of the methods described in the first aspect.
[0015] In a fourth aspect, the present disclosure provides an electronic device, comprising:
[0016] a storage device having a computer program stored thereon;
[0017] A processing device is used to execute the computer program in the storage device to implement the steps of any one of the methods in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale. In the drawings:
[0019] FIG1 is a schematic diagram of a spectrum of a bone conduction microphone sound pickup solution provided according to an exemplary embodiment;
[0020] FIG2 is a flow chart of an audio processing method according to an exemplary embodiment;
[0021] FIG3 is a schematic diagram of a process for generating training data according to an exemplary embodiment;
[0022] FIG4 is a schematic diagram of a beam direction provided according to an exemplary embodiment;
[0023] FIG5 is a schematic diagram of an air conduction signal processing process according to an exemplary embodiment;
[0024] FIG6 is a schematic diagram showing a comparison of the spectra of an air-conduction signal and a directional beam audio signal according to an exemplary embodiment;
[0025] FIG7 is a flow chart of an audio processing method according to an exemplary embodiment;
[0026] FIG8 is a flow chart of an audio processing method according to an exemplary embodiment;
[0027] FIG9 is a schematic diagram showing a comparison of spectra of an air conduction signal and a target enhanced audio signal according to an exemplary embodiment;
[0028] FIG10 is a block diagram of an audio processing apparatus according to an exemplary embodiment; and
[0029] FIG11 is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0030] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0031] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0032] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Definitions of other terms are provided in the following description.
[0033] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0034] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0035] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0036] All actions of acquiring signals, information or data in this disclosure are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0037] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0038] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0039] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0040] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0041] At the same time, it is understood that the data involved in the technical solution of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws, regulations and relevant provisions.
[0042] In related technologies, speech enhancement includes directional speech enhancement solutions, bone conduction microphone sound collection solutions, general speech enhancement solutions, voiceprint speaker separation solutions, etc.
[0043] Among them, directional speech enhancement refers to enhancing the sound in a specific direction in the signal and suppressing the sound in non-specific directions. Generally, a signal beamforming scheme is used, that is, a directional speech enhancement signal is obtained by weighted synthesis of the signals received by a multi-microphone array. This scheme requires multiple air conduction microphones to form an array to achieve relatively good results, but the number of air conduction microphones on wearable devices is relatively small. In addition, this scheme requires a priori assumptions about the signal. It is usually assumed that the noise signal is a uniform Gaussian distribution, but there may be differences in real scenarios, which leads to poor robustness in real scenarios. And for wearable devices worn on the human head, they are easily interfered by the diffraction of the human head, so the directional speech enhancement scheme is not effective in practical applications.
[0044] The bone conduction microphone pickup solution is to receive the signal from the bone conduction microphone and transmit it into the communication link. It should be understood that bone conduction is a sound conduction method, that is, by converting sound into mechanical vibrations of different frequencies, sound waves are transmitted through the human skull, bony labyrinth, inner ear lymph, organ of Corti, auditory nerve, and auditory center. Compared with the classic sound conduction method that generates sound waves through the diaphragm, bone conduction eliminates many steps in the transmission of sound waves, can clearly restore sound in a noisy environment, and the sound waves will not affect others because of diffusion in the air. The bone conduction microphone can receive vibrations caused by the wearer's speech, but the disadvantage is that the received voice frequency is low, generally less than 2kHz or 4kHz, and the high-frequency part is severely lost.
[0045] Referring to Figure 1, the spectrum shows that the signal contains little noise and only the wearer's voice, but the high-frequency signal is severely lost, resulting in a poor listening experience. Although neural networks can be used in related technologies to spread the spectrum of bone conduction microphones, that is, to generate high-frequency signals above 4kHz from low frequencies, on the one hand, the computational complexity is large and not suitable for deployment on wearable devices, and on the other hand, the generated high-frequency signal is not natural enough. In addition, bone conduction microphones have the defect of unreliable sound pickup, and the fit between the wearable device and the human head will affect the microphone's sound pickup.
[0046] General speech enhancement solutions belong to the traditional noise reduction solutions in speech enhancement. They can adopt the Wiener filter or neural network noise reduction solutions in signal processing to extract only the human voice signal, but cannot distinguish whether the person is wearing a wearable device.
[0047] Voiceprint speaker separation requires pre-registration of the wearer's voice, a cumbersome process. In practice, multiple speakers' voices are fed into a speech detection model and a voiceprint model. The resulting voiceprint features are then matched against the pre-registered wearer's voice to determine if the voice is the wearer's. This computationally intensive process is typically performed in the cloud.
[0048] In view of this, the present disclosure provides an audio processing method, device, readable medium and electronic device to solve the above technical problems.
[0049] The following further explains the embodiments of the present disclosure with reference to the accompanying drawings.
[0050] FIG2 is a flow chart of an audio processing method according to an exemplary embodiment of the present disclosure. Referring to FIG2 , the audio processing method includes:
[0051] S201: Acquire a bone conduction signal collected by a bone conduction microphone and an air conduction signal collected by an air conduction microphone.
[0052] S202: Obtaining a directional beam audio feature and a noise audio feature according to the air conduction signal and the first signal processing model, and obtaining a bone conduction audio feature according to the bone conduction signal and the second signal processing model.
[0053] Among them, the first signal processing model is used to extract the directional beam audio features in a specified direction and the omnidirectional noise audio features in the air conduction signal, and the second signal processing model is used to extract the audio features in the bone conduction signal.
[0054] S203: Obtain a target enhanced audio signal in a specified direction based on the air conduction signal, the bone conduction audio feature, the noise audio feature, and the directional beam audio feature.
[0055] By using the above method, by distinguishing the directional beam audio characteristics and noise audio characteristics of the air conduction signal, the audio signal in the specified direction can be enhanced and the audio signal in the non-specified direction can be suppressed. The audio signal is further enhanced by combining the audio characteristics of the bone conduction signal to obtain an enhanced audio signal in the specified direction. There is no need to set up multiple air conduction microphones. It can be applied to wearable devices and can improve the quality of the audio signal, and improve the intelligibility and recognizability of the audio signal during transmission and processing.
[0056] In a possible manner, the first signal processing model is obtained by: obtaining an initial signal processing model, sample audio data and a target simulation function, the target simulation function including a room impulse response function and a directional impulse response function determined based on a simulated omnidirectional air conduction microphone, the directional impulse response function being used to characterize the positive correlation between the signal attenuation of the target beam and the target angle, the target angle representing the angle between the target beam and the beam corresponding to the specified direction; generating sample directional audio data and sample omnidirectional audio data based on the sample audio data and the target simulation function; performing model training on the initial signal processing model based on the sample directional audio data and the sample omnidirectional audio data, and obtaining the first signal processing model after the preset training completion conditions are met.
[0057] For example, a simulation experiment can be performed by simulating an omnidirectional air conduction microphone to determine a room impulse response function (RIR) and a directional impulse response function, thereby obtaining a target simulation function.
[0058] The image source method is used to create a room of random size, fixed microphone placement, and random sound source locations. The reflection of the sound signal from the sound source to the microphone is simulated through mirroring to calculate the energy and phase of the sound signal, thereby determining the room impulse response function. A directional beam, through its directional pattern, can set the signal attenuation at the angle between the beam direction and the specified direction. Generally, the larger the angle, the greater the signal attenuation, thus achieving a directional impulse response function with a directional beam. A directional impulse response function without a directional beam is a directional impulse response function without signal attenuation.
[0059] For example, taking a wearable device as an example, the specified direction can be set to the direction corresponding to the wearer's mouth. In this way, the attenuation of the signal from the direction of the wearer's mouth is smaller, which is equivalent to enhancing the wearer's voice and suppressing the voice of non-wearers.
[0060] For example, referring to FIG3 , sample directional audio data and sample omnidirectional audio data are generated based on the sample audio data and the target simulation function. In order to simulate audio data in different scenarios, the sample audio data may include speech data, speech data with interference, noise data, and the like, which is not limited in the present disclosure. Taking speech data as an example, the directional microphone data of the target speech with direction is determined by the directional impulse response function and the room impulse response function with a directional beam, and the omnidirectional microphone data of the target speech is determined by the directional impulse response function and the room impulse response function without a directional beam. Furthermore, based on the speech data, the speech data with interference, and the noise data, directional microphone data and omnidirectional microphone data with different signal-to-noise ratios are respectively determined, that is, the sample directional audio data and the sample omnidirectional audio data, thereby obtaining paired simulation sample data.
[0061] For example, the initial signal processing model is trained based on the sample directional audio data and the sample omnidirectional audio data, and after the preset training completion conditions are met, the first signal processing model is obtained. Referring to the beam direction diagram shown in Figure 4, 0 degrees is the specified direction, the black solid line is the forming beam direction under the ideal state, and the black dotted line is the actual forming beam direction of the model, that is, the trained first signal processing model can form a directional beam similar to that under the ideal state, and has the ability to separate the audio signal in the specified direction, thereby extracting the directional audio features and the omnidirectional noise audio features. Among them, the preset training completion conditions can be set according to needs, such as meeting the preset model convergence conditions, etc., and the present disclosure does not limit this.
[0062] In a possible manner, obtaining directional beam audio features and noise audio features based on the air conduction signal and the first signal processing model may include: determining the frequency band signal-to-noise ratio based on the air conduction signal, the frequency band signal-to-noise ratio characterizing the signal weight distribution in different directions, and the signal weight in the designated direction is greater than the signal weight in the non-designated direction; based on the air conduction signal, the frequency band signal-to-noise ratio and the first signal processing model, obtaining the directional beam audio features in the designated direction and the omnidirectional noise audio features.
[0063] For example, by performing a short-time Fourier transform (STFT) on the audio signals received from the air conduction microphone in different directions, the transformed audio signals are multiplied by the weight values of the beamforming algorithm (Minimum Variance Distortionless Response, MVDR) to obtain the signal weight distribution in different directions. The MVDR weight value w can be expressed as follows:
[0064] Among them, k represents the audio signal in the kth direction, f represents the frequency, t represents the frame number, R n (f) represents the covariance matrix of the noise, r f Represents a direction vector.
[0065] For example, the pointing direction is the direction corresponding to the wearer's mouth. Since the signal weight in the designated direction is greater than the signal weight in the non-designated direction, the first signal processing model can more accurately extract the audio data in the designated direction based on the frequency band signal-to-noise ratio, thereby facilitating the subsequent extraction of directional beam audio features and noise audio features, thereby improving the accuracy of feature extraction.
[0066] In a possible manner, based on the air conduction signal, the frequency band signal-to-noise ratio and the first signal processing model, the directional beam audio characteristics of the specified direction and the omnidirectional noise audio characteristics are obtained, which can include: performing a frequency domain transformation on the air conduction signal to obtain an air conduction frequency domain signal; and obtaining the directional beam audio characteristics and the noise audio characteristics according to the air conduction frequency domain signal, the frequency band signal-to-noise ratio and the first signal processing model.
[0067] For example, the air conduction signal can be Fourier transformed to convert the air conduction signal into an air conduction frequency domain signal so as to perform frequency domain analysis on the air conduction signal, that is, obtain the directional beam audio characteristics and noise audio characteristics through the first signal processing model.
[0068] In a possible manner, obtaining directional beam audio features and noise audio features based on air conduction frequency domain signals, frequency band signal-to-noise ratio and a first signal processing model may include: frequency grouping the air conduction frequency domain signals to obtain multiple groups of air conduction frequency domain signals; for each group of air conduction frequency domain signals in the multiple groups of air conduction frequency domain signals, inputting the group of air conduction frequency domain signals and the frequency band signal-to-noise ratio into the first signal processing model to obtain directional sub-audio features corresponding to a specified direction and corresponding omnidirectional noise sub-audio features in the group of air conduction frequency domain signals, wherein the model parameters of the first signal processing model corresponding to each group of air conduction frequency domain signals are the same; obtaining directional beam audio features and noise audio features based on the directional sub-audio features and noise sub-audio features corresponding to the multiple groups of air conduction frequency domain signals.
[0069] For example, referring to FIG5 , the air conduction frequency domain signals obtained after Fourier transform of the air conduction signals can be grouped by frequency, for example, into groups of 50 Hz, although this disclosure is not limited thereto. Model processing is then performed on the air conduction frequency domain signals according to the groups, and the model parameters of the first signal processing model corresponding to the air conduction frequency domain signals in different groups are the same. That is, air conduction frequency domain signals of different frequencies share the same model parameters, thereby effectively reducing the amount of computation and parameters. This not only improves model processing efficiency, but also results in a smaller model with lower power consumption, allowing deployment in wearable devices.
[0070] It should be understood that when the frequency band signal-to-noise ratio is determined based on the air conduction signal, Fourier transform is also required. Therefore, the frequency band signal-to-noise ratio can also be determined directly through the air conduction frequency domain signal or the grouped air conduction frequency domain signal. The present disclosure does not limit the calculation method for performing frequency domain transformation on the air conduction signal. For example, the frequency domain transformation can be performed through the Fourier transform calculation formula and its modified calculation formula.
[0071] In a possible manner, obtaining directional beam audio features and noise audio features based on the air conduction frequency domain signal, the frequency band signal-to-noise ratio and the first signal processing model can include: inputting the air conduction frequency domain signal and the frequency band signal-to-noise ratio into the first sub-network of the first signal processing model for feature extraction to obtain omnidirectional time domain audio features and omnidirectional frequency domain audio features; inputting the omnidirectional time domain audio features and omnidirectional frequency domain audio features into the second sub-network of the first signal processing model for time domain modeling to obtain omnidirectional time-frequency audio features; and performing linear classification on the omnidirectional time-frequency audio features to obtain directional beam audio features and noise audio features.
[0072] For example, continuing to refer to Figure 5, the first sub-network may include a convolutional neural network (CNN) and a normalization layer (layernorm), which is used to extract features based on the air conduction frequency domain signal and the frequency band signal-to-noise ratio to obtain omnidirectional time domain audio features and omnidirectional frequency domain audio features. The second sub-network may include a conformer network, which can uniformly model the local features and global features in the speech sequence. In the embodiment of the present disclosure, it is used to perform time domain modeling on the omnidirectional time domain audio features and the omnidirectional frequency domain audio features to obtain omnidirectional time-frequency audio features. Finally, a linear classification layer is used to linearly classify the omnidirectional time-frequency audio features to obtain the time-frequency audio features of the directional beam and the time-frequency audio features of the noise, that is, the directional beam audio features and the noise audio features.
[0073] In this way, the directional beam audio characteristics and noise audio characteristics of the air conduction signal can be obtained through the first signal processing model. It should be understood that the present disclosure does not limit the specific implementation form of the first signal processing model, and it can be adjusted according to model training.
[0074] In a possible manner, obtaining bone conduction audio features based on the bone conduction signal and the second signal processing model can include: inputting the bone conduction signal into the third sub-network of the second signal processing model for feature extraction to obtain bone conduction time domain audio features and bone conduction frequency domain audio features; inputting the bone conduction time domain audio features and bone conduction frequency domain audio features into the fourth sub-network of the second signal processing model for time domain modeling to obtain bone conduction audio features.
[0075] For example, the second signal processing model can be a conditional encoder, including a CNN network for extracting bone conduction time-domain audio features and bone conduction frequency-domain audio features, and a frequency-time-long short-term memory (FT-LSTM) neural network for time-domain modeling, so as to simultaneously scan the frequency and time axes to obtain the time-frequency characteristics of bone conduction audio. It should be understood that the present disclosure does not limit the specific implementation of the second signal processing model, and it can be adjusted according to model training.
[0076] In a possible manner, obtaining a target enhanced audio signal in a specified direction based on air conduction signals, bone conduction audio features, noise audio features and directional beam audio features may include: obtaining a directional beam audio signal based on the air conduction signals and the directional beam audio features; performing feature extraction and time domain modeling on the directional beam audio signal to obtain a target audio feature; performing attention feature fusion on the bone conduction audio features, the noise audio features and the target audio features to obtain a target fusion feature; decoding the target fusion feature to obtain a target fusion audio signal, and obtaining a target enhanced audio signal based on the target fusion audio signal and the directional beam audio signal.
[0077] For example, referring to Figure 5 , the air-conduction signal is masked based on the directional beam audio characteristics to obtain a directional beam audio signal. Referring to Figure 6 , the spectrum of the air-conduction signal with human voice and noise interference is shown on the left, while the spectrum of the directional beam audio signal is shown on the right. This achieves directional sound reception while suppressing sounds outside the non-directional beam.
[0078] For example, referring to Figure 7 , similar to the bone conduction signal processing, the directional beam audio signal can be subjected to feature extraction and time-domain modeling to obtain target audio features. It should be understood that the first audio processing model focuses more on spatial audio features, i.e., directional beam audio features are audio features associated with a specific direction. The CNN used to extract the target audio features here focuses more on features associated with the bone conduction signal and noise audio features, facilitating subsequent attention feature fusion.
[0079] For example, attention fusion is performed on the bone conduction audio features, the noise audio features, and the target audio features to obtain a target fusion feature. This feature focuses more on the corresponding features of the target audio features in the bone conduction audio features. The target fusion feature is then decoded by a decoder to obtain a target fusion audio signal. Finally, further directional enhancement is performed based on the target fusion audio signal and the directional beam audio signal to obtain a target enhanced audio signal.
[0080] It is worth noting that the model used in the embodiments of the present disclosure is relatively small and has low power consumption, so it can be deployed on wearable devices. Taking deployment on a wearable device as an example, referring to Figure 8, the air conduction signal is processed by the first audio processing model to obtain directional beam audio features and noise audio features, and a directional beam audio signal can be obtained. Since the specified direction of the directional beam can be set to match the wearer's mouth, the first audio processing model will form a beam pointing to the mouth in the air conduction microphone, improving the signal-to-noise ratio of the wearer's voice signal, suppressing sounds and noise from other directions, and extracting features from the enhanced directional beam audio signal to obtain target audio features. Combined with the noise audio features and the bone conduction features obtained by the second audio processing model of the bone conduction signal, feature fusion is performed to obtain a target enhanced audio signal that enhances the wearer's voice. No voiceprint registration is required, and the operation is simple. Not only can voice enhancement be performed in a specified direction, but the distortion of the voice can also be reduced, and the model stability is also better.
[0081] For example, referring to Figure 9, comparing the air conduction signal with human voice interference and noise interference and the target enhanced audio signal, the embodiment of the present disclosure can receive voice in a specified direction through a fixed beam, and combine the voice processing unit (Voice Processing Unit) of the bone conduction microphone to obtain the wearer's voice signal for further voice enhancement and solve the noise problem.
[0082] Based on the same concept, the present disclosure further provides an audio processing device. Referring to FIG. 10 , the audio processing device 10 includes:
[0083] An acquisition module 11 is configured to acquire a bone conduction signal acquired by a bone conduction microphone and an air conduction signal acquired by an air conduction microphone;
[0084] a model processing module 12, configured to obtain directional beam audio features and noise audio features based on the air conduction signal and a first signal processing model, and to obtain bone conduction audio features based on the bone conduction signal and a second signal processing model, wherein the first signal processing model is configured to extract directional beam audio features in a specified direction and omnidirectional noise audio features from the air conduction signal, and the second signal processing model is configured to extract audio features from the bone conduction signal;
[0085] The enhancement module 13 is configured to obtain a target enhanced audio signal in the specified direction based on the air conduction signal, the bone conduction audio feature, the noise audio feature, and the directional beam audio feature.
[0086] By using the above-mentioned device, by distinguishing the directional beam audio characteristics and noise audio characteristics of the air conduction signal, the audio signal in the specified direction can be enhanced and the audio signal in the non-specified direction can be suppressed. The audio signal can be further enhanced by combining the audio characteristics of the bone conduction signal to obtain an enhanced audio signal in the specified direction. There is no need to set up multiple air conduction microphones. It can be applied to wearable devices and can improve the quality of the audio signal and improve the comprehensibility and recognizability of the audio signal during transmission and processing.
[0087] Optionally, the model processing module 12 includes:
[0088] a determination module, configured to determine a frequency band signal-to-noise ratio based on the air conduction signal, wherein the frequency band signal-to-noise ratio represents a signal weight distribution in different directions, and the signal weight in the designated direction is greater than the signal weight in the non-designated direction;
[0089] A model processing submodule is used to obtain the directional beam audio characteristics of the specified direction and the omnidirectional noise audio characteristics based on the air conduction signal, the frequency band signal-to-noise ratio and the first signal processing model.
[0090] Optionally, the model processing submodule is used to:
[0091] Performing frequency domain transformation on the gas conduction signal to obtain a gas conduction frequency domain signal;
[0092] The directional beam audio feature and the noise audio feature are obtained according to the air conduction frequency domain signal, the frequency band signal-to-noise ratio and the first signal processing model.
[0093] Optionally, the model processing submodule is used to:
[0094] Frequency grouping the air conduction frequency domain signals to obtain multiple groups of air conduction frequency domain signals;
[0095] For each group of air conduction frequency domain signals in the plurality of groups of air conduction frequency domain signals, inputting the group of air conduction frequency domain signals and the frequency band signal-to-noise ratio into the first signal processing model to obtain a directional sub-audio feature corresponding to the specified direction and a noise sub-audio feature corresponding to the omnidirectional direction in the group of air conduction frequency domain signals, wherein the model parameters of the first signal processing model corresponding to each group of air conduction frequency domain signals are the same;
[0096] The directional beam audio feature and the noise audio feature are obtained according to the directional sub-audio features and the noise sub-audio features corresponding to the multiple groups of air-conducted frequency domain signals.
[0097] Optionally, the model processing submodule is used to:
[0098] Inputting the air conduction frequency domain signal and the frequency band signal-to-noise ratio into the first subnetwork of the first signal processing model for feature extraction to obtain omnidirectional time domain audio features and omnidirectional frequency domain audio features;
[0099] Inputting the omnidirectional time-domain audio feature and the omnidirectional frequency-domain audio feature into the second subnetwork of the first signal processing model for time-domain modeling to obtain omnidirectional time-frequency audio features;
[0100] Linear classification is performed on the omnidirectional time-frequency audio features to obtain the directional beam audio features and the noise audio features.
[0101] Optionally, the first signal processing model is obtained by:
[0102] Obtaining an initial signal processing model, sample audio data, and a target simulation function, wherein the target simulation function includes a room impulse response function and a directional impulse response function determined based on a simulated omnidirectional air conduction microphone, wherein the directional impulse response function is used to characterize a positive correlation between a signal attenuation magnitude of a target beam and a target angle magnitude, wherein the target angle represents an angle between the target beam and a beam corresponding to the specified direction;
[0103] generating sample directional audio data and sample omnidirectional audio data according to the sample audio data and the target simulation function;
[0104] The initial signal processing model is trained according to the sample directional audio data and the sample omnidirectional audio data, and the first signal processing model is obtained after a preset training completion condition is met.
[0105] Optionally, the model processing module 12 is used to:
[0106] Inputting the bone conduction signal into the third sub-network of the second signal processing model for feature extraction to obtain bone conduction time-domain audio features and bone conduction frequency-domain audio features;
[0107] The bone conduction time domain audio feature and the bone conduction frequency domain audio feature are input into the fourth subnetwork of the second signal processing model for time domain modeling to obtain the bone conduction audio feature.
[0108] Optionally, the enhancement module 13 is used to:
[0109] Obtaining a directional beam audio signal according to the air conduction signal and the directional beam audio feature;
[0110] Performing feature extraction and time-domain modeling on the directional beam audio signal to obtain target audio features;
[0111] Performing attention feature fusion on the bone conduction audio feature, the noise audio feature, and the target audio feature to obtain a target fusion feature;
[0112] The target fusion feature is decoded to obtain a target fusion audio signal, and a target enhanced audio signal is obtained according to the target fusion audio signal and the directional beam audio signal.
[0113] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0114] Based on the same concept, an embodiment of the present disclosure further provides a computer-readable medium on which a computer program is stored. When the program is executed by a processing device, the steps of the above-mentioned audio processing method are implemented.
[0115] Based on the same concept, an embodiment of the present disclosure further provides an electronic device, including:
[0116] a storage device having a computer program stored thereon;
[0117] A processing device is used to execute the computer program in the storage device to implement the steps of the above audio processing method.
[0118] Reference is now made to FIG11 , which illustrates a schematic diagram of the structure of an electronic device 1100 suitable for implementing embodiments of the present disclosure. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device illustrated in FIG11 is merely an example and should not limit the functionality or scope of use of embodiments of the present disclosure.
[0119] As shown in FIG11 , the electronic device 1100 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1101, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1102 or a program loaded from a storage device 1108 into a random access memory (RAM) 1103. Various programs and data required for the operation of the electronic device 1100 are also stored in the RAM 1103. The processing device 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0120] Typically, the following devices may be connected to the I / O interface 1105: an input device 1106 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1107 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1108 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1109. The communication device 1109 may allow the electronic device 1100 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 11 illustrates the electronic device 1100 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
[0121] According to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1109, or installed from the storage device 1108, or installed from the ROM 1102. When the computer program is executed by the processing device 1101, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0122] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0123] In some embodiments, communication can be performed using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0124] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0125] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: obtains a bone conduction signal collected by a bone conduction microphone and an air conduction signal collected by an air conduction microphone; obtains a directional beam audio feature and a noise audio feature based on the air conduction signal and a first signal processing model, and obtains a bone conduction audio feature based on the bone conduction signal and a second signal processing model, the first signal processing model is used to extract the directional beam audio feature and omnidirectional noise audio feature in a specified direction from the air conduction signal, and the second signal processing model is used to extract the audio feature in the bone conduction signal; based on the air conduction signal, the bone conduction audio feature, the noise audio feature and the directional beam audio feature, obtains a target enhanced audio signal in the specified direction.
[0126] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0127] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0128] The modules described in the embodiments of the present disclosure may be implemented in software or hardware, wherein the name of a module does not necessarily limit the module itself.
[0129] The functions described above in the present disclosure may be performed at least in part by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0130] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0131] According to one or more embodiments of the present disclosure, Example 1 provides an audio processing method, which includes: acquiring a bone conduction signal collected by a bone conduction microphone and an air conduction signal collected by an air conduction microphone; obtaining a directional beam audio feature and a noise audio feature based on the air conduction signal and a first signal processing model, and obtaining a bone conduction audio feature based on the bone conduction signal and a second signal processing model, the first signal processing model being used to extract the directional beam audio feature and omnidirectional noise audio feature in a specified direction from the air conduction signal, and the second signal processing model being used to extract audio features from the bone conduction signal; obtaining a target enhanced audio signal in the specified direction based on the air conduction signal, the bone conduction audio feature, the noise audio feature and the directional beam audio feature.
[0132] According to one or more embodiments of the present disclosure, Example 2 provides the audio processing method of Example 1, wherein the directional beam audio feature and the noise audio feature are obtained based on the air conduction signal and the first signal processing model, including: determining the frequency band signal-to-noise ratio based on the air conduction signal, the frequency band signal-to-noise ratio characterizing the signal weight distribution in different directions, and the signal weight in the designated direction is greater than the signal weight in the non-designated direction; based on the air conduction signal, the frequency band signal-to-noise ratio and the first signal processing model, obtaining the directional beam audio feature of the designated direction and the omnidirectional noise audio feature.
[0133] According to one or more embodiments of the present disclosure, Example 3 provides the audio processing method of Example 2, which obtains the directional beam audio characteristics of the specified direction and the omnidirectional noise audio characteristics based on the air conduction signal, the frequency band signal-to-noise ratio and the first signal processing model, including: performing frequency domain transformation on the air conduction signal to obtain an air conduction frequency domain signal; and obtaining the directional beam audio characteristics and the noise audio characteristics based on the air conduction frequency domain signal, the frequency band signal-to-noise ratio and the first signal processing model.
[0134] According to one or more embodiments of the present disclosure, Example 4 provides the audio processing method of Example 3, wherein the directional beam audio feature and the noise audio feature are obtained based on the air conduction frequency domain signal, the frequency band signal-to-noise ratio and the first signal processing model, including: frequency grouping the air conduction frequency domain signal to obtain multiple groups of air conduction frequency domain signals; for each group of air conduction frequency domain signals in the multiple groups of air conduction frequency domain signals, inputting the group of air conduction frequency domain signals and the frequency band signal-to-noise ratio into the first signal processing model to obtain the directional sub-audio feature corresponding to the specified direction and the noise sub-audio feature corresponding to the omnidirectional direction in the group of air conduction frequency domain signals, wherein the model parameters of the first signal processing model corresponding to each group of air conduction frequency domain signals are the same; and obtaining the directional beam audio feature and the noise audio feature based on the directional sub-audio feature and the noise sub-audio feature corresponding to the multiple groups of air conduction frequency domain signals.
[0135] According to one or more embodiments of the present disclosure, Example 5 provides the audio processing method of Example 3, wherein the directional beam audio feature and the noise audio feature are obtained based on the air conduction frequency domain signal, the frequency band signal-to-noise ratio and the first signal processing model, including: inputting the air conduction frequency domain signal and the frequency band signal-to-noise ratio into the first sub-network of the first signal processing model for feature extraction to obtain omnidirectional time domain audio features and omnidirectional frequency domain audio features; inputting the omnidirectional time domain audio features and the omnidirectional frequency domain audio features into the second sub-network of the first signal processing model for time domain modeling to obtain omnidirectional time-frequency audio features; and performing linear classification on the omnidirectional time-frequency audio features to obtain the directional beam audio features and the noise audio features.
[0136] According to one or more embodiments of the present disclosure, Example 6 provides the audio processing method described in any one of Examples 1-5, wherein the first signal processing model is obtained in the following manner: obtaining an initial signal processing model, sample audio data and a target simulation function, wherein the target simulation function includes a room impulse response function and a directional impulse response function determined based on a simulated omnidirectional air conduction microphone, wherein the directional impulse response function is used to characterize the positive correlation between the signal attenuation size of the target beam and the target angle size, and the target angle represents the angle between the target beam and the beam corresponding to the specified direction; generating sample directional audio data and sample omnidirectional audio data based on the sample audio data and the target simulation function; performing model training on the initial signal processing model based on the sample directional audio data and the sample omnidirectional audio data, and obtaining the first signal processing model after the preset training completion conditions are met.
[0137] According to one or more embodiments of the present disclosure, Example 7 provides the audio processing method described in any one of Examples 1-5, wherein the bone conduction audio feature is obtained based on the bone conduction signal and the second signal processing model, including: inputting the bone conduction signal into the third subnetwork of the second signal processing model for feature extraction to obtain bone conduction time domain audio features and bone conduction frequency domain audio features; inputting the bone conduction time domain audio features and the bone conduction frequency domain audio features into the fourth subnetwork of the second signal processing model for time domain modeling to obtain the bone conduction audio features.
[0138] According to one or more embodiments of the present disclosure, Example 8 provides the audio processing method described in any one of Examples 1-5, wherein the target enhanced audio signal in the specified direction is obtained based on the air conduction signal, the bone conduction audio feature, the noise audio feature and the directional beam audio feature, including: obtaining a directional beam audio signal based on the air conduction signal and the directional beam audio feature; performing feature extraction and time domain modeling on the directional beam audio signal to obtain a target audio feature; performing attention feature fusion on the bone conduction audio feature, the noise audio feature and the target audio feature to obtain a target fusion feature; decoding the target fusion feature to obtain a target fusion audio signal, and obtaining the target enhanced audio signal based on the target fusion audio signal and the directional beam audio signal.
[0139] According to one or more embodiments of the present disclosure, Example 9 provides an audio processing device, which includes: an acquisition module for acquiring a bone conduction signal collected by a bone conduction microphone and an air conduction signal collected by an air conduction microphone; a model processing module for obtaining directional beam audio features and noise audio features based on the air conduction signal and a first signal processing model, and obtaining bone conduction audio features based on the bone conduction signal and a second signal processing model, the first signal processing model being used to extract the directional beam audio features and omnidirectional noise audio features in a specified direction from the air conduction signal, and the second signal processing model being used to extract audio features from the bone conduction signal; an enhancement module being used to obtain a target enhanced audio signal in the specified direction based on the air conduction signal, the bone conduction audio features, the noise audio features and the directional beam audio features.
[0140] According to one or more embodiments of the present disclosure, Example 10 provides the audio processing device described in Example 9, wherein the model processing module includes: a determination module for determining a frequency band signal-to-noise ratio based on the air conduction signal, wherein the frequency band signal-to-noise ratio characterizes the signal weight distribution in different directions, and the signal weight in the designated direction is greater than the signal weight in the non-designated direction; a model processing sub-module for obtaining the directional beam audio characteristics of the designated direction and the omnidirectional noise audio characteristics based on the air conduction signal, the frequency band signal-to-noise ratio and the first signal processing model.
[0141] According to one or more embodiments of the present disclosure, Example 11 provides the audio processing device described in Example 10, wherein the model processing submodule is used to: perform frequency domain transformation on the air conduction signal to obtain an air conduction frequency domain signal; and obtain the directional beam audio feature and the noise audio feature based on the air conduction frequency domain signal, the frequency band signal-to-noise ratio and the first signal processing model.
[0142] According to one or more embodiments of the present disclosure, Example 12 provides the audio processing device described in Example 11, wherein the model processing submodule is used to: perform frequency grouping on the air conduction frequency domain signals to obtain multiple groups of air conduction frequency domain signals; for each group of air conduction frequency domain signals in the multiple groups of air conduction frequency domain signals, input the group of air conduction frequency domain signals and the frequency band signal-to-noise ratio into the first signal processing model to obtain the directional sub-audio features corresponding to the specified direction and the noise sub-audio features corresponding to the omnidirectional direction in the group of air conduction frequency domain signals, wherein the model parameters of the first signal processing model corresponding to each group of air conduction frequency domain signals are the same; and obtain the directional beam audio features and the noise audio features based on the directional sub-audio features and the noise sub-audio features corresponding to the multiple groups of air conduction frequency domain signals.
[0143] According to one or more embodiments of the present disclosure, Example 13 provides the audio processing device described in Example 11, wherein the model processing submodule is used to: input the air conduction frequency domain signal and the frequency band signal-to-noise ratio into the first subnetwork of the first signal processing model for feature extraction to obtain omnidirectional time domain audio features and omnidirectional frequency domain audio features; input the omnidirectional time domain audio features and the omnidirectional frequency domain audio features into the second subnetwork of the first signal processing model for time domain modeling to obtain omnidirectional time-frequency audio features; perform linear classification on the omnidirectional time-frequency audio features to obtain the directional beam audio features and the noise audio features.
[0144] According to one or more embodiments of the present disclosure, Example 14 provides the audio processing device described in any one of Examples 9-13, wherein the first signal processing model is obtained in the following manner: obtaining an initial signal processing model, sample audio data and a target simulation function, the target simulation function including a room impulse response function and a directional impulse response function determined based on a simulated omnidirectional air conduction microphone, the directional impulse response function being used to characterize the positive correlation between the signal attenuation of the target beam and the target angle, the target angle representing the angle between the target beam and the beam corresponding to the specified direction; generating sample directional audio data and sample omnidirectional audio data based on the sample audio data and the target simulation function; performing model training on the initial signal processing model based on the sample directional audio data and the sample omnidirectional audio data, and obtaining the first signal processing model after the preset training completion conditions are met.
[0145] According to one or more embodiments of the present disclosure, Example 15 provides the audio processing device described in any one of Examples 9-13, wherein the model processing module is used to: input the bone conduction signal into the third subnetwork of the second signal processing model for feature extraction to obtain bone conduction time domain audio features and bone conduction frequency domain audio features; input the bone conduction time domain audio features and the bone conduction frequency domain audio features into the fourth subnetwork of the second signal processing model for time domain modeling to obtain the bone conduction audio features.
[0146] According to one or more embodiments of the present disclosure, Example 16 provides the audio processing device described in any one of Examples 9-13, wherein the enhancement module is used to: obtain a directional beam audio signal based on the air conduction signal and the directional beam audio feature; perform feature extraction and time domain modeling on the directional beam audio signal to obtain a target audio feature; perform attention feature fusion on the bone conduction audio feature, the noise audio feature and the target audio feature to obtain a target fusion feature; decode the target fusion feature to obtain a target fusion audio signal, and obtain a target enhanced audio signal based on the target fusion audio signal and the directional beam audio signal.
[0147] The above description is merely an example of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0148] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0149] Although the present disclosure has been described using language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. Regarding the apparatus in the above-described embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method and will not be elaborated upon here.
Claims
1. An audio processing method, comprising: Acquire a bone conduction signal collected by a bone conduction microphone and an air conduction signal collected by an air conduction microphone; obtaining a directional beam audio feature and a noise audio feature based on the air conduction signal and a first signal processing model, and obtaining a bone conduction audio feature based on the bone conduction signal and a second signal processing model, wherein the first signal processing model is used to extract the directional beam audio feature in a specified direction and the omnidirectional noise audio feature from the air conduction signal, and the second signal processing model is used to extract the audio feature from the bone conduction signal; A target enhanced audio signal in the specified direction is obtained based on the air conduction signal, the bone conduction audio feature, the noise audio feature, and the directional beam audio feature.
2. The audio processing method according to claim 1, wherein: The obtaining of the directional beam audio feature and the noise audio feature according to the air conduction signal and the first signal processing model includes: determining a frequency band signal-to-noise ratio according to the air conduction signal, wherein the frequency band signal-to-noise ratio represents a signal weight distribution in different directions, and the signal weight in the designated direction is greater than the signal weight in the non-designated direction; Based on the air conduction signal, the frequency band signal-to-noise ratio and the first signal processing model, the directional beam audio feature in the specified direction and the omnidirectional noise audio feature are obtained.
3. The audio processing method according to claim 2, wherein: The obtaining, based on the air conduction signal, the frequency band signal-to-noise ratio, and the first signal processing model, the directional beam audio feature in the specified direction and the omnidirectional noise audio feature, includes: Performing frequency domain transformation on the gas conduction signal to obtain a gas conduction frequency domain signal; The directional beam audio feature and the noise audio feature are obtained according to the air conduction frequency domain signal, the frequency band signal-to-noise ratio and the first signal processing model.
4. The audio processing method according to claim 3, wherein: The obtaining, according to the air conduction frequency domain signal, the frequency band signal-to-noise ratio, and the first signal processing model, the directional beam audio feature and the noise audio feature comprises: Frequency grouping the air conduction frequency domain signals to obtain multiple groups of air conduction frequency domain signals; For each group of air conduction frequency domain signals in the plurality of groups of air conduction frequency domain signals, inputting each group of air conduction frequency domain signals and the frequency band signal-to-noise ratio into the first signal processing model to obtain a directional sub-audio feature corresponding to the specified direction and a noise sub-audio feature corresponding to the omnidirectional direction in each group of air conduction frequency domain signals, wherein the model parameters of the first signal processing model corresponding to each group of air conduction frequency domain signals are the same; The directional beam audio feature and the noise audio feature are obtained according to the directional sub-audio features and the noise sub-audio features corresponding to the multiple groups of air-conducted frequency domain signals. The audio processing method according to claim 3 , wherein: The obtaining, according to the air conduction frequency domain signal, the frequency band signal-to-noise ratio, and the first signal processing model, the directional beam audio feature and the noise audio feature comprises: Inputting the air conduction frequency domain signal and the frequency band signal-to-noise ratio into the first subnetwork of the first signal processing model for feature extraction to obtain omnidirectional time domain audio features and omnidirectional frequency domain audio features; Inputting the omnidirectional time-domain audio feature and the omnidirectional frequency-domain audio feature into the second subnetwork of the first signal processing model for time-domain modeling to obtain omnidirectional time-frequency audio features; Linear classification is performed on the omnidirectional time-frequency audio features to obtain the directional beam audio features and the noise audio features.
6. The audio processing method according to any one of claims 1 to 5, wherein: The first signal processing model is obtained by: Obtaining an initial signal processing model, sample audio data, and a target simulation function, wherein the target simulation function includes a room impulse response function and a directional impulse response function determined based on a simulated omnidirectional air conduction microphone, wherein the directional impulse response function is used to characterize a positive correlation between a signal attenuation magnitude of a target beam and a target angle magnitude, wherein the target angle represents an angle between the target beam and a beam corresponding to the specified direction; generating sample directional audio data and sample omnidirectional audio data according to the sample audio data and the target simulation function; The initial signal processing model is trained according to the sample directional audio data and the sample omnidirectional audio data, and the first signal processing model is obtained after a preset training completion condition is met.
7. The audio processing method according to any one of claims 1 to 6, wherein: The obtaining of the bone conduction audio feature according to the bone conduction signal and the second signal processing model includes: Inputting the bone conduction signal into the third sub-network of the second signal processing model for feature extraction to obtain bone conduction time-domain audio features and bone conduction frequency-domain audio features; The bone conduction time domain audio feature and the bone conduction frequency domain audio feature are input into the fourth subnetwork of the second signal processing model for time domain modeling to obtain the bone conduction audio feature.
8. The audio processing method according to any one of claims 1 to 7, wherein: The obtaining, based on the air conduction signal, the bone conduction audio feature, the noise audio feature, and the directional beam audio feature, of the target enhanced audio signal in the specified direction includes: Obtaining a directional beam audio signal according to the air conduction signal and the directional beam audio feature; Performing feature extraction and time-domain modeling on the directional beam audio signal to obtain target audio features; Performing attention feature fusion on the bone conduction audio feature, the noise audio feature, and the target audio feature to obtain a target fusion feature; The target fusion feature is decoded to obtain a target fusion audio signal, and the target enhanced audio signal is obtained according to the target fusion audio signal and the directional beam audio signal.
9. An audio processing device, comprising: an acquisition module, configured to acquire a bone conduction signal acquired by a bone conduction microphone and an air conduction signal acquired by an air conduction microphone; a model processing module configured to obtain directional beam audio features and noise audio features based on the air conduction signal and a first signal processing model, and to obtain bone conduction audio features based on the bone conduction signal and a second signal processing model, wherein the first signal processing model is used to extract directional beam audio features in a specified direction and omnidirectional noise audio features from the air conduction signal, and the second signal processing model is used to extract audio features from the bone conduction signal; The enhancement module is configured to obtain the target enhanced audio signal in the specified direction based on the air conduction signal, the bone conduction audio feature, the noise audio feature and the directional beam audio feature.
10. A computer-readable medium storing a computer program, wherein: When the computer program is executed by a processing device, the audio processing method according to any one of claims 1 to 8 is implemented.
11. An electronic device comprising: a storage device storing at least one computer program; At least one processing device is configured to execute the at least one computer program in the storage device to implement the audio processing method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Voice enhancement method, device, equipment and storage medium
CN109767783A
Multi-channel sound enhancement method and device and electronic equipment
CN112634930A
Voice signal processing method and device, computer equipment and storage medium
CN116030823A
Pickup noise reduction method, control device, pickup equipment and readable storage medium
CN116801147A
Voice activation detecting method of earphones, earphones and storage medium
US20230352038A1