Call noise reduction method, terminal and computer readable storage medium

By separating and processing the audio signals from the headset and the terminal, the problems of unclear and distorted voice in the existing technology are solved, resulting in a clearer call effect.

CN116189701BActive Publication Date: 2025-11-25纳欣科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310187743.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-24
Publication Date
2025-11-25
Estimated Expiration
2043-02-24

AI Technical Summary

Technical Problem

Existing technologies fail to effectively utilize multi-channel radio arrays in call noise reduction processing, resulting in blurred and distorted speech.

Method used

By acquiring the audio signals from the headset and terminal, the voice signal and noise signal are separated and processed, synthesized and noise-reduced to obtain a clear call signal.

Benefits of technology

It achieves clearer and more faithful voice call quality by comprehensively processing the audio signals from the headset and the terminal, thus improving voice quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189701B_ABST
    Figure CN116189701B_ABST
Patent Text Reader

Abstract

The application provides a call noise reduction method, a terminal and a computer readable storage medium. The call noise reduction method comprises the following steps: acquiring a first sound signal collected by a first sound collection module of the terminal and a second sound signal collected by a second sound collection module of a headset during a call, wherein the first sound signal comprises a mixed first voice signal and a first noise signal, and the second sound signal comprises a mixed second voice signal and a second noise signal; separating the first voice signal, the first noise signal, the second voice signal and the second noise signal; processing the first voice signal and the second voice signal to obtain a target voice signal; synthesizing the first noise signal and the second noise signal to obtain a synthesized noise signal; performing noise reduction processing on the synthesized noise signal to obtain a residual noise signal; and synthesizing the target voice signal and the residual noise signal to obtain a call signal. The call noise reduction method provided by the application can realize clearer and more authentic voice calls.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of communication technology, in particular to a call noise reduction method, a terminal and a computer readable storage medium. BACKGROUND

[0002] Currently, the noise reduction processing of the call voice is usually to perform signal quality threshold judgment on the voice signals collected by the microphone array of the earphone and the microphone array of the mobile phone, to perform selection judgment on the microphone array according to the distance between the listener and the microphone array, to select the voice signal of the better one for signal transmission and processing, and to discard the voice signal of the one with lower signal-to-noise ratio. However, this does not completely utilize the complete multi-path microphone array, and may cause the call voice to be blurred and distorted. SUMMARY

[0003] To solve the above technical problems, the present application provides a call noise reduction method, a terminal and a computer readable storage medium, which retains the voice signals collected by the earphone and the terminal, and can realize a clearer and more faithful voice call function.

[0004] The first aspect of the present application provides a call noise reduction method, which comprises the following steps: acquiring a first sound collection signal collected by a first sound collection module of a terminal and a second sound collection signal collected by a second sound collection module of an earphone during a call, the first sound collection signal comprising a mixed first voice signal and a first noise signal, and the second sound collection signal comprising a mixed second voice signal and a second noise signal; separating out the first voice signal, the first noise signal, the second voice signal and the second noise signal; processing the first voice signal and the second voice signal to obtain a target voice signal; synthesizing the first noise signal and the second noise signal to obtain a synthesized noise signal; performing noise reduction processing on the synthesized noise signal to obtain a residual noise signal; and synthesizing the target voice signal and the residual noise signal to obtain a call signal.

[0005] The second aspect of the present application provides a terminal, which comprises a first sound collection module and a processor, the processor being configured to acquire a first sound collection signal collected by the first sound collection module and a second sound collection signal collected by a second sound collection module of an earphone, the second sound collection signal comprising a mixed second voice signal and a second noise signal, and the processor being further configured to separate out the first voice signal, the first noise signal, the second voice signal and the second noise signal, to synthesize the first noise signal and the second noise signal to obtain a synthesized noise signal, to process the first voice signal and the second voice signal to obtain a target voice signal, to perform noise reduction processing on the synthesized noise signal to obtain a residual noise signal, and to synthesize the target voice signal and the residual noise signal to obtain a call signal.

[0006] The third aspect of the present application provides a computer readable storage medium, which stores a computer program for executing the aforementioned call noise reduction method when called by a computer.

[0007] The call noise reduction method, the terminal and the computer readable storage medium provided by the present application can separate the sound signals and noise signals collected by the two sound collection arrays of the earphone and the terminal, perform noise reduction processing on the separated noise signals, and synthesize the separated sound signals and the noise signals after noise reduction processing to obtain a call signal. The sound signals collected by the earphone and the terminal are retained, a clearer and more faithful voice call function can be realized, and a better voice call effect than processing the sound signals collected by the earphone alone or the sound signals collected by the terminal alone can be achieved. BRIEF DESCRIPTION OF DRAWINGS

[0008] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings described below are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0009] Figure 1 The flow chart of the call noise reduction method provided by the embodiments of the present application.

[0010] Figure 2 The structural block diagram of the terminal provided by the embodiments of the present application.

[0011] Main element symbol explanation:

[0012] Terminal 100

[0013] First sound collection module 10

[0014] Processor 20 DETAILED DESCRIPTION

[0015] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0016] In the description of the present application, the terms "first", "second" and the like are used to distinguish different objects, and are not used to describe a specific order, and in addition, the terms "upper", "inner", "outer" and the like indicate the orientation or positional relationship shown in the drawings, and are only used to facilitate the description of the present application and simplify the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application.

[0017] In the description of the present application, unless otherwise explicitly specified and limited, the term "connection" should be understood broadly, for example, it can be fixed connection, or detachable connection, or integral connection; it can be direct connection, or indirect connection through an intermediate medium, or internal communication of two elements; it can be communication connection; it can be electrical connection. For those skilled in the art, the specific meaning of the above-mentioned terms in the present application can be understood according to the specific circumstances.

[0018] Please refer to Figure 1 , Figure 1 The flow chart of the call noise reduction method provided by the embodiment of the present application. The call noise reduction method is applied to a terminal. As shown in Figure 1 The call noise reduction method comprises the following steps:

[0019] S10: Obtain a first sound collection signal collected by a first sound collection module of the terminal during a call and a second sound collection signal collected by a second sound collection module of a headset, the first sound collection signal comprising a mixed first voice signal and a first noise signal, and the second sound collection signal comprising a mixed second voice signal and a second noise signal.

[0020] S20: Separate the first voice signal, the first noise signal, the second voice signal and the second noise signal.

[0021] S30: Process the first voice signal and the second voice signal to obtain a target voice signal.

[0022] S40: Synthesize the first noise signal and the second noise signal to obtain a synthesized noise signal.

[0023] S50: Perform noise reduction processing on the synthesized noise signal to obtain a residual noise signal.

[0024] S60: Synthesize the target voice signal and the residual noise signal to obtain a call signal.

[0025] The call noise reduction method provided by the embodiments of the present application separates the sound signals collected by the sound collection array of the earphone and the sound collection array of the terminal into voice signals and noise signals, performs noise reduction processing on the separated noise signals, and synthesizes the separated voice signals and the noise signals after noise reduction processing to obtain a call signal. The voice signals collected by the earphone and the terminal are retained, clearer and more faithful voice calls can be realized, and better voice call effects than processing the sound signals collected by the earphone alone or the sound signals collected by the terminal alone can be achieved.

[0026] The first sound collection module can be a plurality of microphones arranged on the terminal, and the plurality of microphones can be arranged in an array to collect the first sound signal when a voice call is performed.

[0027] The second sound collection module can be a plurality of microphones arranged on the earphone, and the plurality of microphones can be arranged in an array to collect the second sound signal when a voice call is performed.

[0028] The terminal can be a mobile phone, a tablet, a computer, a wearable device, or other electronic devices with a voice call function. The earphone can be a wireless earphone, for example, a TWS (True Wireless Stereo) earphone. The earphone can be wirelessly connected to the terminal through Bluetooth.

[0029] In some embodiments, the separating the first voice signal, the first noise signal, the second voice signal, and the second noise signal can include: performing separation processing on the first sound signal to obtain the first voice signal and the first noise signal; and performing separation processing on the second sound signal to obtain the second voice signal and the second noise signal.

[0030] The separation processing on the first sound signal to obtain the first voice signal and the first noise signal can include windowing, spectral subtraction, contrast, and prediction of the first sound signal to separate the first voice signal and the first noise signal.

[0031] The separation processing on the second sound signal to obtain the second voice signal and the second noise signal can include windowing, spectral subtraction, contrast, and prediction of the second sound signal to separate the second voice signal and the second noise signal.

[0032] In some embodiments, as Figure 2As shown, the terminal may include a processor 20, which can acquire the first audio signal and the second audio signal, and separate the first voice signal, the first noise signal, the second voice signal and the second noise signal. The processor 20 processes the first voice signal and the second voice signal to obtain a target voice signal, synthesizes the first noise signal and the second noise signal to obtain a synthesized noise signal, performs noise reduction processing on the synthesized noise signal to obtain a residual noise signal, and synthesizes the target voice signal and the residual noise signal to obtain a call signal.

[0033] In some embodiments, processing the first speech signal and the second speech signal to obtain a target speech signal includes: synthesizing the first speech signal and the second speech signal to obtain a synthesized speech signal; and performing enhancement processing on the synthesized speech signal using a preset speech enhancement algorithm to obtain the target speech signal.

[0034] By synthesizing the first and second voice signals, a more complete and realistic voice signal can be obtained, making the voice call clearer and more authentic.

[0035] Specifically, the processor 20 can synthesize the first speech signal and the second speech signal to obtain a synthesized speech signal, and use a preset speech enhancement algorithm to enhance the synthesized speech signal to obtain the target speech signal.

[0036] In some embodiments, the first noise signal and the second noise signal obtained after separation processing may include some noise signals, so that the synthesized speech signal obtained after synthesis also includes some noise signals. By using the preset speech enhancement algorithm to perform speech enhancement processing on the synthesized speech signal, the noise in the synthesized speech signal can be suppressed and reduced, making the obtained target speech signal purer, improving the speech quality, and thus making the voice call function clearer and more faithful.

[0037] In some embodiments, after the first receiving signal collected by the first receiving module of the terminal and the second receiving signal collected by the second receiving module of the earphone are obtained, i.e. after step S10, the talker noise reduction method further comprises: determining a main speech signal and a compensation speech signal according to the signal-to-noise ratio of the first receiving signal and the signal-to-noise ratio of the second receiving signal. The synthesizing the first speech signal and the second speech signal to obtain a synthesized speech signal comprises: extracting a first feature spectrum of the main speech signal and a second feature spectrum of the compensation speech signal; determining a to-be-repaired spectrum segment according to the first feature spectrum; and repairing the to-be-repaired spectrum segment according to the second feature spectrum, thereby obtaining the synthesized speech signal. By repairing the main speech signal according to the compensation speech signal, a more accurate and real synthesized speech signal can be obtained.

[0038] In some embodiments, the first feature spectrum comprises a time domain spectrum and / or a frequency domain spectrum of the main speech signal, and the second feature spectrum comprises a time domain spectrum and / or a frequency domain spectrum of the compensation speech signal.

[0039] In some embodiments, the determining a main speech signal and a compensation speech signal according to the signal-to-noise ratio of the first receiving signal and the signal-to-noise ratio of the second receiving signal can comprise: when the signal-to-noise ratio of the first receiving signal is greater than the signal-to-noise ratio of the second receiving signal, determining the speech signal of the first receiving signal as the main speech signal and the speech signal of the second receiving signal as the compensation speech signal; and when the signal-to-noise ratio of the second receiving signal is greater than the signal-to-noise ratio of the first receiving signal, determining the speech signal of the second receiving signal as the main speech signal and the speech signal of the first receiving signal as the compensation speech signal. That is, the speech signal in the receiving signal with a higher signal-to-noise ratio is taken as the main speech signal, and the speech signal in the receiving signal with a lower signal-to-noise ratio is taken as the compensation speech signal.

[0040] In some embodiments, when the earphone is closer to the sound producing part of the person being received than the terminal, the signal-to-noise ratio of the second receiving signal is higher than that of the first receiving signal.

[0041] In some embodiments, when the terminal is closer to the sound producing part of the person being received than the earphone, the signal-to-noise ratio of the first receiving signal is higher than that of the second receiving signal.

[0042] In some embodiments, the synthesizing the first voice signal and the second voice signal to obtain a synthesized voice signal comprises: extracting a feature spectrum of the first voice signal and a feature spectrum of the second voice signal; and splicing the feature spectrum of the first voice signal and the feature spectrum of the second voice signal to obtain the synthesized voice signal. By splicing the feature spectrum of the first voice signal and the feature spectrum of the second voice signal, a more complete, accurate and real synthesized voice signal can be obtained.

[0043] The feature spectrum of the first voice signal can comprise a time domain spectrum and / or a frequency domain spectrum of the first voice signal, and the feature spectrum of the second voice signal can comprise a time domain spectrum and / or a frequency domain spectrum of the second voice signal.

[0044] In other embodiments, the first voice signal and the second voice signal can be synthesized based on statistical parameters. Specifically, the acoustic features and the duration of the first voice signal and the acoustic features and the duration of the second voice signal are taken as input information in a model training stage, and are trained and modeled to obtain a duration model and an acoustic model. The duration and acoustic parameters are predicted based on the duration model and the acoustic model to obtain prediction information, and the synthesized voice signal is obtained based on the prediction information.

[0045] In some embodiments, the adopting a preset voice enhancement algorithm to perform enhancement processing on the synthesized voice signal to obtain the target voice signal can comprise: adopting an AI (Artificial Intelligence) voice enhancement algorithm based on a neural network to perform voice enhancement processing on the synthesized voice signal to obtain the target voice signal. The processor 20 can adopt a preset voice enhancement algorithm to perform enhancement processing on the synthesized voice signal.

[0046] The AI voice enhancement algorithm based on the neural network can be a voice enhancement algorithm based on a deep neural network (DNN), a voice enhancement algorithm based on a convolutional neural network (CNN), a voice enhancement algorithm based on a recurrent neural network (RNN), etc.

[0047] For example, the neural network-based AI voice enhancement algorithm is a deep neural network-based voice enhancement algorithm. The synthetic voice signal is processed to include: obtaining the feature of the synthetic voice signal; inputting the feature of the synthetic voice signal into a pre-trained neural network model to obtain the pure voice feature; and performing waveform reconstruction on the pure voice feature to obtain the target voice signal. The pre-trained neural network model processes the feature of the synthetic voice signal to filter out noise features to obtain the pure voice feature, so that the synthetic voice signal can be processed for voice enhancement.

[0048] By using the neural network-based AI voice enhancement algorithm for voice enhancement processing, a more pure voice signal can be obtained compared with using a traditional voice enhancement algorithm (for example, spectral subtraction, Wiener filtering, etc.) for enhancement processing.

[0049] In some embodiments, the noise reduction processing of the synthetic noise signal to obtain a residual noise signal includes: using a preset noise reduction algorithm to perform noise reduction processing on the synthetic noise signal to obtain a residual noise signal. The processor 20 can perform noise reduction processing on the synthetic noise signal to obtain a residual noise signal.

[0050] By performing noise reduction processing on the synthetic noise signal, the obtained residual noise signal is cleaner, and the signal-to-noise ratio of the call signal obtained by synthesizing the residual noise signal and the target voice signal is higher, so that the clarity, intelligibility and voice quality of the call signal can be improved.

[0051] The noise reduction processing of the synthetic noise signal using a preset noise reduction algorithm to obtain a residual noise signal can include: using an LMS (Least Mean Square) adaptive filtering algorithm and / or an AI prediction training noise reduction algorithm based on a neural network to perform noise reduction processing on the synthetic noise signal to obtain the residual noise signal.

[0052] The AI prediction training noise reduction algorithm based on a neural network can be an AI prediction training noise reduction algorithm based on a BP (Back Propagation) neural network, an AI prediction training noise reduction algorithm based on an RBF (Radical Basis Function) neural network, etc.

[0053] In some embodiments, the first noise signal and the second noise signal obtained after the separation processing can include partial speech signals, so that the synthesized noise signal also includes partial speech signals, and by performing noise reduction processing on the synthesized noise signal, the noise in the synthesized noise signal can be removed to a large extent and the partial speech signals can be retained, thereby obtaining the residual noise signal. By synthesizing the residual noise signal and the target speech signal to obtain the call signal, the partial speech signals in the residual noise signal can be used to supplement and repair the target speech signal, thereby further improving the completeness and authenticity of the speech in the call signal.

[0054] In some embodiments, before the first speech signal, the first noise signal, the second speech signal, and the second noise signal are separated, the call noise reduction method further includes: performing noise reduction processing on the second received signal to obtain a preprocessed second received signal, and the preprocessed second received signal includes mixed preprocessed second speech signal and preprocessed second noise signal.

[0055] The first speech signal, the first noise signal, the preprocessed second speech signal, and the preprocessed second noise signal are separated, including: separating the first speech signal, the first noise signal, the preprocessed second speech signal, and the preprocessed second noise signal.

[0056] The first speech signal and the second speech signal are processed to obtain a target speech signal, including: processing the first speech signal and the preprocessed second speech signal to obtain a target speech signal.

[0057] The first noise signal and the second noise signal are synthesized to obtain a synthesized noise signal, including: synthesizing the first noise signal and the preprocessed second noise signal to obtain a synthesized noise signal.

[0058] The first speech signal, the first noise signal, the preprocessed second speech signal, and the preprocessed second noise signal are separated, including: performing separation processing on the first received signal to obtain the first speech signal and the first noise signal; and performing separation processing on the preprocessed second received signal to obtain the preprocessed second speech signal and the preprocessed second noise signal.

[0059] The preprocessed second received signal is separated to obtain the preprocessed second speech signal and the preprocessed second noise signal, including: performing windowing, spectral subtraction, contrast, and prediction on the preprocessed second received signal to separate the preprocessed second speech signal and the preprocessed second noise signal.

[0060] The step of processing the first speech signal and the preprocessed second speech signal to obtain the target speech signal may include: synthesizing the first speech signal and the preprocessed second speech signal to obtain a synthesized speech signal; and using the preset speech enhancement algorithm to enhance the synthesized speech signal to obtain the target speech signal.

[0061] Specifically, before separating the speech signal and noise signal in the second radio signal, the second radio signal is first subjected to noise reduction processing to obtain a preprocessed second radio signal. Then, the preprocessed second radio signal is separated, which can improve the signal-to-noise ratio of the preprocessed second radio signal. This is beneficial for the subsequent separation of the preprocessed second speech signal and the preprocessed second noise signal in the preprocessed second radio signal, and makes the separated preprocessed second speech signal purer, thereby improving the signal-to-noise ratio of the synthesized speech signal obtained in the subsequent synthesis.

[0062] The preprocessed second radio signal can be obtained by denoising the second radio signal using traditional speech enhancement algorithms. These traditional speech enhancement algorithms include spectral subtraction and Wiener filtering.

[0063] In some embodiments, the processor 20 can perform noise reduction processing on the second radio signal using a conventional speech enhancement algorithm to obtain the preprocessed second radio signal. In other embodiments, the processor of the headset can perform noise reduction processing on the second radio signal using a conventional speech enhancement algorithm to obtain the preprocessed second radio signal. The headset processor is used to acquire the second radio signal and perform noise reduction processing on the acquired second radio signal to obtain the preprocessed second radio signal. The acquisition of the first radio signal collected by the first radio module of the terminal during a call and the second radio signal collected by the second radio module of the headset may include: acquiring the first radio signal collected by the first radio module of the terminal during a call and the preprocessed second radio signal.

[0064] In some embodiments, before separating the first speech signal, the first noise signal, the second speech signal, and the second noise signal, the call noise reduction method further includes: performing noise reduction processing on the first radio signal to obtain a preprocessed first radio signal, wherein the preprocessed first radio signal includes a mixture of the preprocessed first speech signal and the preprocessed first noise signal.

[0065] The separating the first voice signal, the first noise signal, the preprocessed second voice signal and the preprocessed second noise signal comprises: separating the preprocessed first voice signal, the preprocessed second voice signal, the preprocessed first noise signal and the preprocessed second noise signal.

[0066] The processing the first voice signal and the preprocessed second voice signal to obtain a target voice signal comprises: processing the preprocessed first voice signal and the preprocessed second voice signal to obtain a target voice signal.

[0067] The synthesizing the first noise signal and the preprocessed second noise signal to obtain a synthesized noise signal comprises: synthesizing the preprocessed first noise signal and the preprocessed second noise signal to obtain a synthesized noise signal.

[0068] The separating the first voice signal, the first noise signal, the preprocessed second voice signal and the preprocessed second noise signal comprises: separating the preprocessed first voice signal, the preprocessed second voice signal, the preprocessed first noise signal and the preprocessed second noise signal.

[0069] The separating the preprocessed first voice signal and the preprocessed first noise signal from the preprocessed first sound signal comprises: windowing, spectral subtraction, contrast and prediction of the preprocessed first sound signal to separate the preprocessed first voice signal and the preprocessed first noise signal.

[0070] The processing the preprocessed first voice signal and the preprocessed second voice signal to obtain a target voice signal comprises: synthesizing the preprocessed first voice signal and the preprocessed second voice signal to obtain the synthesized voice signal; and performing enhancement processing on the synthesized voice signal by using the preset voice enhancement algorithm to obtain the target voice signal.

[0071] Before the separation of the first received signal into a speech signal and a noise signal, the first received signal is first subjected to noise reduction processing to obtain a preprocessed first received signal, and then the preprocessed first received signal is subjected to separation processing. This can improve the signal-to-noise ratio of the preprocessed first received signal, is conducive to subsequent separation of a preprocessed first speech signal and a preprocessed first noise signal in the preprocessed first received signal, and makes the preprocessed first speech signal obtained by separation purer, thereby being conducive to improving the signal-to-noise ratio of the synthesized speech signal obtained by subsequent synthesis.

[0072] The preprocessed first received signal can be obtained by subjecting the first received signal to noise reduction processing by using a conventional speech enhancement algorithm. The conventional speech enhancement algorithm includes spectral subtraction, Wiener filtering, etc.

[0073] In some embodiments, the preprocessed first received signal can be obtained by subjecting the first received signal to noise reduction processing by using a conventional speech enhancement algorithm by the processor 20.

[0074] In some embodiments, the synthesizing the target speech signal and the residual noise signal to obtain a call signal includes synthesizing the target speech signal and the residual noise signal and adding a preset noise signal to obtain the call signal. The call signal can be obtained by synthesizing the target speech signal and the residual noise signal and adding a preset noise signal by the processor 20.

[0075] The preset noise signal can be white noise, such as thermal noise, shot noise, and Gaussian white noise.

[0076] The addition of the preset noise signal to obtain the call signal can improve the auditory comfort.

[0077] Please refer to Figure 2 The structure block diagram of the terminal 100 provided by the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, the terminal 100 includes a processor 20, a memory 10, a communication interface 30, and a display 40. Figure 2As shown, the terminal 100 comprises a first receiving module 10 and a processor 20. The first receiving module 10 is configured to collect a first receiving signal, wherein the first receiving signal comprises a mixed first voice signal and a first noise signal. The processor 20 is configured to acquire the first receiving signal collected by the first receiving module 10 and a second receiving signal collected by a second receiving module of a headset, wherein the second receiving signal comprises a mixed second voice signal and a second noise signal, and separate the first voice signal, the first noise signal, the second voice signal and the second noise signal, synthesize the first noise signal and the second noise signal to obtain a synthesized noise signal, process the first voice signal and the second voice signal to obtain a target voice signal, perform noise reduction processing on the synthesized noise signal to obtain a residual noise signal, and synthesize the target voice signal and the residual noise signal to obtain a call signal.

[0078] In the terminal 100, a communication module (not shown) is further included, which is configured to establish a communication connection with the headset, and the processor 20 acquires the second receiving signal collected by the second receiving module of the headset through the communication module. The communication module can be a Bluetooth module or the like.

[0079] The processor 20 can be a central processing unit (CPU), a microcontroller, a single-chip microcomputer, a digital signal processor (DSP) or the like.

[0080] In some embodiments, the processor 20 is configured to synthesize the first voice signal and the second voice signal to obtain a synthesized voice signal, and perform enhancement processing on the synthesized voice signal by using a preset voice enhancement algorithm to obtain the target voice signal.

[0081] In some embodiments, the processor 20 is further configured to determine a main voice signal and a compensation voice signal according to a signal-to-noise ratio of the first receiving signal and a signal-to-noise ratio of the second receiving signal after the first receiving signal and the second receiving signal are acquired. The processor 20 synthesizes the first voice signal and the second voice signal to obtain a synthesized voice signal, which includes extracting a first feature spectrum of the main voice signal and a second feature spectrum of the compensation voice signal, determining a to-be-repaired spectrum segment according to the first feature spectrum, and repairing the to-be-repaired spectrum segment according to the second feature spectrum, so as to obtain the synthesized voice signal.

[0082] In some embodiments, the processor 20 synthesizes the first speech signal and the second speech signal to obtain a synthesized speech signal, including: extracting a feature spectrum of the first speech signal and a feature spectrum of the second speech signal, splicing the feature spectrum of the first speech signal and the feature spectrum of the second speech signal, and obtaining the synthesized speech signal.

[0083] In some embodiments, the processor 20 is configured to perform enhancement processing on the synthesized speech signal by using a neural network-based AI speech enhancement algorithm to obtain the target speech signal.

[0084] In some embodiments, the processor 20 is configured to perform noise reduction processing on the synthesized noise signal by using a preset noise reduction algorithm to obtain a residual noise signal.

[0085] In some embodiments, the processor 20 is configured to perform noise reduction processing on the synthesized noise signal by using an LMS adaptive filtering algorithm and a neural network-based AI prediction and training noise reduction algorithm to obtain a residual noise signal.

[0086] In some embodiments, the processor 20 is further configured to perform noise reduction processing on the second received signal to obtain a preprocessed second received signal before separating the first speech signal, the first noise signal, the second speech signal, and the second noise signal, the preprocessed second received signal including a mixed preprocessed second speech signal and a preprocessed second noise signal, and the processor 20 is further configured to separate the first speech signal, the first noise signal, the preprocessed second speech signal, and the preprocessed second noise signal, process the first speech signal and the preprocessed second speech signal to obtain the target speech signal, and synthesize the first noise signal and the preprocessed second noise signal to obtain the synthesized noise signal.

[0087] In some embodiments, the processor 20 is further configured to perform noise reduction processing on the first received signal to obtain a preprocessed first received signal before separating the first speech signal, the first noise signal, the second speech signal, and the second noise signal, the preprocessed first received signal including a mixed preprocessed first speech signal and a preprocessed first noise signal, and the processor 20 is further configured to separate the preprocessed first speech signal, the preprocessed second speech signal, the preprocessed first noise signal, and the preprocessed second noise signal, process the preprocessed first speech signal and the preprocessed second speech signal to obtain the target speech signal, and synthesize the preprocessed first noise signal and the preprocessed second noise signal to obtain the synthesized noise signal.

[0088] In some embodiments, the processor 20 synthesizes the target speech signal and the residual noise signal to obtain a talk signal, including: synthesizing the target speech signal and the residual noise signal and adding a preset noise signal to obtain the talk signal.

[0089] The terminal 100 can include a memory (not shown in the figure), and the preset noise signal can be pre-stored in the memory. The memory can be a non-volatile memory, such as a flash disk, a read-only memory, etc.

[0090] The terminal 100 corresponds to the aforementioned talk noise reduction method, and more detailed descriptions can be referred to the content of each embodiment of the aforementioned talk noise reduction method. The terminal 100 and the content of the aforementioned talk noise reduction method can also be mutually referred.

[0091] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program. The computer program is called by a processor to be executed, so as to implement the talk noise reduction method provided by any one of the aforementioned embodiments.

[0092] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiments can be completed by programs instructing relevant hardware, and the programs can be stored in a computer readable memory, which can include a flash disk, a read-only memory, a random access memory, a magnetic disk or an optical disk, etc.

[0093] It should be noted that, for the aforementioned method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.

[0094] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0095] The above is the implementation of the embodiments of the present application. It should be noted that, for those skilled in the art, without departing from the principle of the embodiments of the present application, some improvements and refinements can be made, which are also regarded as the protection scope of the present application.

Claims

1. A method for call noise reduction, characterized in that, The call noise reduction method includes: The system acquires a first audio signal collected by the first audio module of the terminal and a second audio signal collected by the second audio module of the headset during a call. The first audio signal includes a mixture of a first speech signal and a first noise signal, and the second audio signal includes a mixture of a second speech signal and a second noise signal. The first speech signal, the first noise signal, the second speech signal, and the second noise signal are separated. The first speech signal and the second speech signal are processed to obtain the target speech signal; The first noise signal and the second noise signal are combined to obtain a composite noise signal; The synthesized noise signal is denoised to obtain a residual noise signal; and The target speech signal and the residual noise signal are synthesized to obtain a call signal; The step of processing the first speech signal and the second speech signal to obtain the target speech signal includes: A synthesized speech signal is obtained by synthesizing the first speech signal and the second speech signal, or by synthesizing the first speech signal and the preprocessed second speech signal, or by synthesizing the preprocessed first speech signal and the preprocessed second speech signal. The synthesized speech signal is enhanced using a preset speech enhancement algorithm to obtain the target speech signal; The preprocessed second speech signal is obtained by noise reduction and separation processing of the second radio signal, and the preprocessed first speech signal is obtained by noise reduction and separation processing of the first radio signal.

2. The call noise reduction method according to claim 1, characterized in that, The noise reduction process for the synthesized noise signal to obtain the residual noise signal includes: The synthetic noise signal is denoised using a preset denoising algorithm to obtain a residual noise signal.

3. The call noise reduction method according to claim 1, characterized in that, Before separating the first speech signal, the first noise signal, the second speech signal, and the second noise signal, the call noise reduction method further includes: The second radio signal is subjected to noise reduction processing to obtain a preprocessed second radio signal, which includes a mixture of the preprocessed second speech signal and the preprocessed second noise signal. The separation of the first speech signal, the first noise signal, the second speech signal, and the second noise signal includes: The first speech signal, the first noise signal, the preprocessed second speech signal, and the preprocessed second noise signal are separated. The process of processing the first speech signal and the second speech signal to obtain the target speech signal includes: The first speech signal and the preprocessed second speech signal are processed to obtain the target speech signal; The step of synthesizing the first noise signal and the second noise signal to obtain a synthesized noise signal includes: The first noise signal and the preprocessed second noise signal are combined to obtain a synthesized noise signal.

4. The call noise reduction method according to claim 1, characterized in that, The process of synthesizing the target speech signal and the residual noise signal to obtain the call signal includes: The target speech signal and the residual noise signal are synthesized and a preset noise signal is added to obtain the call signal.

5. A terminal, characterized in that, The terminal includes: A first radio module is used to collect a first radio signal, the first radio signal including a mixed first speech signal and a first noise signal; The processor is configured to acquire a first audio signal collected by the first audio module and a second audio signal collected by the second audio module of the earphone, wherein the second audio signal includes a mixed second speech signal and a second noise signal; the processor is configured to separate the first speech signal, the first noise signal, the second speech signal, and the second noise signal; synthesize the first noise signal and the second noise signal to obtain a synthesized noise signal; process the first speech signal and the second speech signal to obtain a target speech signal; perform noise reduction processing on the synthesized noise signal to obtain a residual noise signal; and synthesize the target speech signal and the residual noise signal to obtain a call signal. The processor processes the first speech signal and the second speech signal to obtain the target speech signal, including: The processor synthesizes the first speech signal and the second speech signal to obtain a synthesized speech signal, or synthesizes the first speech signal and the preprocessed second speech signal to obtain a synthesized speech signal, or synthesizes the preprocessed first speech signal and the preprocessed second speech signal to obtain a synthesized speech signal, and uses a preset speech enhancement algorithm to enhance the synthesized speech signal to obtain the target speech signal; The preprocessed second speech signal is obtained by noise reduction and separation processing of the second radio signal, and the preprocessed first speech signal is obtained by noise reduction and separation processing of the first radio signal.

6. The terminal according to claim 5, characterized in that, The processor is used to perform noise reduction processing on the synthetic noise signal using a preset noise reduction algorithm to obtain a residual noise signal.

7. The terminal according to claim 5, characterized in that, The processor is further configured to perform noise reduction processing on the second audio signal to obtain a preprocessed second audio signal before separating the first speech signal, the first noise signal, the second speech signal, and the second noise signal. The preprocessed second audio signal includes a mixture of the preprocessed second speech signal and the preprocessed second noise signal. The processor is also configured to separate the first speech signal, the first noise signal, the preprocessed second speech signal, and the preprocessed second noise signal, process the first speech signal and the preprocessed second speech signal to obtain the target speech signal, and synthesize the first noise signal and the preprocessed second noise signal to obtain the synthesized noise signal.

8. The terminal according to claim 5, characterized in that, The processor synthesizes the target speech signal and the residual noise signal to obtain a call signal, including synthesizing the target speech signal and the residual noise signal and adding a preset noise signal to obtain the call signal.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is used by a computer to execute the call noise reduction method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Voice device with noise cancellation and dual-microphone voice system

    CN109215676A

  • Hearing-aid multichannel voice enhancing algorithm based on iterative Wiener filtering

    CN109961799A