Noise reduction method and device, terminal sound transmission method and device, and server noise reduction processing method and device
By training and generating a speech noise reduction program on the server, combined with the collaborative work of the noise reduction device and the terminal, the problems of high hardware cost and inaccurate noise processing in the existing technology are solved, and efficient and real-time speech noise reduction effects are achieved.
Patent Information
- Application Number
- CN202511089936.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-10-10
AI Technical Summary
Existing speech noise reduction technologies have difficulty in efficiently identifying and processing noise in complex environments, with high hardware costs and limited software algorithm effectiveness.
Ambient sound is collected through the noise reduction device, and the neural network model on the server is used for training to generate a speech noise reduction program, which is then downloaded to the terminal and the noise reduction device for execution to achieve accurate speech noise reduction processing.
It improves the clarity of voice signals and the accuracy of noise reduction, avoids delays and network dependence caused by voice upload and download, and ensures the real-time nature of calls.
Smart Images

Figure CN120766699A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of speech noise reduction technology, and in particular to a noise reduction method and device, a terminal sound transmission method and device, and a server noise reduction processing method and device. BACKGROUND
[0002] Speech noise reduction refers to reducing or eliminating environmental noise through technical means to improve the clarity and intelligibility of speech signals. In a noisy environment, environmental noise often obscures the details of speech, making it difficult for listeners to clearly hear the content of the conversation. Effective speech noise reduction technology can significantly improve the clarity of speech, making conversations smoother, especially in remote conferencing, speech recognition systems, mobile phone calls, hearing aids, and other scenarios.
[0003] In existing speech noise reduction technology, multiple microphones are used to form an array to collect speech signals, and then spatial differences are used to enhance speech signals while reducing background noise. However, the hardware cost is high. Software algorithms analyze the noise spectrum characteristics in speech signals and subtract them from the speech signals to reduce background noise, but cannot efficiently identify and process noise in complex environments.
[0004] Therefore, it is necessary to study a speech noise reduction method to adapt to different noise environments. SUMMARY
[0005] To solve the above technical problems, the present application provides a noise reduction device, comprising: a sound signal acquisition module for generating a sound signal based on collected sound information, the sound information including speech and first environmental sound; a memory adapted to store a speech noise reduction program, the speech noise reduction program being generated based on an environmental sound signal, the environmental sound signal being generated based on second environmental sound; and a processor adapted to execute the speech noise reduction program to perform noise reduction processing on the sound signal.
[0006] Optionally, the noise reduction device further comprises a transmission module adapted to upload the sound signal to a terminal, the terminal being adapted for voice communication.
[0007] Optionally, the transmission module is further adapted to download the speech noise reduction program from the terminal.
[0008] Optionally, the transmission module includes a Type-C interface, a Micro-USB interface, a Lightning interface, or a Bluetooth module.
[0009] Optionally, the sound signal acquisition module includes a microphone.
[0010] Optionally, the sound signal acquisition module further includes a single-ended to differential conversion module, adapted to perform differential processing on the sound information to generate the sound signal.
[0011] Optionally, the memory is further adapted to store a general noise reduction program, and the processor is further adapted to execute the general noise reduction program to perform noise reduction processing on the sound signal.
[0012] Optionally, the processor is further adapted to perform equalization processing on the sound signal after noise reduction.
[0013] To solve the above technical problems, an embodiment of the present invention provides a noise reduction method, including: obtaining a first sound signal, the first sound signal is obtained based on collected first sound information, the first sound information includes first ambient sound and speech; and executing a speech noise reduction program to perform noise reduction processing on the first sound signal, the speech noise reduction program is generated based on second sound information, and the second sound information includes second ambient sound.
[0014] Optionally, the noise reduction method further includes: performing equalization processing on the first sound signal after noise reduction.
[0015] Optionally, the noise reduction method further includes: transmitting the first sound signal after noise reduction to a terminal, where the terminal is suitable for conducting a voice call.
[0016] Optionally, the first sound signal is a signal obtained by performing differential processing on the first sound information.
[0017] Optionally, the noise reduction method further includes: acquiring a second sound signal, where the second sound signal is obtained based on the collected second sound information.
[0018] Optionally, the second sound signal is a signal obtained by performing differential processing on the second sound information, and the method further includes: uploading the second sound signal to a terminal, where the terminal is suitable for conducting voice calls.
[0019] Optionally, the noise reduction method further includes: downloading the voice noise reduction program from the terminal, where the voice noise reduction program is generated according to the second sound signal.
[0020] To solve the above technical problems, an embodiment of the present invention provides a terminal sound transmission method, wherein the terminal is suitable for making voice calls, including: the terminal obtains a first sound signal from a noise reduction device, the first sound signal is obtained based on the collected first sound information, the first sound information includes a first ambient sound and voice, the first sound signal is subjected to noise reduction processing, the noise reduction processing is performed according to a voice noise reduction program, the voice noise reduction program is obtained based on a second sound signal, the second sound signal is obtained based on the collected second sound information, the second sound information includes a second ambient sound; and the first sound signal is transmitted to another terminal.
[0021] Optionally, the terminal sound transmission method further includes: the terminal acquiring the second sound signal from the noise reduction device, where the second sound signal is obtained based on the collected second sound information; and transmitting the second sound signal to the server.
[0022] Optionally, the terminal sound transmission method further includes: receiving a user instruction to enter an environmental noise reduction mode.
[0023] Optionally, the terminal sound transmission method further includes: a step of prompting a user after completing the transmission of the second sound signal.
[0024] To solve the above technical problems, an embodiment of the present invention provides a terminal sound transmission device, including: a transmission module, suitable for obtaining a first sound signal from a noise reduction device, and transmitting the first sound signal to another terminal, the first sound signal is obtained based on the collected first sound information, the first sound information includes a first ambient sound and voice, the first sound signal is subjected to noise reduction processing, the noise reduction processing is performed according to a voice noise reduction program, the voice noise reduction program is obtained based on a second sound signal, the second sound signal is obtained based on the collected second sound information, and the second sound information includes a second ambient sound.
[0025] Optionally, the transmission module is further adapted to obtain the second sound signal from the noise reduction device, and transmit the second sound signal to a server.
[0026] Optionally, the terminal sound transmission device further includes: a receiving module, adapted to receive a user instruction to enter an environmental noise reduction mode.
[0027] Optionally, the terminal sound transmission device further includes: a prompting module, adapted to prompt a user that the transmission of the second sound signal is completed after the second sound signal is transmitted.
[0028] To solve the above technical problems, an embodiment of the present invention provides a server noise reduction processing method, comprising: the server obtains a sound signal from a terminal sound transmission device, the sound signal is obtained based on the collected sound information, and the sound information includes ambient sound; based on the sound signal, a neural network model is trained to obtain a speech noise reduction program; and the speech noise reduction program is transmitted to the terminal sound transmission device.
[0029] To solve the above technical problems, an embodiment of the present invention provides a server noise reduction processing device, comprising: a server transmission module, suitable for obtaining a sound signal from a terminal sound transmission device, wherein the sound signal is obtained based on the collected sound information, and the sound information includes ambient sound; and a server processor, suitable for training based on a neural network model and the sound signal to obtain a speech noise reduction program; the server transmission module is also suitable for transmitting the speech noise reduction program to the terminal sound transmission device.
[0030] Optionally, the neural network model is a deep neural network model, a recurrent neural network model or a convolutional neural network model.
[0031] Compared with the prior art, the technical solution of the embodiment of the present invention has the following beneficial effects.
[0032] In an embodiment of the present invention, it does not rely on hardware multi-microphone noise reduction, but first collects ambient sound through a noise reduction device, and then transmits it to a server through a terminal sound transmission device. After obtaining the ambient sound, the server trains according to a neural network model and the ambient sound to obtain a voice noise reduction program. The voice noise reduction program can more accurately match the ambient sound of the user's scene in the subsequent noise reduction processing. The server transmits the voice noise reduction program to the terminal sound transmission device, and then the terminal sound transmission device downloads the voice noise reduction program to the noise reduction device and executes it, making the noise reduction of the noise reduction device more accurate. In addition, compared with the prior art that completes the noise reduction processing through the server, the embodiment of the present invention can complete the noise reduction processing through the noise reduction device, avoiding the delay caused by voice uploading and voice downloading and the dependence on the network, and ensuring the real-time nature of the call. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a flow chart of a noise reduction method according to an embodiment of the present invention;
[0034] Figure 2 1 is a schematic structural diagram of a noise reduction device according to an embodiment of the present invention;
[0035] Figure 3 is a structural diagram of a terminal sound transmission device according to an embodiment of the present invention; and
[0036] Figure 4 It is a structural diagram of a server noise reduction processing device in an embodiment of the present invention. DETAILED DESCRIPTION
[0037] Voice noise reduction technology aims to improve the quality and clarity of voice signals. Multi-microphone noise reduction technology uses multiple microphones to capture both the user's voice and ambient sound. It then uses a noise reduction algorithm to analyze and process these sound signals, generating a sound signal that is in phase with the ambient sound, canceling out the phase and significantly reducing or eliminating the ambient sound. However, due to the need for hardware such as multiple microphones and the corresponding noise reduction algorithm processing chip, multi-microphone noise reduction technology is relatively expensive.
[0038] In this embodiment of the present invention, the model training step is performed on a server, which is typically equipped with high-performance hardware resources such as large-capacity memory, high-performance CPUs and GPUs, and fast storage devices. These resources can support large-scale data processing and complex computing tasks. After training is completed, the speech noise reduction program no longer requires high-performance hardware resources.
[0039] In an embodiment of the present invention, a trained speech noise reduction program is downloaded to a terminal sound transmission device, and then the speech noise reduction program is downloaded to a processor of the noise reduction device through the terminal sound transmission device for execution. In the subsequent speech noise reduction process, no information interaction with the server is required, thereby improving the real-time performance of the speech noise reduction processing.
[0040] In order to make the above-mentioned objects, features and beneficial effects of the present invention more obvious and easy to understand, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0041] Please refer to Figure 1 , Figure 1 This is a flowchart of a noise reduction method 100 according to an embodiment of the present invention. Before performing noise reduction on a sound signal comprising a speech signal and an ambient sound signal, the method provided by this embodiment of the present invention first performs model training based on the current ambient sound to obtain a speech noise reduction program. The obtained speech noise reduction program is then used to perform noise reduction processing on the sound signal. This speech noise reduction program can more accurately match the ambient sound signal of the user's scene. The noise reduction method 100 is described in detail below in conjunction with steps S11 to S31.
[0042] In step S1, the noise reduction device D collects second sound information.
[0043] The second sound information includes a second ambient sound. In some embodiments, the second sound information is the second ambient sound. In order to generate a speech noise reduction program for noise reduction, it is necessary to obtain the second ambient sound for model training.
[0044] In some embodiments, the noise reduction device D collects the second sound information through a microphone (MIC). Specifically, the microphone can be an analog silicon microphone.
[0045] In some embodiments, after receiving a user instruction, the terminal T triggers the noise reduction device D connected thereto to enter an ambient noise reduction mode. After entering the ambient noise reduction mode, the noise reduction device 200 can collect the second sound information. In a specific implementation, the terminal T can be a mobile phone, a tablet, or the like.
[0046] In step S2, the noise reduction device D performs differential processing on the second sound information to generate a second sound signal.
[0047] Specifically, the differential processing can convert a single-ended sound signal into a differential sound signal. The single-ended sound signal usually corresponds to one signal line and one ground (GND) line, while the differential sound signal corresponds to two signal lines, and the voltage amplitudes on the two signal lines are equal but opposite in phase. Through differential signal transmission, the anti-interference ability and stability of the sound signal can be significantly improved, so that the noise reduction device D has strong anti-interference ability, a longer pickup distance, and more stable performance.
[0048] In step S3, the noise reduction device D transmits the second sound signal to the terminal T.
[0049] Specifically, the noise reduction device D can use wired transmission or wireless transmission, for example, wired transmission through a Type-C interface, a Micro-USB interface, or a Lightning interface, or wireless transmission through a Bluetooth module.
[0050] In step S4, after receiving the second sound signal, the terminal T uploads the second sound signal to the server S.
[0051] In the embodiments of the present application, the terminal T uploads the second sound signal to the server S for model training to obtain a voice noise reduction program.
[0052] In some embodiments, after transmitting the second sound signal, the terminal T prompts the user that the second sound signal transmission is complete, for example, by ringing, pop-up window, or the like.
[0053] In step S5, the server S trains a neural network model according to the second sound signal to obtain a voice noise reduction program.
[0054] The voice noise reduction program will be used for noise reduction processing of the noise reduction device D.
[0055] Specifically, the neural network model can be a deep neural network model, a recurrent neural network model or a convolutional neural network model.
[0056] In a specific implementation, the training process of the neural network model may include steps S51 to S58.
[0057] Step S51: Collecting noise data from multiple environments: First, noise data must be collected under a variety of environmental conditions. These environments may include, but are not limited to, city noise, traffic noise, natural sounds, and office background sounds. By deploying high-quality microphones in these diverse environments, audio samples containing a rich set of noise characteristics can be captured.
[0058] Step S52: Dataset Construction and Annotation: The collected raw audio data needs to undergo careful preprocessing to construct an accurately annotated training dataset. This may include steps such as denoising, sound signal segmentation, and feature extraction to ensure that the dataset provides high-quality input for subsequent model training.
[0059] Step S53: Deep Learning Model Design: Next, design and select appropriate deep learning architectures, such as deep neural networks (DNNs), recurrent neural networks (RNNs), and convolutional neural networks (CNNs). These models can learn complex patterns in audio data and effectively distinguish between noise and human voices.
[0060] Step S54: Parallel processing and vectorized operations: To process large datasets and accelerate the model training process, it is crucial to use parallel computing technology and vectorized operations. These technologies can significantly improve computational efficiency, enabling the model to learn and converge quickly.
[0061] Step S55: Model training and optimization: The selected neural network model is trained using the constructed data set. During the training process, network parameters are continuously adjusted through optimization techniques such as backpropagation and gradient descent to minimize prediction errors.
[0062] Step S56: Feature Learning and Noise Reduction Model Construction: The deep learning model learns the characteristic representations of noise and voice, building a noise reduction model that can identify and separate these two sounds. The learned features capture the unique properties of the human voice while suppressing background noise.
[0063] Step S57: Performance Evaluation and Validation: After model training is complete, it needs to be evaluated on an independent validation set to ensure that the model has good generalization capabilities. Evaluation metrics may include signal distortion, signal-to-noise ratio improvement, speech intelligibility, etc.
[0064] Step S58: Integration and deployment of the noise reduction system: Finally, the trained and verified noise reduction model is integrated into the actual audio processing system. In this way, the system can automatically perform noise reduction operations in real-time communication, conference recording, audio monitoring, or any other application scenario that requires clear speech.
[0065] The speech noise reduction program trained by the server S of the embodiment of the application can improve the signal-to-noise ratio of the speech signal, and the improvement amplitude is greater than 0 dB.
[0066] In step S6, the server S transmits the trained speech noise reduction program to the terminal T.
[0067] In step S7, the terminal T transmits the received speech noise reduction program to the noise reduction device D.
[0068] In some embodiments, the noise reduction device D downloads the speech noise reduction program from the terminal T for noise reduction processing of the sound signal.
[0069] In step S8, the noise reduction device D acquires a first sound signal.
[0070] In some embodiments, the first sound signal is obtained according to first sound information collected by the noise reduction device D, and the first sound information includes a first environmental sound and speech. The first sound information is information to be noise reduced.
[0071] The speech is the main carrier of information and is retained after noise reduction processing. The first environmental sound is the environmental noise of the scene where the user is located, and can be weakened after noise reduction processing, thereby increasing the recognition degree of the speech. In a specific implementation, the first sound information can be sound information during a call.
[0072] The second environmental sound and the first environmental sound can be the same environment or similar environment, especially an environment with little difference in noise, such as a train station environment, so that the speech noise reduction program maintains good environmental adaptability.
[0073] In step S9, the noise reduction device D executes the speech noise reduction program to perform noise reduction processing on the first sound signal.
[0074] Since the speech noise reduction program maintains good environmental adaptability, it can more accurately match the environmental sound of the scene where the user is located in the noise reduction processing. The noise reduction device D can obtain more accurate noise reduction effect based on the speech noise reduction program for noise reduction of the first sound signal.
[0075] In some embodiments, if the environment during a call changes, the noise reduction process can be restarted, that is, the noise reduction device D reacquires the ambient sound and uploads it to the terminal T, and the terminal T then uploads it to the server S for model training to obtain a new voice noise reduction program, and the noise reduction device D uses the new noise reduction program to perform noise reduction processing on the sound signal. In a specific implementation, the sensing module in the terminal T can detect changes in the environment in which the call is located, and then notify the noise reduction device D to reacquire the ambient sound. Optionally, after the sensing module in the terminal T detects changes in the environment in which the call is located, it can also issue a prompt to the user, allowing the user to choose whether the noise reduction device D needs to reacquire the ambient sound.
[0076] In some embodiments, a general noise reduction program can be pre-stored in the noise reduction device D. The general noise reduction program is applicable to most environments. The first sound signal can also be denoised by executing the general noise reduction program, but the noise reduction effect may not be as good as the voice noise reduction program trained based on the second sound signal. In a specific implementation, the noise reduction device D can use the most recently acquired noise reduction program for noise reduction processing. Optionally, the noise reduction device D can also receive user instructions and decide whether to use the general noise reduction program or the voice noise reduction program for noise reduction processing based on the user instructions.
[0077] In step S10, the noise reduction device D performs equalization processing on the first sound signal after noise reduction.
[0078] The equalization processing can make the first sound signal after noise reduction smoother and sound more natural.
[0079] In step S11 , the noise reduction device D transmits the first sound signal after noise reduction and equalization processing to the terminal T.
[0080] The terminal T is suitable for conducting voice calls. In some embodiments, after receiving the processed first sound signal, the terminal T transmits the first sound signal to another terminal (the other terminal) with which the terminal is conducting the voice call, so that the user at the other terminal receives a clearer voice signal. For example, transmission methods such as WiFi or mobile communication can be used for transmission.
[0081] The noise reduction method provided by the embodiment of the present invention does not rely on hardware multi-microphone noise reduction, but first collects ambient sound through the noise reduction device, and then transmits it to the server through the terminal. After obtaining the ambient sound, the server is trained according to the neural network model and the ambient sound to obtain a voice noise reduction program. The voice noise reduction program can more accurately match the ambient sound of the user's scene in the subsequent noise reduction processing. The server transmits the voice noise reduction program to the terminal, and then the terminal downloads the voice noise reduction program to the noise reduction device and executes it, so that the noise reduction of the noise reduction device is more accurate. In addition, compared with the prior art that completes the noise reduction processing through the server, the embodiment of the present invention can complete the noise reduction processing through the noise reduction device, avoiding the delay caused by voice uploading and voice downloading and the dependence on the network, and ensuring the real-time nature of the call.
[0082] The embodiment of the present invention further provides a variety of devices corresponding to the above-mentioned noise reduction method 100.
[0083] Please refer to Figure 2 , Figure 2 2 is a schematic diagram of a noise reduction device 200 according to an embodiment of the present invention. The noise reduction device 200 includes: a sound signal acquisition module 201 , a memory 202 , and a processor 203 .
[0084] The sound signal acquisition module 201 generates a sound signal according to the collected sound information, where the sound information includes speech and first ambient sound.
[0085] In a specific implementation, the sound signal acquisition module 201 includes a microphone (MIC). In some embodiments, the microphone can be an analog silicon microphone, which is small in size and can be easily integrated into the noise reduction device 200.
[0086] During use of the noise reduction device 200, the collected sound information includes speech and first ambient sound. The speech, as the primary carrier of information, is preserved after noise reduction processing. The first ambient sound is the ambient noise of the user's scene, which is weakened after noise reduction processing, thereby increasing the intelligibility of the speech.
[0087] The memory 202 stores a speech noise reduction program, which is generated according to an ambient sound signal, and the ambient sound signal is generated according to a second ambient sound.
[0088] Specifically, the memory 202 may include a read-only memory, a random access memory, a magnetic disk or an optical disk, etc. The computer-readable storage medium may also include a non-volatile memory or a non-transitory memory, etc.
[0089] To generate the speech noise reduction program, the sound signal acquisition module 201 is required to obtain a second environmental sound for training. The second environmental sound and the first environmental sound can be the same environment or a similar environment, especially an environment with similar noise levels, such as a train station environment, so that the speech noise reduction program maintains good environmental adaptability. The processor 203 is configured to execute the speech noise reduction program to perform noise reduction processing on the sound signal.
[0090] Specifically, the processor 203 can be a microcontroller unit (MCU) or a central processing unit (CPU). The processor 203 can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0091] Preferably, the processor 203 includes a brain neural network processor core (Brain Neural Processing Unit, BNPU), which is a processor specially designed for deep learning and neural network operations. It can perform deep learning tasks more efficiently and with lower power consumption than other processors.
[0092] In some embodiments, if the environment during a call changes, the noise reduction process can be restarted, that is, the ambient sound is reacquired and uploaded to the terminal, which then uploads the data to the server for model training to obtain a new voice noise reduction program. The noise reduction device 200 uses the new noise reduction program to perform noise reduction on the sound signal. In a specific implementation, the terminal is provided with a sensing module that can detect changes in the environment during a call, and then notify the noise reduction device 200 to reacquire the ambient sound. Optionally, after the sensing module in the terminal detects changes in the environment during a call, it can also issue a prompt to the user, allowing the user to choose whether the noise reduction device 200 needs to reacquire the ambient sound.
[0093] In some embodiments, the memory 202 may also pre-store a general noise reduction program, which is applicable to most environments. The processor 203 is also suitable for executing the general noise reduction program to reduce the noise of the sound signal, but the noise reduction effect may not be as good as using the voice noise reduction program.
[0094] In a specific implementation, the processor 203 may use the most recently acquired noise reduction program to perform noise reduction processing. Optionally, the noise reduction device 200 may further include an instruction receiving module (e.g., a button, not shown) for receiving user instructions. The processor 203 may determine whether to use the general noise reduction program or the voice noise reduction program for noise reduction processing based on the user instructions.
[0095] The processor 203 is further adapted to perform analog-to-digital conversion (ADC) on the sound signal, converting the sound signal from an analog signal to a digital signal to facilitate subsequent signal processing. The processor 203 is further adapted to encode the digital sound signal to achieve data compression. The processor 203 is further adapted to perform digital-to-analog conversion (DAC) on the digital sound signal for subsequent analog channel transmission. The processor 203 is further adapted to perform equalization processing on the sound signal, making the noise-reduced sound signal smoother and more natural-sounding.
[0096] In some embodiments, the noise reduction device 200 further includes a single-ended to differential module 204 adapted to perform differential processing on the sound information.
[0097] The single-ended to differential module 204 is a circuit module that converts single-ended sound signals into differential sound signals. Single-ended sound signals typically have only one signal line and one common ground (GND) line, while differential sound signals use two signal lines, with voltages of equal amplitude but opposite phases. Differential signal transmission significantly improves the sound signal's anti-interference capability and stability, resulting in a stronger anti-interference capability, a longer sound pickup distance, and more stable performance for the noise reduction device.
[0098] Optionally, the single-ended to differential conversion module 204 may also be integrated into the processor 203 .
[0099] In some embodiments, the noise reduction device 200 further includes a transmission module 205 adapted to upload the second ambient sound signal to a terminal for training and generating the speech noise reduction program. The terminal is adapted to conduct a voice call.
[0100] It is understandable that after the training of the speech noise reduction program is completed, the transmission module 205 is further adapted to download the speech noise reduction program from the terminal, and after the speech noise reduction program is downloaded, it covers the general speech noise reduction program.
[0101] The transmission module 205 is further adapted to upload the noise-reduced sound signal to the terminal after noise reduction processing.
[0102] Specifically, the transmission module 205 can use wired transmission, and the transmission module 205 can include, for example, a Type-C interface, a Micro-USB interface, or a Lightning interface. These interfaces can transmit data while also powering the noise reduction device 200. The transmission module 205 can also use wireless transmission, and the transmission module 205 can include, for example, a Bluetooth module. In this case, the noise reduction device 200 is powered by a battery.
[0103] An embodiment of the present invention further provides a noise reduction method, which can be applied to a noise reduction device (such as the noise reduction device 200 ), and includes step S21 and step S22 .
[0104] Step S21: Acquire a first sound signal, where the first sound signal is obtained based on collected first sound information, where the first sound information includes a first ambient sound and speech;
[0105] Step S22: executing a speech noise reduction program to perform noise reduction processing on the first sound signal, wherein the speech noise reduction program is generated according to second sound information, and the second sound information includes a second ambient sound.
[0106] For more information about the working principle and beneficial effects of the noise reduction method in the embodiment of the present invention, please refer to the above description of the noise reduction device 200 and Figure 1 The relevant description of the noise reduction method 100 in FIG. 1 is not repeated here.
[0107] The embodiment of the present invention also provides a terminal sound transmission device. Figure 2 , Figure 2 FIG2 is a schematic diagram of the structure of a terminal sound transmission device 300 according to an embodiment of the present invention. The terminal sound transmission device 300 includes: a transmission module 301 , a receiving module 302 and a prompt module 303 .
[0108] The transmission module 301 is adapted to obtain the first sound signal from the noise reduction device 200 and transmit the first sound signal to another terminal.
[0109] Similarly, the transmission module 301 may include, for example, a Type-C interface, a Micro-USB interface, or a Lightning interface, or may also include a Bluetooth module. After the transmission module 301 obtains the first sound signal, it may transmit the first sound signal to another terminal using a transmission method such as WiFi or mobile communication.
[0110] In a specific implementation, the first sound signal is obtained based on the collected first sound information, the first sound information includes the first ambient sound and voice, the first sound signal is subjected to noise reduction processing, the noise reduction processing is performed according to a voice noise reduction program, the voice noise reduction program is obtained based on the second sound signal, the second sound signal is obtained based on the collected second sound information, and the second sound information includes the second ambient sound.
[0111] As mentioned above, in order to generate the speech noise reduction program, it is necessary to obtain the second environmental sound for training through the sound signal acquisition module 201. The transmission module 301 is also adapted to obtain the second sound signal from the noise reduction device 200 and transmit the second sound signal to the server for model training.
[0112] In some embodiments, the terminal sound transmission device 300 further includes a receiving module 302 configured to receive a user instruction to enter an ambient noise reduction mode. After entering the ambient noise reduction mode, the second ambient sound can be collected by the noise reduction device 200.
[0113] In other embodiments, the terminal sound transmission device 300 further includes a prompting module 303, which is adapted to prompt the user of the completion of the transmission of the second sound signal after the second sound signal is transmitted, for example, by ringing a ringtone, pop-up window, or the like.
[0114] In some embodiments, the terminal sound transmission device 300 further includes a sensing module (not shown) for detecting changes in the call environment. If so, the noise reduction device 200 may be notified to reacquire the ambient sound for model training to obtain a new voice noise reduction program. The device may also prompt the user to select whether the noise reduction device 200 should reacquire the ambient sound.
[0115] In a specific implementation, the terminal sound transmission device 300 can be a smart terminal device such as a mobile phone or a tablet. Various functions of the terminal sound transmission device 300 can be realized by developing an APP application, such as transmission of the first sound signal, receiving user instructions, prompt functions, etc.
[0116] In an embodiment of the present invention, a terminal sound transmission method is further provided. The terminal sound transmission method can be applied to a terminal side (such as the terminal sound transmission device 300 ), and includes step S31 and step S32 .
[0117] Step S31: The terminal obtains a first sound signal from a noise reduction device, the first sound signal being obtained based on collected first sound information, the first sound information including first ambient sound and speech, the first sound signal undergoing noise reduction processing, the noise reduction processing being performed based on a speech noise reduction program, the speech noise reduction program being obtained based on a second sound signal, the second sound signal being obtained based on collected second sound information, the second sound information including second ambient sound;
[0118] Step S32: Transmit the first sound signal to another terminal.
[0119] For more information about the working principle and beneficial effects of the terminal sound transmission method in the embodiment of the present invention, please refer to the above description of the terminal sound transmission device 300 and Figure 1 The relevant description of the noise reduction method 100 in FIG. 1 is not repeated here.
[0120] The embodiment of the present invention also provides a server noise reduction processing device, referring to Figure 3 , Figure 3 FIG4 is a structural diagram of a server noise reduction processing device 400 according to an embodiment of the present invention. The server noise reduction processing device 400 includes: a server transmission module 401 and a server processor 402 .
[0121] The server transmission module 401 is adapted to obtain a sound signal from a terminal sound transmission device, wherein the sound signal is obtained based on collected sound information, and the sound information includes ambient sound.
[0122] The server processor 402 is adapted to perform training based on the neural network model and the sound signal to obtain a speech noise reduction program.
[0123] The server transmission module 401 is further adapted to transmit the speech noise reduction program to the terminal sound transmission device.
[0124] Specifically, the neural network model is a deep neural network model, a recurrent neural network model or a convolutional neural network model.
[0125] The training process of the neural network model may refer to the aforementioned steps S51 to S58.
[0126] In an embodiment of the present invention, a server noise reduction processing method is further provided. The server noise reduction processing method can be applied to a server side (such as the server noise reduction processing device 400 ), and includes steps S41 to S43 .
[0127] Step S41: The server obtains a sound signal from the terminal sound transmission device, wherein the sound signal is obtained based on the collected sound information, and the sound information includes ambient sound;
[0128] Step S42: training the neural network model according to the sound signal to obtain a speech noise reduction program;
[0129] Step S43: transmitting the speech noise reduction program to the terminal sound transmission device.
[0130] For more details about the working principle and benefits of the server noise reduction processing method in the embodiments of the present application, please refer to the relevant description of the noise reduction method 100 in the server noise reduction processing device 400 and Figure 1 in the above, which will not be repeated here.
[0131] The various methods and devices provided in the embodiments of the present application do not rely on hardware multi-microphone noise reduction, but first collect environmental sound through a noise reduction device, and then transmit it to a server through a terminal sound transmission device. The server trains according to a neural network model and the environmental sound after obtaining the environmental sound to obtain a speech noise reduction program. The speech noise reduction program can more accurately match the environmental sound of the scene where the user is in subsequent noise reduction processing. The server transmits the speech noise reduction program to the terminal sound transmission device, and then the terminal sound transmission device downloads the speech noise reduction program to the noise reduction device and executes it, so that the noise reduction of the noise reduction device is more accurate. In addition, compared with the prior art of completing noise reduction processing through the server, the embodiments of the present application can complete noise reduction processing through the noise reduction device, avoiding the delay and dependence on the network caused by speech uploading and downloading, and ensuring the real-time performance of the call.
[0132] It should be understood that the term "and / or" herein merely describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. In addition, the character " / " in this paper represents that the front and rear associated objects are a "or" relationship. As used herein, unless otherwise explicitly stated, the term "or" covers all possible combinations, unless not feasible. For example, if it is stated that a component can include A or B, then unless explicitly stated otherwise or not feasible, the component can include A, or B, or A and B. As a second example, if it is stated that a component can include A, B or C, then unless explicitly stated otherwise or not feasible, the component can include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.
[0133] The "multiple" appearing in the embodiments of the present application means two or more.
[0134] The relational terms herein, such as first, second and the like, are used solely to distinguish one entity or action from another, without necessarily requiring or implying any actual relationship or order between such entities or actions. Moreover, the words "comprises," "has," and "includes" and other similar forms are intended to be equivalent in meaning and be open-ended, such that any item or items following any one of these words is not meant to be an exhaustive listing of possible items, or meant to be limited to only the items listed.
[0135] It should be noted that the sequence of the steps in the embodiments of the present application does not represent the limitation of the execution sequence of the steps.
[0136] Although the present application has been disclosed with reference to the above embodiments, the present application is not limited to the above embodiments. Any person skilled in the art, without departing from the spirit and scope of the present application, can make various modifications and changes, and the scope of protection of the present application should be limited by the scope defined in the claims.
Claims
1. A noise reduction device, characterized in that: include: A sound signal acquisition module, configured to generate a sound signal based on the collected sound information, wherein the sound information includes speech and a first ambient sound; a memory, adapted to store a speech noise reduction program, wherein the speech noise reduction program is generated according to an ambient sound signal, and the ambient sound signal is generated according to a second ambient sound; as well as The processor is adapted to execute the speech noise reduction program to perform noise reduction processing on the sound signal.
2. The noise reduction device according to claim 1, characterized in that Also includes: The transmission module is adapted to upload the sound signal to a terminal, and the terminal is adapted to conduct a voice call.
3. The noise reduction device according to claim 2, characterized in that: The transmission module is further adapted to download the speech noise reduction program from the terminal.
4. The noise reduction device according to claim 2 or 3, characterized in that: The transmission module includes a Type-C interface, a Micro-USB interface, a Lightning interface or a Bluetooth module.
5. The noise reduction device according to claim 1, characterized in that The sound signal acquisition module includes a microphone.
6. The noise reduction device according to claim 5, characterized in that: The sound signal acquisition module further includes a single-ended to differential conversion module, adapted to perform differential processing on the sound information to generate the sound signal.
7. The noise reduction device according to claim 1, characterized in that: The memory is further adapted to store a general noise reduction program, and the processor is further adapted to execute the general noise reduction program to perform noise reduction processing on the sound signal.
8. The noise reduction device according to claim 7, characterized in that: The processor is further adapted to perform equalization processing on the sound signal after noise reduction.
9. A noise reduction method, characterized in that: include: Acquire a first sound signal, where the first sound signal is obtained based on the collected first sound information, where the first sound information includes a first ambient sound and a voice; as well as A speech noise reduction program is executed to perform noise reduction processing on the first sound signal, where the speech noise reduction program is generated according to second sound information, where the second sound information includes a second ambient sound.
10. The noise reduction method according to claim 9, characterized in that: Also includes: Performing equalization processing on the first sound signal after noise reduction.
11. The noise reduction method according to claim 9 or 10, characterized in that: Also includes: The first sound signal after noise reduction is transmitted to a terminal, and the terminal is suitable for conducting a voice call.
12. The noise reduction method according to claim 9, characterized in that: The first sound signal is a signal obtained by performing differential processing on the first sound information.
13. The noise reduction method according to claim 9, characterized in that: Also includes: A second sound signal is obtained, where the second sound signal is obtained based on the collected second sound information.
14. The noise reduction method according to claim 13, characterized in that: The second sound signal is a signal obtained by performing differential processing on the second sound information, and the method further includes: The second sound signal is uploaded to a terminal, where the terminal is suitable for conducting a voice call.
15. The noise reduction method according to claim 14, characterized in that: Also includes: The speech noise reduction program is downloaded from the terminal, and the speech noise reduction program is generated according to the second sound signal.
16. A method for transmitting sound to a terminal, wherein the terminal is suitable for making a voice call, characterized in that: include: The terminal obtains a first sound signal from a noise reduction device, the first sound signal being obtained based on collected first sound information, the first sound information including first ambient sound and speech, the first sound signal undergoing noise reduction processing, the noise reduction processing being performed based on a speech noise reduction program, the speech noise reduction program being obtained based on a second sound signal, the second sound signal being obtained based on collected second sound information, the second sound information including second ambient sound; and The first sound signal is transmitted to another terminal.
17. The terminal sound transmission method according to claim 16, characterized in that: Also includes: The terminal obtains the second sound signal from the noise reduction device, where the second sound signal is obtained based on the collected second sound information; as well as The second sound signal is transmitted to the server.
18. The terminal sound transmission method according to claim 16, characterized in that: Also includes: Receive user instructions to enter the ambient noise reduction mode.
19. The terminal sound transmission method according to claim 17, characterized in that: Also includes: The step of prompting the user after transmitting the second sound signal is completed.
20. A terminal sound transmission device, characterized in that: include: A transmission module is adapted to obtain a first sound signal from a noise reduction device and transmit the first sound signal to another terminal, wherein the first sound signal is obtained based on the collected first sound information, the first sound information includes a first ambient sound and a voice, the first sound signal undergoes noise reduction processing, the noise reduction processing is performed based on a voice noise reduction program, the voice noise reduction program is obtained based on a second sound signal, the second sound signal is obtained based on the collected second sound information, and the second sound information includes a second ambient sound.
21. The terminal sound transmission device according to claim 20, characterized in that: The transmission module is further adapted to obtain the second sound signal from the noise reduction device and transmit the second sound signal to a server.
22. The terminal sound transmission device according to claim 20, characterized in that: Also includes: The receiving module is adapted to receive a user instruction to enter the ambient noise reduction mode.
23. The terminal sound transmission device according to claim 21, characterized in that: Also includes: The prompt module is adapted to prompt the user that the transmission of the second sound signal is completed after the second sound signal is transmitted.
24. A server noise reduction processing method, characterized in that: include: The server obtains a sound signal from the terminal sound transmission device, wherein the sound signal is obtained based on the collected sound information, and the sound information includes ambient sound; Training a neural network model according to the sound signal to obtain a speech noise reduction program; as well as The speech noise reduction program is transmitted to the terminal sound transmission device.
25. A server noise reduction processing device, characterized in that: include: A server transmission module adapted to obtain a sound signal from a terminal sound transmission device, wherein the sound signal is obtained based on collected sound information, wherein the sound information includes ambient sound; as well as A server processor, adapted to perform training based on a neural network model and the sound signal to obtain a speech noise reduction program; The server transmission module is further adapted to transmit the speech noise reduction program to the terminal sound transmission device.
26. The server noise reduction processing device according to claim 25, characterized in that: The neural network model is a deep neural network model, a recurrent neural network model or a convolutional neural network model.