Downlink noise reduction method and device, equipment and storage medium

By determining the noise signal in the call environment for signal noise reduction and adjusting the volume, the problem of users being affected by noise in voice calls is solved, and clear call content and efficient call experience are achieved.

CN120378535APending Publication Date: 2025-07-25GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510420970.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

During the voice call, the user is affected by the other party's environmental noise and his own environmental noise, resulting in being unable to hear clearly or understand the other party's call content. The existing technical solutions cannot effectively reduce the noise in the downlink voice signal.

Method used

By determining the ambient noise signal of the call environment, signal noise reduction is performed according to the noise intensity, and the volume of the audio output device is adjusted to adapt to the noise environment, ensuring the clarity of the downlink voice signal and the appropriate volume.

Benefits of technology

Effectively reduce noise in downlink voice signals, improve user call quality and call efficiency, and enable users to clearly hear each other's voice content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378535A_ABST
    Figure CN120378535A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a downlink noise reduction method and device, equipment and a storage medium, and the method comprises the steps: determining an environment noise signal of a call environment, which is used for representing the environment noise intensity of an answering user side; and performing signal noise reduction processing on the downlink voice signal according to the environmental noise signal, wherein the noise reduction intensity corresponding to the signal noise reduction processing is related to the environmental noise intensity represented by the environmental noise signal. And controlling the audio output device to play the downlink voice signal after noise reduction processing according to the target volume, wherein the target volume and the environmental noise intensity represented by the environmental noise signal are in a positive correlation relationship. According to the noise reduction method provided by the embodiment of the invention, the downlink voice signal can be subjected to noise reduction based on the ambient noise intensity of the answering user side, the noise existing in the downlink voice signal is accurately reduced, and the playing volume of the downlink voice signal is adjusted according to the ambient noise intensity, so that the volume of the played downlink voice signal is proper, and the user experience is improved. Therefore, the user can hear clear voice content, and the call quality and call efficiency of the user are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of communication technologies, including but not limited to a downlink noise reduction method, device, equipment, and storage medium. Background Art

[0002] When a user uses devices such as mobile phones and computers to make a voice call with another user, due to the ambient noise in the environment of the other user and the user's own environment, as well as the noise introduced during the transmission of the voice signal from the other device to the user device, etc., it will affect the voice quality heard by the user, resulting in the situation that the user cannot hear clearly or understand. Summary of the Invention

[0003] In view of this, the downlink noise reduction method, device, equipment, and storage medium provided by the embodiments of the present application can improve the played voice quality, thereby improving the call quality. The downlink noise reduction method, device, equipment, and storage medium provided by the embodiments of the present application are implemented as follows:

[0004] On the one hand, an embodiment of the present application provides a downlink noise reduction method, including:

[0005] Determine the ambient noise signal of the call environment, and the ambient noise signal is used to characterize the ambient noise intensity of the receiving user end;

[0006] Perform signal noise reduction processing on the downlink voice signal according to the ambient noise signal, and the noise reduction strength corresponding to the signal noise reduction processing is related to the ambient noise intensity characterized by the ambient noise signal;

[0007] Control the audio output device to play the downlink voice signal after noise reduction processing according to the target volume, and the target volume has a positive correlation with the ambient noise intensity characterized by the ambient noise signal.

[0008] On the other hand, an embodiment of the present application further provides a downlink noise reduction device, including:

[0009] A signal determination module, configured to determine the ambient noise signal of the call environment, and the ambient noise signal is used to characterize the ambient noise intensity of the receiving user end;

[0010] A signal noise reduction module, configured to perform signal noise reduction processing on the downlink voice signal according to the ambient noise signal, and the noise reduction strength corresponding to the signal noise reduction processing is related to the ambient noise intensity characterized by the ambient noise signal;

[0011] A signal output module, configured to control the audio output device to play the downlink voice signal after noise reduction processing according to the target volume, and the target volume has a positive correlation with the ambient noise intensity characterized by the ambient noise signal.

[0012] The computer device provided by an embodiment of the present application includes a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, the method described in the embodiment of the present application is implemented.

[0013] The computer-readable storage medium provided by an embodiment of the present application stores a computer program thereon, and when the computer program is executed by a processor, the method provided by the embodiment of the present application is implemented.

[0014] The downlink noise reduction method, device, computer device, and computer-readable storage medium provided by the embodiments of the present application determine the environmental noise signal of the call environment, which is used to characterize the environmental noise intensity of the receiving user end. According to the environmental noise signal, signal noise reduction processing is performed on the downlink voice signal. The noise reduction intensity corresponding to the signal noise reduction processing is related to the environmental noise intensity characterized by the environmental noise signal, and the audio output device is controlled to play the downlink voice signal after noise reduction processing according to the target volume. The target volume has a positive correlation with the environmental noise intensity characterized by the environmental noise signal. In this way, the noise reduction method can perform noise reduction on the downlink voice signal based on the environmental noise intensity of the receiving user end, accurately reduce the noise existing in the downlink voice signal, and adjust the playback volume of the downlink voice signal according to the environmental noise intensity, so that the playback volume of the downlink voice signal is appropriate, enabling the user to hear clear voice content and improving the user's call quality and call efficiency. Description of the Drawings

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0016] Figure 1 A schematic diagram of an application scenario of a downlink noise reduction method according to an embodiment of the present application is shown;

[0017] Figure 2 A flowchart of a downlink noise reduction method according to an embodiment of the present application is shown;

[0018] Figure 3 A flowchart of a noise reduction model training process according to an embodiment of the present application is shown;

[0019] Figure 4 A flowchart of a volume determination model training process according to an embodiment of the present application is shown;

[0020] Figure 5 A flowchart of a downlink signal processing model training process according to an embodiment of the present application is shown;

[0021] Figure 6 A flowchart showing another downlink noise reduction method according to an embodiment of the present application;

[0022] Figure 7 A schematic diagram showing a downlink noise reduction device according to an embodiment of the present application;

[0023] Figure 8 A schematic diagram showing an electronic device according to an embodiment of the present application. Detailed implementation manners

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the specific technical solutions of the present application in detail with reference to the accompanying drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.

[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0026] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0027] It should be noted that the terms "first / second / third" related to the embodiments of this application are used to distinguish similar or different objects, and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0028] The downlink noise reduction method of the embodiments of the present application can be executed by an electronic device, which may include but is not limited to mobile phones, wearable devices (such as smart watches, smart bracelets, smart glasses, etc.), tablet computers, laptop computers, vehicle-mounted terminals, PCs (Personal Computers), etc. The functions implemented by this method can be realized by a processor in the electronic device calling program code. Of course, the program code can be stored in a computer storage medium. It can be seen that the electronic device includes at least a processor and a storage medium.

[0029] The downlink noise reduction method of the embodiment of the present application can be used in any voice call scenario. For example, it can include application scenarios where users make voice calls by means of calling through electronic devices with telephone functions such as mobile phones and smart watches. Or, it can also include application scenarios where users make voice calls through the call functions of application programs such as social software and office software deployed on electronic devices such as smart phones, smart watches, tablets, and laptop computers.

[0030] Optionally, the scenarios where users make voice calls can also include scenarios where they make voice calls with a single other user by calling the operator's phone or an online network phone, and scenarios where multiple other users make voice calls together through online group chats, web conferences, etc. via some application programs.

[0031] Figure 1 The schematic diagram of the application scenario of a downlink noise reduction method according to an embodiment of the present application is shown. As Figure 1 shown, during the process of a user making a voice call based on an electronic device, the voice signal of the communication partner user will be mixed with the ambient noise of the other party. At the same time, some other noise signals will also be introduced during the transmission of the voice signal. Therefore, the downlink voice signal received by the user-side electronic device is a downlink voice signal carrying noise, and further, the call voice heard by the user through the audio output device (such as headphones or speakers, etc.) of the electronic device will have the problem of unclear sound.

[0032] In addition, since the user is also affected by the nearby ambient noise when listening to the call voice. Therefore, the call process of the receiving user side will be further affected by the ambient noise interference in the surrounding environment, further affecting the user's acquisition of accurate call content and resulting in the user being unable to clearly hear or understand the call content of the other user.

[0033] In response to the above problems, there are technical solutions in the related art that perform differential operations on the microphone MIC (Microphone Signal) signal of the receiving party's device to perform noise reduction on the downlink voice signal received by the user-side electronic device to improve the sound quality received by the receiving end during the call process. Or, there are technical solutions that perform noise reduction processing on the ambient noise of the receiving party through the ambient noise of the receiving party to improve the sound quality received by the receiving end during the call process. Or, there are technical solutions that intelligently judge the speaker through voiceprint, register with the speaker's voiceprint after judging the speaker and perform noise reduction to improve the sound quality received by the receiving end during the call process. And, there are technical solutions that adjust the playback volume of the audio output device according to the ambient noise of the receiving party during the call process to improve the sound quality received by the receiving end during the call process.

[0034] However, the above technical solutions respectively have defects such as the solution of obtaining the opposite-end MIC signal is not practical in mobile phone calls from the perspective of signal transmission bandwidth and efficiency, the defect of being unable to perform noise reduction processing on the noise in the downlink voice signal, and the defect of being unable to determine the speaker in the case of multiple people speaking, resulting in incorrect noise reduction objects. In the presence of the above-mentioned certain situations, the call quality will still be affected by noise, causing problems for users such as unclear or inaudible voices.

[0035] The noise reduction method of the embodiments of the present application can perform noise reduction on the downlink voice signal based on the environmental noise intensity of the receiving user end, and adjust the playback volume of the downlink voice signal to obtain a downlink voice signal with accurate noise reduction and appropriate volume. This noise reduction method simultaneously considers the environmental noise of the other user and the local environmental noise of the receiving user to improve the quality of the call content finally heard by the user, enabling the user to obtain clear voice information content, and improving the user's call quality and call efficiency.

[0036] Figure 2 The flowchart of a downlink noise reduction method according to an embodiment of the present application is shown. As Figure 2 shown, the operation analysis method of the embodiments of the present application may include the following steps S21-S23.

[0037] For the convenience of description, the downlink noise reduction method of the embodiments of the present application is described with an electronic device as the execution subject. It should be understood that the execution subject of the embodiments of the present application may also be a processor or a chip in the electronic device, and the embodiments of the present application do not make any limitations.

[0038] Step S21: The electronic device determines the environmental noise signal of the call environment.

[0039] In a possible implementation manner, the electronic device determines the environmental noise signal of the call environment before starting the call process or when the call process has already started. Among them, the environmental noise signal is used to characterize the environmental noise intensity of the receiving user end. The environmental noise intensity is used to characterize the environmental noise level of the receiving user end, and can be obtained by calculating the power or voltage of the environmental noise signal.

[0040] In some embodiments, the scenarios before starting a call may include but are not limited to any voice communication request scenario such as the electronic device receiving a phone call request or an online voice request, an online meeting request, a live connection request, etc. sent by an application program. That is, the electronic device can obtain the environmental noise signal of the call environment where the electronic device is currently located in the above scenarios of receiving a voice call request sent by other user devices.

[0041] In some other embodiments, the scenario where the call process has started can be the scenario where the electronic device receives a telephone call request, or the scenario where it receives voice communication requests such as online voice requests, online meeting requests, and live connection requests sent by an application. That is, the electronic device can start obtaining the ambient noise signal of the answering user terminal immediately at the moment when the user accepts the voice request sent by another user. Or, the scenario where the call process has started can also be the scenario where the answering user terminal establishes a communication connection with the requesting electronic device after receiving a telephone call request, or receiving voice communication requests such as online voice requests, online meeting requests, and live connection requests sent by an application. That is, the electronic device can start obtaining the ambient noise signal of the answering user terminal at the moment when it establishes a communication connection with other user devices as described above.

[0042] Optionally, the ambient noise in the embodiments of the present application refers to other sound signals other than the voice signal of the answering user for recording the voice content to be transmitted. Among them, the voice signal of the answering user is the voice signal for recording the user's voice content that the answering user needs to transmit, and the voice signal for recording the audio content played by the receiving user to the other user, etc. Exemplarily, other sound signals other than the voice signal of the answering user can be other voices in the environment, the echo of the voice signal generated by the user in a specific environment, and other environmental sounds such as vehicle sounds and animal calls that can be collected by the electronic device near the answering user.

[0043] In some possible implementation manners, the electronic device can determine the ambient noise signal of the call environment in any way. Optionally, the ambient noise signal can be a sound signal sampled within a period of time. Exemplarily, the electronic device can determine the ambient noise signal through the uplink voice signal segment in the silent state when there is no voice during the call process. Or, it can also process the uplink voice signal of the answering user terminal sampled within a period of time through a trained ambient noise extraction model to obtain the corresponding ambient noise signal, where the ambient noise extraction model can be trained based on the uplink voice signals before and after noise reduction processing during the user's historical voice calls.

[0044] In some embodiments, the embodiments of the present application can also perform noise reduction processing on the uplink voice signal and accurately determine the environmental noise signal based on the uplink voice signal before and after the noise reduction processing. Exemplarily, the electronic device can acquire the uplink voice signal collected by the audio input device, then perform noise reduction processing on the uplink voice signal, and determine the environmental noise signal according to the uplink voice signal after the noise reduction processing. This method can determine the environmental noise signal based on the noise reduction processing of the uplink voice signal. The environmental noise signal can be simply and accurately determined based on the existing noise reduction algorithm of the uplink voice signal without introducing other algorithm calculations, simplifying the acquisition process of the environmental noise signal, improving the acquisition efficiency, and ensuring that the electronic device can obtain an accurate environmental noise signal with relatively low power consumption.

[0045] Optionally, the audio input device in the embodiments of the present application can be any audio acquisition device provided on the electronic device. For example, it can include the microphone MIC (Microphone Signal) of the electronic device. For example, when the electronic device is a smart phone, the audio input device can specifically include the top MIC provided at the top of the phone and the bottom MIC provided at the bottom of the phone. The uplink voice signal is the voice signal collected by the audio input device during a voice call and needs to be transmitted to the other device, which may include the voice signal recording the voice of the answering user, the environmental noise signal, and the echo interference signal brought by the downlink voice signal output by the audio output device provided on the electronic device.

[0046] In some embodiments, based on the signal content included in the uplink voice signal before noise reduction, the electronic device can perform two noise reduction processes to respectively remove the interference signal generated by the sound played by the audio output device and the environmental noise interference signal in it, so as to obtain a voice signal clearly recording the voice of the user. Thus, the electronic device can determine the environmental noise signal through the signal difference obtained by the two noise reduction processes.

[0047] Exemplarily, the electronic device can first perform a first noise reduction process on the uplink voice signal to obtain a first noise reduction signal, and the first noise reduction signal is the signal obtained by removing the interference signal generated by the sound played by the audio output device in the uplink voice signal. Then perform a second noise reduction process on the first noise reduction signal to obtain a second noise reduction signal, and the second noise reduction signal is the signal obtained by removing the environmental noise interference in the first noise reduction signal. Then determine the environmental noise signal according to the difference between the second noise reduction signal and the first noise reduction signal. Among them, the first noise reduction signal still includes the human voice signal of the environmental noise signal, while the second noise reduction signal is the human voice signal after removing the environmental noise signal. Therefore, the above method for determining the environmental noise signal takes into account the echo interference signal in the uplink voice signal and accurately determines the environmental noise signal at the answering user end by taking the difference between the first noise reduction signal and the first noise reduction signal.

[0048] The process of first performing the first noise reduction process and then the second noise reduction process is because in the uplink voice signal collected by the audio input device, in addition to the user's voice, if the audio playback device is playing sound, it may cause the sound played by the audio playback device to be captured by the microphone. In addition, ambient sound generated in the call environment is also collected. Therefore, it is necessary to first perform the first noise reduction process to filter out the noise signal generated by the sound played by the audio playback device included in the uplink voice signal, and then perform the second noise reduction process to filter out the ambient noise signal included in the uplink voice signal.

[0049] In some embodiments, the electronic device in the embodiments of the present application can remove the interference generated by the sound played by the audio output device in the uplink voice signal through an echo cancellation algorithm to implement the first noise reduction process to obtain a first noise reduction signal. The echo cancellation algorithm is used to eliminate the echo generated when the audio output device is captured again by the audio input device during a call or voice interaction. The embodiments of the present application can implement the first noise reduction process through linear echo cancellation algorithms such as the least mean square algorithm and the recursive least squares method, or non-linear echo cancellation algorithms such as deep learning models or piecewise linearization, to avoid the echo interference signal being included in the determined ambient noise signal and affecting the accuracy of the obtained ambient noise signal.

[0050] In some embodiments, the electronic device in the embodiments of the present application can remove the ambient noise interference in the first noise reduction signal through any uplink noise reduction algorithm to implement the second noise reduction process to obtain a second noise reduction signal. The uplink noise reduction algorithm is used to eliminate or suppress the ambient noise in the voice signal to ensure that the voice transmitted to the other party is clear. The embodiments of the present application can implement the second noise reduction process in any way such as spectral subtraction, Wiener filtering, least mean square error estimation, and Gaussian mixture model and deep neural network model to obtain an accurate and clear uplink voice signal.

[0051] Optionally, in the case where the electronic device first performs ambient noise reduction on the uplink voice signal to obtain a first noise reduction signal, and then performs noise reduction to remove the reverberation interference signal to obtain a second noise reduction signal, the electronic device can actually also determine the ambient noise signal of the receiving user end by subtracting the first noise reduction signal from the uplink voice signal before noise reduction. This method of determining the ambient noise signal is simpler in steps than the above method, can further improve the efficiency of obtaining the ambient noise signal, and improve the real-time performance of the entire downlink noise reduction process.

[0052] Step S22: The electronic device performs signal noise reduction processing on the downlink voice signal according to the ambient noise signal.

[0053] In a possible implementation, after the electronic device obtains the ambient noise signal characterizing the ambient noise situation around the answering user, it can perform signal noise reduction processing on the downlink voice signal obtained from the other user terminal based on the ambient noise signal. Among them, the noise reduction intensity used in the signal noise reduction process can be related to the ambient noise intensity characterized by the ambient noise signal, and the noise reduction intensity is the strength of the noise reduction effect applied when performing noise reduction processing on the signal.

[0054] Exemplarily, in the case of a large noise reduction intensity, more noise in the downlink voice signal can be filtered out, improving the clarity of the downlink voice signal, but it may cause distortion of the downlink voice signal. In the case of a small noise reduction intensity, less noise in the downlink voice signal is filtered out, ensuring the authenticity and naturalness of the downlink voice signal, but the clarity of the downlink voice signal may be poor.

[0055] During the process of a user making a voice call, in addition to the voice signal of the other user that the answering user needs to obtain, the answering user is also affected by the noise in the downlink voice signal and the ambient noise of its own environment. Therefore, if the answering user is in an environment with relatively small ambient noise and is relatively quiet, even if there is relatively large noise in the obtained downlink voice signal, the answering user can clearly hear the voice signal of the other user because the overall noise heard by the user is relatively small. On the contrary, if the answering user is in an environment with relatively large ambient noise and is relatively noisy, when there is relatively large noise in the obtained downlink signal, it is difficult for the answering user to clearly hear the voice signal of the other user because the overall noise heard by the user is relatively large. Therefore, the electronic device can perform signal noise reduction processing on the downlink voice signal according to the ambient noise signal of the call environment, so that the noise reduction intensity is adapted to the ambient noise intensity of the call environment, improving the noise reduction accuracy.

[0056] In some embodiments, the noise reduction intensity can be positively correlated with the ambient noise intensity characterized by the ambient noise signal, that is, when the noise at the answering end of the answering user is larger, a stronger noise reduction intensity is used to perform noise reduction on the downlink voice signal, so that the overall noise of the voice signal obtained by the answering user is relatively small, ensuring that the user can hear clearly during the voice call process.

[0057] Optionally, in the embodiments of the present application, the method of noise reduction for the downlink voice signal can be implemented through a deep learning model, that is, the downlink voice signal and the ambient noise signal or the intensity information of the ambient noise signal can be input into the trained deep learning model, and the noise reduction model performs signal noise reduction processing according to the ambient noise signal, and outputs the downlink voice signal after noise reduction with a noise reduction intensity matching the ambient noise intensity. Among them, the intensity information of the ambient noise signal can be the signal intensity of the ambient noise signal.

[0058] Exemplarily, when the deep learning model only considers the environmental noise signal intensity, the training set of the deep learning model can include multiple sets of training samples. Among them, each set of training samples includes a sample environmental noise signal, a noisy sample speech signal, and a noise-free sample speech signal. The noisy sample speech signals corresponding to the same sample speech signal in different training samples are also the same. That is to say, in different training samples including the same sample speech signal, there is only one variable, namely the environmental noise signal, to ensure that the trained model only considers the environmental noise signal variable for signal noise reduction processing.

[0059] In some other embodiments, due to extreme situations where the environmental noise of the receiving party is small but the environmental noise of the calling party is large, or the environmental noise of the receiving party is large but the environmental noise of the calling party is small. Therefore, in the embodiments of the present application, the electronic device can also jointly determine the corresponding noise reduction strength through the environmental noise signal and the downlink speech signal itself, and perform signal noise reduction processing to further improve the signal noise reduction processing effect.

[0060] Exemplarily, in the embodiments of the present application, the electronic device can input the environmental noise signal and the downlink speech signal into a noise reduction model, and the noise reduction model performs signal noise reduction processing on the downlink speech signal according to the environmental noise signal and the downlink speech signal to obtain a noise-reduced downlink speech signal. The noise reduction model has the ability to analyze the noise existing in the downlink speech signal, and can jointly determine the noise reduction strength in combination with the environmental noise and the noise existing in the downlink speech signal. Thus, in the embodiments of the present application, the noise reduction strength can be jointly determined by comprehensively considering the environmental noise signal and the downlink speech signal itself through the trained noise reduction model, and then the signal noise reduction processing is automatically performed, making the noise-reduced downlink speech signal clearer and improving the noise reduction effect.

[0061] Optionally, in the case where the noise reduction model in the embodiments of the present application jointly determines the corresponding noise reduction strength based on the environmental noise signal and the downlink speech signal itself, the noise reduction model can be trained through a first training set including multiple sets of training samples, and each set of training samples includes a sample environmental noise signal, a noisy sample speech signal, and a noise-free sample speech signal. The noisy sample speech signal is obtained by adding random noise to the noise-free sample speech signal. Among them, in different training samples including the same noise-free sample speech signal, at least one of the included sample environmental noise signal and the noisy sample speech signal is different. That is to say, in different training samples in the first training set including the same sample speech signal, there are two variables, namely the sample environmental noise signal and the randomly added noise, to ensure that the trained model can jointly consider the environmental noise signal and the noise in the downlink speech signal itself, and perform signal noise reduction processing on the downlink speech signal.

[0062] In some embodiments, the process of training the noise reduction model based on multiple training samples in the above-mentioned first training set may include: using the noisy sample speech signal and the sample ambient noise signal in each group of training samples as the input of the noise reduction model. The noise reduction model determines the noise reduction intensity according to the noise of the sample speech signal and the ambient noise signal, and denoises the noisy sample speech signal according to this noise reduction intensity to generate a predicted denoised sample speech signal. Then, the model loss is obtained by comparing the denoised sample speech signal with the noise-free sample speech signal in the training samples, and the model parameters of the noise reduction model are adjusted according to the model damage.

[0063] Figure 3 FIG. shows a flowchart of a noise reduction model training process according to an embodiment of the present application. The training process of this noise reduction model can be implemented in an electronic device that executes the downlink noise reduction method of the embodiment of the present application, or in other devices. As Figure 3 shown, in the electronic device for training the noise reduction model, a noise-free sample speech signal is pre-obtained, and then random noise is added to the noise-free sample speech signal by executing step S30 to obtain a noisy sample speech signal. At the same time, the electronic device for training the noise reduction model also executes step S31 to randomly generate a sample ambient noise signal to obtain the sample ambient noise signal.

[0064] After obtaining the noise-free sample speech signal, the noisy sample speech signal, and the sample ambient noise signal, a training sample can be determined according to the noise-free sample speech signal, the noisy sample speech signal, and the randomly generated sample ambient noise signal. Further, the above method can be repeated to obtain multiple training samples to determine a first training set including multiple training samples. After obtaining the first training set, the electronic device for training the noise reduction model can execute step S32 based on the first training set to train the noise reduction model.

[0065] Step S23: The electronic device controls the audio output device to play the downlink speech signal after noise reduction processing at the target volume.

[0066] In a possible implementation manner, after the electronic device obtains the ambient noise signal characterizing the noise situation around the answering user, it can also adjust the volume of the audio output device based on the ambient noise signal, so as to play the downlink speech signal after noise reduction processing at the target volume through the audio output device with the adjusted volume. Among them, the target volume has a positive correlation with the ambient noise intensity characterized by the ambient noise signal. That is, the greater the ambient noise intensity around the answering user, the greater the target volume. The smaller the ambient noise intensity around the answering user, the smaller the target volume.

[0067] During a voice call, after the receiving user performs noise reduction on the downlink voice signal, in addition to the voice signal of the other user that needs to be obtained, it is also affected by the ambient noise of the environment where the user is located. Therefore, if the receiving user is in an environment with relatively low surrounding noise and is relatively quiet, even if the downlink voice signal is played at a relatively low volume, the receiving user can still clearly hear the voice signal of the other user. On the contrary, if the receiving user is in an environment with relatively high surrounding noise and is relatively noisy, when the downlink voice signal is played at a relatively low volume, since the overall noise heard by the user is relatively high and it is difficult for the user to clearly hear the conversation content of the other user, it is necessary to increase the playback volume of the downlink voice signal so that the receiving user can hear the conversation content of the other user clearly.

[0068] Optionally, in the embodiments of the present application, the electronic device can determine the target volume according to the intensity of the ambient noise signal in any way, so as to play the downlink voice signal after noise reduction at the target volume through the audio output device. Among them, the audio output device can include any device on the electronic device that can play audio signals, such as headphones or speakers.

[0069] In some embodiments, a mapping relationship between the ambient noise signal intensity and the volume can be pre-maintained in the electronic device. This mapping relationship can be a mapping table or a linear function. After the electronic device obtains the ambient noise signal, it can determine the signal intensity of the ambient noise signal, and then determine the target volume corresponding to the audio output device according to the ambient noise signal intensity of the ambient noise signal and the preset mapping relationship between the ambient noise signal intensity and the volume. The preset mapping relationship can be determined according to the ambient noise signal intensity and the corresponding audio output device volume during the historical user call process, or during the process of playing audio and video. Thus, the embodiments of the present application can simply and accurately determine the volume that conforms to the user's habit in the corresponding noise environment as the target volume according to the preset mapping relationship, improving the user experience.

[0070] In other embodiments, the electronic device can also determine the target volume through a trained volume determination model. Exemplarily, the electronic device can input the ambient noise signal into the trained volume determination model, and the volume determination model outputs the target volume corresponding to the audio output device according to the ambient noise signal. This method of determining the target volume related to the ambient noise signal through the volume determination model can flexibly and accurately determine the appropriate target volume through the deep learning model, and after the downlink voice signal is played at the target volume by the volume output device, the volume determination model can be optimized after the user manually adjusts the playback volume again, continuously increasing the possibility that the target volume determined by the model conforms to the user's hearing situation.

[0071] Exemplarily, the volume determination model in the embodiments of the present application can be trained by a second training set, where the second training set includes multiple historical volumes corresponding to the answering user terminal and the historical ambient noise signals corresponding to each historical volume. The historical volume is the volume at which the audio output device outputs an audio signal in the noise environment characterized by the corresponding historical ambient noise signal. The historical volume may include the playback volume of the corresponding earphone or speaker during the historical user call, or may also be the playback volume of the corresponding earphone or speaker when the historical user plays voice information, music, video and other data with sound. The historical ambient noise signal corresponding to each historical volume may be the ambient noise acquired by the electronic device when the historical user plays the corresponding call signal or other sound signals at this volume.

[0072] In some possible implementation manners, the electronic device can acquire and store the ambient noise signal of the environment where the user is located while the user uses the audio playback device each time, so as to periodically update and optimize the volume determination model. Optionally, the multiple historical volumes corresponding to the user and the historical ambient noise signals corresponding to each historical volume can be bound to the user, that is, when the electronic device stores the historical volume and the historical ambient noise signal corresponding to the historical volume, the stored data can be bound to the user identifier. Based on this data storage method, the electronic device can train the corresponding volume determination model for each user specifically, so that the target volume output by the volume determination model is more in line with the preferences of the corresponding user.

[0073] Optionally, the storage address of the historical volume and the corresponding historical ambient noise signal used to determine the second training set can be local or in the cloud. Thus, in the case where the user replaces other devices, the model can also be trained according to the data stored in the cloud to obtain a volume determination model that can determine the volume in line with the user's preferences.

[0074] In some embodiments, the volume determination model in the embodiments of the present application can be obtained by specifically training a pre-trained model according to the second training set corresponding to the user, where the pre-trained model can be trained according to multiple historical volumes corresponding to a large number of users and the historical ambient noise signals corresponding to each historical volume. This training method can initially determine the target volume when a large amount of user historical data has not been obtained, and then gradually optimize the model as the user uses it, improving the accuracy of the model output result.

[0075] Figure 4 The flowchart showing a training process of a volume determination model according to an embodiment of the present application is shown. The training process of this volume determination model can be implemented in the electronic device that executes the downlink noise reduction method of the embodiments of the present application, or can be implemented in other devices. As Figure 4As shown, the electronic device for training the volume determination model pre-obtains the historical data of the calling client in advance, and the historical data can be stored locally or in the cloud. Then, data extraction is performed by executing step S40 to obtain the historical volume and historical ambient noise signal during the historical voice signal playback of the user, so as to obtain the second training set. After obtaining the second training set, the electronic device for training the noise reduction model can execute step S41 based on the second training set to train the volume determination model.

[0076] In a possible implementation manner, in the embodiments of the present application, the process of performing noise reduction on the downlink voice signal based on the ambient noise signal and determining the target volume can also be implemented by the same deep learning model. That is, after obtaining the ambient noise signal of the calling client, the electronic device can input the ambient noise signal into the trained downlink signal processing model to perform noise reduction processing on the downlink voice signal and determine the target volume simultaneously. This method can further simplify the number of models deployed in the electronic device, reduce the computational amount, and improve the computational efficiency.

[0077] Figure 5 The flowchart showing the training process of a downlink signal processing model according to an embodiment of the present application is shown. The training process of the downlink signal processing model can be implemented in the electronic device that executes the downlink noise reduction method of the embodiments of the present application, or in other devices. As Figure 5 As shown, the electronic device for training the downlink signal processing model pre-obtains a noise-free sample voice signal in advance, and then adds random noise to the noise-free sample voice signal by executing step S50 to obtain a noisy sample voice signal. At the same time, the electronic device for training the downlink signal processing model also pre-obtains the historical data of the calling client in advance. Then, data extraction is performed by executing step S51 to obtain the historical volume and historical ambient noise signal during the historical voice signal playback of the user.

[0078] After obtaining the noise-free sample voice signal, the noisy sample voice signal, the historical ambient noise signal, and the historical volume, a training sample can be determined according to the noise-free sample voice signal, the noisy sample voice signal, the historical ambient noise signal, and the historical volume. Further, the above method can be repeated to obtain multiple training samples to determine a third training set including multiple training samples. After obtaining the third training set, the electronic device for training the downlink signal processing model can execute step S52 based on the third training set to train the downlink signal processing model.

[0079] In a possible implementation, in the signal obtained after noise reduction processing on the downlink voice signal, the voice content may also have its sound quality reduced due to the impact of noise reduction. Thus, the electronic device can further perform voice signal enhancement processing on the downlink voice signal after noise reduction processing, so as to control the audio output device to play the enhanced downlink voice signal according to the target volume. Optionally, the intensity of voice signal enhancement is related to the ambient noise intensity characterized by the ambient noise signal. That is, a smaller intensity of voice signal enhancement processing is performed when the ambient noise intensity is smaller, and a larger intensity of voice signal enhancement processing is performed when the ambient noise intensity is larger. This voice signal enhancement processing process can be achieved by enhancing signals in specific frequency bands through spectrum processing or balancing volume fluctuations, etc., for enhancing the voice part content in the downlink voice signal after noise reduction processing.

[0080] Figure 6 The flowchart showing another downlink noise reduction method according to an embodiment of the present application is as follows. As Figure 6 shown, during the process of a user making a voice call through an electronic device, the uplink voice signal that needs to be transmitted to the other device at the local end is input by the audio input device. Since the uplink voice signal includes the voice signal recording the speaking voice of the answering user, the ambient noise signal, and the echo interference signal brought by the downlink voice signal output by the audio output device set on the electronic device. The electronic device can execute step S60 to perform first noise reduction processing on the uplink voice signal to obtain a first noise reduction signal. Then execute step S61 to perform second noise reduction processing on the first noise reduction signal to obtain a second noise reduction signal.

[0081] After the electronic device determines the first noise reduction signal and the second noise reduction signal, it executes step S62 to perform a difference calculation on the first noise reduction signal and the second noise reduction signal to obtain the ambient noise signal at the answering user end. Further, the electronic device executes step S63 based on the ambient noise signal to perform signal noise reduction processing on the received downlink voice signal to obtain the downlink voice signal after noise reduction processing. At the same time, the electronic device also executes step S64 based on the ambient noise signal to obtain the target volume that conforms to the user's hearing habit in this ambient noise situation, and controls the audio output device to play the downlink voice signal after noise reduction processing at this target volume, so that the user can obtain clear voice content of the other party during the call.

[0082] Based on the above technical features, the noise reduction method of the embodiment of the present application can perform noise reduction on the downlink voice signal based on the ambient noise intensity at the answering user end, effectively suppressing the noise in the downlink voice signal to obtain an accurately noise-reduced downlink voice signal. At the same time, the playback volume of the downlink voice signal is adjusted based on the ambient noise intensity, so that the downlink voice signal after noise reduction is played at a volume that conforms to the user's hearing in the current ambient noise, enabling the user to obtain clear voice information content from the downlink voice signal and improving the user's call quality and call efficiency.

[0083] It should be understood that although the steps in the above-mentioned flowcharts are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above-mentioned flowcharts may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0084] Based on the foregoing embodiments, an embodiment of the present application provides a downlink noise reduction device. The device includes each module included and each unit included in each module, and can be implemented by a processor; of course, it can also be implemented by specific logic circuits. During implementation, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or the like.

[0085] Figure 7 A schematic diagram showing a downlink noise reduction device 70 according to an embodiment of the present application. As Figure 7 shown, the downlink noise reduction device 70 of the embodiment of the present application includes:

[0086] A signal determination module 71, configured to determine an environmental noise signal of a call environment, and the environmental noise signal is used to characterize the environmental noise intensity of the answering user end;

[0087] A signal noise reduction module 72, configured to perform signal noise reduction processing on the downlink voice signal according to the environmental noise signal, and the noise reduction strength corresponding to the signal noise reduction processing is related to the environmental noise intensity characterized by the environmental noise signal;

[0088] A signal output module 73, configured to control an audio output device to play the downlink voice signal after noise reduction processing according to a target volume, and the target volume has a positive correlation with the environmental noise intensity characterized by the environmental noise signal.

[0089] In a possible implementation manner, the signal determination module 71 is further configured to:

[0090] Obtain an uplink voice signal collected by an audio input device;

[0091] Perform noise reduction processing on the uplink voice signal, and determine the environmental noise signal according to the uplink voice signal after noise reduction processing.

[0092] In a possible implementation, the signal determination module 71 is further configured to:

[0093] Perform first noise reduction processing on the uplink voice signal to obtain a first noise-reduced signal, where the first noise-reduced signal is a signal obtained by removing the interference generated by the sound played by the audio output device from the uplink voice signal;

[0094] Perform second noise reduction processing on the first noise-reduced signal to obtain a second noise-reduced signal, where the second noise-reduced signal is a signal obtained by removing environmental noise interference from the first noise-reduced signal;

[0095] Determine the environmental noise signal according to the difference between the second noise-reduced signal and the first noise-reduced signal.

[0096] In a possible implementation, the signal determination module 71 is further configured to:

[0097] Perform noise reduction processing on the uplink voice signal through an echo cancellation algorithm to obtain a first noise-reduced signal that removes the interference generated by the sound played by the audio output device from the uplink voice signal.

[0098] In a possible implementation, the signal noise reduction module 72 is further configured to:

[0099] Input the environmental noise signal and the downlink voice signal into a noise reduction model, and the noise reduction model performs signal noise reduction processing on the downlink voice signal according to the environmental noise signal and the downlink voice signal to obtain a noise-reduced downlink voice signal.

[0100] In a possible implementation, the noise reduction model is trained through a first training set;

[0101] The first training set includes multiple groups of training samples, and each group of training samples includes a sample environmental noise signal, a noisy sample voice signal, and a noise-free sample voice signal; the noisy sample voice signal is obtained by adding random noise to the noise-free sample voice signal.

[0102] In a possible implementation, before controlling the audio output device to play the noise-reduced downlink voice signal according to the target volume, the device further includes:

[0103] A first volume determination module, configured to determine the target volume corresponding to the audio output device according to the environmental noise signal intensity of the environmental noise signal and the preset mapping relationship between the environmental noise signal intensity and the volume.

[0104] In a possible implementation, before controlling the audio output device to play the noise-reduced downlink voice signal according to the target volume, the device further includes:

[0105] A second volume determination module, configured to input an environmental noise signal into a trained volume determination model, and output, by the volume determination model according to the environmental noise signal, a target volume corresponding to an audio output device.

[0106] In a possible implementation manner, the volume determination model is obtained through a second training set;

[0107] The second training set includes a plurality of historical volumes corresponding to an answering client and historical environmental noise signals corresponding to each historical volume. The historical volume is the volume at which the audio output device outputs an audio signal in a noise environment characterized by the corresponding historical environmental noise signal.

[0108] In a possible implementation manner, after performing signal noise reduction processing on a downlink voice signal according to the environmental noise signal, the apparatus further includes:

[0109] A signal enhancement module, configured to perform voice signal enhancement processing on the noise-reduced downlink voice signal, where the intensity of the voice signal enhancement is related to the environmental noise intensity characterized by the environmental noise signal;

[0110] A signal output module 73, further configured to:

[0111] Control the audio output device to play the enhanced downlink voice signal according to the target volume.

[0112] The description of the above apparatus embodiments is similar to the description of the above method embodiments, and has similar beneficial effects to the method embodiments. For technical details not disclosed in the apparatus embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0113] It should be noted that in the embodiments of the present application Figure 7 The division of the modules of the downlink noise reduction apparatus shown is illustrative, and is only a logical function division. In actual implementation, there may be other division methods. In addition, each functional unit in the embodiments of the present application may be integrated in one processing unit, may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware, or may be implemented in the form of a software functional unit. It may also be implemented in the form of a combination of software and hardware.

[0114] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0115] Figure 8 FIG. shows a schematic diagram of an electronic device according to an embodiment of the present application. As Figure 8 shown, an embodiment of the present application provides an electronic device, which may be a server, and its internal structural diagram may be as Figure 8 shown. The electronic device includes a processor 820, a memory, and a transceiver 840 connected through a system bus 810. Among them, the processor 820 of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium 831 and an internal memory 832. The non-volatile storage medium 831 stores an operating system, a computer program, and a database. The internal memory 832 provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium 831. The database of the electronic device is used to store data. The transceiver 840 of the electronic device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor 820, implements the above method.

[0116] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor 820, it implements the steps in the method provided in the above embodiment.

[0117] An embodiment of the present application provides a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the steps in the method provided in the above method embodiment.

[0118] Those skilled in the art can understand that Figure 8 the structure shown in

[0119] In one embodiment, the job parsing device provided by the present application can be implemented in the form of a computer program, and the computer program can run on an electronic device as shown in Figure 8 . Each program module that makes up the above device can be stored in the memory of the electronic device. The computer program composed of each program module enables the processor 820 to execute the steps in the methods of various embodiments of the present application described in this specification.

[0120] It should be noted here that the descriptions of the above storage medium and device embodiments are similar to those of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium, storage medium and device embodiments of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.

[0121] It should be understood that the phrase "in one embodiment" or "in an embodiment" or "in some embodiments" mentioned throughout the specification means that specific features, structures, or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" or "in some embodiments" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments. The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments, and the same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated herein.

[0122] The term "and / or" in this article is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, object A and / or object B can represent: object A exists alone, object A and object B exist simultaneously, and object B exists alone.

[0123] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the element.

[0124] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the couplings, direct couplings, or communication connections between the various components shown or discussed can be through some interfaces, and the indirect couplings or communication connections of devices or modules can be electrical, mechanical, or other forms.

[0125] The modules described above as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules; they can be located in one place or distributed to multiple network units; some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0126] In addition, in each embodiment of the present application, all the functional modules can be integrated in one processing unit, or each module can be separately used as a unit, or two or more modules can be integrated in one unit; the above-mentioned integrated modules can be implemented in the form of hardware or in the form of a combination of hardware and software functional units.

[0127] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), magnetic disks, or optical disks, etc., which can store program codes.

[0128] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable an electronic device to execute all or part of the methods described in various embodiments of the present application. And the foregoing storage medium includes: removable storage devices, ROM, magnetic disks, or optical disks, etc., which can store program codes.

[0129] The methods disclosed in several method embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments.

[0130] The features disclosed in several product embodiments provided by this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0131] The features disclosed in several method or device embodiments provided by this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0132] As mentioned above, it is only the implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A downlink noise reduction method, characterized in that, The method includes: Determining an environmental noise signal of a call environment, where the environmental noise signal is used to characterize the environmental noise intensity of the receiving user terminal; Performing signal noise reduction processing on a downlink voice signal according to the environmental noise signal, where the noise reduction strength corresponding to the signal noise reduction processing is related to the environmental noise intensity characterized by the environmental noise signal; Controlling an audio output device to play the downlink voice signal after noise reduction processing at a target volume, where the target volume has a positive correlation with the environmental noise intensity characterized by the environmental noise signal.

2. The method according to claim 1, wherein The determining of the environmental noise signal of the call environment includes: Obtaining an uplink voice signal collected by an audio input device; Performing noise reduction processing on the uplink voice signal and determining an environmental noise signal according to the uplink voice signal after noise reduction processing.

3. The method according to claim 2, wherein The performing of noise reduction processing on the uplink voice signal and determining an environmental noise signal according to the uplink voice signal after noise reduction processing includes: Performing a first noise reduction process on the uplink voice signal to obtain a first noise reduction signal, where the first noise reduction signal is a signal obtained by removing the interference generated by the sound played by the audio output device from the uplink voice signal; Performing a second noise reduction process on the first noise reduction signal to obtain a second noise reduction signal, where the second noise reduction signal is a signal obtained by removing environmental noise interference from the first noise reduction signal; Determining an environmental noise signal according to the difference between the second noise reduction signal and the first noise reduction signal.

4. The method according to claim 3, wherein The performing of a first noise reduction process on the uplink voice signal to obtain a first noise reduction signal includes: Performing noise reduction processing on the uplink voice signal through an echo cancellation algorithm to obtain a first noise reduction signal that removes the interference generated by the sound played by the audio output device from the uplink voice signal.

5. The method according to any one of claims 1 to 4, characterized in that The performing of signal noise reduction processing on the downlink voice signal according to the environmental noise signal includes: Inputting the environmental noise signal and the downlink voice signal into a noise reduction model, and through the noise reduction model, performing signal noise reduction processing on the downlink voice signal according to the environmental noise signal and the downlink voice signal to obtain a downlink voice signal after noise reduction.

6. The method according to claim 5, wherein The noise reduction model is trained through a first training set; The first training set includes multiple groups of training samples, and each group of training samples includes a sample environmental noise signal, a sample voice signal with noise, and a sample voice signal without noise; the sample voice signal with noise is obtained by adding random noise to the sample voice signal without noise.

7. The method according to claim 1, characterized in that, Before the controlling of the audio output device to play the downlink voice signal after noise reduction processing at a target volume, the method further includes: Determining the target volume corresponding to the audio output device according to the environmental noise signal intensity of the environmental noise signal and a preset mapping relationship between the environmental noise signal intensity and the volume.

8. The method according to claim 1, wherein Before the controlling of the audio output device to play the downlink voice signal after noise reduction processing at a target volume, the method further includes: Inputting the environmental noise signal into a trained volume determination model, and through the volume determination model, outputting the target volume corresponding to the audio output device according to the environmental noise signal.

9. The method according to claim 8, characterized in that The volume determination model is obtained through a second training set; The second training set includes a plurality of historical volumes corresponding to the answering client and historical ambient noise signals corresponding to each historical volume, where the historical volume is the volume at which the audio output device outputs an audio signal in a noise environment characterized by the corresponding historical ambient noise signal.

10. The method according to any one of claims 1 to 4, 7 to 9, characterized in that, After performing signal noise reduction processing on the downlink voice signal according to the ambient noise signal, the method further includes: Performing voice signal enhancement processing on the downlink voice signal after noise reduction processing, where the intensity of the voice signal enhancement is related to the ambient noise intensity characterized by the ambient noise signal; Controlling the audio output device to play the downlink voice signal after noise reduction processing at a target volume includes: Controlling the audio output device to play the downlink voice signal after enhancement processing at a target volume.

11. A downlink noise reduction device, characterized in that, The device includes: A signal determination module, configured to determine an ambient noise signal of a call environment, where the ambient noise signal is used to characterize the ambient noise intensity of the answering client; A signal noise reduction module, configured to perform signal noise reduction processing on the downlink voice signal according to the ambient noise signal, where the noise reduction strength corresponding to the signal noise reduction processing is related to the ambient noise intensity characterized by the ambient noise signal; A signal output module, configured to control the audio output device to play the downlink voice signal after noise reduction processing at a target volume, where the target volume has a positive correlation with the ambient noise intensity characterized by the ambient noise signal.

12. An electronic device, comprising a memory and a processor, the memory storing a computer program that can run on the processor, characterized in that, When the processor executes the program, the steps of the method according to any one of claims 1 to 10 are implemented.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.