Call audio processing method, device, computer equipment and storage medium

By selecting the call environment noise and masking threshold of the receiving terminal to filter the audio of multi-person calls, the problem of high server bandwidth usage is solved and the call quality of multi-person calls is improved.

CN114067822BActive Publication Date: 2025-09-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010786543.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-07
Publication Date
2025-09-23
Estimated Expiration
2040-08-09

AI Technical Summary

Technical Problem

During a multi-person call, the server needs to forward a large amount of call audio, resulting in high network bandwidth usage and affecting call quality.

Method used

Select a receiving terminal, obtain its call environment noise, determine the masking degree of each call audio based on the masking threshold of the noise and call audio, and filter the call audio based on the masking degree before sending it.

Benefits of technology

Eliminate call audio that is easily masked by the receiving terminal's ambient noise, reduce the amount of audio forwarded by the server, reduce network bandwidth usage, and improve call quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114067822B_ABST
    Figure CN114067822B_ABST
Patent Text Reader

Abstract

The present application relates to a call audio processing method, device, computer equipment and computer-readable storage medium. The method includes: obtaining call audio sent by multiple call member terminals participating in a multi-person call; selecting one of the call member terminals participating in the multi-person call as a receiving terminal, and obtaining the call environment noise of the receiving terminal; determining the masking degree of each call audio according to the call environment noise and the masking threshold of each call audio; the masking degree indicates the degree to which the call audio is masked by the call environment noise; and filtering the call audio according to each masking degree and sending it to the receiving terminal. This method can improve the call quality of multi-person calls.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a call audio processing method, apparatus, computer equipment, and storage medium. Background Art

[0002] With the development of Internet technology, the application of call functions is becoming more and more extensive. For example, it can enable multiple people to talk at the same time and improve communication efficiency.

[0003] During a multi-person call, the sending terminal sends call audio to the server, and the server is responsible for forwarding the received call audio to each receiving terminal. In many cases, multiple sending terminals send call audio at the same time, and the server needs to forward multiple call audios to multiple receiving terminals separately. The more users participate in the call process, the larger the amount of data forwarded by the server, causing the server to occupy more network bandwidth, thereby affecting the call quality during the multi-person call process. Summary of the Invention

[0004] Based on this, it is necessary to provide a call audio processing method, device, computer equipment and storage medium that can improve the call quality of multi-person calls in response to the above technical problems.

[0005] A call audio processing method, the method comprising:

[0006] Obtain call audio sent by multiple call member terminals participating in a multi-person call;

[0007] Selecting one of the call member terminals participating in a multi-person call as a receiving terminal and obtaining the call environment noise of the receiving terminal;

[0008] Determine the masking degree of each call audio based on the call environment noise and the masking threshold of each call audio; the masking degree indicates the degree to which the call audio is masked by the call environment noise;

[0009] The call audio is filtered according to each masking degree and then sent to the receiving terminal.

[0010] A call audio processing device, the device comprising:

[0011] An acquisition module is used to acquire call audio sent by multiple call member terminals participating in a multi-person call;

[0012] The acquisition module is further configured to select one of the call member terminals participating in the multi-person call as a receiving terminal and obtain the call environment noise of the receiving terminal;

[0013] a determination module for determining a masking degree of each call audio according to the call environment noise and a masking threshold of each call audio; the masking degree indicates the degree to which the call audio is masked by the call environment noise;

[0014] The screening module is used to screen the call audio according to each masking degree and then send it to the receiving terminal.

[0015] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are performed:

[0016] Obtain call audio sent by multiple call member terminals participating in a multi-person call;

[0017] Selecting one of the call member terminals participating in a multi-person call as a receiving terminal and obtaining the call environment noise of the receiving terminal;

[0018] Determine the masking degree of each call audio based on the call environment noise and the masking threshold of each call audio; the masking degree indicates the degree to which the call audio is masked by the call environment noise;

[0019] The call audio is filtered according to each masking degree and then sent to the receiving terminal.

[0020] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:

[0021] Obtain call audio sent by multiple call member terminals participating in a multi-person call;

[0022] Selecting one of the call member terminals participating in a multi-person call as a receiving terminal and obtaining the call environment noise of the receiving terminal;

[0023] Determine the masking degree of each call audio based on the call environment noise and the masking threshold of each call audio; the masking degree indicates the degree to which the call audio is masked by the call environment noise;

[0024] The call audio is filtered according to each masking degree and then sent to the receiving terminal.

[0025] The above-mentioned call audio processing method, device, computer equipment and storage medium obtain call audio sent by multiple call member terminals participating in a multi-person call, select one of the call member terminals participating in the multi-person call as a receiving terminal, obtain the call environment noise of the receiving terminal, determine the masking degree of each call audio according to the call environment noise and the masking threshold of each call audio, and filter the call audio according to each masking degree and then send it to the receiving terminal. In this way, the call audio sent by multiple terminals participating in the multi-person call is filtered according to the masking degree, and the call audio that is easily masked by the call environment noise of the receiving terminal can be eliminated, so that the call members of the receiving terminal can hear the received call audio clearly; and the amount of call audio forwarded by the server is reduced, and the network bandwidth occupied by the server is reduced, thereby improving the call quality of the multi-person call process. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 A diagram illustrating an application environment of a call audio processing method in one embodiment;

[0027] Figure 2 This is a diagram of an application environment of a call audio processing method in another embodiment;

[0028] Figure 3 1 is a flow chart of a method for processing call audio in one embodiment;

[0029] Figure 4 is a schematic diagram of critical frequency bands in one embodiment;

[0030] Figure 5 A flowchart of a call audio processing method in another embodiment;

[0031] Figure 6 A flowchart of a call audio processing method according to another embodiment;

[0032] Figure 7 1. A flowchart of a multi-person call process in one embodiment;

[0033] Figure 8 A flowchart of a multi-person call process in another embodiment;

[0034] Figure 9 This is a structural block diagram of a call audio processing device in one embodiment;

[0035] Figure 10 is a diagram of the internal structure of a computer device in one embodiment;

[0036] Figure 11 FIG. 4 is a diagram showing the internal structure of a computer device in another embodiment. DETAILED DESCRIPTION

[0037] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0038] The call audio processing method provided in this application can be applied to Figure 1 In the application environment shown. Figure 1The application environment includes multiple terminals 102 and a server 104. Each terminal 102 is connected to the server 104 via a network, and each terminal 102 can serve as both a transmitter and receiver of call audio. Specifically, at the same time, one or at least two terminals 102 send call audio to the server 104, and the server 104 forwards the received call audio to each terminal 102 participating in the call.

[0039] The terminal 102 may be a mobile phone, a tablet computer, or other devices. The server 104 may be a single server, a server cluster consisting of several servers, or a cloud computing service center.

[0040] Figure 2 This is a schematic diagram of another application environment provided by the embodiment of the present application, see Figure 2 The application environment includes: multiple terminals 202, a first server 204, and a second server 206. The terminals 202 are connected to the first server 204, or the terminals 202 are connected to the second server 206, and the first server 204 is connected to the second server 206.

[0041] The terminal 202 may be a mobile phone, a tablet computer, or other types of devices. The first server 204 and the second server 204 may be a single server, or a server cluster consisting of several servers, or a cloud computing service center.

[0042] For example, when the first terminal and the second terminal are in the same call group, assuming that the first terminal is a sending terminal and the second terminal is a receiving terminal, the first terminal is connected to the first server and the second terminal is connected to the second server, the first server receives the call audio sent by the first terminal and sends the call audio to the second server, and the second server receives the call audio sent by the first server and sends the call audio to the second terminal.

[0043] Cloud technology refers to a hosting technology that unifies hardware, software, network and other resources within a wide area network or local area network to achieve data computing, storage, processing and sharing.

[0044] Cloud technology is a general term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies based on the cloud computing business model. It can form a resource pool that can be used flexibly and conveniently on demand. Cloud computing technology will become a crucial support. Backend services for technical network systems, such as video websites, image websites, and more portals, require extensive computing and storage resources. With the rapid development and application of the internet industry, every item will likely have its own unique identifier, requiring transmission to backend systems for logical processing. Different levels of data will be processed separately, and data from various industries will require a strong system backend, which can only be achieved through cloud computing.

[0045] The method provided in the embodiments of the present application can be applied to voice calls, video calls, or other call scenarios. The voice call or video call can be a VOIP (Voice over Internet Protocol) multi-person conference scenario or other scenarios.

[0046] For example, in a voice call scenario, where multiple terminals exchange voice data, a server uses the method provided in an embodiment of the present application to select target voice data from at least two channels of voice data sent by at least two sending terminals and sends it to a receiving terminal. The receiving terminal decodes the received target voice data and mixes and plays it.

[0047] For example, in a video call scenario, multiple terminals interact with each other to exchange video data, which includes voice data and image data. The server processes the voice data and image data in the video call separately.

[0048] For processing voice data, the server uses the method provided in the embodiment of the present application to select target voice data from at least two channels of voice data sent by at least two sending terminals and send it to the receiving terminal. The receiving terminal decodes the received target voice data and mixes and plays it.

[0049] For image data processing, the server transmits the image data sent by at least two sending terminals to the receiving terminal. The receiving terminal then displays the image data based on the at least two received image data and the image data collected by the terminal. The displayed image data can be a combination of the at least two received image data and the image data collected by the terminal, or it can be a selected image data from the at least two received image data and the image data collected by the terminal based on user operation.

[0050] In one embodiment, Figure 3 As shown in FIG, a method for processing call audio is provided. Figure 1Taking the multiple terminals 102 and the server 104 in the example as an example, the method includes the following steps:

[0051] Step 302: Acquire call audio sent by multiple call member terminals participating in a multi-person call.

[0052] In this embodiment of the present application, at least three terminals are added to the same talk group, and these at least three terminals can conduct a call. During a call, the terminal that sends voice data is called the sending terminal, and the terminal that receives voice data sent by other terminals is called the receiving terminal. Each terminal in the talk group can be either a sending terminal or a receiving terminal.

[0053] The call group can be a voice call group or a video call group, meaning at least three terminals can conduct either a voice call or a video call. During a voice call, at least three terminals need to exchange voice data; during a video call, at least three terminals need to exchange not only voice data but also image data. This embodiment of the application only illustrates the exchange process for voice data.

[0054] A multi-person call is a call involving at least three participants. During a multi-person call, the participating parties collect audio through different terminals, encode it, package it, and send it to the server. The server decodes the encoded data and retrieves the audio from multiple terminals.

[0055] Call audio is the audio signal recorded by the terminal's sound collection device (such as a microphone). Call audio can include the voice data of the call participants, in which case the terminal acts as a transmitter. Alternatively, call audio can exclude the voice data of the call participants, for example, by containing only background noise or silence, in which case the terminal acts as a receiver.

[0056] During a multi-person call, the terminal can collect audio signals in real time. After collecting the audio signals, the terminal can send them to the server immediately, or it can cache the audio signals and then extract the cached audio signals and send them to the server later.

[0057] In this embodiment, obtaining the call audio sent by each terminal participating in a multi-person call may include the following methods:

[0058] (1) Obtain the call audio sent by all terminals participating in a multi-person call.

[0059] That is, in step 302, the server obtains the call audio sent by all terminals participating in the multi-person call.

[0060] For example, a call group includes terminals A, B, C, and D. If the call members corresponding to terminals A, B, and C are speaking, but the call member corresponding to terminal D is silent, the call audio sent by terminals A, B, and C all contains voice data, while the call audio sent by terminal D does not contain voice data. In this case, terminal C is the receiving terminal, and terminals A, B, and D are the sending terminals. The server can obtain the call audio sent by terminals A, B, C, and D.

[0061] (2) Obtain the call audio sent by the sending terminal.

[0062] That is, the call audio obtained by the server in step 302 only includes the call audio sent by the sending terminal, and does not include the call audio sent by the receiving terminal. For example, if terminal C is the receiving terminal and terminals A, B, and D are the sending terminals, the server can obtain the call audio sent by terminals A, B, and D.

[0063] Step 304: Select one of the call member terminals participating in the multi-person call as a receiving terminal, and obtain the call environment noise of the receiving terminal.

[0064] The call environment noise is the noise in the environment where the receiving terminal is located.

[0065] It can be understood that selecting one of the call member terminals participating in the multi-person call as the receiving terminal means processing each call member terminal participating in the multi-person call as a receiving terminal in turn.

[0066] In the embodiments of this application, the background acoustic environment of the receiving terminal affects the listening experience of the call participants at the receiving terminal. For example, when the background acoustic environment of the receiving terminal is noisy, it is more difficult for the call participants at the receiving terminal to hear the voices of multiple parties clearly. Therefore, the noise level of the call environment at the receiving terminal is analyzed.

[0067] In this embodiment, obtaining the call environment noise of the receiving terminal may include the following methods: In one possible implementation, the server obtains the call environment noise based on the call audio of the receiving terminal. The server may simultaneously obtain the call audio sent by the sending terminal and the receiving terminal according to the above method (1), and then obtain the call environment noise based on the call audio of the receiving terminal. The server may also obtain the call audio sent by the sending terminal and the receiving terminal respectively according to the above method (2), and then obtain the call environment noise based on the call audio of the receiving terminal.

[0068] In one possible implementation, the receiving terminal obtains the call environment noise based on the call audio and sends the obtained call environment noise to the server. The server can simultaneously obtain the call audio sent by the sending terminal and the call environment noise sent by the receiving terminal according to the above method (1), or can obtain the call audio sent by the sending terminal and the call environment noise sent by the receiving terminal separately according to the above method (2).

[0069] In one embodiment, a method for obtaining the call environment noise of a receiving terminal includes: performing frequency domain conversion processing on the call audio of the receiving terminal to obtain a power spectrum of the call audio of the receiving terminal in the frequency domain; and determining the call environment noise of the receiving terminal based on the power spectrum.

[0070] Frequency domain conversion is used to convert the audio signal collected by the terminal from the time domain to the frequency domain for call audio analysis. The time domain shows how the audio signal changes over time, while the frequency domain shows how the audio signal changes over frequency.

[0071] Sound signals exist in the form of waves, so they generate power. The power spectrum represents the distribution of signal power in the frequency domain, that is, how the signal power changes with frequency.

[0072] The frequency domain can be divided into multiple frequency ranges of equal length, each of which is called a frequency bin. Each frequency bin has a corresponding power spectrum. The power spectrum reflects the energy of the sound signal at that frequency bin. The higher the power spectrum, the higher the energy of the sound signal at that frequency bin.

[0073] First, the call audio collected by the receiving terminal undergoes frame-based windowing and discrete Fourier transform processing to convert the call audio from the time domain to the frequency domain. Because audio signals are short-term stationary, framing separates the audio signal into at least two segments of equal length, each of which is called a frame. However, segmenting the audio signal can easily lead to spectral leakage. Therefore, windowing is used to make the framed audio signal continuous, reducing spectral leakage and ensuring that each frame exhibits the characteristics of a periodic function.

[0074] A Hamming window can be used to perform frame and window processing on the call audio. The window function can be expressed by the following formula:

[0075]

[0076] Where N is the total number of sample points in a single window, n is the sequence number of each sample point in the window, n∈[0,N-1].

[0077] After the call audio is framed and windowed, a discrete Fourier transform is performed on the call audio. The discrete Fourier transform can be expressed by the following formula:

[0078]

[0079] Where i is the frame number; k is the frequency point number, k = 1, 2, 3, ..., N, where N is the total number of sample points in a single window; X(i, k) is the frequency domain spectrum of the call audio, that is, the spectrum of the i-th frame and the k-th frequency point; n is the number value of each sample point in the window; x(n) is the input sample point value; and j is the complex domain of the Fourier transform.

[0080] Next, after performing frame windowing and discrete Fourier transform on the call audio, the power spectrum of the call audio in the frequency domain is obtained, which can be calculated using the following formula:

[0081] S(i,k)=|X(i,k)| 2 (3)

[0082] Where i is the frame number, k is the frequency number, and S(i,k) is the power spectrum of the call audio in the frequency domain, i.e., the power spectrum of the i-th frame and the k-th frequency.

[0083] In one embodiment, the call environment noise of the call audio of the receiving terminal at each frequency point within a specified frequency band is obtained.

[0084] The designated frequency band range may be a main frequency band range of speech, such as 50 to 3400 Hz.

[0085] In a specific embodiment, a noise estimation algorithm can be used to obtain the call environment noise at each frequency point within a specified frequency band based on the power spectrum of the call audio of the receiving terminal within the specified frequency band. The noise estimation algorithm can be MS (Minimum Statistics), MCRA (Minimum Control Recursion Average), IMCRA (Improved Minimum Control Recursion Average, based on Wiener filtering), or other algorithms.

[0086] Step 306 : Determine the masking degree of each call audio according to the call environment noise and the masking threshold of each call audio.

[0087] When the sound pressure level of one audio signal is greater than that of another, the lower-level signal becomes less audible. This phenomenon is called masking. The masking threshold is the maximum sound pressure level of the masked signal. The masking degree indicates the degree to which the call audio is masked by the ambient noise. The higher the masking degree, the greater the degree to which the call audio is masked by the ambient noise, and the less audible the call audio is.

[0088] In one embodiment, step 306 includes: determining more than one sending terminal from the call member terminals participating in a multi-person call; obtaining the call audio sent by each sending terminal; determining each masking threshold based on the call audio of each sending terminal; and determining the masking degree of the call audio of each sending terminal based on the call environment noise and each masking threshold.

[0089] In the embodiment of the present application, it is analyzed whether each call audio sent by the sending terminal is masked by the call environment noise of the receiving terminal.

[0090] In this embodiment, obtaining the masking threshold of the call audio may include the following methods: In one possible implementation method, the server obtains the masking threshold based on the call audio of the sending terminal. The server may simultaneously obtain the call audio sent by the sending terminal and the receiving terminal (or the call environment noise sent by the receiving terminal) according to the above method (1), and then obtain the masking threshold based on the call audio of the sending terminal. The server may also separately obtain the call audio of the sending terminal and the call audio sent by the receiving terminal (or the call environment noise sent by the receiving terminal) according to the above method (2), and then obtain the masking threshold based on the call audio of the sending terminal.

[0091] In one possible implementation, the sending terminal obtains a masking threshold based on the call audio and sends the obtained masking threshold to the server. The server can simultaneously obtain the masking threshold sent by the sending terminal and the call audio sent by the receiving terminal (or the call environment noise sent by the receiving terminal) according to the above method (1), or can separately obtain the masking threshold sent by the sending terminal and the call audio sent by the receiving terminal (or the call environment noise sent by the receiving terminal) according to the above method (2).

[0092] In one embodiment, the method for obtaining the masking threshold of the call audio of the sending terminal includes: performing frequency domain conversion processing on the call audio of the sending terminal to obtain the power spectrum of the call audio of the sending terminal in the frequency domain; and determining the masking threshold of the call audio of the sending terminal based on the power spectrum.

[0093] The method of obtaining the power spectrum is the same as the method of obtaining the power spectrum in step 304, and will not be repeated here.

[0094] In one embodiment, the masking degree of each call audio is determined based on the call environment noise and the masking threshold of each call audio, including: taking the call audio sent by each sending terminal as the current processing audio in turn; obtaining the masking threshold of each frequency point of the current processing audio within a specified frequency band; and determining the masking degree of the current processing audio based on the call environment noise and the masking threshold of each frequency point.

[0095] The designated frequency band range may be a main frequency band range of speech, such as 50 to 3400 Hz.

[0096] It can be understood that taking the call audio sent by each sending terminal as the currently processed audio in sequence means that the server processes the call audio sent by each sending terminal in sequence.

[0097] In a specific embodiment, a psychoacoustic model can be used to obtain a masking threshold for each frequency point in a specified frequency band based on the power spectrum of the call audio of the transmitting terminal within the specified frequency band. The psychoacoustic model can be a Johnston model, a Terhardt model, or the like.

[0098] In an embodiment of the present application, for each frequency point in a specified frequency band, the magnitude relationship between the corresponding masking threshold and the call environment noise is analyzed, and based on the magnitude relationship, it is determined whether the sound signal of the frequency point is masked by the call environment noise.

[0099] In one embodiment, the masking degree of the currently processed audio is determined based on the call environment noise and masking threshold of each frequency point, including: taking each frequency point as the currently processed frequency point in turn; when the call environment noise of the currently processed frequency point is greater than the masking threshold of the currently processed frequency point, marking the currently processed frequency point as the target frequency point; and determining the masking degree of the currently processed audio based on all target frequencies of the currently processed audio.

[0100] Among them, the call environment noise of the current processing frequency is greater than the masking threshold of the current processing frequency, indicating that the sound signal of the current processing frequency is easily masked by the call environment noise of the receiving terminal.

[0101] Specifically, the server processes each frequency point within the specified frequency band in turn, and determines whether the call environment noise of the currently processed frequency point is greater than the masking threshold of the currently processed frequency point. If so, the currently processed frequency point is marked as the target frequency point, and the next frequency point is processed until the processing of each frequency point within the specified frequency band is completed, and the target frequency point can be selected.

[0102] In a specific embodiment, when the call environment noise of the current processing frequency is greater than the masking threshold of the current processing frequency, the ratio of the call environment noise of the current processing frequency to the masking threshold of the current processing frequency is obtained; when the ratio is greater than the preset ratio, the current processing frequency is marked as the target frequency.

[0103] Among them, when the ratio of the call environment noise of the current processing frequency to the masking threshold of the current processing frequency is greater than the preset ratio, it means that the sound signal of the current processing frequency is basically masked by the call environment noise of the receiving terminal, thereby improving the marking accuracy of the target frequency.

[0104] In a specific embodiment, based on all target frequency points of the currently processed audio, the masking degree of the currently processed audio is determined, including: obtaining the sum of the power spectra of each target frequency point; obtaining the sum of the power spectra of all frequency points of the currently processed audio; determining the proportion of the sum of the power spectra of each target frequency point to the sum of the power spectra of all frequency points; and using the proportion as the masking degree of the currently processed audio.

[0105] The power spectrum reflects the energy of the audio signal at each frequency. The audio signal at the target frequency is easily masked by the ambient noise. The ratio of the sum of the power spectra of each target frequency to the sum of the power spectra of all frequencies reflects the degree to which the call audio is masked by the ambient noise. If the audio signals at most frequencies in the call audio are masked by the ambient noise, it indicates that the call participants at the receiving terminal have difficulty hearing the audio clearly and the audio does not need to be forwarded to the receiving terminal.

[0106] In this embodiment, target frequencies that are easily masked are selected based on the call environment noise and masking threshold of each frequency point, and then the masking degree is determined based on the proportion of the sum of the power spectra of the target frequencies to the sum of the power spectra of all frequencies. In this way, the masking degree of the call audio is accurately calculated, and precise routing of the call audio can be achieved subsequently.

[0107] Step 308: The call audio is filtered according to each masking degree and then sent to the receiving terminal.

[0108] In an embodiment of the present application, when the call audio sent by the sending terminal is easily masked by the call environment noise of the receiving terminal, the call audio can be discarded, and the call audio that is not easily masked by the call environment noise of the receiving terminal can be forwarded to the receiving terminal, so that the call members of the receiving terminal can hear the received call audio clearly, and at the same time, the amount of call audio forwarded by the server is reduced, reducing the network bandwidth occupied by the server.

[0109] Specifically, the server filters the call audios sent by multiple call member terminals participating in a multi-person call according to the masking degree, obtains the target call audio, and sends the target call audio to the receiving terminal.

[0110] The number of target audio calls is no greater than the number of audio calls sent by the multiple call member terminals. The server can select at least two target audio calls from the at least three audio calls participating in the multi-person call, thereby reducing the number of audio calls forwarded by the server and thereby reducing network bandwidth usage and data traffic consumption when sending the audio calls to the receiving terminal.

[0111] In this embodiment, filtering the call audio according to the masking degree may include the following methods:

[0112] (1) The server selects target call audio with a masking degree less than a screening threshold from the call audios sent by multiple call member terminals based on the masking degrees.

[0113] Among them, the masking degree is less than the screening threshold, which means that the call audio is less masked by the call environment noise of the receiving terminal, and the call members of the receiving terminal can hear the forwarded call audio clearly.

[0114] The server can traverse the call audio sent by multiple call member terminals, and determine whether the masking degree of the currently traversed call audio is less than the screening threshold. If so, the currently traversed call audio will be used as the target call audio, and the next call audio will be traversed until the call audio sent by the multiple call member terminals is traversed. At least two target call audios can be selected.

[0115] For example, a call group includes terminals A, B, C, and D. If the call participants corresponding to terminals A, B, and C are speaking, but the call participant corresponding to terminal D is silent, the call audio sent by terminals A, B, and C all contain voice data, while the call audio sent by terminal D does not. Terminal C is then used as the receiving terminal, and terminals A, B, and D as the sending terminals. The server selects the call audio sent by terminals A and B as the target call audio from the three call audio streams sent by terminals A, B, and D. It then sends the call audio sent by terminals A and B to terminal C, but does not send the call audio sent by terminal D to terminal C.

[0116] (2) The server selects a preset number of target call audios whose masking degrees are less than a screening threshold from the call audios sent by multiple call member terminals according to the masking degrees.

[0117] The preset number is not greater than the number of call audios sent by the multiple call member terminals. The preset number is an integer greater than 1 and not greater than the number of call audios sent by the multiple call member terminals.

[0118] The server selects call audio with a masking degree less than a screening threshold from the call audio sent by multiple call member terminals, and selects a preset number of target call audios with the smallest masking degree from the selected call audios, thereby ensuring that the selected target call audios are less affected by the call environment noise.

[0119] The server can sort the selected call audios in order of masking degree from small to large, and select a preset number of call audios before sorting as target call audios.

[0120] For example, the preset number is 2, the masking degree of the call audio sent by terminal A is greater than the masking degree of the call audio sent by terminal B, and the masking degree of the call audio sent by terminal B is greater than the masking degree of the call audio sent by terminal C, then the call audios sent by terminal C and terminal B respectively can be selected as the target call audios.

[0121] In one embodiment, step 308 includes: selecting a preset number of target call audios from the call audio and having a masking degree less than a screening threshold; sending the target call audio to the receiving terminal, and the sent target call audio is used to instruct the receiving terminal to decode the target call audio and then mix and play it; or, performing mixing processing on the target call audio after decoding, and re-encoding the mixed target call audio and sending it to the receiving terminal, and the sent target call audio is used to instruct the receiving terminal to decode the target call audio and then play it.

[0122] The preset number is the number of call audio tracks to be filtered. For example, three call audio tracks may be filtered for mixed playback. The screening threshold is a critical parameter for the degree of masking of the call audio during filtering. In this embodiment, a call audio track will only be selected as a target call audio track if its masking degree is less than the screening threshold.

[0123] Specifically, the server selects candidate audio tracks with a masking degree less than a screening threshold from the audio tracks sent by multiple call participants. It then selects a preset number of target audio tracks from the candidate audio tracks. The server then sends these target audio tracks to the receiving terminal, which decodes them and plays them after mixing the decoded target audio tracks. This reduces the impact of ambient noise on the call.

[0124] The server can also first decode the target call audio, then mix the decoded target call audio and re-encode it. The re-encoded call audio is then sent to the terminal, which can directly decode and play it. This can not only reduce the impact of environmental noise on the call process, but also reduce the workload of the terminal and reduce the impact of the terminal's background data processing on the call process.

[0125] In one embodiment, step 308 includes: selecting a first target call audio whose masking degree is less than a first threshold from the call audio; for candidate call audio whose masking degree is greater than or equal to the screening threshold, when the importance of the candidate call audio is higher than the second threshold, enhancing the candidate call audio to obtain a second target call audio; sending the first target call audio and the second target call audio to the receiving terminal for decoding and mixing for playback; or, mixing the first target call audio and the second target call audio after decoding, and re-encoding the mixed call audio and sending it to the receiving terminal for decoding and playback.

[0126] The importance level indicates the importance of the call audio content. A higher importance level indicates more important call audio content. It is understood that candidate call audio with a masking level greater than or equal to the first threshold can be filtered out to conserve network resources, as it is easily masked by background noise. However, if the candidate call audio is relatively important, filtering it out may cause the receiving terminal to miss important information.

[0127] Specifically, for candidate call audio with a masking degree greater than or equal to the first threshold, the server can evaluate the importance of the candidate call audio and determine whether the importance of the candidate call audio is higher than the second threshold. When the importance of the candidate call audio is higher than the second threshold, it is considered that the candidate call audio is relatively important and should not be filtered out, but it is easily masked by background noise. At this time, the server can enhance the candidate call audio to obtain the second target call audio, and send the first target call audio and the second target call audio to the receiving terminal for mixing and playing after decoding at the receiving terminal; or, mix the first target call audio and the second target call audio after decoding, and re-encode the mixed call audio and send it to the receiving terminal for direct decoding and playing at the receiving terminal.

[0128] The importance of call audio can be assessed through natural language processing or other processing. Natural language processing can be performed using a trained natural language model. Specifically, the natural language model can be an end-to-end model, with audio data or text data converted from audio data as input and importance as output. The natural language model can be a multi-layer network structure, with different network layers performing different processing on the input data and outputting the processing results to the next network layer.

[0129] In one embodiment, the candidate call audio may be enhanced at the volume level. A corresponding relationship may be established between the enhancement intensity and the masking degree of the candidate call audio. For example, if the masking degree of the candidate call audio is high, the enhancement intensity is high.

[0130] For example, let's assume terminal F is the receiving terminal, and terminals A, B, C, D, and E are the sending terminals. From the five audio calls sent by terminals A, B, C, D, and E, the server selects the audio calls sent by terminals A and B whose masking levels are greater than or equal to the screening threshold. Upon analysis, the server finds that the importance of the audio calls sent by terminal A is greater than a second threshold, while the importance of the audio calls sent by terminal B is less than the second threshold. The server then enhances the audio calls sent by terminal A and sends the enhanced audio calls to terminal F, allowing terminal F to clearly hear the audio calls sent by terminal A.

[0131] In this embodiment, when the masking degree of the call audio is greater than or equal to the screening threshold, if the call audio is important, the call audio is enhanced to prevent the receiving terminal from missing important information.

[0132] In one embodiment, step 302 includes: obtaining call audio sent by each call member terminal participating in a multi-person call; preliminarily screening out multiple call audios from the call audio based on audio features of each call audio; and obtaining the preliminarily screened multiple call audios.

[0133] Step 308 includes: performing a secondary screening on the plurality of call audios according to the masking degrees; and sending the secondary screened call audios to a receiving terminal.

[0134] Audio features can include VAD (Voice Activity Detection) information, audio energy, and ambient noise. VAD information indicates whether the call audio contains voice data. Audio energy indicates the presence and volume of sound in the call audio. For example, if the call audio contains only voice data, the louder the voice volume, the higher the audio energy.

[0135] Considering the human ear's limited ability to discern mixed signals from different sound sources simultaneously, for example, the human ear can typically only discern four or fewer sound signals at a time. When four or more sound signals are simultaneously present, the human ear becomes unable to discern the mixed sound signals. Therefore, audio features can be used to initially filter out multiple audio streams from the audio streams sent by multiple call participants. These multiple audio streams are then secondary filtered based on masking, selecting target audio streams with masking below a screening threshold. This reduces the amount of audio forwarded by the server.

[0136] In a specific embodiment, after the terminal collects the call audio, it is encoded and packaged and then sent to the server. The server decodes the received audio encoding data, thereby obtaining the call audio sent by multiple terminals, and extracts features from the call audio to obtain audio features.

[0137] Alternatively, after collecting call audio, the terminal can extract audio features from the audio and send them along with the encoded audio data to the server. This approach uses distributed processing, with audio features processed at the sending terminal and routed on the server, saving computing resources and reducing network bandwidth usage.

[0138] In a specific embodiment, an audio routing strategy is adopted to preliminarily filter out multiple call audios from the call audios sent by multiple call member terminals based on audio features. The audio routing strategy may be an audio routing algorithm.

[0139] In a specific embodiment, the server preliminarily selects at least two channels of call audio from at least three channels of call audio based on audio features, and the number of at least two channels of call audio is no greater than the number of at least three channels of call audio. The server may select at least two channels of call audio containing voice data from the at least three channels of call audio based on VAD information of the at least three channels of call audio. Alternatively, the server may select a preset number of call audios with the highest audio energy from the at least three channels of call audio based on VAD information and audio energy of the at least three channels of call audio. Alternatively, the server may select a preset number of call audios with the highest audio energy and the lowest ambient noise from the at least three channels of call audio based on VAD information, audio energy, and ambient noise of the at least three channels of call audio. The preset number is smaller than the number of at least three channels of call audio, and the preset number is an integer greater than 1 and less than the number of at least three channels of call audio.

[0140] In this embodiment, the server preliminarily screens out multiple call audios from the call audios sent by multiple call member terminals through audio features, and then performs a secondary screening on the multiple call audios according to the masking degree. In this way, the secondary screening of the multiple call audios according to the masking degree can eliminate the call audios masked by the call environment noise of the receiving terminal, ensuring that the call members of the receiving terminal can clearly hear the forwarded call audios; and the preliminarily screening and secondary screening of the call audios sent by multiple call member terminals reduces the number of call audios forwarded by the server, and can reduce the occupied network bandwidth and the consumed data traffic when sending call audios to the receiving terminal, thereby improving the call quality during multi-person calls.

[0141] Based on the above embodiment, filtering call audio according to each masking degree may also include the following methods:

[0142] (3) The server preliminarily screens multiple call audios from the call audios sent by multiple call member terminals based on the audio features; performs secondary screening on the multiple call audios based on each masking degree, and selects target call audios whose masking degree is less than the screening threshold.

[0143] The server can traverse the call audios sent by multiple call member terminals and determine whether the audio features of the currently traversed call audio meet the selection conditions. If so, the currently traversed call audio is selected and the next call audio is traversed until the traversal of the call audios sent by the multiple call member terminals is completed. In this way, multiple call audios can be selected. Afterwards, the server performs a secondary screening process on the multiple call audios according to the masking degrees, which is similar to the above method (1).

[0144] The selection conditions may include: containing voice data; or containing voice data and having high audio energy; or containing voice data, having high audio energy and having little ambient noise, etc.

[0145] For example, with terminal F as the receiving terminal and terminals A, B, C, D, and E as the sending terminals, the server preliminarily selects the call audio sent by terminal A, B, C, D, and E from the five-way call audio sent by terminals A, B, C, D, and E based on the audio features; the server secondarily selects the call audio sent by terminal A and terminal B as the target call audio from the four-way call audio sent by terminals A, B, C, and D based on the masking degree, and subsequently sends the call audio sent by terminals A and B to terminal F, while the call audio sent by terminals C, D, and E will not be sent to terminal F.

[0146] (4) The server preliminarily screens a preset number of call audios from the call audios sent by multiple call member terminals based on the audio features; performs a secondary screening on the preset number of call audios based on each masking degree, and selects target call audios whose masking degree is less than the screening threshold.

[0147] The server selects a preset number of call audios from the call audios sent by multiple call member terminals, and the audio features of the call audios meet the selection conditions. The server can sort the selected call audios according to the audio features, and select the preset number of call audios before sorting. For example, the call audio containing voice data is sorted before the call audio not containing voice data, and among the call audios containing voice data, the call audio with higher audio energy is sorted before the call audio with lower audio energy. Afterwards, the process of the server performing a secondary screening of the multiple call audios according to each masking degree is similar to the above method (1).

[0148] For example, if the preset number is 4, and the call audio sent by terminals A, B, C, D, and E all contain voice data, and the audio energy is ranked from high to low as follows: terminal A, terminal B, terminal C, terminal D, and terminal E, then the call audio sent by terminals A, B, C, and D will be selected. The server then selects the call audio sent by terminals A and B as the target call audio from the four call audio channels sent by terminals A, B, C, and D based on the masking degree. Subsequently, the call audio sent by terminals A and B will be sent to terminal F.

[0149] (5) The server preliminarily screens multiple call audios from the call audios sent by multiple call member terminals based on the audio features; performs secondary screening on the multiple call audios based on each masking degree, and selects a preset number of target call audios whose masking degrees are less than the screening threshold.

[0150] The process of the server preliminarily screening multiple call audios from the call audios sent by multiple call member terminals based on audio features is similar to the above method (3). The process of the server performing secondary screening of the multiple call audios based on each masking degree is similar to the above method (2).

[0151] For example, with terminal F as the receiving terminal and terminals A, B, C, D and E as the sending terminals, the server preliminarily selects the call audios sent by terminal A, B, C, D and E from the five call audios sent by terminals A, B, C, D and E based on the audio features; the preset number is 2, the masking degree of the call audio sent by terminal A is greater than the masking degree of the call audio sent by terminal B, and the masking degree of the call audio sent by terminal B is greater than the masking degree of the call audio sent by terminal C, then the call audios sent by terminals C and B respectively can be selected as the target call audios.

[0152] In one embodiment, the server mixes at least two channels of target call audio, and sends the mixed target call audio to the receiving terminal, which decodes and plays the mixed target call audio.

[0153] Alternatively, the receiving terminal receives at least two channels of target call audio sent by the server, decodes the at least two channels of target call audio, mixes the decoded at least two channels of target call audio, and plays the mixed target call audio. In this way, distributed processing is adopted, the call audio is routed in the server, and the call audio is mixed in the receiving terminal, which saves computing resources and reduces network bandwidth usage.

[0154] In the above-mentioned call audio processing method, call audio sent by multiple call member terminals participating in a multi-person call is obtained, one of the call member terminals participating in the multi-person call is selected as the receiving terminal, the call environment noise of the receiving terminal is obtained, the masking degree of each call audio is determined according to the call environment noise and the masking threshold of each call audio, and the call audio is filtered according to each masking degree and then sent to the receiving terminal. In this way, the call audio sent by multiple terminals participating in the multi-person call is filtered according to the masking degree, and the call audio that is easily masked by the call environment noise of the receiving terminal can be eliminated, so that the call members of the receiving terminal can hear the received call audio clearly; and the amount of call audio forwarded by the server is reduced, and the network bandwidth occupied by the server is reduced, thereby improving the call quality during the multi-person call process.

[0155] In addition, this embodiment can adopt distributed processing, processing the call audio at the sending terminal to obtain the masking threshold, processing the call audio at the receiving terminal to obtain the call environment noise, and routing the call audio in the server. The distributed processing of the call audio by the sending terminal, server and receiving terminal saves computing resources and reduces network bandwidth usage.

[0156] It can be understood that the unmasking degree of each call audio can also be determined based on the call environment noise and the masking threshold of each call audio. The unmasking degree indicates the degree to which the call audio is not masked by the call environment noise. The call audio is filtered according to each unmasking degree and then sent to the receiving terminal.

[0157] In one embodiment, the unmasking degree of the currently processed audio is determined based on the call environment noise and masking threshold of each frequency point, including: taking each frequency point as the currently processed frequency point in turn; when the call environment noise of the currently processed frequency point is less than or equal to the masking threshold of the currently processed frequency point, marking the currently processed frequency point as a reference frequency point; and determining the unmasking degree of the currently processed audio based on all reference frequencies of the currently processed audio.

[0158] Among them, the call environment noise of the current processing frequency is less than or equal to the masking threshold of the current processing frequency, indicating that the sound signal of the current processing frequency is not easily masked by the call environment noise of the receiving terminal.

[0159] Specifically, the server processes each frequency point within the specified frequency band in turn, and determines whether the call environment noise of the currently processed frequency point is less than or equal to the masking threshold of the currently processed frequency point. If so, the currently processed frequency point is marked as the reference frequency point, and the next frequency point is processed until the processing of each frequency point within the specified frequency band is completed, and the reference frequency point can be selected.

[0160] In a specific embodiment, based on all reference frequency points of the currently processed audio, the unmasking degree of the currently processed audio is determined, including: obtaining the sum of the power spectra of each reference frequency point; obtaining the sum of the power spectra of all frequency points of the currently processed audio; determining the ratio of the sum of the power spectra of each reference frequency point to the sum of the power spectra of all frequency points; and using the ratio as the unmasking degree of the currently processed audio.

[0161] The power spectrum reflects the energy of the audio signal at each frequency. The audio signal at the reference frequency is not easily masked by the call environment noise. The ratio of the sum of the power spectra of each reference frequency to the sum of the power spectra of all frequencies reflects the degree to which the call audio is not masked by the call environment noise. If the audio signals at most frequencies in the call audio are not masked by the call environment noise, it means that the call participants at the receiving terminal can hear the call clearly and the call audio can be forwarded to the receiving terminal.

[0162] In one embodiment, the call audio is filtered according to each unmasking degree and then sent to the receiving terminal, including: according to each masking degree, selecting the target call audio with an unmasking degree greater than or equal to the filtering threshold from the call audio sent by multiple call member terminals.

[0163] In this embodiment, the unmasking degree of the call audio is accurately calculated, and accurate routing of the call audio is achieved based on the unmasking degree.

[0164] In one embodiment, the call environment noise of each frequency point of the currently processed audio within a specified frequency band is obtained, including: taking each frequency point as the currently processed frequency point in turn; performing time-frequency domain smoothing on the power spectrum of the currently processed frequency point to obtain a smoothed power spectrum; performing a minimum value search on the smoothed power spectrum through a window function to obtain a local minimum value; determining the probability of speech existence based on the smoothed power spectrum and the local minimum value; and determining the call environment noise of the currently processed frequency point based on the probability of speech existence.

[0165] The call environment noise can be calculated using the following formula:

[0166]

[0167] in, is the call environment noise at the i-th frame and the k-th frequency point; is the probability of speech existence at the kth frequency point in the i-th frame; is the call environment noise at the i-1th frame and the kth frequency point; S(i,k) is the power spectrum of the call audio in the frequency domain.

[0168] Formula (4) shows that to calculate the call environment noise, we need to first calculate the probability of speech presence at the current processing frequency point, and the probability of speech presence is determined by the power spectrum and local minimum of the current processing frequency point. Therefore, in order to obtain the local minimum of the current processing frequency point, we first need to perform time-frequency domain smoothing on the power spectrum of the current processing frequency point, which can be expressed by the following formula:

[0169] Perform frequency domain smoothing on the power spectrum of the current processing frequency point:

[0170]

[0171] Wherein, b(j+w) is the frequency domain smoothing weighting factor group.

[0172] The power spectrum of the current processing frequency point is smoothed in the time domain using a first-order recursive average:

[0173]

[0174] in, is the smoothed power spectrum of the i-th frame and the k-th frequency point; α0 is the time domain smoothing factor.

[0175] Next, the minimum tracking method is used to search for the local minimum, that is, the minimum value of the smoothed power spectrum is searched through the window function in the specified window length L frame, which can be expressed by the following formula:

[0176]

[0177]

[0178] Among them, S min (i,k) is the local minimum value of the i-th frame and the k-th frequency point; S tmp (ik) is the temporary value of the i-th frame and the k-th frequency point.

[0179] Every L frames, the temporary value S tmp (ik) Update as follows:

[0180]

[0181]

[0182] Next, the probability of speech presence is determined based on the smoothed power spectrum and the local minimum. First, the ratio between the smoothed power spectrum and the local minimum is determined, which can be expressed by the following formula:

[0183]

[0184] The probability of the existence of instant speech is determined based on the relationship between the ratio and the preset ratio, which can be expressed by the following formula:

[0185]

[0186] Here, p(i,k) is the probability of the existence of instant speech at the i-th frame and the k-th frequency point. p(i,k)=1 means that there is speech data at the i-th frame and the k-th frequency point, and p(i,k)=0 means that there is no speech data at the i-th frame and the k-th frequency point.

[0187] The probability of speech presence is determined based on the instant speech presence probability, which can be calculated using the following formula:

[0188]

[0189] in, is the probability of speech existence at the kth frequency point in the i-th frame; α p is the smoothing coefficient; is the probability of speech existence in the i-1th frame and the kth frequency point; p(i,k) is the probability of instant speech existence in the i-1th frame and the kth frequency point.

[0190] Next, the call environment noise at the current processing frequency is calculated based on the speech existence probability and the above formula (4).

[0191] In this embodiment, the call environment noise of each frequency point of the call audio within the specified frequency band is obtained, and each frequency point of the specified frequency band is subsequently analyzed to improve the accuracy of the audio signal analysis.

[0192] In one embodiment, obtaining a masking threshold value of each frequency point of the currently processed audio within a specified frequency band includes: taking each frequency point as the currently processed frequency point in turn; determining the critical frequency band to which the currently processed frequency point belongs; obtaining a global masking threshold value of the critical frequency band; and determining a masking threshold value of the currently processed frequency point based on the global masking threshold value.

[0193] Among them, in the spectrum of audio signals, the human ear perceives audio signals of different frequencies differently. In order to uniformly measure sound frequencies from the perspective of human ear perception, critical frequency bands are introduced, and sound frequencies with the same degree of human ear perception are classified into the same critical frequency band. Figure 4 , Figure 4 FIG. 4 is a schematic diagram of critical frequency bands in an embodiment, which can be divided into 24 critical frequency bands.

[0194] First, determine the critical frequency band to which the current processing frequency belongs, that is, determine the critical frequency band number corresponding to the current processing frequency, which can be calculated using the following formula:

[0195] z(f)=13*arctan(0.76*f)+3.5*arctan(f / 7.5) 2 (14)

[0196] Where f is the frequency point; z(f) is the critical frequency band number corresponding to the frequency point f.

[0197] Next, a global masking threshold of the critical band is obtained, including: obtaining a critical band power spectrum of the critical band; expanding the critical band power spectrum by an expansion function to obtain an expanded power spectrum; and determining the global masking threshold according to the expanded power spectrum.

[0198] Specifically, first obtain the critical band power spectrum of the critical band. The critical band power spectrum is determined by the sum of the power spectrum values ​​of each frequency point in the critical band and can be calculated using the following formula:

[0199]

[0200] Where B(i,z) is the critical band power spectrum of the i-th frame and the z-th critical band; b2(m) and b1(m) are the frequency range limits of the critical band; and P(i,l) is the power spectrum value of the i-th frame and the l-th frequency point.

[0201] The human ear has different perception abilities for different critical frequency bands, resulting in inter-band masking effects. The strength of this effect is related to the distance between the critical frequency bands; generally speaking, the greater the distance, the smaller the effect. To account for the inter-band influence, a spread function is used to expand the critical frequency band power spectrum. The spread function is related to the distance between the critical frequency bands and can be expressed as follows:

[0202]

[0203] Here, δz is the distance between bands. For example, δz = mn, where m and n are the critical band numbers.

[0204] The critical band power spectrum is expanded by the expansion function to obtain the expanded power spectrum, which can be expressed by the following formula:

[0205] C(i,z)=B(i,z)×SF(δz) (17)

[0206] Where C(i,z) is the extended power spectrum of the i-th frame and the z-th critical band; SF(δz) is the expansion function.

[0207] The global masking threshold under the extended power spectrum can be calculated using the following formula:

[0208]

[0209] The global masking threshold obtained from formula (18) is for the extended power spectrum. It is necessary to normalize the result and reconvert it to the critical band. For example, T(i,z) can be divided by the power gain and then frequency-domain expanded to obtain the global masking threshold for the critical band.

[0210] However, the absolute hearing threshold of the critical frequency band also needs to be considered. The absolute hearing threshold is the statistically determined audible sound intensity of pure speech at a specific frequency in the absence of ambient noise. Frequencies in the sound signal with sound intensity below the absolute hearing threshold are imperceptible to the human ear.

[0211] The absolute hearing threshold of the critical band can be calculated using the following formula:

[0212] T abs (z)=3.64*(btof(z)) -0.8 -6.5exp((btof(z))-3.3) 2 +10 -3 (btof(z)) 4 (19)

[0213] Where btof(z) is the center frequency corresponding to the zth critical frequency band, which can be obtained by Figure 4 Obtained by looking up the table.

[0214] The maximum value of the calculated absolute hearing threshold of the critical band and the absolute hearing threshold is used as the final global masking threshold of the critical band, which can be expressed by the following formula:

[0215] T'(i,z)=max(T(i,z),T abs (z)) (20)

[0216] Next, the masking threshold of the current processing frequency point is determined based on the global masking threshold. It can be calculated using the following formula:

[0217] P mask (i,f)=10 0.05*(T ' (i,z(f))) (twenty one)

[0218] Among them, P mask (i,f) is the masking threshold of the i-th frame and the f-th frequency point; T'(i,z(f) is the global masking threshold of the i-th frame and the z(f)-th critical frequency band.

[0219] In this embodiment, the masking threshold of each frequency point of the call audio within the specified frequency band is obtained, and each frequency point of the specified frequency band is subsequently analyzed to improve the accuracy of the audio signal analysis.

[0220] The present application also provides an application scenario, which applies the above-mentioned call audio processing method. The application scenario is as follows: at least three terminals join the same call group, and the at least three terminals can make calls. The call group can be a voice call group or a video call group, that is, at least three terminals can make voice calls or video calls. The voice call or video call can be a VOIP (Voice over Internet Protocol) multi-person conference scenario or other scenarios. During a call, the terminal that sends voice data is the sending terminal, and the terminal that receives voice data sent by other terminals is the receiving terminal. Each terminal in the call group can be either a sending terminal or a receiving terminal.

[0221] In one possible implementation, Figure 5 As shown, the application of the call audio processing method in this application scenario is as follows:

[0222] Step 502: Acquire call audio sent by multiple call member terminals participating in a multi-person call.

[0223] Step 504 : Perform frequency domain conversion processing on each call audio to obtain a power spectrum of each call audio in the frequency domain.

[0224] In step 506, one of the call member terminals participating in the multi-person call is selected as the receiving terminal, and the call environment noise at each frequency point of the call audio of the receiving terminal within the specified frequency band is determined based on the power spectrum of the call audio sent by the receiving terminal.

[0225] In step 508, more than one sending terminal is determined from the terminals of the call members participating in the multi-person call, and the call audio sent by each sending terminal is used as the current processing audio in turn. Based on the power spectrum of the current processing audio, the masking threshold of each frequency point of the current processing audio within the specified frequency band is determined.

[0226] In step 510, each frequency point is taken as the current processing frequency point in turn. When the call environment noise of the current processing frequency point is greater than the masking threshold of the current processing frequency point, the ratio of the call environment noise of the current processing frequency point to the masking threshold of the current processing frequency point is obtained. When the ratio is greater than the preset ratio, the current processing frequency point is marked as the target frequency point.

[0227] Step 512: Obtain the sum of the power spectrum values ​​of each target frequency point, obtain the sum of the power spectrum values ​​of all frequency points of the currently processed audio, determine the ratio of the sum of the power spectrum values ​​of each target frequency point to the sum of the power spectrum values ​​of all frequency points, and use the ratio as the masking degree of the currently processed audio.

[0228] Step 514 : Select a preset number of target call audios with a masking degree less than a screening threshold from the call audios sent by the sending terminal, and send the target call audios to the receiving terminal.

[0229] Among them, the target call audio is used to instruct the receiving terminal to decode the target call audio and then mix and play it; or, the target call audio is decoded and then mixed, and the mixed target call audio is re-encoded and sent to the receiving terminal. The sent target call audio is used to instruct the receiving terminal to decode the target call audio and then play it.

[0230] Specifically, refer to Figure 7 , Figure 7 This is a flowchart of a multi-person call process in one embodiment. As can be seen, the server obtains a masking threshold based on the call audio sent by the sending terminal, obtains the call environment noise based on the call audio sent by the receiving terminal, performs psychoacoustic masking analysis on the masking threshold and the call environment noise, selects target call audio from the call audio sent by the sending terminal, and sends the target call audio to the receiving terminal. The receiving terminal decodes the target call audio and mixes it for playback.

[0231] In this embodiment, the call audio sent by multiple terminals participating in a multi-person call is filtered according to the masking degree, and the call audio that is easily masked by the call environment noise of the receiving terminal can be eliminated, so that the call members of the receiving terminal can hear the received call audio clearly; and the amount of call audio forwarded by the server is reduced, and the network bandwidth occupied by the server is reduced, thereby improving the call quality during the multi-person call process.

[0232] In one possible implementation, Figure 6 As shown, the application of the call audio processing method in this application scenario is as follows:

[0233] Step 602: obtain the call audio sent by each sending terminal participating in a multi-person call, and preliminarily screen out multiple call audios from the call audio sent by the sending terminal based on the audio features of the call audio sent by each sending terminal, and use the preliminarily screened multiple call audios as the current processing audio in turn.

[0234] Step 604 : Perform frequency domain conversion on the currently processed audio to obtain a power spectrum of the currently processed audio in the frequency domain. Based on the power spectrum of the currently processed audio, determine the masking threshold of each frequency point of the currently processed audio within the specified frequency band.

[0235] In step 606, one of the call member terminals participating in the multi-person call is selected as the receiving terminal, the call audio sent by the receiving terminal is obtained, and the call audio of the receiving terminal is converted into a frequency domain to obtain the power spectrum of the call audio of the receiving terminal in the frequency domain.

[0236] Step 608 : Determine the call environment noise at each frequency point of the call audio of the receiving terminal within a specified frequency band according to the power spectrum of the call audio of the receiving terminal in the frequency domain.

[0237] In step 610, each frequency point is taken as the current processing frequency point in turn. When the call environment noise of the current processing frequency point is greater than the masking threshold of the current processing frequency point, the ratio of the call environment noise of the current processing frequency point to the masking threshold of the current processing frequency point is obtained. When the ratio is greater than the preset ratio, the current processing frequency point is marked as the target frequency point.

[0238] Step 612: Obtain the sum of the power spectrum values ​​of each target frequency point, obtain the sum of the power spectrum values ​​of all frequency points of the currently processed audio, determine the ratio of the sum of the power spectrum values ​​of each target frequency point to the sum of the power spectrum values ​​of all frequency points, and use the ratio as the masking degree of the currently processed audio.

[0239] Step 614 , a preset number of target call audios with a masking degree less than a screening threshold are selected from the multiple call audios initially screened, and the target call audios are sent to a receiving terminal.

[0240] Among them, the target call audio is used to instruct the receiving terminal to decode the target call audio and then mix and play it; or, the target call audio is decoded and then mixed, and the mixed target call audio is re-encoded and sent to the receiving terminal. The sent target call audio is used to instruct the receiving terminal to decode the target call audio and then play it.

[0241] Specifically, refer to Figure 8 , Figure 8 This is a flowchart of a multi-person call process in another embodiment. As can be seen, the server extracts audio features from the call audio sent by the sending terminal and preliminarily selects multiple call audio tracks based on the audio features. It then determines a masking threshold based on the multiple call audio tracks, obtains the call environment noise based on the call audio sent by the receiving terminal, performs psychoacoustic masking analysis on the masking threshold and the call environment noise, and then reselects a target call audio track from the preliminarily selected multiple call audio tracks. The target call audio track is then sent to the receiving terminal, which decodes the target call audio track and mixes it for playback.

[0242] In this embodiment, the server preliminarily screens out multiple call audios from the call audios sent by multiple call member terminals through audio features, and then performs a secondary screening on the multiple call audios according to the masking degree. In this way, the secondary screening of the multiple call audios according to the masking degree can eliminate the call audios masked by the call environment noise of the receiving terminal, ensuring that the call members of the receiving terminal can clearly hear the forwarded call audios; and the preliminarily screening and secondary screening of the call audios sent by multiple call member terminals reduces the number of call audios forwarded by the server, and can reduce the occupied network bandwidth and the consumed data traffic when sending call audios to the receiving terminal, thereby improving the call quality during multi-person calls.

[0243] It should be understood that although Figure 3 、 Figure 5 、 Figure 6 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 3 、 Figure 5 、 Figure 6 At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The order of execution of these steps or stages is not necessarily one by one, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.

[0244] In one embodiment, Figure 9As shown, a call audio processing device is provided. The device can be a software module or a hardware module, or a combination of the two to form a part of a computer device. The device specifically includes: an acquisition module 902, a determination module 904 and a screening module 906, wherein:

[0245] An acquisition module 902 is configured to acquire call audio sent by multiple call member terminals participating in a multi-person call;

[0246] The acquisition module 902 is further configured to select one of the call member terminals participating in the multi-person call as a receiving terminal and obtain the call environment noise of the receiving terminal;

[0247] Determination module 904, configured to determine the masking degree of each call audio according to the call environment noise and the masking threshold of each call audio; the masking degree indicates the degree to which the call audio is masked by the call environment noise;

[0248] The screening module 906 is configured to screen the call audio according to the masking degrees and then send the filtered audio to the receiving terminal.

[0249] In one embodiment, the acquisition module 902 is also used to: select one of the call member terminals participating in a multi-person call as a receiving terminal, and determine the call audio sent by the receiving terminal; perform frequency domain conversion processing on the call audio to obtain the power spectrum of the call audio in the frequency domain; and determine the call environment noise of the receiving terminal based on the power spectrum.

[0250] In one embodiment, the determination module 904 is also used to: take each call audio as the currently processed audio in turn; obtain the masking threshold of each frequency point of the currently processed audio within the specified frequency band; obtain the call environment noise of each frequency point of the currently processed audio within the specified frequency band; and determine the masking degree of the currently processed audio based on the call environment noise and the masking threshold of each frequency point.

[0251] In one embodiment, the determination module 904 is further used to: take each frequency point as the current processing frequency point in turn; mark the current processing frequency point as the target frequency point when the call environment noise of the current processing frequency point is greater than the masking threshold of the current processing frequency point; and determine the masking degree of the current processing audio based on all target frequencies of the currently processed audio.

[0252] In one embodiment, the determination module 904 is further used to: obtain the sum of the power spectrum values ​​of each target frequency point; obtain the sum of the power spectrum values ​​of all frequency points of the currently processed audio; determine the proportion of the sum of the power spectrum values ​​of each target frequency point to the sum of the power spectrum values ​​of all frequency points; and use the proportion as the masking degree of the currently processed audio.

[0253] In one embodiment, the determination module 904 is further used to: obtain the ratio of the call environment noise of the current processing frequency point to the masking threshold of the current processing frequency point when the call environment noise of the current processing frequency point is greater than the masking threshold of the current processing frequency point; when the ratio is greater than the preset ratio, mark the current processing frequency point as the target frequency point.

[0254] In one embodiment, the determination module 904 is further used to: take each frequency point as the current processing frequency point in turn; determine the critical frequency band to which the current processing frequency point belongs; obtain the global masking threshold of the critical frequency band; and determine the masking threshold of the current processing frequency point based on the global masking threshold.

[0255] In one embodiment, the determination module 904 is further configured to: obtain a critical band power spectrum of the critical band; expand the critical band power spectrum using an expansion function to obtain an expanded power spectrum; and determine a global masking threshold according to the expanded power spectrum.

[0256] In one embodiment, the determination module 904 is also used to: take each frequency point as the current processing frequency point in turn; perform time-frequency domain smoothing on the power spectrum of the current processing frequency point to obtain a smoothed power spectrum; perform minimum value search on the smoothed power spectrum through a window function to obtain a local minimum value; determine the probability of speech existence based on the smoothed power spectrum and the local minimum value; and determine the call environment noise of the current processing frequency point based on the probability of speech existence.

[0257] In one embodiment, the determination module 904 is further configured to: determine a ratio between a smoothed power spectrum and a local minimum; determine an immediate speech presence probability based on a magnitude relationship between the ratio and a preset ratio; and determine a speech presence probability based on the immediate speech presence probability.

[0258] In one embodiment, the screening module 906 is also used to: select a preset number of target call audios with a masking degree less than a screening threshold from the call audio; send the target call audio to the receiving terminal, and the sent target call audio is used to instruct the receiving terminal to decode the target call audio and then mix it for playback; or, perform mixing processing on the target call audio after decoding, and re-encode the mixed target call audio and send it to the receiving terminal, and the sent target call audio is used to instruct the receiving terminal to decode the target call audio and then play it.

[0259] In one embodiment, the acquisition module 902 is further used to: obtain the call audio sent by the terminal of each call member participating in a multi-person call; preliminarily screen out multiple call audios from the call audio according to the audio characteristics of each call audio; obtain the multiple call audios preliminarily screened out; the screening module 906 is further used to: perform secondary screening on the multiple call audios according to each masking degree; and send the secondary screened call audios to the receiving terminal.

[0260] In one embodiment, the screening module 906 is also used to: select a first target call audio with a masking degree less than a first threshold from the call audio; for a candidate call audio with a masking degree greater than or equal to the screening threshold, when the importance of the candidate call audio is higher than a second threshold, enhance the candidate call audio to obtain a second target call audio; send the first target call audio and the second target call audio to the receiving terminal for decoding and mixing for playback; or, perform mixing processing on the first target call audio and the second target call audio after decoding, and re-encode the mixed call audio and send it to the receiving terminal for decoding and playback.

[0261] In the above-mentioned call audio processing device, call audio sent by multiple call member terminals participating in a multi-person call is obtained, one of the call member terminals participating in the multi-person call is selected as the receiving terminal, the call environment noise of the receiving terminal is obtained, the masking degree of each call audio is determined according to the call environment noise and the masking threshold of each call audio, and the call audio is filtered according to each masking degree and then sent to the receiving terminal. In this way, the call audio sent by multiple terminals participating in the multi-person call is filtered according to the masking degree, and the call audio that is easily masked by the call environment noise of the receiving terminal can be eliminated, so that the call members of the receiving terminal can hear the received call audio clearly; and the amount of call audio forwarded by the server is reduced, and the network bandwidth occupied by the server is reduced, thereby improving the call quality during the multi-person call process.

[0262] In addition, this embodiment can adopt distributed processing, processing the call audio at the sending terminal to obtain the masking threshold, processing the call audio at the receiving terminal to obtain the call environment noise, and routing the call audio in the server. The distributed processing of the call audio by the sending terminal, server and receiving terminal saves computing resources and reduces network bandwidth usage.

[0263] For the specific definition of the call audio processing device, please refer to the definition of the call audio processing method above and will not be repeated here. Each module in the above-mentioned call audio processing device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.

[0264] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 10As shown. The computer device includes a processor, a memory and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store call audio processing data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a call audio processing method is implemented.

[0265] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 11 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a call audio processing method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0266] For example, this embodiment can adopt distributed processing, process the call audio at the sending terminal to obtain the masking threshold, process the call audio at the receiving terminal to obtain the call environment noise, and perform routing processing on the call audio in the server. The distributed processing of the call audio by the sending terminal, server and receiving terminal saves computing resources and reduces network bandwidth usage.

[0267] Those skilled in the art will understand that Figure 10 、 Figure 11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0268] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0269] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.

[0270] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of each of the above-described method embodiments.

[0271] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0272] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0273] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A call audio processing method, characterized in that: The method comprises: Obtain call audio sent by multiple call member terminals participating in a multi-person call; Selecting one of the call member terminals participating in the multi-person call as a receiving terminal, and obtaining the call environment noise of the receiving terminal based on the call audio of the receiving terminal; Determining more than one sending terminal from the terminals of the call members participating in the multi-person call, and preliminarily screening a plurality of call audios from the call audios sent by each of the sending terminals based on audio features of the call audios sent by each of the sending terminals, wherein the plurality of call audios are a preset number of call audios with the maximum audio energy and the least ambient noise among the call audios containing voice data; Determining a masking degree of each of the preliminarily screened call audios based on the call environment noise and a masking threshold of each of the preliminarily screened call audios; the masking degree indicates the degree to which the call audio is masked by the call environment noise; Selecting a preset number of first target call audios whose masking degree is less than a screening threshold from the multiple call audios preliminarily screened; For candidate call audio whose masking degree is greater than or equal to the screening threshold, when the importance of the candidate call audio is higher than the second threshold, the candidate call audio is enhanced to obtain the second target call audio according to the correspondence between the enhancement processing intensity and the masking degree of the candidate call audio, and the first target call audio and the second target call audio are sent to the receiving terminal.

2. The method according to claim 1, characterized in that The selecting one of the call member terminals participating in the multi-person call as a receiving terminal, and obtaining the call environment noise of the receiving terminal according to the call audio of the receiving terminal, includes: Selecting one of the call member terminals participating in the multi-person call as the receiving terminal, and determining the call audio sent by the receiving terminal; Performing frequency domain conversion processing on the call audio of the receiving terminal to obtain a power spectrum of the call audio of the receiving terminal in the frequency domain; The call environment noise of the receiving terminal is determined according to the power spectrum.

3. The method according to claim 1, characterized in that The selecting one of the call member terminals participating in the multi-person call as a receiving terminal, and obtaining the call environment noise of the receiving terminal according to the call audio of the receiving terminal, includes: Selecting one of the call member terminals participating in the multi-person call as the receiving terminal, and determining the call audio sent by the receiving terminal; Acquire the call environment noise of the call audio of the receiving terminal at each frequency point within a specified frequency band; The determining, based on the call environment noise and the masking threshold of each of the call audios preliminarily screened out, a masking degree of each of the call audios preliminarily screened out includes: The call audios initially screened are sequentially used as current processing audios; Obtaining the masking threshold of each frequency point of the currently processed audio within the specified frequency band; The masking degree of the currently processed audio is determined according to the call environment noise at each frequency point and the masking threshold.

4. The method according to claim 3, characterized in that The determining the masking degree of the currently processed audio according to the call environment noise at each frequency point and the masking threshold includes: Each frequency point is used as the current processing frequency point in turn; When the call environment noise of the currently processed frequency point is greater than the masking threshold of the currently processed frequency point, marking the currently processed frequency point as a target frequency point; The masking degree of the currently processed audio is determined based on all target frequency points of the currently processed audio.

5. The method according to claim 4, characterized in that The determining, based on all target frequency points of the currently processed audio, a masking degree of the currently processed audio, includes: Obtaining the sum of the power spectra of each of the target frequency points; Obtain the sum of the power spectra of all frequency points of the currently processed audio; Determining the proportion of the sum of the power spectra of each target frequency point to the sum of the power spectra of all frequency points; The proportion is used as the masking degree of the currently processed audio.

6. The method according to claim 4, characterized in that The method further comprises: When the call environment noise of the current processing frequency is greater than the masking threshold of the current processing frequency, obtaining a ratio of the call environment noise of the current processing frequency to the masking threshold of the current processing frequency; When the ratio is greater than a preset ratio, the current processing frequency point is marked as the target frequency point.

7. The method according to claim 3, characterized in that The obtaining of the masking threshold of each frequency point of the currently processed audio within the specified frequency band includes: Each frequency point is used as the current processing frequency point in turn; Determining the critical frequency band to which the currently processed frequency point belongs; Obtaining a global masking threshold of the critical frequency band; The masking threshold of the current processing frequency point is determined according to the global masking threshold.

8. The method according to claim 7, characterized in that The obtaining of the global masking threshold of the critical band includes: Obtaining a critical band power spectrum of the critical band; Expanding the critical band power spectrum by using an expansion function to obtain an expanded power spectrum; The global masking threshold is determined according to the extended power spectrum.

9. The method according to claim 3, characterized in that The acquiring of the call environment noise of the call audio of the receiving terminal at each frequency point within the specified frequency band includes: Each frequency point is used as the current processing frequency point in turn; Performing time-frequency domain smoothing on the power spectrum of the current processing frequency point to obtain a smoothed power spectrum; Performing minimum search on the smoothed power spectrum through a window function to obtain a local minimum; Determining a probability of speech existence according to the smoothed power spectrum and the local minimum; The call environment noise of the current processing frequency point is determined according to the speech existence probability.

10. The method according to claim 9, characterized in that Determining the speech presence probability according to the smoothed power spectrum and the local minimum value includes: determining a ratio between the smoothed power spectrum and the local minimum; Determining the probability of the immediate speech existence according to the magnitude relationship between the ratio and a preset ratio; The voice presence probability is determined based on the immediate voice presence probability.

11. The method according to claim 1, wherein The sending the first target call audio and the second target call audio to the receiving terminal includes: Sending the first target call audio and the second target call audio to the receiving terminal, where the sent first target call audio and the second target call audio are used to instruct the receiving terminal to decode the target call audio and mix and play it; or The first target call audio and the second target call audio are decoded and mixed, and the mixed target call audio is re-encoded and sent to the receiving terminal. The sent mixed target call audio is used to instruct the receiving terminal to decode and play the mixed target call audio.

12. The method according to claim 1, characterized in that The multi-person call includes a voice call and a video call.

13. The method according to claim 1, wherein The plurality of call member terminals include at least three call member terminals.

14. A call audio processing device, characterized in that: The device comprises: An acquisition module is used to acquire call audio sent by multiple call member terminals participating in a multi-person call; The acquisition module is further configured to select one of the call member terminals participating in the multi-person call as a receiving terminal, and acquire the call environment noise of the receiving terminal based on the call audio of the receiving terminal; The acquisition module is further configured to determine more than one sending terminal from the terminals of the call members participating in the multi-person call, and preliminarily screen out a plurality of call audios from the call audios sent by each of the sending terminals based on the audio characteristics of the call audios sent by each of the sending terminals, wherein the plurality of call audios are a preset number of call audios with the maximum audio energy and the least ambient noise among the call audios containing voice data; a determination module, configured to determine a masking degree of each of the preliminarily screened call audios based on the call environment noise and a masking threshold of each of the preliminarily screened call audios; the masking degree indicating the degree to which the call audio is masked by the call environment noise; A screening module is used to select a preset number of first target call audios with a masking degree less than a screening threshold from the multiple call audios preliminarily screened, and for candidate call audios with a masking degree greater than or equal to the screening threshold, when the importance of the candidate call audio is higher than a second threshold, the candidate call audio is enhanced to obtain a second target call audio based on the correspondence between the enhancement processing intensity and the masking degree of the candidate call audio, and the first target call audio and the second target call audio are sent to the receiving terminal.

15. The device according to claim 14, characterized in that The acquisition module is also used to select one of the call member terminals participating in the multi-person call as the receiving terminal, determine the call audio sent by the receiving terminal; perform frequency domain conversion processing on the call audio of the receiving terminal to obtain the power spectrum of the call audio of the receiving terminal in the frequency domain; and determine the call environment noise of the receiving terminal based on the power spectrum.

16. The device according to claim 14, characterized in that The determining module is further configured to select one of the call member terminals participating in the multi-person call as the receiving terminal, determine the call audio sent by the receiving terminal; and obtain the call environment noise of each frequency point of the call audio of the receiving terminal within a specified frequency band; The call audios initially screened are sequentially used as current processing audios; Obtaining a masking threshold of each frequency point of the currently processed audio within the specified frequency band; and determining a masking degree of the currently processed audio according to the call environment noise of each frequency point and the masking threshold.

17. The device according to claim 16, characterized in that The determination module is further configured to sequentially use each frequency point as a current processing frequency point; when the call environment noise of the current processing frequency point is greater than the masking threshold of the current processing frequency point, mark the current processing frequency point as a target frequency point; The masking degree of the currently processed audio is determined based on all target frequency points of the currently processed audio.

18. The device according to claim 17, characterized in that The determination module is also used to obtain the sum of the power spectra of each target frequency point; obtain the sum of the power spectra of all frequency points of the currently processed audio; determine the ratio of the sum of the power spectra of each target frequency point to the sum of the power spectra of all frequency points; and use the ratio as the masking degree of the currently processed audio.

19. The device according to claim 17, characterized in that The determining module is further configured to obtain a ratio of the call environment noise of the current processing frequency to the masking threshold of the current processing frequency when the call environment noise of the current processing frequency is greater than the masking threshold of the current processing frequency; When the ratio is greater than a preset ratio, the current processing frequency point is marked as the target frequency point.

20. The device according to claim 16, characterized in that The determination module is further configured to sequentially use each frequency point as a current processing frequency point; determine a critical frequency band to which the current processing frequency point belongs; and obtain a global masking threshold value of the critical frequency band; The masking threshold of the current processing frequency point is determined according to the global masking threshold.

21. The device according to claim 20, characterized in that The determination module is further configured to obtain a critical band power spectrum of the critical band; expand the critical band power spectrum using an expansion function to obtain an expanded power spectrum; and determine the global masking threshold according to the expanded power spectrum.

22. The device according to claim 16, characterized in that The determination module is also used to take each frequency point as the current processing frequency point in turn; perform time-frequency domain smoothing on the power spectrum of the current processing frequency point to obtain a smoothed power spectrum; perform minimum value search on the smoothed power spectrum through a window function to obtain a local minimum value; determine the probability of speech existence based on the smoothed power spectrum and the local minimum value; and determine the call environment noise of the current processing frequency point based on the speech existence probability.

23. The device according to claim 22, characterized in that The determination module is further configured to determine a ratio between the smoothed power spectrum and the local minimum; determine an immediate speech presence probability based on a magnitude relationship between the ratio and a preset ratio; and determine the speech presence probability based on the immediate speech presence probability.

24. The device according to claim 14, characterized in that The screening module is further configured to send the first target call audio and the second target call audio to the receiving terminal, and the sent first target call audio and the second target call audio are used to instruct the receiving terminal to decode the target call audio and mix and play it; Alternatively, the first target call audio and the second target call audio are decoded and mixed, and the mixed target call audio is re-encoded and sent to the receiving terminal. The sent mixed target call audio is used to instruct the receiving terminal to decode and play the mixed target call audio.

25. The device according to claim 14, characterized in that The multi-person call includes a voice call and a video call.

26. The device according to claim 14, characterized in that The plurality of call member terminals include at least three call member terminals.

27. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 13 are implemented.

28. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.

29. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.

Citation Information

Patent Citations

  • Method and device for regulating and controlling coding parameters, equipment and storage medium

    CN110265046A

  • Call audio mixing processing method and device, storage medium and computer equipment

    CN111048119A

  • Call method, device and system, server and storage medium

    CN111049848A