Audio adjustment method and device, computer equipment and storage medium
By collecting environmental audio and obtaining voice signals, combining the audio to be played and the intent of the local object, the volume is automatically adjusted, which solves the problem of limited applicable scenarios for volume adjustment in voice interaction, and improves the accuracy and user experience of audio adjustment.
Patent Information
- Application Number
- CN202311555198.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-20
AI Technical Summary
In voice interaction scenarios, due to device differences, environmental differences or distance changes, the sound received by one end of the other end is too small or too large, which affects the user experience. The existing technology has limited applicable scenarios, and the user needs to manually adjust the target volume gain, which has a low user experience.
By collecting the ambient audio from the local end, a voice signal is obtained, and based on the voice signal, the audio to be played and the object intention of the local end, the target volume gain of the audio to be played is automatically adjusted to achieve audio adjustment.
This method can automatically adjust the volume in any scenario, improve the accuracy and user satisfaction of audio adjustments, and is suitable for a variety of environments and user intentions.
Smart Images

Figure CN120020948A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio processing, and more particularly, to an audio adjustment method, apparatus, computer device, and storage medium. Background Art
[0002] With the popularity of audio and video, real-time voice interaction can be achieved through the network, which greatly facilitates communication between users and improves the information transmission efficiency. However, in the voice interaction scenario, due to device differences, environmental differences, or distance changes, etc., there may be a phenomenon that the sound of the other end received by one end (which can be regarded as the local end) is too small or too large, affecting the user experience. Therefore, it is necessary to adjust the audio.
[0003] Currently, in the voice interaction scenario, mainly the audio data collected from the local end or the other end is processed by AGC (Automatic Gain Control) to obtain the target volume gain, so as to achieve audio adjustment.
[0004] However, due to the limited applicable scenarios of the above method, in some scenarios, users need to manually adjust the target volume gain, resulting in a low user experience. Summary of the Invention
[0005] The main object of the present application is to provide an audio adjustment method, apparatus, computer device, and storage medium, which can be applicable to any scenario, and can automatically adjust the target volume gain according to the environmental audio and the local end object intention, improving the accuracy of audio adjustment and user satisfaction.
[0006] To achieve the above object, in a first aspect, the present application provides an audio adjustment method, including:
[0007] Collect the environmental audio of the local end;
[0008] Based on the environmental audio, obtain the first voice signal;
[0009] Based on the first voice signal, the audio to be played, and the local end object intention, determine the target volume gain of the audio to be played, where the audio to be played is the audio stream of the other end in voice communication with the local end, and the local end object intention is used to represent the desired volume adjustment trend of the object at the local end;
[0010] Adjust the audio to be played according to the target audio gain to obtain the adjusted audio to be played, and play the adjusted audio to be played at the local end.
[0011] In an embodiment, based on the environmental audio, obtaining the first voice signal includes:
[0012] Process the environmental audio to obtain the environmental audio signal;
[0013] Detect whether there is a first voice signal in the ambient audio signal;
[0014] If so, extract the first voice signal from the ambient audio signal.
[0015] In an embodiment, detecting whether there is a first voice signal in the ambient audio signal includes:
[0016] Calculate the energy ratio of the ambient audio signal;
[0017] Based on the magnitude relationship between the energy ratio and a first preset threshold, detect whether there is a first voice signal in the ambient audio signal.
[0018] In an embodiment, calculating the energy ratio of the ambient audio signal includes:
[0019] Obtain the frequency-domain signal corresponding to the ambient audio signal, and calculate the total energy value of the frequency-domain signal;
[0020] Obtain the target frequency-domain signal in the frequency-domain signal, and calculate the total energy value of the target frequency-domain signal, where the target frequency-domain signal is a signal with a preset frequency;
[0021] Divide the total energy value of the target frequency-domain signal by the total energy value of the frequency-domain signal to obtain the energy ratio of the ambient audio signal.
[0022] In an embodiment, determining the target volume gain of the audio to be played based on the first voice signal, the audio to be played, and the intention of the local object includes:
[0023] Calculate the target energy value of the first voice signal;
[0024] Obtain the audio signal to be played corresponding to the audio to be played, and detect whether there is a second voice signal in the audio signal to be played;
[0025] If so, calculate the target energy value of the second voice signal;
[0026] Based on the target energy value of the first voice signal, the target energy value of the second voice signal, and the intention of the local object, determine the target volume gain of the audio to be played;
[0027] Wherein, the first voice signal is the voice signal of the local object, and the second voice signal is the voice signal of the remote object.
[0028] In an embodiment, calculating the target energy value of the first voice signal includes:
[0029] Calculate the initial energy value of the first voice signal;
[0030] Obtain the first preset volume gain;
[0031] Divide the initial energy value of the first voice signal by the first preset volume gain to obtain the target energy value of the first voice signal.
[0032] In an embodiment, calculating the target energy value of the second voice signal includes:
[0033] Calculate the initial energy value of the second voice signal;
[0034] Obtain the second preset volume gain, where the second preset volume gain corresponds to the initial energy value of the second voice signal;
[0035] Add the initial energy value of the second voice signal and the second preset volume gain to obtain the target energy value of the second voice signal.
[0036] In an embodiment, determining the target volume gain of the audio to be played based on the target energy value of the first voice signal, the target energy value of the second voice signal, and the local object intention includes:
[0037] Compare the target energy value of the first voice signal and the target energy value of the second voice signal to obtain the initial volume gain of the audio to be played;
[0038] Obtain the local object intention;
[0039] Determine the target volume gain of the audio to be played based on the local object intention and the initial volume gain.
[0040] In an embodiment, comparing the target energy value of the first voice signal and the target energy value of the second voice signal to obtain the initial volume gain of the audio to be played includes:
[0041] Subtract the target energy value of the second voice signal from the target energy value of the first voice signal to obtain a difference value;
[0042] Use the difference value as the initial volume gain of the audio to be played.
[0043] In an embodiment, obtaining the local object intention includes:
[0044] Obtain the volume trend of the volume button and the playback volume trend of the application within a preset time;
[0045] Determine the local object intention based on the volume trend of the volume button and the playback volume trend of the application.
[0046] In an embodiment, determining the local object intention based on the volume trend of the volume button and the playback volume trend of the application includes:
[0047] If both the volume trend of the volume button and the playback volume trend of the application are increasing, the local object intends to increase the volume of the audio to be played;
[0048] If both the volume trend of the volume button and the playback volume trend of the application are decreasing, the local object intends to decrease the volume of the audio to be played;
[0049] If both the volume trend of the volume button and the playback volume trend of the application remain unchanged, the local object intends to keep the volume of the audio to be played.
[0050] In one embodiment, determining the target volume gain of the audio to be played based on the local object intention and the initial volume gain includes:
[0051] If the local object intention is to keep the volume of the audio to be played and the initial volume gain is greater than the second preset threshold, adjust the initial volume gain and use the adjusted initial volume gain as the target volume gain of the audio to be played;
[0052] If the local object intention is to increase the volume of the audio to be played and the initial volume gain is less than or equal to the second preset threshold, set the target volume gain of the audio to be played to zero;
[0053] If the local object intention is to decrease the volume of the audio to be played and the initial volume gain is less than or equal to the second preset threshold, set the target volume gain of the audio to be played to zero.
[0054] In a second aspect, an audio adjustment device provided by an embodiment of the present application includes:
[0055] An acquisition module for acquiring the ambient audio of the local end;
[0056] A voice acquisition module for acquiring a first voice signal based on the ambient audio;
[0057] A gain determination module for determining the target volume gain of the audio to be played based on the first voice signal, the audio to be played, and the local object intention, where the audio to be played is the audio stream of the peer end in voice communication with the local end, and the local object intention is used to represent the desired volume adjustment trend of the object at the local end;
[0058] An audio adjustment module for adjusting the audio to be played according to the target audio gain to obtain the adjusted audio to be played, so as to play the adjusted audio to be played at the local end.
[0059] In one embodiment, the voice acquisition module is further configured to process the ambient audio to obtain an ambient audio signal;
[0060] Detect whether there is a first voice signal in the ambient audio signal;
[0061] If it exists, extract the first voice signal from the environmental audio signal.
[0062] In an embodiment, the voice acquisition module is further configured to calculate the energy ratio of the environmental audio signal;
[0063] Based on the magnitude of the energy ratio and a first preset threshold, detect whether a first voice signal exists in the environmental audio signal.
[0064] In an embodiment, the voice acquisition module is further configured to obtain the frequency-domain signal corresponding to the environmental audio signal and calculate the total energy value of the frequency-domain signal;
[0065] Obtain the target frequency-domain signal in the frequency-domain signal and calculate the total energy value of the target frequency-domain signal, where the target frequency-domain signal is a signal with a preset frequency;
[0066] Divide the total energy value of the target frequency-domain signal by the total energy value of the frequency-domain signal to obtain the energy ratio of the environmental audio signal.
[0067] In an embodiment, the gain determination module is further configured to calculate the target energy value of the first voice signal;
[0068] Obtain the audio signal to be played corresponding to the audio to be played and detect whether a second voice signal exists in the audio signal to be played;
[0069] If it exists, calculate the target energy value of the second voice signal;
[0070] Based on the target energy value of the first voice signal, the target energy value of the second voice signal, and the local object intention, determine the target volume gain of the audio to be played;
[0071] Wherein, the first voice signal is the voice signal of the local object, and the second voice signal is the voice signal of the remote object.
[0072] In an embodiment, the gain determination module is further configured to calculate the initial energy value of the first voice signal;
[0073] Obtain the first preset volume gain;
[0074] Divide the initial energy value of the first voice signal by the first preset volume gain to obtain the target energy value of the first voice signal.
[0075] In an embodiment, the gain determination module is further configured to calculate the initial energy value of the second voice signal;
[0076] Obtain the second preset volume gain, where the second preset volume gain corresponds to the initial energy value of the second voice signal;
[0077] Add the initial energy value of the second voice signal to the second preset volume gain to obtain the target energy value of the second voice signal.
[0078] In one embodiment, the gain determination module is further configured to compare the target energy value of the first voice signal with the target energy value of the second voice signal to obtain the initial volume gain of the audio to be played;
[0079] Obtain the intention of the local object;
[0080] Based on the intention of the local object and the initial volume gain, determine the target volume gain of the audio to be played.
[0081] In one embodiment, the gain determination module is further configured to subtract the target energy value of the second voice signal from the target energy value of the first voice signal to obtain a difference value;
[0082] Use the difference value as the initial volume gain of the audio to be played.
[0083] In one embodiment, the gain determination module is further configured to obtain the volume trend of the volume button and the playback volume trend of the application within a preset time;
[0084] Based on the volume trend of the volume button and the playback volume trend of the application, determine the intention of the local object.
[0085] In one embodiment, the gain determination module is further configured to, if both the volume trend of the volume button and the playback volume trend of the application are rising, the intention of the local object is to increase the volume of the audio to be played;
[0086] If both the volume trend of the volume button and the playback volume trend of the application are falling, the intention of the local object is to decrease the volume of the audio to be played;
[0087] If both the volume trend of the volume button and the playback volume trend of the application remain unchanged, the intention of the local object is to keep the volume of the audio to be played.
[0088] In one embodiment, the gain determination module is further configured to, if the intention of the local object is to keep the volume of the audio to be played and the initial volume gain is greater than the second preset threshold, adjust the initial volume gain and use the adjusted initial volume gain as the target volume gain of the audio to be played;
[0089] If the intention of the local object is to increase the volume of the audio to be played and the initial volume gain is less than or equal to the second preset threshold, set the target volume gain of the audio to be played to zero;
[0090] If the intention of the local object is to decrease the volume of the audio to be played and the initial volume gain is less than or equal to the second preset threshold, set the target volume gain of the audio to be played to zero.
[0091] In a third aspect, an embodiment of the present application provides a device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of any of the above methods are implemented.
[0092] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0093] An embodiment of the present application provides an audio adjustment method, device, computer device, and storage medium, including: first collecting the ambient audio at the local end, then obtaining a first voice signal based on the ambient audio, then determining the target volume gain of the audio to be played based on the first voice signal, the audio to be played, and the local object intention, and finally adjusting the audio to be played according to the target audio gain to obtain the adjusted audio to be played, so as to play the adjusted audio to be played at the local end. The present application obtains a first voice signal from the ambient audio to automatically match the environment where the local object is located with the target volume gain, expands the applicable scenarios of audio adjustment, enables more comprehensive consideration of environmental factors during audio adjustment, and thus improves the accuracy and adaptability of audio adjustment; in addition, adding the local object intention in obtaining the target volume gain realizes precise gain adjustment of the playing volume of the audio to be played, enables more comprehensive consideration of the user's needs and intentions during audio adjustment, and further improves the accuracy of audio adjustment and user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0094] The drawings constituting a part of the present application are used to provide a further understanding of the present application, making other features, objectives, and advantages of the present application more obvious. The schematic embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0095] Figure 1 is a schematic flowchart of an audio processing method provided by an embodiment of the present application;
[0096] Figure 2 is a schematic flowchart of another audio processing method provided by an embodiment of the present application;
[0097] Figure 3 is a schematic diagram of an application scenario of an audio adjustment method provided by an embodiment of the present application;
[0098] Figure 4 is a schematic flowchart of an audio adjustment method provided by an embodiment of the present application;
[0099] Figure 5It is a schematic flowchart of a voice detection method provided by an embodiment of the present application;
[0100] Figure 6 It is a schematic diagram of the numerical form corresponding to the frame audio signal provided by an embodiment of the present application;
[0101] Figure 7 It is a schematic flowchart of a method for determining voice energy value provided by an embodiment of the present application;
[0102] Figure 8 It is a schematic diagram of an audio frame structure provided by an embodiment of the present application;
[0103] Figure 9 It is a schematic flowchart of a method for determining the corresponding intention at the local end provided by an embodiment of the present application;
[0104] Figure 10 It is a schematic flowchart of a method for determining the adjusted audio to be played provided by an embodiment of the present application;
[0105] Figure 11 It is a schematic diagram of the structure of an audio adjustment device provided by an embodiment of the present application;
[0106] Figure 12 It is a schematic diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0107] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are only a part rather than all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0108] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein.
[0109] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the various processes do not mean the order of execution, and the execution order of the various processes should be determined by their functions and internal logics, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0110] It should be understood that in this application, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0111] It should be understood that in this application, "a plurality of" means two or more. "And / or" is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "Including A, B, and C" and "including A, B, C" mean that all of A, B, and C are included. "Including A, B, or C" means including any one of A, B, and C. "Including A, B, and / or C" means including any one or any two or all three of A, B, and C.
[0112] It should be understood that in this application, "B corresponding to A", "B corresponding to A relatively", "A corresponding to B relatively", or "B corresponding to A relatively" means that B is associated with A, and B can be determined according to A. Determining B according to A does not mean determining B only according to A, but also B can be determined according to A and / or other information. The matching of A and B means that the similarity between A and B is greater than or equal to a preset threshold.
[0113] Depending on the context, as used herein, "if" can be interpreted as "when", "while", "in response to determining", or "in response to detecting".
[0114] The data involved in this application can be data authorized by testers or fully authorized by all parties. The collection, dissemination, use, etc. of the data all comply with the requirements of relevant laws, regulations, and standards in relevant countries and regions. The implementation manners / embodiments of this application can be combined with each other.
[0115] The embodiments of this application can be applied to various scenarios such as video, voice interaction, and co-hosting.
[0116] The technical solutions of this application will be described in detail below with specific embodiments. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0117] To make the purpose, technical solutions, and advantages of this application clearer, the professional terms in this application are explained as follows:
[0118] Loudness / Energy Value: It is a description of the sound intensity. In one embodiment, for example, in a scenario where the acquisition depth is 16 bits, the acquired audio will be described by a series of 16-bit numbers, with a maximum of 32767 and a minimum of 0. The magnitude of this value can also be regarded as loudness. The larger the value, the greater the loudness; conversely, the smaller the value.
[0119] AEC (Acoustic Echo Cancellation): During a voice connection, the received far-end or near-end signal is used as a reference signal to estimate the signal after being reflected by the microphone and the room, obtaining a predicted signal. Then, the predicted signal is subtracted from the signal collected by the microphone, leaving only the locally collected sound.
[0120] ANS (Automatic Noise Suppression): By identifying the noise in the collected signal and filtering out the noise from the collected signal.
[0121] AGC: During the audio acquisition process, due to differences in the microphone itself, changes in the distance between the user and the microphone, and differences in the system acquisition gain set by the user, the acquired audio data may not be fixed, which may cause the played audio to be too low or too high, or even fluctuate, greatly affecting the listener's experience. At this time, audio automatic gain needs to be added. From a technical perspective, there are two ways of audio gain. One is analog gain, which is to adjust the system acquisition gain to make the acquired audio data reach the target value; the other is digital gain, which directly adjusts the acquired audio data to reach the target value. In addition, in actual engineering, analog gain and digital gain can be combined and, combined with the adaptive ability, form audio gain based on feedback.
[0122] Vad (Voice activity detection): It is used to detect whether a certain segment of speech contains a voice signal, where the voice signal refers to the signal corresponding to the sound made by a person. Voice activity detection generally uses an energy-based voice detection method. Since the energy of the voice signal is mainly distributed below 2 kHz, while the noise is above 2 kHz, by extracting the audio energy below 2 kHz and then comparing it with a threshold, it is possible to identify whether a voice signal is contained.
[0123] Next, the solution of this application will be described through specific embodiments with reference to the accompanying drawings.
[0124] With the popularity of audio and video, real-time voice interaction can be achieved through the network, which greatly facilitates the communication between users and improves the information transmission efficiency. However, in the voice interaction scenario, due to device differences, environmental differences, or distance changes, etc., there will be a phenomenon that the sound received by one end (which can be regarded as the local end) from the other end is too small or too large, affecting the user experience. Therefore, it is necessary to adjust the audio.
[0125] Currently, in the voice interaction scenario, mainly the audio data collected at the local end or the other end is processed by AGC to obtain the target volume gain, so as to achieve audio adjustment.
[0126] In the Figure 1 shown co-hosting scenario, including the remote end and the proximal end. Here, taking the proximal end as the local end and the remote end as the other end as an example, the process of obtaining the target volume gain by processing the audio data collected at the proximal end by AGC to achieve audio adjustment is described as follows:
[0127] When the remote end and the proximal end are in a co-hosting, the proximal end collects the voice signal through the audio acquisition module. The collected voice signal first passes through the AEC / ANS module, which eliminates the echo and noise of the voice signal to obtain the voice signal after elimination. Then the voice signal after elimination passes through the AGC module, which adjusts the voice signal after elimination and the original acquisition gain of the proximal system (such as the microphone) according to the target gain to obtain the acquired signal after gain. Among them, the intensity of the acquired signal after gain is obtained by amplifying or reducing the intensity of the voice signal according to the target gain.
[0128] The acquired signal after gain is successively uplinked to the server through the encoding module and the uplink module, and then successively passes through the server, the downlink module, and the decoding module to generate the audio to be played. The audio to be played is directly played through the playback module configured at the remote end. Among them, the sound intensity of the audio to be played is equal to the intensity of the acquired signal after gain.
[0129] While the remote end receives and plays the audio to be played from the proximal end, the remote end also collects the voice signal, processes and uploads it, and then it is pulled by the proximal end to form the data to be played, and is played through the playback module configured at the proximal end. It should be noted that the process of the remote end collecting the voice signal, processing and uploading it, and then being pulled by the proximal end to form the data to be played, and being played through the playback module configured at the proximal end to achieve audio adjustment is similar to the process of the proximal end collecting the voice signal, processing and uploading it, and then being pulled by the remote end to form the data to be played, and being played through the playback module configured at the remote end to achieve audio adjustment, which will not be elaborated here.
[0130] Through the above interaction process between the remote end and the proximal end, the co-hosting between the remote end and the proximal end is achieved.
[0131] However, this solution also has its drawbacks. Traditional gain mainly relies on the acquisition gain parameters of the system and the audio energy value for gain. At this time, the target volume gain may not meet the customer's needs.
[0132] For example, when the user is in a relatively noisy environment, if the target volume gain in a quiet environment is used for adjustment, it will be found that the user still cannot clearly hear the audio sent by the other end; while if the user is in a relatively quiet environment and the target volume gain in a noisy environment is used for adjustment, it will cause the sound played at the local end to sound too loud, affecting the listening experience.
[0133] At this time, it can only rely on the user to manually adjust the target volume gain, which is very inconvenient to use and the user experience is low.
[0134] In addition, the above method mainly performs gain on the voice signal collected by the other end. The target volume gain refers to the system acquisition gain and the acquisition signal energy value of the other end, and does not consider the environmental factors at the local end (such as the playback end), and cannot be applied to different environmental scenarios, so that the audio after gain may not meet the needs of the local end.
[0135] In summary, due to the limited applicable scenarios of the above method, in some scenarios, the user needs to manually adjust the target volume gain, and the user experience is low. Therefore, this application provides an audio adjustment method.
[0136] This solution adds the playback AGC module 201 and the playback AGC module 202 as shown in Figure 2 at the local end (near end) and the other end (far end) for automatically adjusting the audio to be played according to the collected environmental audio.
[0137] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the application scenario of an audio adjustment method provided by an embodiment of this application.
[0138] As shown in Figure 3 , the first client 302, the second client 306 communicate with the server 304. Among them, the first client 302 is used as the local end and the second client 306 is used as the other end. The first client 302 collects environmental audio and then sends the environmental audio to the server 304. The server 304 first obtains the first voice signal based on the environmental audio, and then determines the target volume gain of the audio to be played based on the first voice signal, the audio to be played, and the local object intention. The audio to be played is the audio stream of the second client 306 that communicates with the first client 302. The local object intention is used to represent the desired volume adjustment trend of the object of the first client 302. Then, the audio to be played is adjusted according to the target audio gain to obtain the adjusted audio to be played, so as to play the adjusted audio to be played on the first client 302.
[0139] Among them, the first client 302, the second client 306 and the server 304 can communicate through any communication means, including but not limited to network communication, and the above network can include but not limited to: wired network, wireless network. Among them, the wired network includes: local area network, metropolitan area network and wide area network, and the wireless network includes: Bluetooth, WIFI, and other networks that implement wireless communication. Among them, the first client 302 and the second client 306 include at least one of the following: mobile phone (such as Android mobile phone, iOS mobile phone, etc.), laptop computer, tablet computer, handheld computer, mobile Internet device (Mobile Internet Devices, MID), PAD, desktop computer, smart TV, etc. Among them, the server 304 can be an on-site server or a remote server. Among them, both the on-site server and the remote server can be implemented by an independent server or a service cluster composed of multiple servers. The above is only an example, and this embodiment does not make any limitation on this.
[0140] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of an audio adjustment method provided by an embodiment of the present application. As Figure 4 shown, this method is applied to Figure 3 the server 304 shown as follows, and includes the following steps:
[0141] Step S401: Collect the ambient audio of the local end.
[0142] Among them, the ambient audio of the local end is all audio data included in the environment where the local end is located, such as voice data, noise data, whistle sound data, etc.
[0143] When the local end plays the to-be-played audio sent by the peer end, in order to make the intensity of the audio played by the local end meet the user's requirements, considering all the audio data included in the environment where the local end is located, the intensity of the audio played by the local end is adjusted along with the ambient audio of the local end, which broadens the applicable scenarios of audio adjustment and improves the user experience.
[0144] For example, when the local end plays the to-be-played audio sent by the peer end, if the user is in a noisy environment and wants to easily hear the to-be-played audio sent by the peer end, then the intensity of the audio played by the local end needs to be greater. And since the intensity of the audio played by the local end is adjusted along with the ambient audio of the local end, the intensity of the ambient audio collected by the local end needs to be greater, that is, the user's speaking voice is adjusted, that is, the user's speaking voice needs to be greater.
[0145] If the user is in a relatively quiet environment, in order to avoid disturbing others, the intensity of the audio played at the local end needs to be lower. And the intensity of the audio played at the local end is adjusted according to the ambient audio at the local end, so the intensity of the ambient audio collected at the local end needs to be relatively low, that is, the user's speaking voice is adjusted, that is, the user's speaking voice needs to be lower.
[0146] Step S402: Obtain a first voice signal based on the ambient audio.
[0147] Among them, the first voice signal is the voice signal of the local object. The local object is an object at the local end in scenarios such as co-hosting, live interaction, online meetings, video calls, etc. The object includes but is not limited to people, robots, animals, etc.
[0148] Steps for obtaining the first voice signal based on the ambient audio: First, the ambient audio needs to be processed to obtain an ambient audio signal, and then it is detected whether there is a first voice signal in the ambient audio signal. If there is a first voice signal in the ambient audio signal, the first voice signal is extracted from the ambient audio signal.
[0149] Among them, processing the ambient audio to obtain an ambient audio signal mainly removes the noise in the ambient audio. Among them, for noise removal, a noise reduction algorithm can be used. The noise reduction algorithm includes but is not limited to frequency domain noise reduction algorithms, time domain noise reduction algorithms, wavelet domain noise reduction algorithms, deep learning-based noise reduction algorithms, active noise reduction methods, passive noise reduction methods, etc.
[0150] For example, after obtaining the ambient audio of the current environment, noise data can be removed by using an active noise reduction or passive noise reduction method. Specifically, when using the passive noise reduction method to remove noise data in this step, noise reduction can be performed by using a noise reduction circuit, that is, the parameters of the noise reduction circuit are set according to the noise data, and the ambient audio is input into the noise reduction circuit with the parameters set to remove the noise data.
[0151] Optionally, when using the active noise reduction method to remove the target noise data, the parameters of the filter are set according to the noise data, and the ambient audio is input into the filter with the parameters set to remove the noise data.
[0152] Optionally, the noise can also be classified by level, and corresponding processing is performed on noise data of different levels.
[0153] It should be noted that the more the number of noise level divisions, the more frequent the volume gain adjustment, and the smoother the change in the volume of the external target audio. However, the increase and decrease of the volume will appear relatively slow and not very timely. On the contrary, the fewer the number of divisions, the rarer the volume gain adjustment, and the volume change will respond very quickly, but the change in the volume of the external target audio will feel abrupt. For example, when the device is in a similar place such as an indoor shopping mall where the noise type and noise level are basically constant or change very slowly, since the noise in such places does not suddenly appear or disappear, the number of noise level divisions can be set to be larger, making the adjustment of the external volume appear smooth. When the device is in an outdoor place, since there will be some high-intensity sudden noises in such places, such as passing vehicles, etc., the number of noise level divisions can be set to be smaller, so that the external sound can quickly respond to the change of the sudden noise. When a strong noise suddenly occurs, the volume of the external audio increases accordingly, so that the user will not fail to hear the content of the external audio clearly and lose the listening information when the noise suddenly occurs; when the strong noise suddenly disappears, the volume of the external audio decreases accordingly, so that the user will not feel that the external audio is too loud and harsh in a relatively quiet environment after the noise has disappeared.
[0154] After obtaining the environmental audio signal, it is necessary to detect whether there is a first voice signal in the environmental audio signal. Specifically, first calculate the energy proportion of the environmental audio signal, and then based on the magnitude relationship between the energy proportion and the first preset threshold, detect whether there is a first voice signal in the environmental audio signal. Among them, the first voice signal refers to the signal corresponding to the sound made by people in the environmental audio signal.
[0155] Among them, to calculate the energy proportion of the environmental audio signal, it is necessary to first obtain the frequency-domain signal corresponding to the environmental audio signal, calculate the total energy value of the frequency-domain signal, then obtain the target frequency-domain signal in the frequency-domain signal, and calculate the total energy value of the target frequency-domain signal. Among them, the target frequency-domain signal is a signal with a preset frequency. Then divide the total energy value of the target frequency-domain signal by the total energy value of the frequency-domain signal to obtain the energy proportion of the environmental audio signal.
[0156] Exemplarily, as Figure 5 shown, assuming the preset frequency is 2 kHz, after step S501 obtains the environmental audio signal, step S502 is executed, that is, the VAD (Voice Activity Detection) method is used to detect whether there is a first voice signal in the environmental audio signal.
[0157] The detection process of step S502 can be specifically as follows: First, convert each data frame in the environmental audio signal into a frequency-domain signal, then calculate the total energy value of the target frequency-domain signal below 2 kHz and the total energy value of the frequency-domain signal, and then obtain the energy ratio of the environmental audio signal through the quotient of the total energy value of the target frequency-domain signal and the total energy value of the frequency-domain signal. Finally, compare the energy ratio with the first preset threshold, and authenticate the frequency-domain signal corresponding to the energy ratio greater than the first preset threshold as the first voice signal.
[0158] For the calculation of the energy value of a certain frame signal of the frequency-domain signal and the target frequency-domain signal, generally, the maximum data value in the frame signal is obtained as the energy value of the frame signal. For example, if a certain frame signal is an audio signal, the maximum data value in the audio signal is the energy value of the frame audio signal, where the audio format is as Figure 6 shown. The frame audio signal can be regarded as a string of 16-bit numerical values, and then the maximum value is used as the energy value of the frame audio signal.
[0159] After calculating the energy value of a certain frame signal, the energy values of all frames in the frequency-domain signal or the target frequency-domain signal can be summed up to obtain the total energy values of the frequency-domain signal and the target frequency-domain signal.
[0160] Step S403: Determine the target volume gain of the audio to be played based on the first voice signal, the audio to be played, and the local object intention.
[0161] Among them, the object can be a person, object, etc. with thinking intention.
[0162] Among them, the audio to be played is the audio stream of the peer end in voice communication with the local end, and the local object intention is used to represent the volume adjustment trend expected by the local object.
[0163] To determine the target volume gain of the audio to be played based on the first voice signal, the audio to be played, and the local object intention, it is necessary to first calculate the target energy value of the first voice signal, then obtain the audio signal to be played corresponding to the audio to be played, and detect whether there is a second voice signal in the audio signal to be played. If so, calculate the target energy value of the second voice signal, and then determine the target volume gain of the audio to be played based on the target energy value of the first voice signal, the target energy value of the second voice signal, and the local object intention, where the first voice signal is the voice signal of the local object and the second voice signal is the voice signal of the peer object.
[0164] Combined with Figure 7, to calculate the target energy value of the first voice signal, it is necessary to first execute step S701: calculate the initial energy value of the first voice signal, then execute step S702: obtain the first preset volume gain, and then execute step S703: calculate the target energy value of the first voice signal. In an embodiment, the initial energy value of the first voice signal can be divided by the first preset volume gain to obtain the target energy value of the first voice signal. Among them, the first preset volume gain can be set by the system or can be set by the user through operating the device.
[0165] The user can set the first preset volume gain through the local device, that is, open the settings option of the local device, select the audio input settings from the settings option, display the audio input settings interface, and complete the setting of the first preset volume gain by adjusting the volume gain in the audio input settings interface.
[0166] Since some devices can set the volume gain, there is a difference between the target energy value of the first voice signal and the initial energy value of the first voice signal. If only the initial energy value of the first voice signal is referred to, misjudgment may occur. Therefore, it is necessary to obtain the target energy value of the first voice signal through the initial energy value of the first voice signal and the first preset volume gain, and the obtained target energy value of the first voice signal is more accurate.
[0167] To obtain the target energy value of the first voice signal through the initial energy value of the first voice signal and the first preset volume gain, the following formula can be used for calculation:
[0168] Initial energy value of the first voice signal / First preset volume gain = Target energy value of the first voice signal.
[0169] After calculating the target energy value of the first voice signal by the above method, the target energy value of the first voice signal can be stored in a data structure containing audio signals as shown in Figure 8 . The audio signal is stored in this data structure in the form of an audio frame structure, and this audio frame structure includes audio data, data duration, sampling rate, number of channels, whether it contains speech, and energy value, etc.
[0170] To calculate the target energy value of the first voice signal, it is also necessary to detect whether there is a second voice signal in the audio signal to be played. If it exists, calculate the target energy value of the second voice signal. Among them, calculating the target energy value of the second voice signal includes first calculating the initial energy value of the second voice signal and obtaining the second preset volume gain, and then adding the initial energy value of the second voice signal and the second preset volume gain to obtain the target energy value of the second voice signal. Among them, the second preset volume gain corresponds to the initial energy value of the second voice signal.
[0171] Since each frame signal in the voice signal corresponds to a gain, each frame signal in the second voice signal also corresponds to a gain (i.e., the second preset volume gain). Then, to calculate the target energy value of the second voice signal, the initial energy value of the second voice signal can be summed with the second preset volume gain.
[0172] Specifically, to obtain the target energy value of the second voice signal through the initial energy value of the second voice signal and the second preset volume gain, the following formula can be used for calculation:
[0173] Initial energy value of the second voice signal + Second preset volume gain = Target energy value of the second voice signal.
[0174] It should be noted that the methods for calculating the initial energy value of the first voice signal and the initial energy value of the second voice signal are similar. The maximum data value corresponding to the frame signal in the voice signal can be used as the initial energy value of this voice signal, which will not be elaborated here.
[0175] After calculating the target energy value of the first voice signal and the target energy value of the second voice signal, it is also necessary to obtain the intention of the local object. That is, based on the target energy value of the first voice signal, the target energy value of the second voice signal, and the intention of the local object, the target volume gain of the audio to be played can be determined.
[0176] In an embodiment, to determine the target volume gain of the audio to be played based on the target energy value of the first voice signal, the target energy value of the second voice signal, and the intention of the local object, it is necessary to first compare the target energy value of the first voice signal and the target energy value of the second voice signal to obtain the initial volume gain of the audio to be played, and obtain the intention of the local object. Then, based on the intention of the local object and the initial volume gain, the target volume gain of the audio to be played is determined.
[0177] Among them, comparing the target energy value of the first voice signal and the target energy value of the second voice signal to obtain the initial volume gain of the audio to be played includes: first taking the difference between the target energy value of the first voice signal and the target energy value of the second voice signal to obtain a difference value, and then using the difference value as the initial volume gain of the audio to be played.
[0178] Combined with Figure 9 , obtaining the intention of the local object includes: first obtaining the volume trend of the volume button within a preset time through step S901 and the volume trend of the application's playback volume through step S902, and then executing step S903, that is, determining the intention of the local object based on the volume trend of the volume button and the volume trend of the application's playback volume.
[0179] When the playback volume fails to meet the user's requirements, the volume will be adjusted. The volume button can only adjust the volume of the same type as the current volume, and cannot guarantee that the playback volume of the current application can be adjusted. Moreover, adjusting the system volume will also cause the volume of other applications to change synchronously. Therefore, adjusting the volume through the volume button may not be accurate, but it can express the user's intention regarding the volume, such as the desire to increase, decrease the volume, or keep it unchanged. The playback volume of an application is similar. Since the playback volume of an application is a software volume, the adjustment range is limited and may not necessarily meet the user's requirements, but it can also reflect the user's intention regarding the volume.
[0180] Therefore, this application combines the volume trend of the volume button and the playback volume trend of the application to determine the intention of the local object, that is, to determine whether the user hopes to increase, decrease, or keep the playback volume unchanged.
[0181] Specifically, based on the volume trend of the volume button and the playback volume trend of the application, determining the intention of the local object includes: if both the volume trend of the volume button and the playback volume trend of the application are rising, the intention of the local object is to increase the volume of the audio to be played; if both the volume trend of the volume button and the playback volume trend of the application are falling, the intention of the local object is to decrease the volume of the audio to be played; if both the volume trend of the volume button and the playback volume trend of the application remain unchanged, the intention of the local object is to keep the volume of the audio to be played.
[0182] For example, in addition to determining the intention of the local object based on the volume trend of the volume button and the playback volume trend of the application, the acquisition of the intention of the local object can also be analyzed and inferred through various other methods. For example, the local object can express its demand or preference for the volume through voice instructions or gesture instructions.
[0183] For example, based on the preset keywords detected in the first voice signal, determine the intention of the local object, where the preset keywords represent volume adjustment instruction commands. Specifically, the preset keywords in the first voice signal can be detected, such as "increase volume" or "decrease volume". When these preset keywords are detected, it can be inferred that the local object hopes to adjust the volume. This detection of preset keywords can be implemented based on speech recognition technology. By training a model to recognize specific voice commands, the target volume gain can be automatically adjusted in subsequent steps.
[0184] For example, based on a preset gesture detected in the video frame of the current communication, determine the intention of the local object. This preset gesture represents a volume adjustment instruction. Specifically, if the object enables the video function during a co-hosting or video call, the gestures of the object in the video frame can be matched with the preset gestures. For example, the preset gestures may include swiping left or down on the screen to indicate a decrease in volume, and the preset gestures also include swiping right or up on the screen to indicate an increase in volume. By recognizing these preset gestures, the intention of the local object can be inferred, and thus the target volume gain can be automatically adjusted in subsequent steps.
[0185] The automatic adjustment of the target volume gain can be achieved in the following ways:
[0186] After a gesture of swiping right or up is displayed on the screen, it can be inferred that the intention of the local object is to increase the volume. If the user clicks on the gesture of swiping right or up, a volume gain setting page is displayed, and the automatically adjusted target volume gain is shown on this page.
[0187] If the automatically adjusted target volume gain meets the requirements of the local object, the automatic adjustment of the volume gain is completed; if the automatically adjusted target volume gain does not meet the requirements of the local object, the target volume gain can be adjusted manually on the volume gain setting page until it meets the requirements of the local object.
[0188] In addition, multiple signals can be combined to obtain the intention of the local object. For example, in addition to the above methods of volume trend, voice indication, and gesture indication, the intention of the local object can also be inferred by analyzing the operation behavior of the local object. If the local object frequently operates at a certain volume level, it can be considered that this volume level is the preferred volume of the local object, and it is automatically adjusted to this volume level.
[0189] After obtaining the intention of the local object, the target volume gain of the audio to be played can be determined based on the intention of the local object and the initial volume gain, including: if the intention of the local object is to keep the volume of the audio to be played, and the initial volume gain is greater than the second preset threshold, adjust the initial volume gain, and use the adjusted initial volume gain as the target volume gain of the audio to be played; if the intention of the local object is to increase the volume of the audio to be played and the initial volume gain is less than or equal to the second preset threshold, set the target volume gain of the audio to be played to zero; if the intention of the local object is to decrease the volume of the audio to be played, and the initial volume gain is less than or equal to the second preset threshold, set the target volume gain of the audio to be played to zero. Among them, the second preset threshold can be set according to specific circumstances and is not specifically limited here.
[0190] Among them, the initial volume gain is set by the local device and can be directly read.
[0191] The setting of the initial volume gain can be set through the local device, that is, open the setting option of the local device, select the audio input setting from the setting option, display the audio input setting interface, and directly read the volume gain displayed in the audio input setting interface as the initial volume gain.
[0192] The volume gain in the audio input setting interface can also be adjusted, that is, directly input or adjust the required volume gain in the audio input setting interface.
[0193] Step S404: Adjust the audio to be played according to the target audio gain to obtain the adjusted audio to be played, so as to play the adjusted audio to be played locally.
[0194] Combined with Figure 10 , in step S1001 of the present application, the first voice signal, the second voice signal, the local object intention, and the initial volume gain are combined to estimate the volume gain, so as to obtain the target volume gain of the audio to be played, and then the audio to be played is obtained, so as to adjust the audio to be played according to the target audio gain executed by step S1002 to obtain the adjusted audio to be played.
[0195] Among them, for adjusting the audio to be played according to the target audio gain, the audio to be played can be processed by AGC according to the target audio gain, so as to obtain the adjusted audio to be played.
[0196] After obtaining the adjusted audio to be played, the local device can play the adjusted audio to be played through the player.
[0197] It should be noted that the above embodiments of the present application can also be applied to the peer end, so that the adjusted audio to be played by the local end and the peer end matches the corresponding environment and the user volume, and at the same time realizes the accurate gain adjustment of the playing volume.
[0198] An embodiment of the present application provides an audio adjustment method, including: first collecting the ambient audio of the local end, then obtaining a first voice signal based on the ambient audio, then determining the target volume gain of the audio to be played based on the first voice signal, the audio to be played, and the intention of the local end object, and finally adjusting the audio to be played according to the target audio gain to obtain the adjusted audio to be played, so as to play the adjusted audio to be played at the local end. The present application obtains the first voice signal from the ambient audio, automatically matches the environment where the local end object is located with the target volume gain, expands the applicable scenarios of audio adjustment, enables more comprehensive consideration of environmental factors when performing audio adjustment, and thus improves the accuracy and adaptability of audio adjustment; in addition, adding the intention of the local end object when obtaining the target volume gain realizes precise gain adjustment of the playback volume of the audio to be played, enables more comprehensive consideration of the user's needs and intentions when performing audio adjustment, and further improves the accuracy of audio adjustment and user satisfaction.
[0199] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0200] The following is an apparatus embodiment of the present application. For the details not described in detail herein, reference may be made to the corresponding method embodiment above.
[0201] Figure 11 The structure diagram of an audio adjustment apparatus provided by an embodiment of the present application is shown. For the sake of convenience of description, only the parts related to the embodiment of the present application are shown. An audio adjustment apparatus includes a collection module 1101, a voice acquisition module 1102, a gain determination module 1103, and an audio adjustment module 1104, specifically as follows:
[0202] The collection module 1101 is configured to collect the ambient audio of the local end;
[0203] The voice acquisition module 1102 is configured to obtain a first voice signal based on the ambient audio;
[0204] The gain determination module 1103 is configured to determine the target volume gain of the audio to be played based on the first voice signal, the audio to be played, and the intention of the local end object, where the audio to be played is the audio stream of the peer end for voice communication with the local end, and the intention of the local end object is used to represent the expected volume adjustment trend of the object at the local end;
[0205] The audio adjustment module 1104 is configured to adjust the audio to be played according to the target audio gain to obtain the adjusted audio to be played, so as to play the adjusted audio to be played at the local end.
[0206] In an embodiment, the voice acquisition module 1102 is further configured to process the ambient audio to obtain an ambient audio signal;
[0207] Detect whether a first voice signal exists in the ambient audio signal;
[0208] If it exists, extract the first voice signal from the ambient audio signal.
[0209] In an embodiment, the voice acquisition module 1102 is further configured to calculate the energy proportion of the ambient audio signal;
[0210] Based on the size of the energy proportion and a first preset threshold, detect whether a first voice signal exists in the ambient audio signal.
[0211] In an embodiment, the voice acquisition module 1102 is further configured to obtain the frequency-domain signal corresponding to the ambient audio signal and calculate the total energy value of the frequency-domain signal;
[0212] Obtain the target frequency-domain signal in the frequency-domain signal and calculate the total energy value of the target frequency-domain signal, where the target frequency-domain signal is a signal with a preset frequency;
[0213] Divide the total energy value of the target frequency-domain signal by the total energy value of the frequency-domain signal to obtain the energy proportion of the ambient audio signal.
[0214] In an embodiment, the gain determination module 1103 is further configured to calculate the target energy value of the first voice signal;
[0215] Obtain the to-be-played audio signal corresponding to the to-be-played audio and detect whether a second voice signal exists in the to-be-played audio signal;
[0216] If it exists, calculate the target energy value of the second voice signal;
[0217] Based on the target energy value of the first voice signal, the target energy value of the second voice signal, and the local object intention, determine the target volume gain of the to-be-played audio;
[0218] Wherein, the first voice signal is the voice signal of the local object, and the second voice signal is the voice signal of the remote object.
[0219] In an embodiment, the gain determination module 1103 is further configured to calculate the initial energy value of the first voice signal;
[0220] Obtain the first preset volume gain;
[0221] Divide the initial energy value of the first voice signal by the first preset volume gain to obtain the target energy value of the first voice signal.
[0222] In one embodiment, the gain determination module 1103 is further configured to calculate the initial energy value of the second voice signal;
[0223] Obtain a second preset volume gain, where the second preset volume gain corresponds to the initial energy value of the second voice signal;
[0224] Sum the initial energy value of the second voice signal and the second preset volume gain to obtain the target energy value of the second voice signal.
[0225] In one embodiment, the gain determination module 1103 is further configured to compare the target energy value of the first voice signal with the target energy value of the second voice signal to obtain the initial volume gain of the audio to be played;
[0226] Obtain the intention of the local object;
[0227] Based on the intention of the local object and the initial volume gain, determine the target volume gain of the audio to be played.
[0228] In one embodiment, the gain determination module 1103 is further configured to subtract the target energy value of the second voice signal from the target energy value of the first voice signal to obtain a difference;
[0229] Use the difference as the initial volume gain of the audio to be played.
[0230] In one embodiment, the gain determination module 1103 is further configured to obtain the volume trend of the volume button and the playback volume trend of the application within a preset time;
[0231] Based on the volume trend of the volume button and the playback volume trend of the application, determine the intention of the local object.
[0232] In one embodiment, if both the volume trend of the volume button and the playback volume trend of the application are rising, the gain determination module 1103 is further configured to determine that the intention of the local object is to increase the volume of the audio to be played;
[0233] If both the volume trend of the volume button and the playback volume trend of the application are falling, the intention of the local object is to decrease the volume of the audio to be played;
[0234] If both the volume trend of the volume button and the playback volume trend of the application remain unchanged, the intention of the local object is to keep the volume of the audio to be played.
[0235] In one embodiment, if the intention of the local object is to keep the volume of the audio to be played and the initial volume gain is greater than the second preset threshold, the gain determination module 1103 is further configured to adjust the initial volume gain and use the adjusted initial volume gain as the target volume gain of the audio to be played;
[0236] If the intention of the local object is to increase the volume of the audio to be played and the initial volume gain is less than or equal to the second preset threshold, set the target volume gain of the audio to be played to zero;
[0237] If the intention of the local object is to decrease the volume of the audio to be played and the initial volume gain is less than or equal to the second preset threshold, set the target volume gain of the audio to be played to zero.
[0238] This application Figure 12 provides a schematic diagram of a computer device. As Figure 12 shown, the computer device 12 of this embodiment includes: a processor 1201, a memory 1202, and a computer program 1203 stored in the memory 1202 and executable on the processor 1201. When the processor 1201 executes the computer program 1203, it implements the steps in the above-mentioned embodiments of various audio adjustment methods, such as Figure 4 the steps 401 to 404 shown. Alternatively, when the processor 1201 executes the computer program 1203, it implements the functions of each module / unit in the above-mentioned embodiments of various audio adjustment devices, such as Figure 10 the functions of the modules / units 1001 to 1004 shown.
[0239] This application also provides a readable storage medium. The readable storage medium stores a computer program, and when the computer program is executed by a processor, it is used to implement the audio adjustment methods provided by the above various embodiments.
[0240] Among them, the readable storage medium can be a computer storage medium or a communication medium. The communication medium includes any medium that facilitates the transmission of a computer program from one place to another. The computer storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer. For example, the readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuits (ASIC). In addition, the ASIC can be located in a user device. Of course, the processor and the readable storage medium can also exist as discrete components in a communication device. The readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0241] The present application also provides a program product, which includes execution instructions stored in a readable storage medium. At least one processor of the device can read the execution instructions from the readable storage medium, and the execution of the execution instructions by at least one processor enables the device to implement the audio adjustment method provided by the above various embodiments.
[0242] In the embodiment of the above device, it should be understood that the processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the present application can be directly embodied as being completed by the execution of a hardware processor, or can be completed by a combination of hardware and software modules in the processor.
[0243] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. An audio adjustment method, characterized in that: include: Collect the ambient audio of the local end; Based on the ambient audio, obtaining a first voice signal; Determining a target volume gain of the audio to be played based on the first voice signal, the audio to be played, and the local object intention, wherein the audio to be played is an audio stream of a peer end that performs voice communication with the local end, and the local object intention is used to characterize a volume adjustment trend expected by the object of the local end; The audio to be played is adjusted according to the target audio gain to obtain the adjusted audio to be played, and the adjusted audio to be played is played on the local end.
2. The audio adjustment method according to claim 1, characterized in that: The acquiring a first voice signal based on the ambient audio includes: Processing the ambient audio to obtain an ambient audio signal; Detecting whether there is a first voice signal in the ambient audio signal; If present, the first speech signal is extracted from the ambient audio signal.
3. The audio adjustment method according to claim 2, characterized in that: The detecting whether there is a first voice signal in the ambient audio signal includes: Calculating the energy ratio of the ambient audio signal; Based on the energy ratio to a first preset threshold, it is detected whether a first speech signal exists in the ambient audio signal.
4. The audio adjustment method according to claim 3, characterized in that: The calculating the energy proportion of the ambient audio signal includes: Obtaining a frequency domain signal corresponding to the ambient audio signal, and calculating a total energy value of the frequency domain signal; Acquire a target frequency domain signal in the frequency domain signal, and calculate a total energy value of the target frequency domain signal, wherein the target frequency domain signal is a signal of a preset frequency; The total energy value of the target frequency domain signal is divided by the total energy value of the frequency domain signal to obtain the energy proportion of the ambient audio signal.
5. The audio adjustment method according to claim 1, characterized in that: The step of determining a target volume gain of the audio to be played based on the first voice signal, the audio to be played, and the local object intention includes: Calculating a target energy value of the first speech signal; Acquire an audio signal to be played corresponding to the audio to be played, and detect whether a second voice signal exists in the audio signal to be played; If so, calculating a target energy value of the second speech signal; Determining a target volume gain of the audio to be played based on a target energy value of the first voice signal, a target energy value of the second voice signal, and the local object intention; The first voice signal is a voice signal of a local object, and the second voice signal is a voice signal of a remote object.
6. The audio adjustment method according to claim 5, characterized in that: The calculating a target energy value of the first speech signal includes: Calculating an initial energy value of the first speech signal; Obtaining a first preset volume gain; The target energy value of the first speech signal is obtained by dividing the initial energy value of the first speech signal by the first preset volume gain.
7. The audio adjustment method according to claim 5, characterized in that: The calculating a target energy value of the second speech signal comprises: Calculating an initial energy value of the second speech signal; Acquire a second preset volume gain, wherein the second preset volume gain corresponds to an initial energy value of the second speech signal; The initial energy value of the second speech signal is summed with the second preset volume gain to obtain a target energy value of the second speech signal.
8. The audio adjustment method according to claim 5, characterized in that: The determining, based on the target energy value of the first voice signal, the target energy value of the second voice signal, and the local object intention, a target volume gain of the audio to be played includes: Compare the target energy value of the first speech signal and the target energy value of the second speech signal to obtain an initial volume gain of the audio to be played; Obtaining the local object intention; Based on the local object intention and the initial volume gain, a target volume gain of the audio to be played is determined.
9. The audio adjustment method according to claim 8, characterized in that: The initial volume gain of the audio to be played obtained by comparing the target energy value of the first voice signal and the target energy value of the second voice signal includes: Subtracting a target energy value of the first speech signal from a target energy value of the second speech signal to obtain a difference; The difference is used as the initial volume gain of the audio to be played.
10. The audio adjustment method according to claim 8, characterized in that: The acquiring the local object intention includes: Get the volume trend of the volume button and the playback volume trend of the application within a preset time; The local object intention is determined based on the volume trend of the volume button and the playback volume trend of the application.
11. The audio adjustment method according to claim 10, characterized in that: The determining the local object intention based on the volume trend of the volume button and the playback volume trend of the application program includes: If the volume trend of the volume button and the playback volume trend of the application are both increasing, the local object intends to increase the volume of the audio to be played; If the volume trend of the volume button and the playback volume trend of the application are both decreasing, the local object intends to reduce the volume of the audio to be played; If the volume trend of the volume button and the playback volume trend of the application program remain unchanged, the local object intends to maintain the volume of the audio to be played.
12. The audio adjustment method according to claim 11, characterized in that: The determining, based on the local object intention and the initial volume gain, a target volume gain of the audio to be played, includes: If the local object intends to maintain the volume of the audio to be played, and the initial volume gain is greater than a second preset threshold, adjusting the initial volume gain, and using the adjusted initial volume gain as the target volume gain of the audio to be played; If the local object intends to increase the volume of the audio to be played and the initial volume gain is less than or equal to the second preset threshold, setting the target volume gain of the audio to be played to zero; If the local object intends to reduce the volume of the audio to be played, and the initial volume gain is less than or equal to the second preset threshold, the target volume gain of the audio to be played is set to zero.
13. An audio adjustment device, characterized in that: include: The acquisition module is used to collect the ambient audio of the local end; A human voice acquisition module, used to acquire a first voice signal based on the environmental audio; a gain determination module, configured to determine a target volume gain of the audio to be played based on the first voice signal, the audio to be played, and the object intention of the local end, wherein the audio to be played is an audio stream of the other end that performs voice communication with the local end, and the object intention of the local end is used to characterize the volume adjustment trend expected by the object of the local end; The audio adjustment module is used to adjust the audio to be played according to the target audio gain to obtain the adjusted audio to be played, and play the adjusted audio to be played on the local end.
14. A computer device, characterized in that: comprising a memory, and one or more processors communicatively connected to the memory; The memory stores instructions that can be executed by the one or more processors, and the instructions are executed by the one or more processors to enable the one or more processors to implement the audio adjustment method as described in claims 1 to 12.
15. A computer-readable storage medium, characterized in that: The invention comprises a program or an instruction, and when the program or the instruction is executed on a computer, the audio adjustment method according to claims 1 to 12 is implemented.