Audio processing method, apparatus, computing device, and medium

By acquiring the current and historical energy distribution information of audio frames, the target audio adjustment information is determined, which solves the problem of differences in listening among users, realizes adaptive personalized audio adjustment, and improves the audio listening experience.

CN115713942BActive Publication Date: 2026-04-24HANGZHOU NETEASE ZHIQI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU NETEASE ZHIQI TECH CO LTD
Filing Date
2022-11-01
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Different users have different needs for the volume and timbre of the headphone's downstream audio, and existing technology lacks adaptive and personalized audio adjustment methods.

Method used

By acquiring the current and historical energy distribution information of audio frames, the target audio adjustment information is determined, and the energy values ​​of subsequent audio frames in the audio sequence are adjusted accordingly to achieve adaptive personalized audio adjustment.

Benefits of technology

It offers adaptive and personalized audio adjustment methods to meet the listening needs of different users and improve the audio listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115713942B_ABST
    Figure CN115713942B_ABST
Patent Text Reader

Abstract

The present disclosure provides an audio processing method, device, computing equipment and medium. In response to a volume adjustment operation for a to-be-processed audio sequence, target audio adjustment information is determined based on current energy distribution information and historical energy distribution information corresponding to a target audio frame in the to-be-processed audio sequence, so as to implement audio processing based on the target audio adjustment information. Since the current energy distribution information corresponds to an audio frame that the user is currently listening to, and the historical energy distribution information corresponds to an audio frame that the user has listened to in the past, both of which have been perceived, accepted and approved by the user, the target audio adjustment information is determined based on the current energy distribution information and the historical energy distribution information, and the adjustment of the audio sequence is implemented based on the target audio adjustment information, which can meet the personalized listening needs of the user, and thus the scheme provided by the present disclosure can provide an adaptive personalized audio adjustment mode for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure relate to the field of multimedia data processing technology, and more specifically, the embodiments of this disclosure relate to an audio processing method, apparatus, computing device, and medium. Background Technology

[0002] This section is intended to provide background or context for the embodiments of this disclosure as set forth in the claims. The description herein is not intended to be a prior art simply because it is included in this section.

[0003] With the continuous development of computer technology and network transmission technology, various types of applications such as online meetings, online education, and interactive entertainment have emerged. In order to obtain an immersive audio listening experience, more and more users tend to use headphones to listen to audio.

[0004] However, due to differences in listening comfort and preferences among different users, their needs for headphone downlink volume and timbre will vary. Therefore, there is an urgent need for an audio processing method to provide users with adaptive and personalized audio adjustment. Summary of the Invention

[0005] In view of the fact that different users have different needs for the volume and timbre of the downlink audio in the related technology, the embodiments of this disclosure provide at least one audio processing method, apparatus, computing device and medium.

[0006] In a first aspect of this disclosure, an audio processing method is provided, the method comprising:

[0007] In response to a volume adjustment operation on the audio sequence to be processed, the current energy distribution information and historical energy distribution information of the target audio frame in the audio sequence to be processed are obtained. The target audio frame is the audio frame located at a preset frame number among the audio frames after volume adjustment. The current energy distribution information is used to indicate the energy value of the target audio frame in multiple frequency sub-bands, and the historical energy distribution information is used to indicate the energy value of the previous audio frame of the target audio frame in multiple frequency sub-bands.

[0008] Based on the current energy distribution information and historical energy distribution information corresponding to the target audio frame, the target audio adjustment information is determined. The target audio adjustment information is used to adjust the energy value of any audio frame located after the target audio frame at multiple frequency points.

[0009] Based on the target audio adjustment information, the energy values ​​of each audio frame following the target audio frame in the audio sequence to be processed are adjusted at multiple frequency points.

[0010] In one embodiment of this disclosure, target audio adjustment information is determined based on the current energy distribution information and historical energy distribution information corresponding to the target audio frame, including:

[0011] Based on the current energy distribution information and the historical energy distribution information, the first energy distribution difference information is determined. The first energy distribution difference information is used to indicate the energy difference between the current energy distribution information and the historical energy distribution information in multiple frequency domain sub-bands.

[0012] Based on historical audio adjustment information and first energy distribution difference information, target audio adjustment information is determined. Historical audio adjustment information refers to the audio adjustment information determined in response to the previous volume adjustment operation on the audio sequence to be processed.

[0013] In one embodiment of this disclosure, determining target audio adjustment information based on historical audio adjustment information and first energy distribution difference information includes:

[0014] The first energy distribution difference information is weighted using a first set parameter, and the target audio adjustment information is determined based on the weighted first energy distribution difference information and historical audio adjustment information.

[0015] In one embodiment of this disclosure, the first energy distribution difference information is weighted using a first set parameter, and the target audio adjustment information is determined based on the weighted first energy distribution difference information and historical audio adjustment information, including:

[0016] Using the first set parameter as the weight of the first energy distribution difference information, the first energy distribution difference information is weighted to obtain the weighted first energy distribution information.

[0017] The sum of the weighted first energy distribution information and the historical audio adjustment information is determined as the target audio adjustment information.

[0018] In one embodiment of this disclosure, before obtaining the current energy distribution information and historical energy distribution information corresponding to the target audio frame in the audio sequence to be processed in response to a volume adjustment operation on the audio sequence to be processed, the method further includes:

[0019] For each audio frame received in the audio sequence to be processed, the energy value of the audio frame in each frequency subband and the total energy value of the audio frame are determined.

[0020] If the total energy value of the received audio frame is greater than the set energy threshold, the current energy distribution information of the audio frame is determined based on the historical energy distribution information of the audio frame and the energy value of the audio frame in multiple frequency sub-bands.

[0021] In one embodiment of this disclosure, determining the energy value of an audio frame in each frequency subband and the overall energy value of the audio frame includes:

[0022] Determine the signal energy value of the audio frame in each frequency subband;

[0023] The total energy value of the audio frame is obtained by summing the signal energy values ​​of each frequency sub-band.

[0024] In one embodiment of this disclosure, determining the current energy distribution information of an audio frame based on its historical energy distribution information and energy values ​​across multiple frequency subbands includes:

[0025] The energy values ​​of the audio frame in multiple frequency sub-bands are weighted by the second set parameters and the energy values ​​in the corresponding frequency sub-bands indicated by the historical energy distribution information, and the first energy distribution information is determined based on the result of the weighting process.

[0026] The current energy distribution information is determined based on the energy values ​​in multiple frequency sub-bands indicated by the first energy distribution information and the sum of the energy values ​​in multiple frequency sub-bands indicated by the first energy distribution information.

[0027] The second setting parameter is determined based on the sampling rate and preset duration of the audio sequence to be processed. The preset duration is the duration required to smooth the spectrum of each audio frame in the audio sequence to be processed.

[0028] In one embodiment of this disclosure, the energy values ​​of an audio frame in multiple frequency sub-bands are weighted with the energy values ​​in the corresponding frequency sub-bands indicated by historical energy distribution information using a second set parameter, and the first energy distribution information is determined based on the result of the weighting process, including:

[0029] The second set parameter is used as the weight of the energy value of the audio frame in multiple frequency domain sub-bands, and the energy value of the audio frame in multiple frequency domain sub-bands is weighted. The difference between the set parameter value and the second set parameter is used as the weight of the historical energy distribution information, and the historical energy distribution information is weighted.

[0030] The sum of the results obtained from the two weighted processing is determined as the first energy distribution information.

[0031] In one embodiment of this disclosure, in response to a volume adjustment operation on an audio sequence to be processed, current energy distribution information and historical energy distribution information corresponding to a target audio frame in the audio sequence to be processed are obtained, including:

[0032] In response to a volume adjustment operation on the audio sequence to be processed, if the total energy value of the received audio frame is greater than a set energy threshold, the step of determining the current energy distribution information of the audio frame based on the historical energy distribution information corresponding to the audio frame and the energy value of the audio frame in multiple frequency sub-bands is performed.

[0033] If the target audio frame is received and its current energy distribution information has been determined, the step of obtaining the current energy distribution information and historical energy distribution information of the target audio frame in the audio sequence to be processed is executed.

[0034] In one embodiment of this disclosure, before determining the current energy distribution information of an audio frame based on its historical energy distribution information and energy values ​​in multiple frequency subbands when the total energy value of the received audio frame exceeds a set energy threshold, the method further includes:

[0035] If no volume adjustment operation is detected for the audio sequence to be processed, historical audio adjustment information is determined as the target audio adjustment information.

[0036] In one embodiment of this disclosure, without performing volume adjustment operations on the audio sequence to be processed, vector curves with values ​​of 1 for multiple frequency domain sub-bands are used as historical audio adjustment information.

[0037] In one embodiment of this disclosure, the method further includes:

[0038] The noise energy value of the target audio frame is obtained. The noise energy value is determined by the sending end of the audio sequence to be processed based on the target audio frame.

[0039] Before obtaining the current energy distribution information and historical energy distribution information corresponding to the target audio frame in the audio sequence to be processed in response to a volume adjustment operation on the audio sequence to be processed, the method further includes:

[0040] If the system volume level has not reached the maximum volume level and the noise energy value is not greater than the first energy threshold, the following steps are performed in response to the volume adjustment operation for the audio sequence to be processed: obtaining the current energy distribution information and historical energy distribution information corresponding to the target audio frame in the audio sequence to be processed.

[0041] In one embodiment of this disclosure, the method further includes:

[0042] When the system volume level has not reached the maximum volume level and the noise energy value is greater than the first energy threshold but less than the second energy threshold, the second energy distribution difference information is determined based on the current energy distribution information of the target audio frame and the noise energy distribution information of the noise signal included in the target audio frame. The noise energy distribution information is used to indicate the energy value of the noise signal in multiple frequency domain sub-bands, and the second energy distribution difference information is used to indicate the energy difference between the energy distribution information and the noise energy distribution information in multiple frequency domain sub-bands.

[0043] Based on the second energy distribution difference information, the overall energy value of the target audio frame, the noise energy value, and the first energy distribution information, the target audio adjustment information is determined. The first energy distribution information is determined based on the weighted processing result of the energy values ​​of the audio frame in multiple frequency sub-bands and the energy values ​​in the corresponding frequency sub-bands indicated by the historical energy distribution information.

[0044] In one embodiment of this disclosure, the method further includes:

[0045] If the system volume level does not reach the maximum volume level and the noise energy value is not less than the second energy threshold and not greater than the third energy threshold, the target energy difference is determined based on the total energy value and noise energy value of the target audio frame. The total energy value of the target audio frame is the sum of the signal energy values ​​of the target audio frame in each frequency domain sub-band.

[0046] Based on the target energy difference, the overall energy value of the target audio frame, the noise energy value, the first energy distribution information, and the first energy adjustment information, the target audio adjustment information is determined. The first energy adjustment information is a first preset energy value used to raise the energy value of the target audio frame at each frequency point.

[0047] In one embodiment of this disclosure, the method further includes:

[0048] When the system volume level has reached the maximum volume level and the noise energy value is not less than the third energy threshold, the second energy distribution difference information is determined based on the current energy distribution information of the target audio frame and the noise energy distribution information of the noise signal included in the target audio frame.

[0049] Based on the second energy distribution difference information, the overall energy value of the target audio frame, the noise energy value, the first energy distribution information, and the second energy adjustment information, the target audio adjustment information is determined. The second energy adjustment information is a second preset energy value used to raise the energy value of the audio frame at each frequency point.

[0050] In one embodiment of this disclosure, after adjusting the energy values ​​of audio frames following the target audio frame in the audio sequence to be processed at multiple frequency points based on target audio adjustment information, the method further includes at least one of the following:

[0051] If the energy value of any audio frame at any frequency point is greater than the first set threshold after adjustment, the energy value of the audio frame at the frequency point is adjusted so that the energy value of the audio frame at the frequency point is less than or equal to the first set threshold.

[0052] If the energy value of any adjusted audio frame at any frequency point is less than the second set threshold, the energy value of the audio frame at the frequency point is adjusted so that the energy value of the audio frame at the frequency point is greater than or equal to the second set threshold.

[0053] In a second aspect of this disclosure, an audio processing apparatus is provided, the apparatus comprising:

[0054] The acquisition module is used to respond to the volume adjustment operation for the audio sequence to be processed, and to acquire the current energy distribution information and historical energy distribution information of the target audio frame in the audio sequence to be processed. The target audio frame is the audio frame located at a preset frame number among the audio frames after volume adjustment. The current energy distribution information is used to indicate the energy value of the target audio frame in multiple frequency domain sub-bands, and the historical energy distribution information is used to indicate the energy value of the previous audio frame of the target audio frame in multiple frequency domain sub-bands.

[0055] The determination module is used to determine the target audio adjustment information based on the current energy distribution information and historical energy distribution information corresponding to the target audio frame. The target audio adjustment information is used to adjust the energy value of any audio frame located after the target audio frame at multiple frequency points.

[0056] The adjustment module is used to adjust the energy values ​​of each audio frame in the audio sequence to be processed at multiple frequency points based on the target audio adjustment information.

[0057] In one embodiment of this disclosure, the determining module, when determining target audio adjustment information based on the current energy distribution information and historical energy distribution information corresponding to the target audio frame, is configured to:

[0058] Based on the current energy distribution information and the historical energy distribution information, the first energy distribution difference information is determined. The first energy distribution difference information is used to indicate the energy difference between the current energy distribution information and the historical energy distribution information in multiple frequency domain sub-bands.

[0059] Based on historical audio adjustment information and first energy distribution difference information, target audio adjustment information is determined. Historical audio adjustment information refers to the audio adjustment information determined in response to the previous volume adjustment operation on the audio sequence to be processed.

[0060] In one embodiment of this disclosure, the determining module, when determining target audio adjustment information based on historical audio adjustment information and first energy distribution difference information, is configured to:

[0061] The first energy distribution difference information is weighted using a first set parameter, and the target audio adjustment information is determined based on the weighted first energy distribution difference information and historical audio adjustment information.

[0062] In one embodiment of this disclosure, the determining module, when weighting the first energy distribution difference information with a first set parameter and determining the target audio adjustment information based on the weighted first energy distribution difference information and historical audio adjustment information, is configured to:

[0063] Using the first set parameter as the weight of the first energy distribution difference information, the first energy distribution difference information is weighted to obtain the weighted first energy distribution information.

[0064] The sum of the weighted first energy distribution information and the historical audio adjustment information is determined as the target audio adjustment information.

[0065] In one embodiment of this disclosure, the determining module is further configured to determine the energy value of the audio frame in each frequency domain subband and the total energy value of the audio frame for each audio frame received in the audio sequence to be processed.

[0066] The determination module is also used to determine the current energy distribution information of the audio frame based on the historical energy distribution information corresponding to the audio frame and the energy value of the audio frame in multiple frequency sub-bands when the total energy value of the received audio frame is greater than a set energy threshold.

[0067] In one embodiment of this disclosure, the determining module, when determining the energy value of the audio frame in each frequency domain subband and the overall energy value of the audio frame, is configured to:

[0068] Determine the signal energy value of the audio frame in each frequency subband;

[0069] The total energy value of the audio frame is obtained by summing the signal energy values ​​of each frequency sub-band.

[0070] In one embodiment of this disclosure, the determining module, when determining the current energy distribution information of an audio frame based on the historical energy distribution information corresponding to the audio frame and the energy values ​​of the audio frame in multiple frequency sub-bands, is configured to:

[0071] The energy values ​​of the audio frame in multiple frequency sub-bands are weighted by the second set parameters and the energy values ​​in the corresponding frequency sub-bands indicated by the historical energy distribution information, and the first energy distribution information is determined based on the result of the weighting process.

[0072] The current energy distribution information is determined based on the energy values ​​in multiple frequency sub-bands indicated by the first energy distribution information and the sum of the energy values ​​in multiple frequency sub-bands indicated by the first energy distribution information.

[0073] The second setting parameter is determined based on the sampling rate and preset duration of the audio sequence to be processed. The preset duration is the duration required to smooth the spectrum of each audio frame in the audio sequence to be processed.

[0074] In one embodiment of this disclosure, the determining module, when performing weighted processing on the energy values ​​of the audio frame in multiple frequency sub-bands and the energy values ​​in the corresponding frequency sub-bands indicated by historical energy distribution information using a second set parameter, and determining the first energy distribution information based on the result of the weighted processing, is configured to:

[0075] The second set parameter is used as the weight of the energy value of the audio frame in multiple frequency domain sub-bands, and the energy value of the audio frame in multiple frequency domain sub-bands is weighted. The difference between the set parameter value and the second set parameter is used as the weight of the historical energy distribution information, and the historical energy distribution information is weighted.

[0076] The sum of the results obtained from the two weighted processing is determined as the first energy distribution information.

[0077] In one embodiment of this disclosure, the acquisition module, when acquiring current energy distribution information and historical energy distribution information corresponding to a target audio frame in the audio sequence to be processed in response to a volume adjustment operation for the audio sequence to be processed, is configured to:

[0078] In response to a volume adjustment operation on the audio sequence to be processed, if the total energy value of the received audio frame is greater than a set energy threshold, the step of determining the current energy distribution information of the audio frame based on the historical energy distribution information corresponding to the audio frame and the energy value of the audio frame in multiple frequency sub-bands is performed.

[0079] If the target audio frame is received and its current energy distribution information has been determined, the step of obtaining the current energy distribution information and historical energy distribution information of the target audio frame in the audio sequence to be processed is executed.

[0080] In one embodiment of this disclosure, the determining module is further configured to determine historical audio adjustment information as target audio adjustment information when no volume adjustment operation for the audio sequence to be processed is detected.

[0081] In one embodiment of this disclosure, without performing volume adjustment operations on the audio sequence to be processed, vector curves with values ​​of 1 for multiple frequency domain sub-bands are used as historical audio adjustment information.

[0082] In one embodiment of this disclosure, the acquisition module is further configured to acquire the noise energy value of the target audio frame, wherein the noise energy value is determined by the transmitting end of the audio sequence to be processed based on the target audio frame;

[0083] The acquisition module is further configured to perform a step of acquiring the current energy distribution information and historical energy distribution information corresponding to the target audio frame in the audio sequence to be processed, in response to a volume adjustment operation for the audio sequence to be processed, when the system volume level has not reached the maximum volume level and the noise energy value is not greater than the first energy threshold.

[0084] In one embodiment of this disclosure, the determining module is further configured to determine second energy distribution difference information based on the current energy distribution information of the target audio frame and the noise energy distribution information of the noise signal included in the target audio frame when the system volume level has not reached the maximum volume level and the noise energy value is greater than the first energy threshold and less than the second energy threshold. The noise energy distribution information is used to indicate the energy value of the noise signal in multiple frequency domain sub-bands, and the second energy distribution difference information is used to indicate the energy difference between the energy distribution information and the noise energy distribution information in multiple frequency domain sub-bands.

[0085] The determination module is also used to determine target audio adjustment information based on the second energy distribution difference information, the overall energy value of the target audio frame, the noise energy value, and the first energy distribution information. The first energy distribution information is determined based on the weighted processing result of the energy values ​​of the audio frame in multiple frequency sub-bands and the energy values ​​in the corresponding frequency sub-bands indicated by the historical energy distribution information.

[0086] In one embodiment of this disclosure, the determining module is further configured to determine a target energy difference based on the total energy value and noise energy value of the target audio frame when the system volume level has not reached the maximum volume level and the noise energy value is not less than the second energy threshold and not greater than the third energy threshold. The total energy value of the target audio frame is the sum of the signal energy values ​​of the target audio frame in each frequency domain sub-band.

[0087] The determining module is also used to determine target audio adjustment information based on the target energy difference, the overall energy value of the target audio frame, the noise energy value, the first energy distribution information, and the first energy adjustment information. The first energy adjustment information is a first preset energy value used to raise the energy value of the target audio frame at each frequency point.

[0088] In one embodiment of this disclosure, the determining module is further configured to determine second energy distribution difference information based on the current energy distribution information of the target audio frame and the noise energy distribution information of the noise signal included in the target audio frame when the system volume level has reached the maximum volume level and the noise energy value is not less than the third energy threshold.

[0089] The determination module is also used to determine the target audio adjustment information based on the second energy distribution difference information, the overall energy value of the target audio frame, the noise energy value, the first energy distribution information, and the second energy adjustment information. The second energy adjustment information is a second preset energy value used to raise the energy value of the audio frame at each frequency point.

[0090] In one embodiment of this disclosure, the adjustment module is further configured to adjust the energy value of the audio frame at any frequency point when the energy value of any adjusted audio frame at any frequency point is greater than the first set threshold, so that the energy value of the audio frame at the frequency point is less than or equal to the first set threshold.

[0091] The adjustment module is also used to adjust the energy value of the audio frame at any frequency point when the energy value of any adjusted audio frame at any frequency point is less than the second set threshold, so that the energy value of the audio frame at the frequency point is greater than or equal to the second set threshold.

[0092] In a third aspect of the present disclosure, a computing device is provided, the computing device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to perform the operations performed by the audio processing method provided in the first aspect and any embodiment of the first aspect.

[0093] In a third aspect of this disclosure, a computer-readable storage medium is provided, on which a program is stored, the program being executed by a processor as performed by the audio processing method provided in the first aspect and any embodiment of the first aspect.

[0094] In a fifth aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, performs the operations performed by the audio processing method provided in the first aspect and any embodiment of the first aspect.

[0095] This disclosure, in response to a volume adjustment operation on an audio sequence to be processed, obtains the current energy distribution information and historical energy distribution information corresponding to the target audio frame in the audio sequence to be processed. Based on the current and historical energy distribution information of the target audio frame, target audio adjustment information is determined. This allows for the adjustment of the energy values ​​of each audio frame following the target audio frame in the audio sequence to be processed at multiple frequency points. Since the current energy distribution information corresponds to the audio frame the user is currently listening to, and the historical energy distribution information corresponds to audio frames the user has previously listened to, both of which are perceived and accepted by the user (if not accepted, the user would inevitably make adjustments). Therefore, determining the target audio adjustment information using the current and historical energy distribution information, and adjusting the audio sequence based on the target audio adjustment information, can meet the user's personalized listening needs. Consequently, the solution provided by this disclosure can offer users an adaptive and personalized audio adjustment method. Attached Figure Description

[0096] The above and other objects, features, and advantages of this disclosure will become readily apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. Several embodiments of this disclosure are illustrated in the drawings by way of example and not limitation, in which:

[0097] Figure 1 This is a flowchart illustrating an audio processing method according to an exemplary embodiment of the present disclosure;

[0098] Figure 2 This disclosure illustrates a time-domain waveform of an audio frame according to an exemplary embodiment;

[0099] Figure 3 This disclosure illustrates a spectrogram of an audio frame according to an exemplary embodiment;

[0100] Figure 4 This disclosure illustrates a subband spectrogram of an audio frame according to an exemplary embodiment;

[0101] Figure 5 This disclosure is a schematic flowchart illustrating an audio processing procedure based on the total energy value of an audio frame according to an exemplary embodiment.

[0102] Figure 6 This is a schematic flowchart illustrating an audio processing procedure according to an exemplary embodiment of the present disclosure;

[0103] Figure 7This is a schematic flowchart illustrating an audio adjustment process under different circumstances according to an exemplary embodiment of the present disclosure;

[0104] Figure 8 This is a block diagram illustrating an audio processing apparatus according to an exemplary embodiment of the present disclosure;

[0105] Figure 9 This is a schematic diagram illustrating a computer-readable storage medium according to an exemplary embodiment of the present disclosure;

[0106] Figure 10 This is a schematic diagram of the structure of a computing device according to an exemplary embodiment of the present disclosure;

[0107] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation

[0108] The principles and spirit of this disclosure will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are given merely to enable those skilled in the art to better understand and implement this disclosure, and are not intended to limit the scope of this disclosure in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.

[0109] Those skilled in the art will recognize that embodiments of this disclosure can be implemented as a system, apparatus, device, method, or computer program product. Therefore, this disclosure can be specifically implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0110] According to embodiments of this disclosure, an audio processing method is proposed for processing any audio frame of an audio sequence to be processed, thereby achieving real-time audio processing of each audio frame and optimizing the playback effect of the audio frame.

[0111] The above-mentioned audio processing method can be executed by a computing device, which can be a terminal device, such as a desktop computer, portable computer, smartphone, tablet computer, smartwatch, Moving Picture Experts Group Audio Layer III (MP3) player, Moving Picture Experts Group Audio Layer IV (MP4) player, etc. Alternatively, the computing device can be headphones, such as wired headphones, wireless headphones (such as wireless Bluetooth headphones), over-ear headphones, etc. Alternatively, the computing device can be a hearing aid. This disclosure does not limit the type of computing device.

[0112] Taking a computing device as an example of a terminal device, the terminal device can receive audio sequences sent by other devices in real time and treat the received audio sequences as audio sequences to be processed. Upon receiving each audio frame from the audio sequence to be processed, the terminal device can process the received audio frame using the solution provided in this disclosure, thereby achieving real-time processing of each audio frame in the audio sequence to be processed. It should be noted that after processing the audio frame, it can be played back. Optionally, the computing device can directly play the audio frame using its built-in audio playback component, or the computing device can be connected to headphones to play the audio frame through headphones.

[0113] For example, taking headphones as a computing device, the headphones can communicate with the terminal device via wired or wireless connections. When the terminal device receives an audio sequence from another device, it can send the received audio sequence to the connected headphones. The headphones can then use the received audio sequence as a processing sequence. Upon receiving each audio frame from the processing sequence, the headphones can process the received audio frame using the scheme provided in this disclosure, achieving real-time processing of each audio frame in the processing sequence. It should be noted that after processing the audio frames, the headphones can then play the audio frames.

[0114] It should be noted that regardless of the implementation scenario described above, both the transmitting and receiving ends of the audio sequence are involved. For the transmitting end, the audio signal collected by the audio input device can be called the near-end signal, which, after processing, can be sent as an uplink signal to the other end or multiple ends of the call. For the receiving end, the audio signal received through the network can be called the far-end signal, which can be called the downlink signal.

[0115] It should be noted that the above application scenarios are shown only to facilitate understanding of the spirit and principles of this disclosure. The embodiments of this disclosure are not limited in any way; on the contrary, the embodiments of this disclosure can be applied to any applicable scenario. For example, the embodiments of this disclosure can also be applied to music scenarios such as listening to songs, and this disclosure does not limit this application.

[0116] The following section, in conjunction with the above descriptions of application scenarios, provides further information. Figure 1 This describes an audio processing method provided according to exemplary embodiments of the present disclosure.

[0117] See Figure 1 , Figure 1 This is a flowchart illustrating an audio processing method according to an exemplary embodiment of the present disclosure, the method comprising:

[0118] S101. In response to the volume adjustment operation for the audio sequence to be processed, obtain the current energy distribution information and historical energy distribution information corresponding to the target audio frame in the audio sequence to be processed. The target audio frame is the audio frame located at a preset frame number among the audio frames after volume adjustment. The current energy distribution information is used to indicate the energy value of the target audio frame in multiple frequency sub-bands, and the historical energy distribution information is used to indicate the energy value of the previous audio frame of the target audio frame in multiple frequency sub-bands.

[0119] Optionally, the volume adjustment operation may include an increase volume operation and a decrease volume operation. For example, a computing device may be equipped with an increase volume button and a decrease volume button, and the user can trigger an increase volume operation by pressing the increase volume button and a decrease volume operation by pressing the decrease volume button.

[0120] It should be noted that since the reception of each audio frame in the audio sequence to be processed is carried out in real time, when the user triggers a volume adjustment operation, the computing device can respond to the user's volume adjustment operation and adjust the system volume of the computing device so that the system volume used when playing the audio frame received after the volume adjustment operation is triggered is the system volume adjusted according to the volume adjustment operation.

[0121] Additionally, it should be noted that the frequency domain sub-bands can be pre-divided according to the sampling rate of the audio sequence to be processed. Optionally, the number of frequency domain sub-bands to be divided can be determined based on the sampling rate of the audio sequence to be processed. Then, according to the determined number of sub-bands and the effective bandwidth (or simply bandwidth) of the audio sequence to be processed, the audio sequence to be processed is divided into multiple frequency domain sub-bands that meet the determined number of sub-bands. Furthermore, the bandwidth of each frequency domain sub-band is consistent, or in other words, the number of frequency points included in each frequency domain sub-band is consistent.

[0122] Generally speaking, the difference between the number of sub-bands and the sampling rate of the audio sequence to be processed is a set value. Let the sampling rate of the audio sequence to be processed be fHz, then the number of sub-bands can be [fn, f+n], where n is the set value. The set value can be any positive integer value, and this disclosure does not limit it.

[0123] Optionally, the setting value can be 4. For example, for an audio sequence with a sampling rate of 16 kHz, the effective bandwidth is generally 8 kHz, and the audio sequence can be divided into 12 to 20 frequency domain sub-bands; for another example, for an audio sequence with a sampling rate of 48 kHz, the effective bandwidth is generally 24 kHz, and the audio sequence can be divided into 20 to 28 frequency domain sub-bands.

[0124] The above are only two exemplary methods for dividing frequency domain sub-bands. In many other possible implementations, frequency domain sub-bands can be divided according to actual technical requirements, and this disclosure does not limit this.

[0125] An audio sequence includes multiple audio frames, and an audio frame includes multiple frequency points. By dividing the audio frame into multiple frequency domain sub-bands, the multiple frequency points corresponding to the audio frame can be divided into frequency domain sub-bands of the corresponding frequencies. Optionally, the number of frequency points included in different frequency domain sub-bands can be the same or different, and this disclosure does not limit this.

[0126] Optionally, the energy distribution information can be a sub-band spectrum curve. It should be noted that for any audio frame, a Fast Fourier Transform (FFT) can be used to analyze the audio frame to obtain the signal energy magnitude of each frequency (frequency point) component. Then, based on the divided frequency domain sub-bands, the energy of the frequency points belonging to the same frequency domain sub-band is summed to obtain the signal energy magnitude corresponding to each frequency domain sub-band. This signal energy magnitude can generally be represented by a curve, which is the sub-band spectrum curve (or sub-band energy distribution curve). In this sub-band spectrum curve, the horizontal axis can be the frequency range corresponding to the frequency domain sub-band (such as 1-2kHz, 2-3kHz, 3-4kHz, ...), and the vertical axis can be the signal energy value corresponding to the frequency domain sub-band; or, the horizontal axis of the sub-band spectrum curve can be the frequency domain sub-band identifier, and the vertical axis can be the signal energy value corresponding to the frequency domain sub-band. Optionally, the divided frequency domain sub-bands can be numbered, and the numbering result can be used as the frequency domain sub-band identifier. For example, the frequency domain sub-band corresponding to 1-2kHz can be numbered as 1, the frequency domain sub-band corresponding to 2-3kHz can be numbered as 2, the frequency domain sub-band corresponding to 3-4kHz can be numbered as 3, and so on. The number obtained by numbering can be used as the frequency domain sub-band identifier.

[0127] To facilitate understanding, the following example, an audio frame, will be used to illustrate the process of dividing the frequency domain into subbands. See [link / reference]. Figure 2 , Figure 2 This disclosure illustrates a time-domain waveform of an audio frame according to an exemplary embodiment, sampled at an 8kHz sampling rate. Figure 2 By sampling the audio frames shown, we can obtain the following: Figure 3 The spectrum shown is Figure 3 This disclosure illustrates a spectrogram of an audio frame according to an exemplary embodiment, see [link to relevant documentation]. Figure 3 That is, as Figure 2 The spectrum diagram (or spectrum curve) of the time-domain signal shown is used to indicate the amplitude corresponding to each frequency point.

[0128] Furthermore, it can be based on, for example Figure 3 The frequency spectrum diagram shown is used to divide the frequency domain into sub-bands. Figure 3 The spectrum shown corresponds to a sampling rate of 8kHz, which can... Figure 3 The audio frame shown is divided into 8 frequency sub-bands, see [link / reference]. Figure 4 , Figure 4 This disclosure illustrates a subband spectrogram of an audio frame according to an exemplary embodiment, such as... Figure 4 As shown, these eight frequency sub-bands are 0-1kHz, 1-2kHz, 2-3kHz, 3-4kHz, 4-5kHz, 5-6kHz, 6-7kHz, and 7-8kHz. Figure 4 The image shown is a sub-band spectrum diagram (or sub-band spectrum curve). The sub-band spectrum curve uses the frequency range corresponding to the frequency domain sub-band as the horizontal axis and the signal energy magnitude corresponding to the frequency domain sub-band as the vertical axis to indicate the amplitude of each frequency domain sub-band. Each frequency domain sub-band includes multiple frequency points whose corresponding frequency values ​​are located within the frequency range of the frequency domain sub-band.

[0129] S102. Based on the current energy distribution information and historical energy distribution information corresponding to the target audio frame, determine the target audio adjustment information. The target audio adjustment information is used to adjust the energy value of any audio frame located after the target audio frame at multiple frequency points.

[0130] It should be noted that by adjusting the energy values ​​of an audio frame at multiple frequency points, for example, by adjusting the energy values ​​at different frequency points by different magnitudes, the timbre of the audio frame can be adjusted.

[0131] S103. Based on the target audio adjustment information, adjust the energy values ​​of each audio frame in the audio sequence to be processed at multiple frequency points after the target audio frame.

[0132] Since the current energy distribution information corresponds to the audio frame the user is currently listening to, and the historical energy distribution information corresponds to the audio frames the user has previously listened to, both the current and previously listened audio frames have been perceived and accepted by the user (if not accepted or accepted, the user will inevitably make adjustments), determining the target audio adjustment information through the current and historical energy distribution information, and adjusting the audio sequence based on the target audio adjustment information, can meet the user's personalized listening needs. Thus, the solution provided in this disclosure can provide users with an adaptive personalized audio adjustment method.

[0133] After introducing the basic principles of this disclosure, various non-limiting embodiments of this disclosure will be described in detail below.

[0134] It should be noted that, taking real-time communication scenarios as an example, users can make calls with others through real-time communication software. During a call, the computing device can receive audio frames sent by other devices in real time. The multiple audio frames received during each call can be used as an audio sequence to be processed, so that each call process corresponds to an audio sequence to be processed.

[0135] Taking the processing of any audio sequence as an example, each time the computing device receives an audio frame in the audio sequence, it can determine the energy value of the audio frame in each frequency sub-band and the total energy value of the audio frame. Based on the energy value of the audio frame in each frequency sub-band and the total energy value of the audio frame, the computing device can process the audio frame accordingly.

[0136] In one possible implementation, determining the energy value of the audio frame in each frequency subband and the overall energy value of the audio frame can be achieved as follows:

[0137] The signal energy value of the audio frame in each frequency domain sub-band is determined, and then the signal energy values ​​of the audio frame in each frequency domain sub-band are summed to obtain the total energy value of the audio frame.

[0138] For example, for any frequency sub-band, the energy values ​​of multiple frequency points included in the frequency sub-band can be summed to obtain the signal energy value of the audio frame in that frequency sub-band. This process can be repeated to obtain the signal energy value of the audio frame in each frequency sub-band. Then, the signal energy values ​​in each frequency sub-band can be summed to obtain the total energy value of the audio frame.

[0139] After determining the overall energy value of the audio frame through the above process, it can be determined whether the audio frame is a valid audio signal based on the determined overall energy value. A valid audio signal is one that includes valid human voices, not just ambient noise.

[0140] If the total energy value of the received audio frame is less than or equal to the set energy threshold, it can be determined that there is no valid speech signal (such as human voice signal, background music signal, etc.) in the audio frame, and it may only contain some environmental noise or no sound signal at all. In this case, there is no need to process the audio frame, so as to avoid unnecessary waste of computing resources for the computing device by processing invalid audio frames. Moreover, it can also avoid errors in the subsequent audio adjustment process caused by processing these audio frames that do not have valid speech signals, thereby ensuring the audio processing effect.

[0141] If the total energy value of the received audio frames is greater than the set energy threshold, it can be determined that there is a valid speech signal in the received audio frames, and thus processing can be performed based on the received audio frames.

[0142] See Figure 5 , Figure 5 This disclosure is a schematic flowchart illustrating an audio processing procedure based on the total energy value of an audio frame, according to an exemplary embodiment. Figure 5 As shown, when an audio frame is received from the audio sequence to be processed, it is determined whether the total energy value of the audio frame is greater than the set energy threshold. If the total energy value of the audio frame is greater than the set energy threshold, the energy distribution information of the audio frame can be analyzed to determine the target audio adjustment information. If the total energy of the audio frame is less than or equal to the set energy threshold, there is no need to process the currently received audio frame, and the next received audio frame can be processed instead.

[0143] Optionally, if the total energy value of the received audio frame is greater than a set energy threshold, the current energy distribution information of the audio frame can be determined based on the historical energy distribution information corresponding to the audio frame and the energy value of the audio frame in multiple frequency sub-bands, so as to carry out subsequent processing based on the determined current energy distribution information.

[0144] Among them, the historical energy distribution information is used to indicate the energy value of the previous audio frame in multiple frequency sub-bands of the audio frame currently being processed. Optionally, the historical energy distribution information can be a historical sub-band spectrum curve.

[0145] In one possible implementation, when determining the current energy distribution information of an audio frame based on its historical energy distribution information and the energy values ​​of the audio frame across multiple frequency subbands, it can be achieved in the following way:

[0146] The energy values ​​of the audio frame in multiple frequency sub-bands are weighted by the second set parameters and the energy values ​​in the corresponding frequency sub-bands indicated by the historical energy distribution information. The first energy distribution information is determined based on the result of the weighting process. The current energy distribution information is determined based on the energy values ​​in multiple frequency sub-bands indicated by the first energy distribution information and the sum of the energy values ​​in multiple frequency sub-bands indicated by the first energy distribution information.

[0147] Optionally, the second set parameter can be used as the weight of the energy value of the audio frame in multiple frequency domain sub-bands, and the energy value of the audio frame in multiple frequency domain sub-bands can be weighted. The difference between the set parameter value and the second set parameter can be used as the weight of the historical energy distribution information, and the historical energy distribution information can be weighted. The sum of the results obtained from the two weighting processes is determined as the first energy distribution information.

[0148] For example, the first energy distribution information can be determined using the following formula (1):

[0149] BandE{k,n}=(1-α)×Band{k,n-1}+α×E{k,n} (1)

[0150] Where BandE{k,n} represents the first energy distribution information, BandE{k,n-1} represents the historical energy distribution information, E{k,n} represents the energy value of the audio frame in multiple frequency sub-bands, k represents any frequency point included in the frequency sub-band, n represents the audio frame identifier, α represents the second setting parameter, and 1 is the setting parameter value.

[0151] It should be noted that the second setting parameter can be determined based on the sampling rate and preset duration of the audio sequence to be processed. The preset duration can be the duration required for smoothing the spectrum of each audio frame in the audio sequence to be processed. Optionally, the preset duration can be a duration value pre-set by relevant technicians according to technical requirements, or the preset duration can be a duration value adaptively determined by the computing device based on previous audio processing processes.

[0152] For example, the second setting parameter can be determined by the following formula (2):

[0153]

[0154] Where α represents the second set parameter, f s The value represents the sampling rate, and T represents the preset duration.

[0155] After determining the first energy distribution information through the above process, the current energy distribution information can be determined based on the energy values ​​of multiple frequency domain sub-bands indicated by the first energy distribution information and the sum of the energy values ​​of multiple frequency domain sub-bands indicated by the first energy distribution information.

[0156] In one possible implementation, the ratio of the energy values ​​in the multiple frequency sub-bands indicated by the first energy distribution information and the sum of the energy values ​​in the multiple frequency sub-bands indicated by the first energy distribution information can be used as the current energy distribution information.

[0157] Optionally, the energy values ​​on multiple frequency sub-bands indicated by the first energy distribution information can be summed to obtain the sum of energy values ​​on multiple frequency sub-bands indicated by the first energy distribution information, thereby determining the ratio of the energy value on each frequency sub-band to the sum of the energy values, so as to obtain the current energy distribution information.

[0158] For example, the current energy distribution information can be determined using the following formula (3):

[0159]

[0160] in, The current energy distribution information is represented by BandE{k, n}, which represents the first energy distribution information. k represents any frequency point included in the frequency domain sub-band, and n represents the audio frame identifier.

[0161] It should be noted that the above process determines the ratio of the energy value in each frequency sub-band to the sum of the energy values ​​in multiple frequency sub-bands. This can achieve a processing effect similar to normalization, eliminating the influence of energy magnitude on the waveform of energy distribution information and retaining only the influence of timbre on the waveform of energy distribution information. This allows the waveforms of two audio frames with different energy magnitudes to be unified into the same energy range, ensuring the smooth progress of subsequent audio processing.

[0162] Through the above process, the current energy distribution information of each received audio frame can be determined and stored. This information is then used to determine the current energy distribution information of the next audio frame. The determined current energy distribution information of the audio frames at this point constitutes the historical energy distribution information. Therefore, in S101, when a volume adjustment operation is detected for the audio sequence to be processed, the current energy distribution information and historical energy distribution information corresponding to the target audio frame in the audio sequence to be processed can be obtained.

[0163] It should be noted that since the process of determining the energy distribution information of the audio frame is performed in real time, for S101, when responding to the volume adjustment operation for the audio sequence to be processed, the current energy distribution information and historical energy distribution information corresponding to the target audio frame in the audio sequence to be processed can be obtained in the following way:

[0164] In response to a volume adjustment operation on the audio sequence to be processed, if the total energy value of the received audio frame is greater than a set energy threshold, the current energy distribution information of the audio frame is determined based on the historical energy distribution information of the audio frame and the energy value of the audio frame in multiple frequency sub-bands; if the target audio frame is received and the current energy distribution information of the target audio frame has been determined, the current energy distribution information and historical energy distribution information of the target audio frame in the audio sequence to be processed are obtained.

[0165] In other words, upon detecting a volume adjustment operation, the computing device further determines the overall energy value of the received audio frame. Based on this overall energy value, it determines whether a valid speech signal exists in the received audio frame. If the overall energy value is greater than a set energy threshold, it can be determined that a valid speech signal exists in the received audio frame, thus determining the current energy distribution information of the audio frame. The process of determining the overall energy value of the received audio frame and the current energy distribution information of the audio frame can be found in the above embodiments and will not be repeated here.

[0166] The aforementioned energy threshold can be any value, for example, -45dB. Optionally, the energy threshold can also be other values, and this disclosure does not limit the specific value of the energy threshold. This energy threshold can be obtained through offline analysis of a large amount of data to determine whether the downlink signal received in the current time period meets the energy requirements, thereby eliminating the influence of some downlink silent segments or segments with low energy, reducing the waste of computing resources, reducing the processing pressure on the computing device, and ensuring the processing efficiency of the computing device.

[0167] It should be noted that when a volume adjustment operation is detected, the computing device will determine the total energy value of the received audio frames in real time, and determine the current energy distribution information of audio frames whose total energy value is greater than the set energy threshold, until the current energy distribution information of the target audio frame is determined. Then, through S102, the target audio adjustment information is determined based on the current energy distribution information and historical energy distribution information corresponding to the target audio frame.

[0168] The target audio frame can be an audio frame located at a preset frame number among the volume-adjusted audio frames. When the user triggers a volume adjustment operation, the computing device can respond to the volume adjustment operation by adjusting the volume of the audio frames received after the moment the volume adjustment operation is triggered. Therefore, the audio frames received after the moment the volume adjustment operation is triggered are the volume-adjusted audio frames, and the target audio frame is the audio frame located at the preset frame number among these volume-adjusted audio frames.

[0169] Optionally, the audio frame located at the preset frame number can be the audio frame corresponding to the moment 60 seconds after the moment the volume adjustment operation is triggered, or the audio frame located at the preset frame number can be the 6000th audio frame received after the moment the volume adjustment operation is triggered. This disclosure does not limit which specific frame the target audio frame is.

[0170] Since a period of time has passed between the moment the target audio frame is received and the moment the volume adjustment operation is triggered, and the user has not triggered any further volume adjustment operation during this period, it means that the volume of the target audio frame meets the user's listening needs. Therefore, by using the current energy distribution information and historical energy distribution information corresponding to the target audio frame as the basis for determining the target audio adjustment information, it can be ensured that the determined target audio adjustment information is more in line with the user's listening needs.

[0171] In some embodiments, when determining the target audio adjustment information based on the current energy distribution information and historical energy distribution information corresponding to the target audio frame in S102, the following steps can be taken:

[0172] S1021. Based on the current energy distribution information and the historical energy distribution information, determine the first energy distribution difference information. The first energy distribution difference information is used to indicate the energy difference between the current energy distribution information and the historical energy distribution information in multiple frequency domain sub-bands.

[0173] In one possible implementation, the difference between the energy values ​​in the corresponding frequency sub-bands of the current energy distribution information and the historical energy distribution information can be determined to obtain the energy difference between the current energy distribution information and the historical energy distribution information in multiple frequency sub-bands, which serves as the first energy distribution difference information.

[0174] For example, the difference in the first energy distribution can be determined using the following formula (4):

[0175]

[0176] in, This indicates information about the differences in the first energy distribution. This indicates the current energy distribution information corresponding to the target audio frame. This represents the historical energy distribution information corresponding to the target audio frame, where k represents any frequency point included in the frequency domain sub-band, and n represents the audio frame identifier.

[0177] S1022. Based on historical audio adjustment information and first energy distribution difference information, determine target audio adjustment information. Historical audio adjustment information is audio adjustment information determined in response to the previous volume adjustment operation for the audio sequence to be processed.

[0178] In one possible implementation, the first energy distribution difference information can be weighted using a first set parameter, and the target audio adjustment information can be determined based on the weighted first energy distribution difference information and historical audio adjustment information.

[0179] Optionally, the first set parameter can be used as the weight of the first energy distribution difference information to perform weighted processing on the first energy distribution difference information to obtain the weighted first energy distribution information; thereby determining the sum of the weighted first energy distribution information and the historical audio adjustment information as the target audio adjustment information.

[0180] For example, the target audio adjustment information can be determined using the following formula (5):

[0181]

[0182] Where EQ{k, l} represents the target audio adjustment information, and EQ{k, l-1} represents the historical audio adjustment information. β represents the first energy distribution difference information, k represents any frequency point included in the frequency domain sub-band, and l represents the number of times the volume adjustment operation is triggered.

[0183] It should be noted that both the target audio adjustment information and the historical audio adjustment information can be a vector curve. Therefore, the audio adjustment information can also be called a tuning curve. Different frequency points have different values ​​on the tuning curve, so when processing audio frames based on audio adjustment information, the energy values ​​of different frequency points can be adjusted by different amplitudes, thereby achieving timbre adjustment of the audio frame.

[0184] Optionally, the audio adjustment information can be an equalizer (EQ) curve, which can adjust the energy level of each frequency component of the audio signal to achieve timbre adjustment.

[0185] It is important to emphasize that each time a user's volume adjustment operation is detected, a target audio adjustment information is determined in response to the user's volume adjustment operation, and the determined target audio adjustment information is stored as historical audio adjustment information so that it can be retrieved later.

[0186] Additionally, it should be noted that when a volume adjustment operation is detected for the first time, the computing device has not previously determined audio adjustment information. In order to ensure that the target audio adjustment information can be determined, relevant technicians can configure a default initial audio adjustment information, so that the initial audio adjustment information can be used as historical audio adjustment information when the user has not yet performed a volume adjustment operation on the audio sequence to be processed.

[0187] Optionally, the default initial audio adjustment information can be a vector curve with a value of 1 for multiple frequency domain sub-bands. That is, if no volume adjustment operation is performed on the audio sequence to be processed, the vector curve with a value of 1 for multiple frequency domain sub-bands can be used as historical audio adjustment information.

[0188] Taking the default initial audio adjustment information as a vector curve with values ​​of 1 for multiple frequency sub-bands as an example, the initial audio adjustment information can be expressed as the vector curve shown in the following formula (6):

[0189] EQ(k, 0) = 1 (6)

[0190] Where EQ(k, 0) represents the initial audio adjustment information, and the fact that multiple frequency domain sub-bands all take the value of 1 indicates that no adjustment is made to the spectral distribution of the audio frame in the initial stage.

[0191] The above process mainly describes how to determine the target audio information when a volume adjustment operation is detected. When no volume adjustment operation for the audio sequence to be processed is detected, the historical audio adjustment information can be directly determined as the target audio adjustment information. That is, when no volume adjustment operation is detected, there is no need to re-determine the audio adjustment information, but the audio adjustment information determined when the volume adjustment operation was detected last time can be directly used.

[0192] Additionally, it should be noted that after a volume adjustment operation is detected, the computing device does not determine the target audio information directly based on the first audio frame received after the volume adjustment operation is detected. Instead, it determines the target audio adjustment information based on the target audio frame. There are several audio frames between the target audio frame and the first audio frame received after the volume adjustment operation is detected. For these multiple audio frames, the audio adjustment information determined when the volume adjustment operation was detected last time (i.e., the historical audio adjustment information) can still be used as the target audio adjustment information.

[0193] The solution provided in this disclosure can, when a user adjusts the system volume, obtain target audio adjustment information that better matches the user's preferred timbre and frequency response by statistically analyzing the energy distribution information of received audio frames over a long period. This target audio adjustment information is then used to perform audio processing on the audio sequence to be processed. Furthermore, determining the target audio adjustment information through long-term statistical analysis of energy distribution information also contributes to the stability of the timbre.

[0194] The above embodiments mainly introduce the process of audio processing of the audio sequence to be processed based on the user's volume adjustment operation. During the audio acquisition process, some environmental noise will inevitably be collected. This environmental noise may lead to a decrease in audio intelligibility (audio intelligibility is a measure of speech comprehension ability under given conditions, which can generally be quantified by calculating the number of correctly recognized words or phonemes). At this time, it is necessary to increase the system volume of the computing device to ensure that the user can hear a clearer speech signal.

[0195] In response to users' need for higher downlink listening volume in noisy scenarios, this patent designs the following scheme to adjust the downlink signal volume and timbre, which mainly relies on the noise level estimated by the uplink processing algorithm on this end.

[0196] In some embodiments, the computing device may acquire the noise energy value of the target audio frame, which is determined by the transmitter of the audio sequence to be processed based on the target audio frame.

[0197] For example, after acquiring the audio signal and generating the corresponding audio frame, the transmitting end of the audio sequence to be processed can estimate the noise energy value of the audio frame through an uplink algorithm, and send the estimated noise energy value and the generated audio frame together to the computing device that is the receiving end of the audio sequence to be processed.

[0198] The uplink algorithm refers to a series of audio processing algorithms that process the audio signals acquired by the transmitting device, including but not limited to acoustic echo cancellation, noise suppression, automatic gain control (AGC), and noise estimation algorithms. Optionally, the noise estimation algorithm can be any of the following: Minimum Statistics, Improved Minimum Controlled Regressive Averaging, or a deep neural network-based noise estimation algorithm. Alternatively, the noise estimation algorithm can be other types of algorithms, which are not limited in this disclosure.

[0199] Optionally, the noise signal in the target audio frame can be determined by a noise estimation algorithm to obtain the initial noise energy distribution information of the noise signal. Then, the energy values ​​of multiple frequency sub-bands indicated by the initial noise energy distribution information are summed to obtain the sum of the energy values ​​of multiple frequency sub-bands indicated by the initial noise energy distribution information as the noise energy value.

[0200] Once the noise energy value of the target audio frame has been obtained, the target audio adjustment information can be determined by taking appropriate measures based on the obtained noise energy value.

[0201] In some embodiments, when the system volume level has not reached the maximum volume level and the noise energy value is not greater than the first energy threshold, the computing device can determine the target audio adjustment information through the audio processing procedure provided in the above embodiments.

[0202] To facilitate understanding, the audio processing process described in the above embodiments will be explained below with a specific implementation scheme.

[0203] See Figure 6 , Figure 6 This disclosure is a schematic flowchart illustrating an audio processing procedure according to an exemplary embodiment, such as... Figure 6 As shown, the target audio adjustment information can be determined through the following steps, thereby enabling audio adjustment of the audio sequence to be processed.

[0204] S601. Receive an audio frame from the audio sequence to be processed and obtain the noise energy value of the audio frame.

[0205] S602. When the system volume level has not reached the maximum volume level and the noise energy value is not greater than the first energy threshold, determine the energy value of the audio frame in each frequency sub-band and the total energy value of the audio frame.

[0206] The system volume, also known as the device volume of the computing device, can be adjusted using the volume up and volume down buttons located on the computing device.

[0207] Generally speaking, the system volume of a computing device can be divided into 15 levels. When the system volume reaches the maximum level, that is, when the device volume of the computing device has reached the 15th level.

[0208] Optionally, the first energy threshold can be any value, and this disclosure does not limit the specific value of the first energy threshold.

[0209] S603. If the total energy value of the audio frame is greater than the set energy threshold, determine the current energy distribution information of the audio frame based on the historical energy distribution information corresponding to the audio frame and the energy value of the audio frame in multiple frequency sub-bands.

[0210] S604. In response to the volume adjustment operation for the audio sequence to be processed, continue to process the received audio frame according to the steps of S601 to S603 above until the target audio frame is received and the current energy distribution information of the target audio frame is determined.

[0211] S605. Based on the current energy distribution information and historical energy distribution information corresponding to the target audio frame, determine the target audio adjustment information.

[0212] In addition, if the system volume level has not reached the maximum volume level and the noise energy value is not greater than the first energy threshold, and no volume adjustment operation for the audio sequence to be processed is detected, then the historical audio adjustment information can be directly determined as the target audio adjustment information.

[0213] S606. Based on the target audio adjustment information, adjust the energy values ​​of each audio frame in the audio sequence to be processed at multiple frequency points after the target audio frame.

[0214] like Figure 6 The above is only one possible implementation method under the condition that the system volume level has not reached the maximum volume level and the noise energy value is not greater than the first energy threshold. It does not constitute a limitation on the specific implementation method. In more possible implementation methods, the order of each step can be adjusted as needed.

[0215] The above describes the process of determining the target audio adjustment information only when the system volume level has not reached the maximum volume level and the noise energy value is not greater than the first energy threshold. For the process of determining the target audio adjustment information in other possible cases, please refer to the following embodiments.

[0216] In other embodiments, when the system volume level has not reached the maximum volume level and the noise energy value is greater than the first energy threshold but less than the second energy threshold, the second energy distribution difference information is determined based on the current energy distribution information of the target audio frame and the noise energy distribution information of the noise signal included in the target audio frame; thereby, the target audio adjustment information is determined based on the second energy distribution difference information, the total energy value of the target audio frame and the noise energy value.

[0217] The noise energy distribution information can be used to indicate the energy values ​​of the noise signal across multiple frequency sub-bands, and the second energy distribution difference information can be used to indicate the energy difference between the energy distribution information and the noise energy distribution information across multiple frequency sub-bands. It should be noted that when acquiring the noise energy distribution information, the sub-band spectrum curve of the noise signal can be used as the initial noise energy distribution information. The ratio of the energy values ​​indicated by the initial noise energy distribution information across multiple frequency sub-bands to the sum of the energy values ​​indicated by the initial noise energy distribution information across multiple frequency sub-bands is then used as the noise energy distribution information.

[0218] Optionally, the energy values ​​on multiple frequency sub-bands indicated by the initial noise energy distribution information can be summed to obtain the sum of energy values ​​on multiple frequency sub-bands indicated by the initial noise energy distribution information, thereby determining the ratio of the energy value on each frequency sub-band indicated by the noise energy distribution information to the sum of the energy values, so as to obtain the noise energy distribution information.

[0219] The second energy threshold can be any value. This disclosure does not limit the specific value of the second energy threshold, but only requires that the second energy threshold is greater than the first energy threshold.

[0220] In one possible implementation, when determining the second energy distribution difference information based on the current energy distribution information of the target audio frame and the noise energy distribution information of the noise signal included in the target audio frame, the difference between the energy values ​​in the corresponding frequency domain sub-bands of the current energy distribution information and the noise energy distribution information can be determined to obtain the energy difference between the current energy distribution information and the noise energy distribution information in multiple frequency domain sub-bands, which serves as the second energy distribution difference information.

[0221] When determining the target audio adjustment information based on the second energy distribution difference information, the total energy value of the target audio frame, the noise energy value, and the first energy distribution information, the first target parameter can be determined based on the second energy distribution difference information, the total energy value of the target audio frame, the noise energy value, and the third setting parameter. Thus, the target audio adjustment information can be determined based on the first target parameter, the total energy value of the target audio frame, the first energy distribution information, and the fourth setting parameter.

[0222] Optionally, when determining the first target parameter based on the second energy distribution difference information, the total energy value of the target audio frame, the noise energy value, and the third set parameter, it can be achieved by the following formula (7):

[0223]

[0224] in, This indicates the current energy distribution information. Represents noise energy distribution information, BandE s (k, n) represents the first energy distribution information. BandE represents the total energy value of the target audio frame. n (k, n) represents the initial noise energy distribution information. The noise energy value is represented by x(k,n), the first target parameter is represented by x(k,n), the predefinedSNR is represented by the third setting parameter, k represents any frequency point included in the frequency domain sub-band, and n represents the audio frame identifier.

[0225] It should be noted that when determining the overall energy value of the target audio frame, the energy values ​​of multiple frequency sub-bands indicated by the first energy distribution information can be summed to obtain the sum of the energy values ​​of multiple frequency sub-bands indicated by the first energy distribution information as the overall energy value of the target audio frame. The process for determining the noise energy value can be found in the above embodiments, and will not be repeated here.

[0226] Additionally, it should be noted that the third setting parameter can be a scalar data, such as a fixed value; or it can be a vector data, where the value of the third setting parameter can differ for different frequency sub-bands, thus making it a curve representing the variation of different values ​​for different frequency sub-bands. Optionally, the value of the third setting parameter can be determined experimentally by relevant technical personnel. For example, they can conduct experiments based on the human ear's auditory characteristics curve to determine the signal energy value most suitable for human auditory characteristics in different frequency sub-bands. The value of the third setting parameter can then be determined based on this signal energy value, ensuring that the target audio adjustment information determined based on the third setting parameter provides the user with the best listening experience.

[0227] Optionally, when determining the first target parameter through the above formula (7), the first target parameter can be solved under preset constraints. Optionally, the preset constraints can be that the value of the determined first target parameter is within a set range. For example, the preset constraints can be a≤x(k,n)≤b, where a and b can be arbitrary values, and it is only necessary to ensure that the value of a is less than b.

[0228] By setting constraints for the solution process of the first objective parameter, the timbre of the adjusted audio frame will not change abruptly with the previous audio frame, thereby improving the audio quality of the audio sequence to be processed.

[0229] After determining the first target parameter value through the above process, the target audio adjustment information can be determined using the following formula (8) based on the first target parameter, the overall energy value of the target audio frame, the first energy distribution information, and the fourth setting parameter:

[0230]

[0231] Where α represents the target audio adjustment information, x(k,n) represents the first target parameter, and BandE s (k, n) represents the first energy distribution information. The total energy value of the target audio frame is represented by , c represents the fourth setting parameter, k represents any frequency point included in the frequency domain sub-band, and n represents the audio frame identifier.

[0232] Optionally, the fourth setting parameter can take any value. For example, the fourth setting parameter can be 1 or slightly greater than 1. This disclosure does not limit the specific value of the fourth setting parameter.

[0233] In other embodiments, when the system volume level has not reached the maximum volume level and the noise energy value is not less than the second energy threshold, the target energy difference is determined based on the total energy value and noise energy value of the target audio frame; thereby, the target audio adjustment information is determined based on the target energy difference, the total energy value of the target audio frame, the noise energy value, the first energy distribution information, and the first energy adjustment information.

[0234] The process of determining the overall energy value and noise energy value of the target audio frame can be found in the above embodiments and will not be repeated here. For the process of determining the target energy difference based on the overall energy value and noise energy value of the target audio frame, the difference between the overall energy value and the noise energy value of the target audio frame can be used as the target energy difference.

[0235] When determining the target audio adjustment information based on the target energy difference, the overall energy value of the target audio frame, the noise energy value, the first energy distribution information, and the first energy adjustment information, the first target parameter can be determined based on the overall energy value, the overall energy value of the target audio frame, the noise energy value, and the third setting parameter. Then, the target audio adjustment information can be determined based on the first target parameter, the overall energy value of the target audio frame, the first energy distribution information, the first energy adjustment information, and the fourth setting parameter.

[0236] Among them, when determining the first target parameter based on the target energy difference, the total energy value of the target audio frame, the noise energy value, and the third set parameter, it can be achieved by the following formula (9):

[0237]

[0238] Among them, BandE s (k, n) represents the first energy distribution information. BandE represents the total energy value of the target audio frame. n (k, n) represents the initial noise energy distribution information. Indicates the noise energy value. The target energy difference is represented by , predefinedSNR is represented by the third setting parameter, k is represented by any frequency point included in the frequency domain subband, and n is represented by the audio frame identifier.

[0239] For the process of determining the total energy value and noise energy value of the target audio frame, and for the introduction of the third setting parameter, please refer to the above embodiments. In addition, when solving the first target parameter by formula (9), constraints can also be set for the solution process of the first target parameter, which will not be elaborated here.

[0240] After determining the first target parameter value through the above process, the target audio adjustment information can be determined using the following formula (10), based on the first target parameter, the overall energy value of the target audio frame, the first energy distribution information, the first energy adjustment information, and the fourth setting parameter:

[0241]

[0242] Where E1 represents the first energy adjustment information, α represents the target audio adjustment information, x(k, n) represents the first target parameter, and BandE s (k, n) represents the first energy distribution information. The total energy value of the target audio frame is represented by , c represents the fourth setting parameter, k represents any frequency point included in the frequency domain sub-band, and n represents the audio frame identifier.

[0243] Optionally, the fourth setting parameter can take any value. For example, the fourth setting parameter can be 1 or slightly greater than 1. This disclosure does not limit the specific value of the fourth setting parameter.

[0244] It should be noted that the first energy adjustment information is a first preset energy value used to raise the energy value of the target audio frame at each frequency point. Optionally, the first preset energy value can be any value, for example, the first preset energy value can be 3dB, or the first preset energy value can be other values. This disclosure does not limit the specific value of the first preset energy value.

[0245] In other embodiments, when the system volume level has reached the maximum volume level and the noise energy value is not less than and not greater than the third energy threshold, the second energy distribution difference information is determined based on the current energy distribution information of the target audio frame and the noise energy distribution information of the noise signal included in the target audio frame; thereby, the target audio adjustment information is determined based on the second energy distribution difference information, the overall energy value of the target audio frame, the noise energy value, the first energy distribution information and the second energy adjustment information.

[0246] The third energy threshold can be any value, and this disclosure does not limit the specific value of the third energy threshold, only requiring that the third energy threshold is greater than the second energy threshold. The process of determining the overall energy value, noise energy value, and second energy distribution difference information of the target audio frame can be found in the above embodiments, and will not be repeated here.

[0247] When determining the target audio adjustment information based on the second energy distribution difference information, the overall energy value of the target audio frame, the noise energy value, the first energy distribution information, and the second energy adjustment information, the first target parameter can be determined based on the second energy distribution difference information, the overall energy value of the target audio frame, the noise energy value, and the third setting parameter. Thus, the target audio adjustment information can be determined based on the first target parameter, the overall energy value of the target audio frame, the first energy distribution information, the second energy adjustment information, and the fourth setting parameter.

[0248] The process of determining the first target parameter based on the second energy distribution difference information, the total energy value of the target audio frame, the noise energy value, and the third set parameter can be found in the above embodiments and will not be repeated here.

[0249] After determining the first target parameter value, the target audio adjustment information can be determined using the following formula (11), based on the first target parameter, the overall energy value of the target audio frame, the first energy distribution information, the first energy adjustment information, and the fourth setting parameter:

[0250]

[0251] Where E2 represents the second energy adjustment information, α represents the target audio adjustment information, x(k,n) represents the first target parameter, and BandE s (k, n) represents the first energy distribution information. The total energy value of the target audio frame is represented by , c represents the fourth setting parameter, k represents any frequency point included in the frequency domain sub-band, and n represents the audio frame identifier.

[0252] Optionally, the fourth setting parameter can take any value. For example, the fourth setting parameter can be 1 or slightly greater than 1. This disclosure does not limit the specific value of the fourth setting parameter.

[0253] It should be noted that the second energy adjustment information is a second preset energy value used to raise the energy value of the audio frame at each frequency point. Optionally, the second preset energy value can be any value, and this disclosure does not limit the specific value of the second preset energy value.

[0254] To facilitate understanding, the following will be explained... Figure 7 The flowchart shown illustrates the audio processing methods in different situations. (See also...) Figure 7 , Figure 7 This disclosure is a flowchart illustrating an audio adjustment process under different circumstances according to an exemplary embodiment, such as... Figure 7 As shown, upon receiving a downlink signal and its noise energy value, if the system volume level has not reached the maximum volume level and the noise energy value is not greater than the first energy threshold, target audio adjustment information is determined based on the current energy distribution information and historical energy distribution information corresponding to the target audio frame. If the system volume level has not reached the maximum volume level and the noise energy value is greater than the first energy threshold but less than the second energy threshold, target audio adjustment information is determined based on the second energy distribution difference information, the overall energy value of the target audio frame, and the noise energy value. If the system volume level has not reached the maximum volume level and the noise energy value is not less than the second energy threshold, target audio adjustment information is determined based on the target energy difference, the overall energy value of the target audio frame, the noise energy value, the first energy distribution information, and the first energy adjustment information. If the system volume level has reached the maximum volume level and the noise energy value is not less than the third energy threshold and not greater than the third energy threshold, target audio adjustment information is determined based on the second energy distribution difference information, the overall energy value of the target audio frame, the noise energy value, the first energy distribution information, and the second energy adjustment information. Once the target audio adjustment information is determined using any of the methods described above, audio processing can be performed on the audio sequence to be processed based on the target audio adjustment information.

[0255] After determining the audio adjustment information through any of the methods provided in the above embodiments, the energy values ​​of each audio frame in the audio sequence to be processed at multiple frequency points can be adjusted based on the determined audio adjustment information through S103.

[0256] In one possible implementation, when adjusting the energy values ​​of each audio frame following the target audio frame in the audio sequence to be processed at multiple frequency points based on the target audio adjustment information, the value of the target audio adjustment information can be used as the weight of the energy value at the corresponding frequency point to achieve the adjustment of the energy values ​​of each audio frame following the target audio frame at multiple frequency points.

[0257] Since the target audio adjustment information is determined by integrating multiple aspects such as the overall energy value of the audio frame, the noise energy value, the current energy distribution information, and the historical energy distribution information, the target audio adjustment information can more comprehensively meet the user's various listening needs. As a result, the audio sequence processed based on the target audio adjustment information also better meets the user's listening needs, thus improving the audio processing effect of the audio sequence to be processed.

[0258] Optionally, after adjusting the energy values ​​of audio frames located after the target audio frame in the audio sequence to be processed at multiple frequency points based on the target audio adjustment information, anti-clipping processing can also be applied to the adjusted audio frames.

[0259] In one possible implementation, if the energy value of any adjusted audio frame at any frequency point is greater than a first set threshold, the energy value of the audio frame at the frequency point is adjusted so that the energy value of the audio frame at the frequency point is less than or equal to the first set threshold.

[0260] In another possible implementation, if the energy value of any adjusted audio frame at any frequency point is less than the second set threshold, the energy value of the audio frame at the frequency point is adjusted so that the energy value of the audio frame at the frequency point is greater than or equal to the second set threshold.

[0261] The first and second set thresholds can both be any values. This disclosure does not limit the values ​​of the first and second set thresholds, but only requires that the second set threshold is less than the first set threshold.

[0262] The above process can control the energy value of each frequency point in each audio frame within the range formed by the second set threshold and the second set threshold. When the energy value is within the range formed by the second set threshold and the second set threshold, the audio effect is better, thereby further improving the audio processing effect of the audio sequence to be processed.

[0263] For example, based on the target audio adjustment information, after adjusting the energy values ​​of any audio frame following the target audio frame in the audio sequence to be processed at multiple frequency points, the adjusted energy value of the audio frame at each frequency point can be obtained. When the energy value at any frequency point is greater than a first set threshold, the energy value at that frequency point can be adjusted to the first set threshold so that the energy value at that frequency point is within the value range formed by the second set threshold and the second set threshold. In addition, when the energy value at any frequency point is less than the second set threshold, the energy value at that frequency point can be adjusted to the second set threshold so that the energy value at that frequency point is within the value range formed by the second set threshold and the second set threshold.

[0264] The solution provided in this disclosure can automatically select and determine the audio adjustment information based on the external environment (mainly noise level) and the user's operation on the system volume. This ensures that the method for determining the audio adjustment information changes accordingly when the external environment or system volume changes. Based on this, this disclosure can save the intermediate variables and historical results generated in the ongoing processing scheme under various conditions with different system volume levels and noise energy values. This allows for the direct acquisition of relevant intermediate variables and historical results after subsequent adjustments to the system volume level or changes in noise energy values, thus ensuring a better and more continuous user experience.

[0265] Optionally, the audio processing functions provided in this disclosure can be provided through software, for example, through a real-time communication-based service software. Alternatively, the software can be deployed on a computing device so that the computing device can perform audio processing on the audio sequence to be processed using the audio processing methods provided in this disclosure.

[0266] The software deployed on the computing device can provide various types of function buttons for users to set and adjust the corresponding functions according to their own needs. For example, the software can provide a button for users to set whether to enable the audio adaptive adjustment function. When the user confirms that the audio adaptive adjustment function is enabled by using this button, the internal algorithm of the software can perform relevant processing through calculation and optimization to ensure that when the user plays audio, it can provide a better listening experience and intelligibility based on the user's settings and habits.

[0267] Optionally, to ensure the security and flexibility of the software during use, a user identifier (ID) can be configured for each user. Users can log in to the software based on their user ID, ensuring that the same device and the same software can provide functions for different users. This allows the same device and the same software to provide corresponding hearing assistance / hearing optimization effects based on different user settings.

[0268] After introducing the audio processing method according to the exemplary embodiments of the present disclosure, the structure of the audio processing apparatus and the computing device and computer-readable storage medium for implementing the audio processing method according to the exemplary embodiments of the present disclosure will be described next.

[0269] See Figure 8 , Figure 8 This is a block diagram illustrating an audio processing apparatus according to an exemplary embodiment of the present disclosure, the apparatus comprising:

[0270] The acquisition module 801 is used to acquire the current energy distribution information and historical energy distribution information of the target audio frame in the audio sequence to be processed in response to the volume adjustment operation for the audio sequence to be processed. The target audio frame is the audio frame located at a preset frame number among the audio frames after volume adjustment. The current energy distribution information is used to indicate the energy value of the target audio frame in multiple frequency sub-bands, and the historical energy distribution information is used to indicate the energy value of the previous audio frame of the target audio frame in multiple frequency sub-bands.

[0271] The determining module 802 is used to determine target audio adjustment information based on the current energy distribution information and historical energy distribution information corresponding to the target audio frame. The target audio adjustment information is used to adjust the energy value of any audio frame located after the target audio frame at multiple frequency points.

[0272] The adjustment module 808 is used to adjust the energy values ​​of each audio frame in the audio sequence to be processed at multiple frequency points based on the target audio adjustment information.

[0273] In one embodiment of this disclosure, the determining module 802, when determining target audio adjustment information based on the current energy distribution information and historical energy distribution information corresponding to the target audio frame, is configured to:

[0274] Based on the current energy distribution information and the historical energy distribution information, the first energy distribution difference information is determined. The first energy distribution difference information is used to indicate the energy difference between the current energy distribution information and the historical energy distribution information in multiple frequency domain sub-bands.

[0275] Based on historical audio adjustment information and first energy distribution difference information, target audio adjustment information is determined. Historical audio adjustment information refers to the audio adjustment information determined in response to the previous volume adjustment operation on the audio sequence to be processed.

[0276] In one embodiment of this disclosure, the determining module 802, when determining target audio adjustment information based on historical audio adjustment information and first energy distribution difference information, is configured to:

[0277] The first energy distribution difference information is weighted using a first set parameter, and the target audio adjustment information is determined based on the weighted first energy distribution difference information and historical audio adjustment information.

[0278] In one embodiment of this disclosure, the determining module 802, when weighting the first energy distribution difference information with a first set parameter and determining the target audio adjustment information based on the weighted first energy distribution difference information and historical audio adjustment information, is configured to:

[0279] Using the first set parameter as the weight of the first energy distribution difference information, the first energy distribution difference information is weighted to obtain the weighted first energy distribution information.

[0280] The sum of the weighted first energy distribution information and the historical audio adjustment information is determined as the target audio adjustment information.

[0281] In one embodiment of this disclosure, the determining module 802 is further configured to determine the energy value of the audio frame in each frequency domain sub-band and the total energy value of the audio frame for each audio frame received in the audio sequence to be processed.

[0282] The determining module 802 is further configured to determine the current energy distribution information of the audio frame based on the historical energy distribution information corresponding to the audio frame and the energy value of the audio frame in multiple frequency sub-bands when the total energy value of the received audio frame is greater than a set energy threshold.

[0283] In one embodiment of this disclosure, the determining module 802, when determining the energy value of the audio frame in each frequency domain subband and the overall energy value of the audio frame, is configured to:

[0284] Determine the signal energy value of the audio frame in each frequency subband;

[0285] The total energy value of the audio frame is obtained by summing the signal energy values ​​of each frequency sub-band.

[0286] In one embodiment of this disclosure, the determining module 802, when determining the current energy distribution information of an audio frame based on the historical energy distribution information corresponding to the audio frame and the energy values ​​of the audio frame in multiple frequency sub-bands, is configured to:

[0287] The energy values ​​of the audio frame in multiple frequency sub-bands are weighted by the second set parameters and the energy values ​​in the corresponding frequency sub-bands indicated by the historical energy distribution information, and the first energy distribution information is determined based on the result of the weighting process.

[0288] The current energy distribution information is determined based on the energy values ​​in multiple frequency sub-bands indicated by the first energy distribution information and the sum of the energy values ​​in multiple frequency sub-bands indicated by the first energy distribution information.

[0289] The second setting parameter is determined based on the sampling rate and preset duration of the audio sequence to be processed. The preset duration is the duration required to smooth the spectrum of each audio frame in the audio sequence to be processed.

[0290] In one embodiment of this disclosure, the determining module 802, when performing weighted processing on the energy values ​​of the audio frame in multiple frequency sub-bands and the energy values ​​in the corresponding frequency sub-bands indicated by historical energy distribution information using a second set parameter, and determining the first energy distribution information based on the result of the weighted processing, is configured to:

[0291] The second set parameter is used as the weight of the energy value of the audio frame in multiple frequency domain sub-bands, and the energy value of the audio frame in multiple frequency domain sub-bands is weighted. The difference between the set parameter value and the second set parameter is used as the weight of the historical energy distribution information, and the historical energy distribution information is weighted.

[0292] The sum of the results obtained from the two weighted processing is determined as the first energy distribution information.

[0293] In one embodiment of this disclosure, the acquisition module 801, when acquiring current energy distribution information and historical energy distribution information corresponding to a target audio frame in the audio sequence to be processed in response to a volume adjustment operation for the audio sequence to be processed, is configured to:

[0294] In response to a volume adjustment operation on the audio sequence to be processed, if the total energy value of the received audio frame is greater than a set energy threshold, the step of determining the current energy distribution information of the audio frame based on the historical energy distribution information corresponding to the audio frame and the energy value of the audio frame in multiple frequency sub-bands is performed.

[0295] If the target audio frame is received and its current energy distribution information has been determined, the step of obtaining the current energy distribution information and historical energy distribution information of the target audio frame in the audio sequence to be processed is executed.

[0296] In one embodiment of this disclosure, the determining module 802 is further configured to determine historical audio adjustment information as target audio adjustment information when no volume adjustment operation for the audio sequence to be processed is detected.

[0297] In one embodiment of this disclosure, without performing volume adjustment operations on the audio sequence to be processed, vector curves with values ​​of 1 for multiple frequency domain sub-bands are used as historical audio adjustment information.

[0298] In one embodiment of this disclosure, the acquisition module 801 is further configured to acquire the noise energy value of the target audio frame, wherein the noise energy value is determined by the transmitting end of the audio sequence to be processed based on the target audio frame;

[0299] The acquisition module 801 is further configured to perform a step of acquiring the current energy distribution information and historical energy distribution information corresponding to the target audio frame in the audio sequence to be processed, in response to a volume adjustment operation for the audio sequence to be processed, when the system volume level has not reached the maximum volume level and the noise energy value is not greater than the first energy threshold.

[0300] In one embodiment of this disclosure, the determining module 802 is further configured to determine second energy distribution difference information based on the current energy distribution information of the target audio frame and the noise energy distribution information of the noise signal included in the target audio frame when the system volume level has not reached the maximum volume level and the noise energy value is greater than the first energy threshold and less than the second energy threshold. The noise energy distribution information is used to indicate the energy value of the noise signal in multiple frequency domain sub-bands, and the second energy distribution difference information is used to indicate the energy difference between the energy distribution information and the noise energy distribution information in multiple frequency domain sub-bands.

[0301] The determining module 802 is further configured to determine target audio adjustment information based on the second energy distribution difference information, the overall energy value of the target audio frame, the noise energy value, and the first energy distribution information. The first energy distribution information is determined based on the weighted processing result of the energy values ​​of the audio frame in multiple frequency sub-bands and the energy values ​​in the corresponding frequency sub-bands indicated by the historical energy distribution information.

[0302] In one embodiment of this disclosure, the determining module 802 is further configured to determine a target energy difference based on the total energy value and noise energy value of the target audio frame when the system volume level has not reached the maximum volume level and the noise energy value is not less than the second energy threshold and not greater than the third energy threshold. The total energy value of the target audio frame is the sum of the signal energy values ​​of the target audio frame in each frequency domain sub-band.

[0303] The determining module 802 is further configured to determine target audio adjustment information based on the target energy difference, the overall energy value of the target audio frame, the noise energy value, the first energy distribution information, and the first energy adjustment information. The first energy adjustment information is a first preset energy value used to raise the energy value of the target audio frame at each frequency point.

[0304] In one embodiment of this disclosure, the determining module 802 is further configured to determine second energy distribution difference information based on the current energy distribution information of the target audio frame and the noise energy distribution information of the noise signal included in the target audio frame when the system volume level has reached the maximum volume level and the noise energy value is not less than the third energy threshold.

[0305] The determination module 802 is also used to determine the target audio adjustment information based on the second energy distribution difference information, the overall energy value of the target audio frame, the noise energy value, the first energy distribution information, and the second energy adjustment information. The second energy adjustment information is a second preset energy value used to raise the energy value of the audio frame at each frequency point.

[0306] In one embodiment of this disclosure, the adjustment module 808 is further configured to adjust the energy value of the audio frame at any frequency point when the energy value of any adjusted audio frame at any frequency point is greater than the first set threshold, so that the energy value of the audio frame at the frequency point is less than or equal to the first set threshold.

[0307] The adjustment module 808 is also used to adjust the energy value of the audio frame at any frequency point when the energy value of any adjusted audio frame at any frequency point is less than the second set threshold, so that the energy value of the audio frame at the frequency point is greater than or equal to the second set threshold.

[0308] It should be noted that although several modules of the audio processing device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules described above can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules. Furthermore, it should be noted that since the device embodiments and method embodiments correspond, the specific implementation process of the functions and roles of each module in the device embodiments can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0309] This disclosure also provides a computer-readable storage medium. Figure 9 This is a schematic diagram illustrating a computer-readable storage medium according to an exemplary embodiment of the present disclosure, such as... Figure 9 As shown, the storage medium stores a computer program 901, which, when executed by a processor, can perform the audio processing method provided in any embodiment of this disclosure.

[0310] This disclosure also provides a computing device, which may include a memory and a processor. The memory stores computer instructions that can run on the processor, and the processor, when executing the computer instructions, implements the audio processing method provided in any embodiment of this disclosure. See also Figure 10 , Figure 10 This is a schematic diagram of the structure of a computing device according to an exemplary embodiment of the present disclosure. The computing device 1000 may include, but is not limited to, a processor 1010, a memory 1020, and a bus 1030 connecting different system components (including the memory 1020 and the processor 1010).

[0311] The memory 1020 stores computer instructions that can be executed by the processor 1010, enabling the processor 1010 to perform the audio processing method provided in any embodiment of this disclosure. The memory 1020 may include a random access memory (RAM) 1021, a cache memory 1022, and / or a read-only memory (ROM) 1023. The memory 1020 may also include a program tool 1025 having a set of program modules 1024, including but not limited to: an operating system, one or more application programs, other program modules, and program data. One or more combinations of these program modules may include an implementation of a network environment.

[0312] Bus 1030 may include, for example, a data bus, an address bus, and a control bus. The computing device 1000 can also communicate with external devices 1050 via I / O interface 1040, such as a keyboard or a Bluetooth device. The computing device 1000 can also communicate with one or more networks via network adapter 1060, such as a local area network (LAN), a wide area network (WAN), or a public network. Figure 10 As shown, the network adapter 1060 can also communicate with other modules of the computing device 1000 via the bus 1030.

[0313] This disclosure also provides a computer program product, which includes a computer program that, when executed by the processor 1010 of the computing device 1000, can implement the audio processing method provided in any embodiment of this disclosure.

[0314] Furthermore, although the operations of the methods disclosed herein are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0315] While the spirit and principles of this disclosure have been described with reference to several specific embodiments, it should be understood that this disclosure is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for convenience of expression. This disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. An audio processing method, characterized in that, The method includes: In response to a volume adjustment operation on an audio sequence to be processed, current energy distribution information and historical energy distribution information corresponding to a target audio frame in the audio sequence to be processed are obtained. The target audio frame is an audio frame located at a preset frame number among the audio frames after volume adjustment. The current energy distribution information is used to indicate the energy value of the target audio frame in multiple frequency sub-bands, and the historical energy distribution information is used to indicate the energy value of the previous audio frame of the target audio frame in multiple frequency sub-bands. Based on the current energy distribution information and historical energy distribution information corresponding to the target audio frame, a first energy distribution difference information is determined. The first energy distribution difference information is used to indicate the energy difference between the current energy distribution information and the historical energy distribution information in multiple frequency domain sub-bands. Based on historical audio adjustment information and the first energy distribution difference information, target audio adjustment information is determined. The historical audio adjustment information is the audio adjustment information determined in response to the previous volume adjustment operation for the audio sequence to be processed. The target audio adjustment information is used to adjust the energy value of any audio frame located after the target audio frame at multiple frequency points. Based on the target audio adjustment information, the energy values ​​of each audio frame following the target audio frame in the audio sequence to be processed are adjusted at multiple frequency points.

2. The method according to claim 1, characterized in that, The step of determining the target audio adjustment information based on historical audio adjustment information and the first energy distribution difference information includes: The first energy distribution difference information is weighted using a first set parameter, and the target audio adjustment information is determined based on the weighted first energy distribution difference information and the historical audio adjustment information.

3. The method according to claim 2, characterized in that, The step of weighting the first energy distribution difference information with a first set parameter, and determining the target audio adjustment information based on the weighted first energy distribution difference information and the historical audio adjustment information, includes: Using the first set parameter as the weight of the first energy distribution difference information, the first energy distribution difference information is weighted to obtain the weighted first energy distribution information. The sum of the weighted first energy distribution information and the historical audio adjustment information is determined as the target audio adjustment information.

4. The method according to claim 1, characterized in that, Before obtaining the current energy distribution information and historical energy distribution information corresponding to the target audio frame in the audio sequence to be processed in response to a volume adjustment operation on the audio sequence to be processed, the method further includes: For each audio frame received in the audio sequence to be processed, the energy value of the audio frame in each frequency sub-band and the total energy value of the audio frame are determined. If the total energy value of the received audio frame is greater than a set energy threshold, the current energy distribution information of the audio frame is determined based on the historical energy distribution information corresponding to the audio frame and the energy value of the audio frame in multiple frequency sub-bands.

5. The method according to claim 4, characterized in that, The step of determining the current energy distribution information of the audio frame based on the historical energy distribution information corresponding to the audio frame and the energy values ​​of the audio frame in multiple frequency sub-bands includes: The energy values ​​of the audio frame in multiple frequency sub-bands are weighted by the second set parameters and the energy values ​​in the corresponding frequency sub-bands indicated by the historical energy distribution information, and the first energy distribution information is determined based on the result of the weighting process. The current energy distribution information is determined based on the energy values ​​in multiple frequency sub-bands indicated by the first energy distribution information and the sum of the energy values ​​in multiple frequency sub-bands indicated by the first energy distribution information. The second setting parameter is determined based on the sampling rate and preset duration of the audio sequence to be processed. The preset duration is the duration required to smooth the spectrum of each audio frame in the audio sequence to be processed.

6. The method according to claim 5, characterized in that, The step of weighting the energy values ​​of the audio frame in multiple frequency sub-bands with the energy values ​​in the corresponding frequency sub-bands indicated by the historical energy distribution information using a second set parameter, and determining the first energy distribution information based on the result of the weighting process, includes: The second set parameter is used as the weight of the energy value of the audio frame in multiple frequency sub-bands, and the energy value of the audio frame in multiple frequency sub-bands is weighted. The difference between the set parameter value and the second set parameter is used as the weight of the historical energy distribution information, and the historical energy distribution information is weighted. The sum of the results obtained from the two weighted processing is determined as the first energy distribution information.

7. The method according to claim 4, characterized in that, The step of responding to a volume adjustment operation on the audio sequence to be processed, and obtaining the current energy distribution information and historical energy distribution information corresponding to the target audio frame in the audio sequence to be processed, includes: In response to a volume adjustment operation on an audio sequence to be processed, if the total energy value of the received audio frame is greater than a set energy threshold, the step of determining the current energy distribution information of the audio frame based on the historical energy distribution information corresponding to the audio frame and the energy value of the audio frame in multiple frequency sub-bands is performed. Upon receiving the target audio frame and determining its current energy distribution information, the step of obtaining the current energy distribution information and historical energy distribution information corresponding to the target audio frame in the audio sequence to be processed is performed.

8. The method according to claim 4, characterized in that, Before determining the current energy distribution information of the audio frame based on the historical energy distribution information corresponding to the audio frame and the energy values ​​of the audio frame in multiple frequency sub-bands when the total energy value of the received audio frame is greater than a set energy threshold, the method further includes: If no volume adjustment operation is detected for the audio sequence to be processed, the historical audio adjustment information is determined as the target audio adjustment information.

9. The method according to claim 8, characterized in that, Without performing volume adjustment operations on the audio sequence to be processed, the vector curves with values ​​of 1 for multiple frequency domain sub-bands are used as the historical audio adjustment information.

10. The method according to claim 1, characterized in that, The method further includes: The noise energy value of the target audio frame is obtained, and the noise energy value is determined by the transmitting end of the audio sequence to be processed based on the target audio frame; Before obtaining the current energy distribution information and historical energy distribution information corresponding to the target audio frame in the audio sequence to be processed in response to a volume adjustment operation on the audio sequence to be processed, the method further includes: If the system volume level has not reached the maximum volume level and the noise energy value is not greater than the first energy threshold, the following steps are performed in response to the volume adjustment operation for the audio sequence to be processed: obtaining the current energy distribution information and historical energy distribution information corresponding to the target audio frame in the audio sequence to be processed.

11. The method according to claim 10, characterized in that, The method further includes: When the system volume level has not reached the maximum volume level and the noise energy value is greater than the first energy threshold and less than the second energy threshold, a second energy distribution difference information is determined based on the current energy distribution information of the target audio frame and the noise energy distribution information of the noise signal included in the target audio frame. The noise energy distribution information is used to indicate the energy value of the noise signal in multiple frequency domain sub-bands, and the second energy distribution difference information is used to indicate the energy difference between the energy distribution information and the noise energy distribution information in multiple frequency domain sub-bands. Based on the second energy distribution difference information, the overall energy value of the target audio frame, the noise energy value, and the first energy distribution information, the target audio adjustment information is determined. The first energy distribution information is determined based on the weighted processing result of the energy values ​​of the audio frame in multiple frequency sub-bands and the energy values ​​in the corresponding frequency sub-bands indicated by the historical energy distribution information.

12. The method according to claim 11, characterized in that, The method further includes: If the system volume level does not reach the maximum volume level, and the noise energy value is not less than the second energy threshold and not greater than the third energy threshold, a target energy difference is determined based on the total energy value of the target audio frame and the noise energy value. The total energy value of the target audio frame is the sum of the signal energy values ​​of the target audio frame in each frequency domain sub-band. Based on the target energy difference, the overall energy value of the target audio frame, the noise energy value, the first energy distribution information, and the first energy adjustment information, the target audio adjustment information is determined. The first energy adjustment information is a first preset energy value used to raise the energy value of the target audio frame at each frequency point.

13. The method according to claim 11, characterized in that, The method further includes: When the system volume level has reached the maximum volume level and the noise energy value is greater than the third energy threshold, the second energy distribution difference information is determined based on the current energy distribution information of the target audio frame and the noise energy distribution information of the noise signal included in the target audio frame. Based on the second energy distribution difference information, the overall energy value of the target audio frame, the noise energy value, the first energy distribution information, and the second energy adjustment information, the target audio adjustment information is determined. The second energy adjustment information is a second preset energy value used to raise the energy value of the audio frame at each frequency point.

14. The method according to any one of claims 1 to 13, characterized in that, After adjusting the energy values ​​of audio frames following the target audio frame in the audio sequence to be processed at multiple frequency points based on the target audio adjustment information, the method further includes at least one of the following: If the energy value of any adjusted audio frame at any frequency point is greater than the first set threshold, the energy value of the audio frame at that frequency point is adjusted so that the energy value of the audio frame at that frequency point is less than or equal to the first set threshold. If the energy value of any adjusted audio frame at any frequency point is less than the second set threshold, the energy value of the audio frame at that frequency point is adjusted so that the energy value of the audio frame at that frequency point is greater than or equal to the second set threshold.

15. An audio processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire the current energy distribution information and historical energy distribution information of the target audio frame in the audio sequence to be processed in response to the volume adjustment operation for the audio sequence to be processed. The target audio frame is the audio frame located at a preset frame number among the audio frames after volume adjustment. The current energy distribution information is used to indicate the energy value of the target audio frame in multiple frequency sub-bands, and the historical energy distribution information is used to indicate the energy value of the previous audio frame of the target audio frame in multiple frequency sub-bands. The determining module is used to determine first energy distribution difference information based on the current energy distribution information and historical energy distribution information corresponding to the target audio frame. The first energy distribution difference information is used to indicate the energy difference between the current energy distribution information and the historical energy distribution information in multiple frequency domain sub-bands. The determining module is further configured to determine target audio adjustment information based on historical audio adjustment information and the first energy distribution difference information. The historical audio adjustment information is the audio adjustment information determined in response to the previous volume adjustment operation for the audio sequence to be processed. The target audio adjustment information is used to adjust the energy value of any audio frame located after the target audio frame at multiple frequency points. The adjustment module is used to adjust the energy values ​​of each audio frame located after the target audio frame in the audio sequence to be processed at multiple frequency points based on the target audio adjustment information.

16. The apparatus according to claim 15, characterized in that, The determining module, when determining the target audio adjustment information based on historical audio adjustment information and the first energy distribution difference information, is used for: The first energy distribution difference information is weighted using a first set parameter, and the target audio adjustment information is determined based on the weighted first energy distribution difference information and the historical audio adjustment information.

17. The apparatus according to claim 16, characterized in that, The determining module, when weighting the first energy distribution difference information with a first set parameter and determining the target audio adjustment information based on the weighted first energy distribution difference information and the historical audio adjustment information, is configured to: Using the first set parameter as the weight of the first energy distribution difference information, the first energy distribution difference information is weighted to obtain the weighted first energy distribution information. The sum of the weighted first energy distribution information and the historical audio adjustment information is determined as the target audio adjustment information.

18. The apparatus according to claim 15, characterized in that, The determining module is further configured to determine the energy value of the audio frame in each frequency domain sub-band and the total energy value of the audio frame for each audio frame received in the audio sequence to be processed. The determining module is further configured to determine the current energy distribution information of the audio frame based on the historical energy distribution information corresponding to the audio frame and the energy value of the audio frame in multiple frequency sub-bands when the total energy value of the received audio frame is greater than a set energy threshold.

19. The apparatus according to claim 18, characterized in that, The determining module, when determining the current energy distribution information of the audio frame based on the historical energy distribution information corresponding to the audio frame and the energy values ​​of the audio frame in multiple frequency sub-bands, is used to: The energy values ​​of the audio frame in multiple frequency sub-bands are weighted by the second set parameters and the energy values ​​in the corresponding frequency sub-bands indicated by the historical energy distribution information, and the first energy distribution information is determined based on the result of the weighting process. The current energy distribution information is determined based on the energy values ​​in multiple frequency sub-bands indicated by the first energy distribution information and the sum of the energy values ​​in multiple frequency sub-bands indicated by the first energy distribution information. The second setting parameter is determined based on the sampling rate and preset duration of the audio sequence to be processed. The preset duration is the duration required to smooth the spectrum of each audio frame in the audio sequence to be processed.

20. The apparatus according to claim 19, characterized in that, The determining module, when performing weighted processing on the energy values ​​of the audio frame in multiple frequency sub-bands and the energy values ​​in the corresponding frequency sub-bands indicated by the historical energy distribution information using a second set parameter, and determining the first energy distribution information based on the result of the weighted processing, is configured to: The second set parameter is used as the weight of the energy value of the audio frame in multiple frequency sub-bands, and the energy value of the audio frame in multiple frequency sub-bands is weighted. The difference between the set parameter value and the second set parameter is used as the weight of the historical energy distribution information, and the historical energy distribution information is weighted. The sum of the results obtained from the two weighted processing is determined as the first energy distribution information.

21. The apparatus according to claim 18, characterized in that, The acquisition module, when acquiring the current energy distribution information and historical energy distribution information corresponding to the target audio frame in the audio sequence to be processed in response to a volume adjustment operation on the audio sequence to be processed, is used for: In response to a volume adjustment operation on an audio sequence to be processed, if the total energy value of the received audio frame is greater than a set energy threshold, the step of determining the current energy distribution information of the audio frame based on the historical energy distribution information corresponding to the audio frame and the energy value of the audio frame in multiple frequency sub-bands is performed. Upon receiving the target audio frame and determining its current energy distribution information, the step of obtaining the current energy distribution information and historical energy distribution information corresponding to the target audio frame in the audio sequence to be processed is performed.

22. The apparatus according to claim 18, characterized in that, The determining module is further configured to determine the historical audio adjustment information as the target audio adjustment information when no volume adjustment operation for the audio sequence to be processed is detected.

23. The apparatus according to claim 22, characterized in that, Without adjusting the volume of the audio sequence to be processed, the vector curves with values ​​of 1 for multiple frequency sub-bands are used as the historical audio adjustment information.

24. The apparatus according to claim 15, characterized in that, The acquisition module is further configured to acquire the noise energy value of the target audio frame, wherein the noise energy value is determined by the transmitting end of the audio sequence to be processed based on the target audio frame; The acquisition module is further configured to perform a step of acquiring the current energy distribution information and historical energy distribution information corresponding to the target audio frame in the audio sequence to be processed, in response to a volume adjustment operation for the audio sequence to be processed, when the system volume level has not reached the maximum volume level and the noise energy value is not greater than the first energy threshold.

25. The apparatus according to claim 24, characterized in that, The determining module is further configured to, when the system volume level has not reached the maximum volume level and the noise energy value is greater than the first energy threshold and less than the second energy threshold, determine second energy distribution difference information based on the current energy distribution information of the target audio frame and the noise energy distribution information of the noise signal included in the target audio frame. The noise energy distribution information is used to indicate the energy value of the noise signal in multiple frequency domain sub-bands, and the second energy distribution difference information is used to indicate the energy difference between the energy distribution information and the noise energy distribution information in multiple frequency domain sub-bands. The determining module is further configured to determine the target audio adjustment information based on the second energy distribution difference information, the overall energy value of the target audio frame, the noise energy value, and the first energy distribution information. The first energy distribution information is determined based on the weighted processing result of the energy values ​​of the audio frame in multiple frequency sub-bands and the energy values ​​in the corresponding frequency sub-bands indicated by the historical energy distribution information.

26. The apparatus according to claim 25, characterized in that, The determining module is further configured to determine a target energy difference based on the total energy value of the target audio frame and the noise energy value when the system volume level has not reached the maximum volume level and the noise energy value is not less than the second energy threshold and not greater than the third energy threshold. The total energy value of the target audio frame is the sum of the signal energy values ​​of the target audio frame in each frequency domain sub-band. The determining module is further configured to determine the target audio adjustment information based on the target energy difference, the overall energy value of the target audio frame, the noise energy value, the first energy distribution information, and the first energy adjustment information, wherein the first energy adjustment information is a first preset energy value used to raise the energy value of the target audio frame at each frequency point.

27. The apparatus according to claim 25, characterized in that, The determining module is further configured to determine the second energy distribution difference information based on the current energy distribution information of the target audio frame and the noise energy distribution information of the noise signal included in the target audio frame, when the system volume level has reached the maximum volume level and the noise energy value is not less than the third energy threshold. The determining module is further configured to determine the target audio adjustment information based on the second energy distribution difference information, the overall energy value of the target audio frame, the noise energy value, the first energy distribution information, and the second energy adjustment information, wherein the second energy adjustment information is a second preset energy value used to raise the energy value of the audio frame at each frequency point.

28. The apparatus according to any one of claims 15 to 27, characterized in that, The adjustment module is further configured to adjust the energy value of the audio frame at any frequency point when the energy value of any adjusted audio frame at any frequency point is greater than the first set threshold, so that the energy value of the audio frame at the frequency point is less than or equal to the first set threshold. The adjustment module is further configured to adjust the energy value of the audio frame at any frequency point when the energy value of any adjusted audio frame at any frequency point is less than the second set threshold, so that the energy value of the audio frame at the frequency point is greater than or equal to the second set threshold.

29. A computing device, characterized in that, The computing device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to perform the operations performed by the audio processing method as described in any one of claims 1 to 14.

30. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that is executed by a processor as the audio processing method as described in any one of claims 1 to 14.

Citation Information

Patent Citations

  • Processing method, device, apparatus and storage medium for volume adjustment

    CN109240637A

  • Earphone, automatic volume adjustment control module and method thereof and storage medium

    CN110806850A