Audio processing method, electronic device, and storage medium

By classifying and weighting audio, the high-frequency and low-frequency parts of the audio can be flexibly repaired, solving audio quality issues caused by low-bitrate encoding and improving the listening experience and repair effect after audio decoding.

WO2025199960A1PCT designated stage Publication Date: 2025-10-02AAC ACOUSTIC TECH (SHANGHAI) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/084857
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-29
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

In existing audio coding and decoding technologies, low-bitrate encoding results in poor audio quality, especially after the high-frequency and low-frequency information are lost, which impairs the hearing experience and cannot meet the user's playback effect requirements.

Method used

Classify the audio according to its content, determine the high-frequency and low-frequency repair weights, and flexibly repair the high-frequency and low-frequency parts of the audio through amplitude superposition and phase update to avoid excessive high-frequency restoration.

Benefits of technology

It improves the listening quality of audio after decoding, restores the high-frequency and low-frequency information in low-bitrate encoded audio, and achieves better audio repair effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024084857_02102025_PF_FP_ABST
    Figure CN2024084857_02102025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to the technical field of audios and videos, and disclose an audio processing method, an electronic device, and a storage medium. The audio processing method comprises: classifying an audio on the basis of the content of the audio; determining a high-frequency restoration weight and a low-frequency restoration weight on the basis of the classification result; on the basis of the high-frequency restoration weight and the low-frequency restoration weight, performing amplitude superposition on the audio having undergone frequency band extension and the audio having undergone low-frequency restoration; and updating a phase having a frequency higher than a cut-off frequency in the amplitude superposition result into a corresponding low-frequency phase having a frequency lower than the cut-off frequency, so as to obtain a restored audio. The present application is at least conducive to improving the audio quality.
Need to check novelty before this filing date? Find Prior Art

Description

Audio processing method, electronic device and storage medium Technical Field

[0001] The embodiments of the present application relate to the field of audio and video technology, and in particular to an audio processing method, electronic device, and storage medium. Background Art

[0002] In the digital media world, audio and video data is represented and stored in digital form. To achieve efficient storage and transmission, audio and video data must be encoded and compressed. This process converts the original audio and video data into a compressed bitstream, reducing data size and improving transmission efficiency. Correspondingly, before audio playback, decoding is required to restore the encoded data to the original audio and video signal.

[0003] However, the audio quality obtained by decoding is currently poor, and the playback effect is not satisfactory to users. Summary of the Invention

[0004] The embodiments of the present application provide an audio processing method, an electronic device, and a storage medium, which are at least beneficial to improving audio quality.

[0005] According to some embodiments of the present application, on the one hand, the embodiments of the present application provide an audio processing method, including: classifying the audio according to the content of the audio; determining the high-frequency repair weight and the low-frequency repair weight according to the classification result; performing amplitude superposition on the audio after frequency band expansion and the audio after low-frequency repair according to the high-frequency repair weight and the low-frequency repair weight; updating the phase above the cutoff frequency in the result of the amplitude superposition to the corresponding low-frequency phase below the cutoff frequency to obtain the repaired audio.

[0006] In some embodiments, when the audio category is mixed sound, determining the high-frequency repair weight and the low-frequency repair weight according to the classification result includes: determining two groups of the high-frequency repair weights and the low-frequency repair weights; performing amplitude superposition on the audio after frequency band expansion and the audio after low-frequency repair according to the high-frequency repair weights and the low-frequency repair weights, including: separating the audio to obtain foreground sound and background sound; performing amplitude superposition on the foreground sound after frequency band expansion and the foreground sound after low-frequency repair according to one group of the high-frequency repair weights and the low-frequency repair weights, and performing amplitude superposition on the background sound after frequency band expansion and the background sound after low-frequency repair according to another group of the high-frequency repair weights and the low-frequency repair weights.

[0007] In some embodiments, the value of the high-frequency repair weight used for superimposing the foreground sound after frequency band expansion and the foreground sound after low-frequency repair approaches the left boundary of the value range of the high-frequency repair weight, and the value of the low-frequency repair weight approaches the right boundary of the value range of the low-frequency repair weight; and / or, the value of the high-frequency repair weight used for superimposing the background sound after frequency band expansion and the background sound after low-frequency repair approaches the right boundary of the value range of the high-frequency repair weight, and the value of the low-frequency repair weight approaches the left boundary of the value range of the low-frequency repair weight.

[0008] In some embodiments, when the audio category is music, the value of the high-frequency repair weight approaches the right boundary of the value range of the high-frequency repair weight, and the value of the low-frequency repair weight approaches the left boundary of the value range of the low-frequency repair weight; or, when the audio category is human voice, the value of the high-frequency repair weight approaches the left boundary of the value range of the high-frequency repair weight, and the value of the low-frequency repair weight approaches the right boundary of the value range of the low-frequency repair weight; or, when the audio category is noise, the value of the high-frequency repair weight approaches the left boundary of the value range of the high-frequency repair weight, and the value of the low-frequency repair weight approaches the left boundary of the value range of the low-frequency repair weight.

[0009] In some embodiments, before amplitude superposition is performed on the audio after frequency band expansion and the audio after low-frequency repair according to the high-frequency repair weight and the low-frequency repair weight, the method further includes: frequency band expansion of the audio according to the encoding method and bit rate used in audio encoding, as well as a frequency band extension model; and low-frequency repair of the audio according to the encoding method and bit rate used in audio encoding, as well as a low-frequency repair model.

[0010] In some embodiments, classifying the audio according to the content of the audio includes classifying the audio according to the content of historical frames and / or current frames in the audio.

[0011] In some embodiments, updating the phase in the result of the amplitude superposition that is higher than the cutoff frequency to a corresponding low-frequency phase that is lower than the cutoff frequency includes: determining the phase corresponding to the result of the amplitude superposition according to the following expression: ,in, represents the phase corresponding to the result of the amplitude superposition, represents the phase of the audio frequency, is the cutoff frequency of the audio.

[0012] In some embodiments, superimposing the band-extended audio and the low-frequency repaired audio according to the high-frequency repair weight and the low-frequency repair weight includes superimposing using the following expression: ,in, represents the result of amplitude superposition, represents the amplitude of the audio after frequency band expansion, Indicates the amplitude of the audio after low-frequency repair, represents the high-frequency restoration weight, represents the low-frequency restoration weight, Indicates the audio The amplitude of 、 The value range of is (0, 1).

[0013] According to some embodiments of the present application, on the other hand, the embodiments of the present application further provide an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the audio processing method as described in any embodiment of the present application.

[0014] According to some embodiments of the present application, on the other hand, embodiments of the present application further provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the audio processing method as described in any embodiment of the present application.

[0015] The technical solution provided by the embodiments of the present application has at least the following advantages:

[0016] By classifying the audio according to its content, the appropriate high-frequency and low-frequency restoration weights can be accurately determined using the classification results. This allows the amplitude of the band-expanded audio and the low-frequency restoration audio to be superimposed based on the accurate high-frequency and low-frequency restoration weights, enabling more accurate and independent control of the restoration effects of the high-frequency and low-frequency parts of the audio. Furthermore, the phase of the audio above the cutoff frequency is updated to the corresponding low-frequency phase below the cutoff frequency, avoiding over-restoration of the high-frequency part of the audio. This is conducive to achieving better audio restoration effects and improving audio quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0018] FIG1 is an audio spectrum before encoding provided in this application;

[0019] FIG2 is an encoded audio spectrum provided in this application;

[0020] FIG3 is another encoded audio spectrum provided in this application;

[0021] FIG4 is another encoded audio spectrum provided in this application;

[0022] FIG5 is a flowchart of an audio method provided in an embodiment of the present application;

[0023] FIG6 is a second flowchart of the audio method provided in an embodiment of the present application;

[0024] FIG7 is a third flowchart of the audio method provided in an embodiment of the present application;

[0025] FIG8 is a schematic diagram of a flow chart of separating foreground sound and background sound involved in an audio method provided in an embodiment of the present application;

[0026] FIG9 is a fourth flowchart of the audio method provided in an embodiment of the present application;

[0027] FIG10 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. Modes for Carrying Out the Invention

[0028] As can be seen from the background technology, the current audio coding and decoding technology has the problem that the quality of the decoded audio is poor and needs to be improved urgently.

[0029] After analysis, the above problems are caused at least by the fact that, due to storage space or transmission bandwidth limitations, audio is often encoded at a low bitrate when encoding and decoding. Existing encoding schemes selectively ignore some information in the audio at low bitrates, which changes the listening experience of the audio. In particular, as the encoding bitrate decreases, more information is lost, and accordingly, the audio quality becomes worse. Among them, the loss of listening experience after the sound source is encoded at a low bitrate can be divided into the following two categories:

[0030] 1. High-frequency information. Specifically, in low-bitrate encoding, in order to reduce the size of the encoded file, high-frequency information is often discarded, which results in only low-frequency parts in the decoded audio, which often sounds rough and muffled. For example, comparing the audio spectrum before encoding shown in Figure 1 and the audio spectrum after 64kbps MP3 encoding shown in Figure 2, it can be seen that the high-frequency parts above 10kHz in the encoded audio are completely lost, and the mid-to-high-frequency parts between 6kHz and 10kHz are also greatly lost. The horizontal axes of Figures 1 and 2 are both the timestamps of the audio, and the vertical axes are the frequencies.

[0031] 2. Low-frequency information. During the audio encoding process, a quantizer is used to quantize the original audio signal. When the encoding bit rate is low, the quantizer's precision is often set very low. This makes it difficult for the quantized dynamic range to match the actual signal, causing some signal frequencies to be quantized to 0 or 1, a phenomenon known as birdies. Birdies cause spectral gaps or islands in the spectrum, which in turn affects the listening experience of the decoded audio. Furthermore, low-bitrate encoding often takes into account the human ear's auditory masking effect. Due to this masking effect, the human ear is less sensitive to certain audio information, so discarding this information has a smaller impact on the listening experience. Even after discarding this information, spectral gaps and islands will still occur. However, since the masking effect of hearing varies from person to person, and sensitivity to discarded information varies, the listening experience of low-bitrate audio can still be perceived as reduced due to missing information. For example, comparing the audio spectrum before encoding shown in Figure 1 and the audio spectrum after 64kbps MP3 encoding shown in Figures 3 and 4, the encoded audio spectrum is no longer completely continuous, but has gaps, namely the aforementioned spectral gaps or spectral islands.

[0032] To this end, the embodiments of the present application provide an audio processing method, an electronic device, and a storage medium, which flexibly determine high-frequency repair weights and low-frequency repair weights for the audio based on the content of the audio, so as to control the high-frequency repair and low-frequency repair effects of the audio, so that the final repair effect of the audio can be adapted to its content, thereby better restoring the missing high-frequency information, masking information, and spatial information in the low-bitrate audio, and improving the listening experience after decoding the low-bitrate encoded audio.

[0033] To make the purpose, technical solutions, and advantages of the embodiments of the present application more clear, each embodiment of the present application will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will appreciate that many technical details are provided in each embodiment of the present application to help readers better understand the present application. However, even without these technical details and various variations and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented.

[0034] The following embodiments are divided for the convenience of description and should not constitute any limitation on the specific implementation of the present application. The various embodiments can be combined with each other and referenced to each other without contradiction.

[0035] In one aspect, embodiments of the present application provide an audio processing method that can be applied to any electronic device, such as a mobile phone, a computer, a music player, etc. In some embodiments, the process of the audio processing method is shown in FIG5 , and includes the following steps:

[0036] Step 501: classify the audio according to the content of the audio.

[0037] Step 502: Determine the high-frequency restoration weight and the low-frequency restoration weight according to the classification result.

[0038] Step 503 : performing amplitude superposition on the band-expanded audio and the low-frequency repaired audio according to the high-frequency repair weight and the low-frequency repair weight.

[0039] Step 504: Update the phase above the cutoff frequency in the amplitude superposition result to the corresponding low-frequency phase below the cutoff frequency to obtain the repaired audio.

[0040] In this way, the audio is classified according to its content, so that the appropriate high-frequency and low-frequency repair weights can be accurately determined using the classification results. Thus, the audio after band expansion and the audio after low-frequency repair are amplitude-superimposed based on the accurate high-frequency and low-frequency repair weights, making it possible to more accurately and independently control the repair effects of the high-frequency and low-frequency parts of the audio. In addition, the phase of the audio above the cutoff frequency is updated to the corresponding low-frequency phase below the cutoff frequency, avoiding excessive restoration of the high-frequency part of the audio. This is conducive to achieving better audio repair effects and improving audio quality.

[0041] To facilitate those skilled in the art to better understand the embodiment shown in FIG5 , it will be described below.

[0042] Regarding step 501, it is understandable that the audio is composed of different audio frames. Therefore, in some examples, all audio frames included in the audio can be classified, so that the audio processing can be completed at one time, which is conducive to more efficient completion of the audio processing. However, for a specific frame, the accuracy of its classification will be reduced. Therefore, in other examples, the audio can be processed frame by frame, or multiple consecutive frames, or multiple consecutive frames with basically unchanged content, can be classified and subsequently processed as a whole to improve the processing accuracy of each audio frame, which is conducive to further improving the processing effect of each frame of audio and achieving better audio quality.

[0043] Furthermore, it is also understood that when the audio is composed of different audio frames, the audio frames contained in the audio generally have content connections, that is, the previous audio frame can also provide a reference for the classification of the subsequent audio frames. Based on this, in some embodiments, as shown in FIG6 , audio classification according to the content of the audio can be achieved by the following steps:

[0044] Step 5011: classify the audio according to the content of the historical frames and / or the current frame in the audio.

[0045] That is to say, historical frames can also be referenced during classification, such as the previous frame or multiple frames before the current frame. In this way, the category of the audio frame can be estimated in advance using the reference provided by the historical frame, so that the subsequent processing of the audio frame can be started more efficiently, which is beneficial for application in scenarios such as audio streams that have high requirements for real-time audio processing, and can reduce audio processing delays and improve user experience. In the case of using the current frame and historical frames for classification at the same time, since the classification refers to more information, the accuracy of the classification can be improved, thereby ensuring the accuracy of subsequent audio processing, which is beneficial to improving the audio repair effect and thus improving the audio quality.

[0046] It should be noted that this embodiment does not limit the classification method. It can be direct content classification, or it can be sound source switching recognition, or a combination of the two, etc., which will not be described in detail here.

[0047] It should also be noted that the embodiments of this application do not limit the various situations and specific number of audio categories. It is understandable that it can be any classification or category that is conducive to distinguishing the proportion of high-frequency content and low-frequency content in the audio. In this way, the subsequent steps can determine the high-frequency repair weight and low-frequency repair weight based on the proportion of high-frequency content and low-frequency content reflected by the classification.

[0048] In some embodiments, the audio category may include at least one of the following: music, human voice, noise, and mixed sound, etc. For example, based on the current audio content and the previous audio content, the audio may be divided into four categories: pure music, pure human voice, noise, and mixed sound.

[0049] It's understandable that different audio content will have different high-frequency and low-frequency characteristics, and the corresponding loss information will also vary. Therefore, different audio should have different emphases on high-frequency and low-frequency restoration. For example, music often has more original high-frequency components. Therefore, after encoding at a low bitrate, the high-frequency components in the original music will be lost, seriously affecting the listening experience. In other words, restoration should focus on using high-frequency extension to increase the high-frequency components, which will be more conducive to improving the listening quality. The energy and effective information of the human voice are mainly concentrated in the mid- and low-frequency ranges. Therefore, the main loss after encoding is the information loss caused by the Birdies phenomenon caused by quantization. In other words, restoration should focus on suppressing the Birdies phenomenon through low-frequency restoration, which will be more conducive to improving the listening quality. Noise generally has no practical significance for restoration, so neither high nor low frequencies are necessary.

[0050] Based on this, in some embodiments, for step 502, when the audio category is music, the high-frequency repair weight can be set to a number close to the right boundary of the high-frequency repair weight range, and the low-frequency repair weight can be set to a number close to the left boundary of the low-frequency repair weight range. In this way, the fact that high-frequency content in music generally has a large proportion and the resulting loss mainly comes from the high-frequency part can be fully considered. High-frequency repair is performed to the greatest extent to achieve the best repair effect, while low-frequency repair is performed to the minimum extent to highlight the high-frequency content, so that the music can have a better high-frequency playback effect, thereby presenting a better listening experience.

[0051] In some embodiments, when the audio is classified as human voice, the high-frequency repair weight can be set to a number close to the left boundary of the high-frequency repair weight range, and the low-frequency repair weight can be set to a number close to the right boundary of the low-frequency repair weight range. This fully considers the characteristics of human voices, which are mainly low- and mid-frequency, and maximizes the repair of the lost low-frequency portion to achieve the best repair effect, while minimizing high-frequency repair to prevent high-frequency content from interfering with low-frequency content. This ensures better playback quality of human voice content and presents a better listening experience.

[0052] In some embodiments, when the audio is classified as noise, the high-frequency repair weight can be set to a number close to the left edge of the high-frequency repair weight range, and the low-frequency repair weight can be set to a number close to the left edge of the low-frequency repair weight range. This allows noise repair to follow the same processing flow as music, vocals, and other repairs, avoiding meaningless repair operations, reducing resource waste, and preventing incorrect audio adjustments, such as adjustments to white noise for sleep.

[0053] Of course, in some cases, when the audio category is noise, no repair is required, which will not be discussed here.

[0054] Similarly, mixed audio—audio resulting from various combinations of different sounds, such as music and vocals, or vocals and noise—is a problem. Background audio typically consists of continuous, stable sounds, such as ambient sounds or quiet musical accompaniment, while foreground audio originates from prominent, direct sources, including human voices, singing, and large musical instruments. Therefore, band-stretching the foreground audio can easily lead to distortion and increased auditory roughness, impacting the listening experience. Therefore, in some examples, when the audio is classified as mixed audio, it can be considered a superposition of different content. Specifically, in a mixed audio, the foreground sound is primarily transient, while the background sound is primarily steady-state. Transient sounds are not suitable for high-frequency extension and restoration, as this can easily introduce noise. Meanwhile, the birds-of-the-dark effect in steady-state sounds is less impactful, resulting in a minimal improvement in low-frequency restoration. In other words, the foreground and background sounds in a mixed audio have different characteristics, requiring different emphasis on high-frequency extension and low-frequency restoration, making the same restoration process inappropriate. Therefore, in order to better meet the restoration needs of different contents in the mixed sound, the mixed sound can be separated, and then the separated foreground sound and background sound can be processed separately to avoid mutual influence between the restoration of the two.

[0055] Based on this, in some embodiments, as shown in FIG7 , when the audio category is mixed sound, determining the high-frequency repair weight and the low-frequency repair weight according to the classification result can be achieved by the following steps:

[0056] Step 5021: Determine two sets of high-frequency repair weights and low-frequency repair weights.

[0057] Accordingly, according to the high-frequency repair weight and the low-frequency repair weight, the amplitude of the band-expanded audio and the low-frequency repaired audio are superimposed, which can be achieved by the following steps:

[0058] Step 5031: Separate the audio to obtain foreground sound and background sound.

[0059] Step 5032: Based on a set of high-frequency repair weights and low-frequency repair weights, the amplitudes of the foreground sound after band expansion and the foreground sound after low-frequency repair are superimposed, and based on another set of high-frequency repair weights and low-frequency repair weights, the amplitudes of the background sound after band expansion and the background sound after low-frequency repair are superimposed.

[0060] In this way, by separating the audio, we can obtain foreground sounds mainly composed of mid- and low-frequency sounds and background sounds mainly composed of high-frequency sounds, so that we can restore the corresponding loss characteristics and avoid mutual interference between the repair of mid- and low-frequency content and high-frequency content, which is beneficial to improving the audio repair effect and further improving the audio quality.

[0061] It should be noted that the embodiments of the present application do not limit the method of separating audio. In some embodiments, the foreground sound can be mainly divided into two parts: one is the transient sound (Transient Signal) in the audio, and the other is the tonal signal (Tonal Signal) caused by voice or musical instruments. That is to say, the foreground sound can be removed from the signal by attenuating the above two parts separately in the original audio to obtain the background sound, thereby realizing the separation of foreground sound and background sound.

[0062] In some embodiments, as shown in Figure 8, after the audio signal X(k) undergoes transient attenuation and tonal attenuation, the spectrum is combined to produce the background signal. Furthermore, the foreground sound can be obtained by removing the background sound from the audio signal. Assuming the signal gain due to transient attenuation is G_tran and the signal gain due to tonal attenuation is G_tona, the signal gain of the background sound relative to the audio signal is: G = min⁡(G_tran, G_tona). That is, the signal spectrum of the background sound is: |B(k)| = |X(k)| * G. Correspondingly, the signal spectrum of the foreground sound is |F(k)| = |X(k)| - |B(k)|.

[0063] As for the two groups of high-frequency repair weights and low-frequency repair weights in the above-mentioned embodiment, the corresponding high-frequency repair weights and low-frequency repair weights can be flexibly set according to the different characteristics of the repair requirements of different contents in the mixed sound mentioned above. That is, in some embodiments, the value of the high-frequency repair weight used for the superposition of the foreground sound after frequency band expansion and the foreground sound after low-frequency repair approaches the left boundary of the value range of the high-frequency repair weight, and the value of the low-frequency repair weight approaches the right boundary of the value range of the low-frequency repair weight, so as to fit the mid- and low-frequency characteristics of the foreground sound, which is conducive to achieving a better repair effect. In some embodiments, the value of the high-frequency repair weight used for the superposition of the background sound after frequency band expansion and the background sound after low-frequency repair approaches the right boundary of the value range of the high-frequency repair weight, and the value of the low-frequency repair weight approaches the left boundary of the value range of the low-frequency repair weight, so as to fit the high-frequency characteristics of the background sound, which is conducive to achieving a better repair effect.

[0064] Regarding step 503, the embodiment of the present application does not limit the audio Blind Bandwidth Extension (BWE) method and Low-frequency Restoration (LFR) method. It can be any solution that achieves the corresponding effect.

[0065] In some embodiments, as shown in FIG9 , before amplitude superposition is performed on the band-extended audio and the low-frequency repaired audio based on the high-frequency repair weight and the low-frequency repair weight, the audio processing method further includes the following steps:

[0066] Step 505 : performing frequency band expansion on the audio according to the encoding method and bit rate used in audio encoding, as well as the frequency band expansion model.

[0067] Step 506: Perform low-frequency repair on the audio according to the encoding method and bit rate used in audio encoding, as well as the low-frequency repair model.

[0068] That is, as shown in Figure 10, two models (i.e., a frequency band extension model and a low-frequency repair extension model) can be used to implement the relevant repairs respectively, where X is the audio, and the prior information is the encoding method and bit rate used in audio encoding. The encoding methods include MP3, Advanced Audio Coding (AAC), Opus, etc., and the bit rates include 64kbps, 96kbps, 128kbps, etc.

[0069] It should be noted that the reason why the model uses encoding mode and bit rate as prior information as input is mainly because different encoding modes and bit rates generally have different cutoff frequencies and degrees of low-frequency loss. Therefore, by providing encoding mode and bit rate as prior information, the model can select more accurate parameters or configurations for audio repair, which is conducive to more accurate repair and further improving the audio repair effect, that is, improving the audio quality.

[0070] To help those skilled in the art better understand the amplitude superposition solution, the following will illustrate step 503 and step 504. It should be emphasized that the following description is only an example and does not mean that step 503 and step 504 can only be implemented in the following manner.

[0071] In some embodiments, the band-extended audio and the low-frequency repaired audio are superimposed based on the high-frequency repair weight and the low-frequency repair weight, which is achieved by the following expression:

[0072] ,

[0073] in, represents the result of amplitude superposition, Indicates the amplitude of the audio after band expansion. Indicates the amplitude of the audio after low-frequency repair. represents the high-frequency repair weight, represents the low-frequency repair weight, Indicates the amplitude of audio X.

[0074] In some embodiments, 、 The value range of is set to (0, 1). Of course, in other embodiments, it can also be set according to the needs 、 The specific value range of will not be described in detail here.

[0075] In this way, the amplitude of the audio is also introduced into the superimposed amplitude, and 、 The value range of is set to (0, 1) to achieve moderate repair of the audio, avoiding excessive adjustment of the repaired signal and insufficient amplitude of the repaired signal.

[0076] Among them, for the above given 、 The value range of the audio category can be set to music. Take a value close to 1, The value is close to 0; the audio can be classified as voice Take a value close to 0, The value is close to 1; the audio category can be classified as noise Take a value close to 0, The value is close to 0.

[0077] In some embodiments, the phase above the cutoff frequency in the result of the amplitude superposition is modified to a low-frequency phase below the cutoff frequency, which is achieved according to the following expression:

[0078] ,

[0079] in, Indicates the phase corresponding to the result of amplitude superposition, Indicates the phase of the audio, is the cutoff frequency of the audio.

[0080] Therefore, the signal obtained based on the amplitude obtained by the above superposition and the determined phase can be transformed into the time domain signal y through inverse Fourier transform, that is, the processed high-quality audio.

[0081] The steps of the various methods above are divided only for the purpose of clear description. During implementation, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this patent. Adding insignificant modifications or introducing insignificant designs to the algorithm or process without changing the core design of the algorithm and process are all within the scope of protection of this patent.

[0082] On the other hand, an embodiment of the present application further provides an electronic device, as shown in Figure 10, comprising: at least one processor 1001; and a memory 1002 communicatively connected to the at least one processor 1001; wherein the memory 1002 stores instructions that can be executed by the at least one processor 1001, and the instructions are executed by the at least one processor 1001 so that the at least one processor 1001 can execute the audio processing method described in any of the above method embodiments.

[0083] The memory 1002 and processor 1001 are connected using a bus. The bus may include any number of interconnected buses and bridges, connecting various circuits of one or more processors 1001 and memory 1002. The bus may also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits. These are all well known in the art and are therefore not described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver may be a single component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor 1001 is transmitted over a wireless medium via an antenna. Furthermore, the antenna receives data and transmits it to the processor 1001.

[0084] The processor 1001 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory 1002 can be used to store data used by the processor 1001 when performing operations.

[0085] Another aspect of the present application further provides a computer-readable storage medium storing a computer program that implements the above method embodiment when executed by a processor.

[0086] That is, those skilled in the art will understand that all or part of the steps in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a program. The program is stored in a storage medium and includes a number of instructions for causing a device (which may be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps in the methods described in the various embodiments of this application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0087] Those skilled in the art will appreciate that the above embodiments are specific embodiments for implementing the present application, and that in actual applications, various changes may be made thereto in form and detail without departing from the spirit and scope of the present application.

Claims

1. An audio processing method, comprising: Classifying the audio according to the content of the audio; According to the classification results, the high-frequency repair weight and the low-frequency repair weight are determined; Performing amplitude superposition on the audio after frequency band expansion and the audio after low-frequency repair according to the high-frequency repair weight and the low-frequency repair weight; The phase above the cutoff frequency in the result of the amplitude superposition is updated to the corresponding low-frequency phase below the cutoff frequency to obtain the repaired audio.

2. The audio processing method according to claim 1, wherein: In a case where the audio category is mixed sound, determining the high-frequency restoration weight and the low-frequency restoration weight according to the classification result includes: Determining two groups of the high-frequency repair weights and the low-frequency repair weights; The step of performing amplitude superposition on the audio after frequency band expansion and the audio after low-frequency repair according to the high-frequency repair weight and the low-frequency repair weight includes: Separating the audio to obtain foreground sound and background sound; According to a set of high-frequency repair weights and the low-frequency repair weights, the amplitudes of the foreground sound after frequency band expansion and the foreground sound after low-frequency repair are superimposed, and according to another set of high-frequency repair weights and the low-frequency repair weights, the amplitudes of the background sound after frequency band expansion and the background sound after low-frequency repair are superimposed.

3. The audio processing method according to claim 2, wherein: The value of the high-frequency repair weight used for superimposing the foreground sound after frequency band expansion and the foreground sound after low-frequency repair approaches the left boundary of the value range of the high-frequency repair weight, and the value of the low-frequency repair weight approaches the right boundary of the value range of the low-frequency repair weight; and / or, the value of the high-frequency repair weight used for superimposing the background sound after frequency band expansion and the background sound after low-frequency repair approaches the right boundary of the value range of the high-frequency repair weight, and the value of the low-frequency repair weight approaches the left boundary of the value range of the low-frequency repair weight.

4. The audio processing method according to any one of claims 1 to 3, wherein: When the audio category is music, the value of the high-frequency repair weight approaches the right boundary of the value range of the high-frequency repair weight, and the value of the low-frequency repair weight approaches the left boundary of the value range of the low-frequency repair weight; or, When the audio category is human voice, the value of the high-frequency repair weight approaches the left boundary of the value range of the high-frequency repair weight, and the value of the low-frequency repair weight approaches the right boundary of the value range of the low-frequency repair weight; or, When the audio category is noise, the value of the high-frequency repair weight approaches the left boundary of the value range of the high-frequency repair weight, and the value of the low-frequency repair weight approaches the left boundary of the value range of the low-frequency repair weight.

5. The audio processing method according to any one of claims 1 to 4, wherein: Before performing amplitude superposition on the band-expanded audio and the low-frequency repaired audio according to the high-frequency repair weight and the low-frequency repair weight, the method further includes: Performing frequency band expansion on the audio according to the encoding mode and bit rate used in audio encoding, and a frequency band expansion model; The audio is repaired at a low frequency according to the encoding mode and bit rate used in audio encoding, and the low frequency repair model.

6. The audio processing method according to any one of claims 1 to 5, wherein: Classifying the audio according to the content of the audio includes: The audio is classified according to content of historical frames and / or current frames in the audio.

7. The audio processing method according to any one of claims 1 to 6, wherein: The updating of the phase higher than the cutoff frequency in the result of the amplitude superposition to the corresponding low-frequency phase lower than the cutoff frequency includes: The phase corresponding to the result of the amplitude superposition is determined according to the following expression: , in, represents the phase corresponding to the result of the amplitude superposition, represents the phase of the audio frequency, is the cutoff frequency of the audio.

8. The audio processing method according to any one of claims 1 to 7, wherein: The step of superimposing the audio after frequency band expansion and the audio after low-frequency repair according to the high-frequency repair weight and the low-frequency repair weight includes: The superposition is done by the following expression: , in, represents the result of amplitude superposition, represents the amplitude of the audio after frequency band expansion, Indicates the amplitude of the audio after low-frequency repair, represents the high-frequency restoration weight, represents the low-frequency restoration weight, Indicates the audio The amplitude of 、 The value range of is (0, 1).

9. An electronic device comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the audio processing method according to any one of claims 1 to 8. 10 . A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the audio processing method according to claim 1 is implemented.

Citation Information

Patent Citations

  • Method and apparatus of band spreading

    CN108172239A

  • Frequency band expansion method and device, electronic equipment and computer readable storage medium

    CN110556123A

  • Method and apparatus for high frequency decoding for bandwidth extension

    CN111312277A

  • Bandwidth expansion method and device, medium and equipment

    CN116189693A

  • Multi-band limiter, sound recording apparatus, and program storage medium

    US20170117864A1