Audio frequency processing methods, apparatus, electronic equipment, computer-readable storage media, and computer program products
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2023-04-27
- Publication Date
- 2026-08-04
AI Technical Summary
【0014】 本願実施例は以下の有益な効果を有している。
Smart Images

Figure 0007900494000028 
Figure 0007900494000029 
Figure 0007900494000030
Abstract
Description
[Technical Field]
[0001] Cross-citation of related applications This application is filed based on a Chinese patent application with application number 202210681060.9 and filing date June 15, 2022, and seeks priority from the Chinese patent application. The entire contents of the Chinese patent application are incorporated herein by reference.
[0002] This application relates to audio frequency processing technology, and more particularly to audio frequency processing methods, apparatus, electronic devices, computer-readable storage media, and computer program products. [Background technology]
[0003] In ultra-wideband audio frequency coding scenarios, according to the auditory mechanisms of the human ear and psychoacoustic models, users are typically more sensitive to the low-frequency portions of a signal than to the high-frequency portions. During coding and decoding processes, more code rate is allocated to the low-frequency portions of the signal compared to the high-frequency portions. However, this does not mean abandoning the high-frequency portions; the absence of high-frequency portions will still affect subjective auditory perception.
[0004] Therefore, in ultra-wideband audio frequency coding scenarios, it is necessary to code and decode high-frequency signals. However, regarding the technical problem of how to achieve highly efficient coding and decoding of high-frequency signals at extremely low coding rates, relevant technologies have not yet provided an effective solution. [Overview of the project] [Problems that the invention aims to solve]
[0005] This embodiment provides an audio frequency processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, and makes it possible to improve the quality of audio frequencies obtained by decoding in a subsequent process by coding spectral flatness information during coding to enhance the perfection of coding in the high-frequency portion. [Means for solving the problem]
[0006] The technical solution in the embodiment of this application is realized as follows.
[0007] In this embodiment, a method for processing audio frequencies is provided, the method being performed by an electronic device. The process involves filtering the audio frequency signal to obtain low-frequency and high-frequency signals. The low-frequency signal is coded to obtain a coded stream of the low-frequency signal, The low-frequency signal is subjected to frequency domain conversion processing to obtain a low-frequency spectrum, and the high-frequency signal is subjected to frequency domain conversion processing to obtain a high-frequency spectrum. The low-frequency spectrum and the high-frequency spectrum are subjected to spectral envelope extraction processing to obtain spectral envelope information of the audio frequency signal, and the high-frequency spectrum is subjected to spectral flatness extraction processing to obtain spectral flatness information of the high-frequency spectrum. The method includes quantizing and coding the spectral flatness information of the high-frequency spectrum and the spectral envelope information of the audio frequency signal to obtain a bandwidth-extended code stream of the audio frequency signal, and constructing a coding code stream of the audio frequency signal using the bandwidth-extended code stream and the code stream of the low-frequency signal.
[0008] This embodiment provides an audio frequency processing device. A band division module is arranged to filter audio frequency signals to obtain low-frequency and high-frequency signals, A coding module arranged to perform coding processing on the low-frequency signal to obtain a code stream of the low-frequency signal; A frequency-domain conversion module arranged to perform frequency-domain conversion processing on the low-frequency signal to obtain a low-frequency spectrum, and to perform frequency-domain conversion processing on the high-frequency signal to obtain a high-frequency spectrum; An extraction module arranged to perform spectrum envelope extraction processing on the low-frequency spectrum and the high-frequency spectrum to obtain spectrum envelope information of the audio frequency signal, and to perform spectrum flatness extraction processing on the high-frequency spectrum to obtain spectrum flatness information of the high-frequency spectrum; A quantization module arranged to perform quantization coding processing on the spectrum flatness information of the high-frequency spectrum and the spectrum envelope information of the audio frequency signal to obtain a bandwidth expansion code stream of the audio frequency signal, and to constitute a coding code stream of the audio frequency signal by the bandwidth expansion code stream and the code stream of the low-frequency signal.
[0009] An embodiment of the present application provides an audio frequency processing method, which is executed by an electronic device. Performing decomposition processing on the coding code stream to obtain a bandwidth expansion code stream and the code stream of the low-frequency signal; Performing decoding processing on the code stream of the low-frequency signal to obtain a low-frequency signal, and performing frequency-domain conversion processing on the low-frequency signal to obtain a low-frequency spectrum of the low-frequency signal; Performing inverse quantization processing on the bandwidth expansion code stream to obtain spectrum flatness information and spectrum envelope information; Performing high-frequency spectrum reconstruction processing based on the spectrum flatness information, the spectrum envelope information, and the low-frequency spectrum to obtain a high-frequency spectrum; Performing time-domain conversion processing on the high-frequency spectrum to obtain a high-frequency signal, and performing synthesis processing on the low-frequency signal and the high-frequency signal to obtain an audio frequency signal corresponding to the coding code stream.
[0010] Embodiments of the present application provide an audio frequency processing apparatus, A decomposition module arranged to perform decomposition processing on the coding code stream to obtain a bandwidth expansion code stream and a code stream of the low-frequency signal; A core module arranged to perform decoding processing on the code stream of the low-frequency signal to obtain a low-frequency signal, and performing frequency-domain conversion processing on the low-frequency signal to obtain a low-frequency spectrum of the low-frequency signal; An inverse quantization module arranged to perform inverse quantization processing on the bandwidth expansion code stream to obtain spectral flatness information and spectral envelope information; A reconstruction module arranged to perform high-frequency spectrum reconstruction processing based on the spectral flatness information, the spectral envelope information, and the low-frequency spectrum to obtain a high-frequency spectrum; A time-domain conversion module arranged to perform time-domain conversion processing on the high-frequency spectrum to obtain a high-frequency signal, and performing synthesis processing on the low-frequency signal and the high-frequency signal to obtain an audio frequency signal corresponding to the coding code stream.
[0011] Embodiments of the present application provide an electronic device, A memory for storing computer-executable instructions; A processor that, when executing the computer-executable instructions stored in the memory, realizes the audio frequency processing method provided by the embodiments of the present application.
[0012] Embodiments of the present application provide a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, realize the audio frequency processing method provided by the embodiments of the present application.
[0013] This embodiment provides a computer program product that includes a computer-executable instruction, and when the computer-executable instruction is executed by a processor, the audio frequency processing method provided by this embodiment is realized.
[0014] The embodiments of this invention have the following beneficial effects.
[0015] By filtering the audio frequency signal to obtain low-frequency signals and high-frequency signals, coding the low-frequency signals to obtain a code stream of the low-frequency signals, extracting spectral envelope information of the audio frequency signal and spectral flatness information of the high-frequency signal from the low-frequency spectrum of the low-frequency signal and the high-frequency spectrum of the high-frequency signal, and quantizing and coding the spectral flatness information and spectral envelope information to obtain a bandwidth-extended code stream of the audio frequency signal, and combining this with the code stream of the low-frequency signal to form a coding code stream of the audio frequency signal, effective coding of the high-frequency signal can be achieved via the spectral envelope information. The spectral flatness information helps in the recovery of the high-frequency signal and has a supplementary effect on the spectral envelope information, so that the perfection of coding the high-frequency portion is ultimately improved and the quality of the audio frequency obtained in the subsequent decoding process can be improved. [Brief explanation of the drawing]
[0016] [Figure 1] This is a schematic diagram of the structure of the audio frequency processing system provided in this embodiment. [Figure 2A-2B] This is a schematic diagram of the structure of the electronic device provided in the present embodiment. [Figure 3A-3D] This is a schematic flowchart of the audio frequency processing method provided in the embodiment of the present invention. [Figure 4] This is a schematic diagram of the bandwidth expansion coding in the audio frequency processing method provided in the embodiment of the present invention. [Figure 5]This is a schematic diagram of the bandwidth expansion decoding in the audio frequency processing method provided in the embodiment of the present invention. [Figure 6] This is a schematic diagram of the bandwidth expansion decoding in the audio frequency processing method provided in the embodiment of the present invention. [Figure 7] This is a schematic diagram of the coding in the audio frequency processing method provided in the embodiment of the present invention. [Figure 8] This is a schematic diagram of the decoding in the audio frequency processing method provided in the embodiment of the present invention. [Figure 9] This is a schematic diagram of the spectrum provided by the embodiment of the present invention. [Modes for carrying out the invention]
[0017] To further clarify the purpose, technical proposal and advantages of this application, the attached drawings will be used to describe the application in more detail below. However, the embodiments described should not be considered limitations to this application, and all other embodiments that can be obtained by a person skilled in the art without creative work are all within the scope of protection of this application.
[0018] In the following description, the “series of embodiments” refers to a subset of all possible embodiments, but it should be understood that the “series of embodiments” can be the same subset or different subsets of all possible embodiments, and can be combined with each other as long as they do not conflict.
[0019] In the following description, the terms "first / second / third" are merely used to distinguish similar objects and do not indicate a specific order of arrangement of the objects. "First / second / third" may be interchangeable with any particular order or sequence where permitted, and the embodiments of the present application described herein may be carried out in an order other than that shown or described herein.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same implications as those commonly understood by those skilled in the art. The terms used herein are solely for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0021] Before describing the embodiments of this application in more detail, we will explain the nouns and terms used in these embodiments, and the following interpretations apply to the nouns and terms used in these embodiments.
[0022] 1) Bandwidth Extension (BWE): Also known as bandwidth duplication, this is one of the typical techniques in the field of audio frequency coding. Bandwidth extension is a type of parameter coding technique that effectively increases the bandwidth at the receiving terminal, thereby improving the quality of the audio frequency signal (also called the "audio signal"). Users can intuitively perceive a more vibrant tone, higher volume, and better speech clarity.
[0023] 2) Quadrature Mirror Filters (QMF): These include analysis-synthesis filter pairs. The analysis filters are used to decompose the subband signals, reduce the signal bandwidth, and allow each subband signal to be processed smoothly through the channel. The synthesis filters are used to synthesize each subband signal recovered at the decoding terminal, reconstructing the original audio frequency signal through methods such as zero interpolation and bandpass filtering.
[0024] 4) Modified Discrete Cosine Transform (MDCT): A type of linear overlapping orthogonal transform that uses time-domain aliasing cancellation techniques and includes a 50% time-domain overlapping window. It effectively overcomes edge effects within the overlapping window without degrading coding performance, thereby effectively removing periodic noise generated by edge effects.
[0025] 5) Spectral Band Replication (SBR): This is a technique for improving source coding systems. It is achieved by reducing the spectral bandwidth at the coding terminal and replicating the corresponding audio frequencies at the decoding terminal. This allows for a significant reduction in the coding bitrate while maintaining equivalent sound quality perception.
[0026] 7) Neural Networks (NN): These are mathematical algorithmic models that mimic the behavioral characteristics of animal neural networks and perform distributed and parallel information processing. This type of network relies on the complexity of the system and achieves its information processing objectives by adjusting the interconnection relationships between the large number of nodes within it.
[0027] In ultra-wideband audio frequency coding scenarios, according to the auditory mechanisms of the human ear and psychoacoustic models, users are typically more sensitive to the low-frequency portions of a signal than to the high-frequency portions. During coding and decoding processes, more code rate is allocated to the low-frequency portions of the signal compared to the high-frequency portions. However, this does not mean abandoning the high-frequency portions; their absence would affect subjective auditory perception. Therefore, high-frequency signals need to be coded and decoded in ultra-wideband audio frequency coding scenarios.
[0028] In related technologies, parameterized representation can be performed for high-frequency signals, and the high-frequency portion of the audio frequency signal is reconstructed in the decoding terminal using these parameters and the corresponding low-frequency portion of the audio frequency signal. In related technologies, when parameterized coding is performed for high-frequency signals, only the spectral envelope information of the high-frequency signal is considered, and it is not possible to perform coding with stronger characterization for high-frequency signals. Therefore, when implementing the embodiments of the present invention, the applicant found that for the bandwidth extension proposal of the non-AI audio codec application related technology, there is a certain error in the results obtained by decoding, and the error is due to the relatively large degree that the coding and decoding of high-frequency signals is not very accurate. For the bandwidth extension proposal of the AI audio codec application related technology, the model of the low-frequency signal is constructed via a neural network, and the error associated with the results obtained after code transmission is significantly different from the error in the results obtained by the non-AI audio codec, and there is an even larger error in the results obtained by decoding, that is, the error due to the inaccurate coding and decoding of high-frequency signals is even more pronounced, and there is clear noise in the high-frequency portion reconstructed in the decoding terminal.
[0029] This embodiment provides an audio frequency processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. By coding spectral flatness information during coding, the perfection of coding in the high-frequency portion can be improved, thereby improving the quality of the audio frequency obtained by the subsequent decoding process. The following describes an exemplary application of the electronic device provided by this embodiment. The electronic device provided by this embodiment can be implemented as a terminal, as a server, or through the cooperation of a terminal and a server. The following description will use as an example a case in which a terminal and a server cooperate to implement the audio frequency processing method provided by this embodiment.
[0030] Referring to Figure 1, Figure 1 is a schematic diagram of the structure of the voice frequency decoding system 100 provided by the embodiment of the present invention. In order to support voice applications, as shown in Figure 1, the voice frequency decoding system 100 includes a server 200, a network 300, a first terminal 400 (i.e., a coding terminal), and a second terminal 500 (i.e., a decoding terminal), and the network 300 can be a local area network, a wide area network, or a combination of both.
[0031] In a series of embodiments, a client 410 is operated on a first terminal 400, and the client 410 can be various types of clients, such as an instant messaging client, a network conferencing client, a live streaming client, a browser, etc. In response to a voice frequency collection command issued by the sender (for example, the initiator of a network conference, the main caster, the speaker of a voice call, etc.), the client 410 activates the microphone provided on the first terminal 400 to collect a voice frequency signal, and filters the collected voice frequency signal to obtain low-frequency and high-frequency signals. The frequency of the low-frequency signal is lower than the frequency of the high-frequency signal. The low-frequency signal is coded to obtain a code stream of the low-frequency signal. The low-frequency signal is then subjected to frequency domain conversion to obtain a low-frequency spectrum, and the high-frequency signal is subjected to frequency domain conversion to obtain a high-frequency spectrum. The low-frequency spectrum and the high-frequency spectrum are subjected to spectral envelope extraction to obtain spectral envelope information of the audio frequency signal. The high-frequency spectrum is subjected to spectral flatness extraction to obtain spectral flatness information of the high-frequency spectrum. The spectral flatness information of the high-frequency spectrum and the spectral envelope information of the audio frequency signal are subjected to quantization coding to obtain a bandwidth-expanded code stream of the audio frequency signal. The bandwidth-expanded code stream and the code stream of the low-frequency signal constitute the coded code stream of the audio frequency signal. Next, the client 410 transmits the coded code stream to the server 200 via the network 300, and the server 200 can transmit the code stream to a second terminal 500 associated with the receiving side (e.g., participants in a network conference, an audience, or a receiver of a voice call).After receiving the coding code stream transmitted by the server 200, the client 510 (e.g., an instant messaging client, a network conferencing client, a live streaming client, a browser, etc.) decomposes the coding code stream to obtain a bandwidth-extended code stream and the code stream of the low-frequency signal, decodes the code stream of the low-frequency signal to obtain a low-frequency signal, and performs frequency-domain conversion on the low-frequency signal to obtain the low-frequency spectrum of the low-frequency signal, inverse quantization on the bandwidth-extended code stream to obtain spectral flatness information and spectral envelope information, and performs high-frequency spectrum reconstruction based on the spectral flatness information, spectral envelope information and the low-frequency spectrum to obtain a high-frequency spectrum, the frequency of the high-frequency spectrum being higher than the frequency of the low-frequency spectrum, and performs time-domain conversion on the high-frequency spectrum to obtain a high-frequency signal, and then synthesizes the low-frequency signal and the high-frequency signal to obtain an audio frequency signal corresponding to the coding code stream.
[0032] The voice frequency processing method provided in this embodiment can be widely applied to various different types of voice communication applications, such as voice calls conducted via instant messaging clients, voice calls within game applications, and voice calls in network conferencing clients.
[0033] For example, considering a network conference scenario, which is an important part of online business, a network conference participant's voice acquisition device (e.g., microphone) needs to collect the speaker's voice signal and then transmit the collected voice signal to other participants in the network conference. This process involves the transmission and reproduction of the voice signal among multiple participants. In this scenario, applying the voice frequency processing method provided in the embodiment of this application can code and decode the voice signal of the network conference, making the coding and decoding of high-frequency signals within the voice signal more efficient and accurate, thereby improving the quality of voice communication in the network conference.
[0034] In another series of embodiments, the embodiments of the present invention can be realized via Cloud Technology, which refers to hosting technology that unifies system sources such as hardware, software, and networks within a wide area network or local area network to enable data computation, storage, processing, and sharing.
[0035] Cloud technology is a general term encompassing network technology, information technology, integration technology, management platform technology, and application technology, all applied based on the cloud computing business model. It allows for the configuration of resource pools, enabling their use as needed, offering flexibility and convenience. Cloud computing technology should provide crucial support. Interoperability between the 200 servers mentioned above can be achieved through cloud technology.
[0036] For example, the server 200 shown in Figure 1 can be an independent physical server, a server group or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud memory, network services, cloud communications, middleware services, domain services, security services, CDNs, and big data and artificial intelligence platforms. The first terminal 400 and second terminal 500 shown in Figure 1 can be, but are not limited to, smartphones, tablet PCs, laptop computers, desktop computers, smart speakers, smartwatches, voice recognition devices, smart home appliances, in-car terminals, aircraft, etc. The terminals (e.g., the first terminal 400 and the second terminal 500) and the server 200 can be connected directly or indirectly by wired or wireless communication, and are not limited in this embodiment.
[0037] In a series of embodiments, a terminal (e.g., a second terminal 500) or a server 200 can further implement the voice frequency processing method provided by the embodiments of this invention by operating a computer program. For example, the computer program can be a native program or software module within an operating system, a native application program (APP), that is, a program that can only be operated after being implemented in an operating system, such as a live streaming APP, a network conferencing APP, or an instant messaging APP, or a mini-program, that is, a program that can be operated after being downloaded to a browser environment. In short, the above computer program can be an application program, module, or plug-in of any form.
[0038] Referring to Figure 2A, which is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention, the first terminal 400 shown in Figure 2A includes at least one processor 410, memory 450, at least one network interface 420, and a user interface 430. Each component of the first terminal is coupled together via a bus system 440. As can be understood, the bus system 440 is used to realize connection communication between these components. The bus system 440 includes a data bus, as well as a power bus, a control bus, and a status signal bus. However, for clarity, in Figure 2A, all types of buses are referred to as the bus system 440.
[0039] The processor 410 can be an integrated circuit chip having signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, where the general-purpose processor can be a microprocessor or any ordinary processor.
[0040] The user interface 430 includes one or more output devices 431 capable of presenting media content, one or more speakers, and / or one or more visual displays. The user interface 430 further includes one or more input devices 432, which include user interface components that support user input, such as a keyboard, mouse, microphone, touch panel, camera, and other input buttons and control materials.
[0041] The memory 450 may be removable, non-removable, or a combination of both. Exemplary hardware devices include solid memory, hard disk drives, optical disk drives, and the like. The memory 450 may include one or more storage devices whose physical location is spaced apart from the processor 410.
[0042] The memory 450 includes volatile memory or non-volatile memory, and may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random-access memory (RAM). The memory 450 described in this embodiment is intended to include any suitable type of memory.
[0043] In a series of embodiments, the memory 450 can store data and support various operations, and examples of such data include programs, modules, and data structures or subsets or supersets, which are described illustratively below.
[0044] The operating system 451 includes system programs for handling various basic system services and executing hardware-related tasks, such as a framework layer, core-based layer, and drive layer, and is used to implement various basic business operations and process hardware-based tasks.
[0045] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB).
[0046] The presentation module 453 is used to enable information to be presented via output devices 431 (e.g., displays, speakers, etc.) associated with one or more user interfaces 430 (e.g., user interfaces for operating peripheral devices and displaying content and information).
[0047] The input processing module 454 is used to detect one or more user inputs or interactions from one or more input devices 432 and to translate the detected inputs or interactions.
[0048] In a series of embodiments, the apparatus provided by the present embodiment can be implemented using a software method. Figure 2A shows an audio frequency processing device 455 stored in memory 450, which can be software in the form of a program or plug-in, and includes the following software modules: a band division module 4551, a core module 4552, a frequency domain conversion module 4553, an extraction module 4554, and a quantization module 4555. Since these modules are logical, they can be arbitrarily combined or further decomposed depending on the function to be implemented. The function of each module will be described in the following text.
[0049] Referring to Figure 2B, which is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention, the second terminal 500 shown in Figure 2B includes at least one processor 510, memory 550, at least one network interface 520 and a user interface 530. Each component of the second terminal 500 is coupled together via a bus system 540. The user interface 530 includes one or more output devices 531 capable of presenting media content. The user interface 530 further includes one or more input devices 532. The memory 550 further includes an operating system 551, a network communication module 552, a presentation module 553 and an input processing module 554.
[0050] In a series of embodiments, the apparatus provided by the present embodiment can be implemented using a software method. Figure 2B shows an audio frequency processing device 555 stored in memory 550, which can be software in the form of a program or plug-in, and includes the following software modules: a decomposition module 5551, a decoding module 5552, an inverse quantization module 5553, a reconstruction module 5554, and a time-domain conversion module 5555. Since these modules are logical, they can be arbitrarily combined or further decomposed depending on the function to be implemented. The function of each module will be described in the following text.
[0051] In the following, the audio frequency decoding method provided by this embodiment will be described from the perspective of the interaction between the first terminal device (i.e., the coding terminal), the server, and the second terminal device (i.e., the decoding terminal).
[0052] Referring to Figure 3A, which is a schematic flowchart of the audio frequency processing method provided by the embodiment of the present invention, the steps shown in Figure 3A will be explained in combination, but the steps shown in Figure 3A are performed by the first terminal device (i.e., the coding terminal).
[0053] What needs to be explained here is that the steps performed by the terminal device are specifically performed by a client running on the terminal device. For simplicity, this application does not make a specific distinction between the terminal device and the client running on the terminal device. Furthermore, it should be explained that the audio frequency decoding method provided in this embodiment can be executed by various types of computer programs running on the terminal device, but is not limited to the client running on the terminal device. It can also be the operating system, software module, script, and miniprogram mentioned in the above text. Therefore, the examples of clients in the following text should not be considered as limitations on this embodiment.
[0054] In step 101, the audio frequency signal is filtered to obtain low-frequency and high-frequency signals.
[0055] As an example, the frequency of a low-frequency signal is lower than the frequency of a high-frequency signal, a high-frequency signal refers to a signal with a frequency exceeding 3 megahertz, and a low-frequency signal refers to a signal with a frequency between 30 kilohertz and 300 kilohertz. The audio frequency signal is an ultra-wideband signal with a sampling rate of 32 kilohertz (kHz). A 32kHz sampling rate means sampling 32,000 times per second to obtain 32,000 sampling points. The audio frequency signal is divided into frames according to a frame length of 640 points, with 640 sampling points forming one frame. Each frame has a frame length of 640 points and a frame duration of 0.02 seconds. The audio frequency signal is filtered through a group of QMF band-divided filters to obtain a high-frequency portion (high-frequency signal) with a frame length of 320 points and a low-frequency portion (low-frequency signal) with a frame length of 320 points.
[0056] In step 102, the low-frequency signal is coded to obtain a coded stream of the low-frequency signal.
[0057] In the series of embodiments, step 102 involves coding a low-frequency signal to obtain a code stream of the low-frequency signal. However, this can be achieved by the following technical proposal, which involves performing a feature extraction process on the low-frequency signal to obtain a first feature of the low-frequency signal, and then performing a quantization coding process on the first feature to obtain a code stream of the low-frequency signal of the audio frequency signal.
[0058] As an example, coding is used to compress low-frequency signals and retain the information contained within them. Coding can be a conventional coding technique, or it can be a coding technique based on deep learning. For example, coding for low-frequency signals can be implemented via a Penguins speech engine based on deep learning. Feature extraction is performed on the low-frequency signal via a neural network model to obtain a feature vector (first feature), where the data dimension of the feature vector is smaller than the data dimension of the low-frequency signal. Vector quantization or scalar quantization is performed on the feature vector corresponding to the low-frequency signal to obtain an index value, and entropy coding is performed on the index value to obtain a code stream of the low-frequency signal.
[0059] As an example, the full dynamic range of a low-frequency signal's feature vector is divided into multiple intervals, each interval having a representative value. When feature vectors within an interval are replaced with this representative value during quantization, the resulting feature vector is one-dimensional, hence it is called standard quantization. Vector quantization shows an extension and expansion of scalar quantization. A vector is constructed from multiple scalar data, and in vector quantization, quantization is performed on the vector. This involves dividing the vector space into multiple sub-regions, searching for a representative vector in each sub-region, and replacing the vectors within the sub-region with this representative vector during quantization.
[0060] In this embodiment, a feature vector (first feature) with dimensions significantly smaller than the original signal is generated via a neural network model, and then the effect of low-code-rate coding can be achieved through techniques such as entropic coding.
[0061] In step 103, a low-frequency signal is subjected to frequency domain conversion to obtain a low-frequency spectrum, and a high-frequency signal is subjected to frequency domain conversion to obtain a high-frequency spectrum.
[0062] For example, the frequency domain conversion process can be MDCT processing. MDCT processing can be performed on a high-frequency signal to obtain a high-frequency spectrum (including multiple spectral coefficients). Alternatively, DCT processing can be performed on a high-frequency signal to obtain a high-frequency spectrum (including multiple spectral coefficients). Similarly, MDCT processing can be performed on a low-frequency signal to obtain a low-frequency spectrum (including multiple spectral coefficients). Alternatively, DCT processing can be performed on a low-frequency signal to obtain a low-frequency spectrum (including multiple spectral coefficients).
[0063] In step 104, the low-frequency spectrum and high-frequency spectrum are subjected to spectral envelope extraction processing to obtain spectral envelope information of the audio frequency signal, and the high-frequency spectrum is subjected to flatness extraction processing to obtain spectral flatness information of the high-frequency spectrum.
[0064] In the following, we will first interpret the origin of the spectrum. Audio frequency signals are time-domain signals, where the horizontal axis of a sequence signal represents time and the vertical axis represents amplitude. For example, if the amplitude (which can indicate volume) of each audio frequency frame is different, it is not possible to embody the rule of change that accompanies changes in frequency over time. Therefore, it is necessary to perform frequency-domain analysis (e.g., Fourier transform) on a time-domain signal to obtain a spectrogram. A spectrogram can display the spectral range of an audio frequency signal. Specifically, a time-domain signal can be decomposed into a single DC component (i.e., a single constant) and the sum of multiple sinusoidal signals. Each sinusoidal component has its own frequency and amplitude value. Therefore, if we plot the amplitude values of the above several sinusoidal signals on their corresponding frequencies, with frequency values on the horizontal axis and amplitude values on the vertical axis, we obtain a signal amplitude distribution diagram, which is the spectrogram shown in Figure 9.
[0065] As an example, the spectrum shown in Figure 9 has a frequency range of 0 Hz to 10,000 Hz. From Figure 9, it can be seen that the amplitude increases or decreases with the change in frequency. The curve formed by connecting the highest points (peaks in the spectrogram) before each amplitude decrease in the spectrogram is called the spectral envelope. Spectral envelope extraction is the process of extracting information shown by the low-frequency spectral envelope from the low-frequency spectrum, i.e., low-frequency spectral envelope information, and extracting information shown by the high-frequency spectral envelope from the high-frequency spectrum, i.e., high-frequency envelope information. The low-frequency spectral envelope here is the curve formed by connecting the highest points before each amplitude decrease in the spectrogram corresponding to the low-frequency section in Figure 9, and the high-frequency spectral envelope here is the curve formed by connecting the highest points before each amplitude decrease in the spectrogram corresponding to the high-frequency section in Figure 9.
[0066] In a series of embodiments, obtaining spectral envelope information of an audio frequency signal by performing spectral envelope extraction processing on low-frequency and high-frequency spectra in step 104 can be achieved by the following technical proposal, namely, performing spectral envelope extraction processing on low-frequency spectra to obtain low-frequency spectral envelope information of low-frequency spectra, performing spectral envelope extraction processing on high-frequency spectra to obtain high-frequency spectral envelope information of high-frequency spectra, and constructing spectral envelope information of an audio frequency signal using the low-frequency spectral envelope information and high-frequency spectral envelope information.
[0067] In this embodiment, by extracting spectral envelope information, coding of the energies of low-frequency and high-frequency signals is realized during bandwidth expansion coding, improving the effectiveness of coding and thereby obtaining a better recovery effect during the subsequent decoding process.
[0068] In the following sections, we will introduce specific embodiments for obtaining low-frequency spectral envelope information by performing spectral envelope extraction on low-frequency spectra, and specific embodiments for obtaining high-frequency spectral envelope information by performing spectral envelope extraction on high-frequency spectra.
[0069] In a series of embodiments, obtaining high-frequency spectral envelope information of a high-frequency spectrum by performing spectral envelope extraction processing on the high-frequency spectrum can be achieved by the following technical proposal, which involves obtaining second fusion configuration data of the high-frequency spectrum, the second fusion configuration data including the spectral ordinal number of each second spectral line combination, extracting spectral coefficients corresponding to each spectral ordinal number of the second spectral line combination from the high-frequency spectrum for each second spectral line combination, squaring the spectral coefficients of each spectral ordinal number to obtain the second squared spectral coefficient of each spectral ordinal number, if there are multiple spectral ordinal numbers of the second spectral line combination, adding the second squared spectral coefficients of the multiple spectral ordinal numbers to obtain a second addition result, logarithmically scaling the second addition result to obtain second fusion spectral envelope information corresponding to the second spectral line combination, and generating high-frequency spectral envelope information based on the second fusion spectral envelope information of at least one second spectral line combination.
[0070] As an example, the second fusion configuration data is stored locally on the terminal or on a server in the form of a data table, so that the first terminal device can easily read it directly from the terminal or retrieve it from the server. An example of the second fusion configuration data is shown in Table 1. From Table 1, it can be seen that the second fusion configuration data includes four second spectral line combinations, with the spectral line ordinal numbers of the first second spectral line combination being 0 to 19, the second second spectral line combination being 20 to 54, the third second spectral line combination being 55 to 89, and the fourth second spectral line combination being 90 to 130, and each spectral line has its own spectral coefficient.
[0071] [Table 1]
[0072] As an example, the spectral coefficients of the high-frequency spectrum are merged based on Table 1, and the spectral envelope information of each second spectral line combination in the high-frequency spectrum is extracted as shown in formula (1).
[0073]
number
[0074] In the formula, m k represents the spectral coefficients of a spectrum with spectral line ordinal number k in the high-frequency spectrum (obtained after MDCT transformation), and i represents the spectral envelope ordinal number (ordinal number of the second spectral line combination). For example, if i is 1,
[0075]
number
[0076] are m0, m1, …, m 19 Each is squared, and the squared results are added together. Spec_env(i) represents the second fused spectral envelope information of the second spectral line combination where the spectral envelope ordinal number is i.
[0077] In the following, the second spectral line combination with a spectral envelope ordinal number of 1 will be described in detail as an example. From Table 1, it can be read that the number of the second spectral line combination is 4. For the second spectral line combination A with a spectral envelope ordinal number of 1, the spectral coefficients of each spectral line ordinal corresponding to the second spectral line combination A are extracted from the high-frequency spectrum, that is, from the spectral coefficient m0 of the spectrum with a spectral line ordinal number of 0 to the spectral coefficient m 19 up to the spectrum with a spectral line ordinal number of 19. The spectral coefficients of each spectral line ordinal are squared to obtain the second squared spectral coefficients m k 2 , for example, m0 2 , m1 2 , …, m 19 2 are obtained. Since the number of spectral line ordinals of the second spectral line combination A is 20, when there are multiple second spectral line ordinals, the second squared spectral coefficients of the multiple spectral line ordinals are added together to obtain the second addition result
[0078]
Number
[0079] Obtaining this corresponds to adding 20 second spectral squarening coefficients and logarithmically scaling the second addition result to obtain the second fused spectral envelope information Spec_env(1) corresponding to the second spectral line combination. If the spectral line ordinal number of the second spectral line combination B is 1, the second squared spectral coefficient corresponding to the unique spectral line ordinal number is logarithmically scaling to obtain the second fused spectral envelope information corresponding to the second spectral line combination. If there are multiple second spectral line combinations, the high-frequency spectral envelope information is constructed using the second fused spectral envelope information Spec_env(i) of the multiple second spectral line combinations, or if there is only one second spectral line combination, the second fused spectral envelope information Spec_env(i) of that second spectral line combination is used as the high-frequency spectral envelope information.
[0080] In this embodiment, when performing spectral envelope information fusion processing based on second fusion configuration data for the spectrum of a high-frequency signal, the second fusion configuration data is used to indicate which spectral lines need to be fused. This is based on the critical band in the psychoacoustic model as a theoretical foundation, and in specific experiments, it is obtained by comprehensively considering the quality of the BWE and the code rate. The critical band is a result obtained from experiments based on psychoacoustics, and specifically reflects the conversion between physical-mechanical stimulation and neuroelectrical stimulation in the cochlea of the human inner ear. The neuroelectrical stimulation converted by the human ear is consistent for a certain frequency and other frequencies within a specific range near it, as well as for pure tone speech frequency signals. This indicates that it is not necessary to achieve excessively high frequency domain resolution by using an excessively high code rate. Through multiple experimental measurements, it was found that relatively good code rates and speech frequency quality can be achieved when the energy envelope selection range for the high-frequency portion is the second fusion configuration data.
[0081] In a series of embodiments, obtaining low-frequency spectral envelope information of a low-frequency spectrum by performing spectral envelope extraction processing on the low-frequency spectrum is achieved by the following technical proposal: obtaining first fused configuration data of the low-frequency spectrum, the first fused configuration data containing the spectral ordinal number of each first spectral line combination, extracting spectral coefficients corresponding to each spectral ordinal number of the first spectral line combination from the low-frequency spectrum for each first spectral line combination, squaring the spectral coefficients of each spectral ordinal number to obtain the first squared spectral coefficient of each spectral ordinal number, if there are multiple spectral ordinal numbers of the first spectral line combination, adding the first squared spectral coefficients of the multiple spectral ordinal numbers to obtain a first addition result, logarithmically scaling the first addition result to obtain first fused spectral envelope information corresponding to the first spectral line combination, and generating low-frequency spectral envelope information based on the first fused spectral envelope information of at least one first spectral line combination.
[0082] As an example, the first fusion configuration data can be stored locally on the terminal or on a server in the form of a data table, and can be read directly from the terminal or retrieved from the server by the first terminal device. An example of the first fusion configuration data is shown in Table 2, from which it can be seen that the first fusion configuration data includes one first spectral line combination, and the spectral line ordinal number of the first spectral line combination is 80 to 150, with each spectral line having its own spectral coefficient.
[0083] [Table 2]
[0084] As an example, the spectral coefficients of the low-frequency spectrum are merged based on Table 2, and the spectral envelope information for each first spectral line combination in the low-frequency spectrum is extracted as shown in formula (2).
[0085]
number
[0086] In the formula, m k represents the spectral coefficients of a spectrum with spectral line ordinal number k in the low-frequency spectrum (obtained after MDCT transformation), and i represents the spectral envelope ordinal number (ordinal number of the first spectral line combination). For example, if i is 1,
[0087]
number
[0088] is, m 80 , m 81 , ..., m 149 Each of these is squared, and the results of the squares are added together to obtain Spec_env(i), which represents the first fused spectral envelope information of the first spectral line combination whose spectral envelope ordinal number is i.
[0089] In the following, we will explain in detail using the first spectral line combination with a spectral envelope ordinal number of 1 as an example. From Table 2, we can see that the quantity of the first spectral line combination is 1. For the first spectral line combination A with a spectral envelope ordinal number of 1, we extract the spectral coefficients for each spectral line ordinal number corresponding to the first spectral line combination A from the low-frequency spectrum, that is, the spectral coefficient m of the spectrum with a spectral line ordinal number of 80. 80 From the spectral coefficient m of the spectrum with spectral ordinal number 150 150 Extract up to and then square the spectral coefficients of each spectral ordinal number to obtain the first squared spectral coefficient m of each spectral ordinal number.k 2 For example, m 80 2 , m 81 2 , ..., m 150 2 Since the number of spectral ordinal numbers in the first spectral line combination A is 71, if there are multiple spectral ordinal numbers in the first spectral line combination A, the first squared spectral coefficients of the multiple spectral ordinal numbers are added together to obtain the first addition result.
[0090]
number
[0091] Obtaining this corresponds to adding 71 first spectral squared coefficients and logarithmically scaling the first addition result to obtain the first fused spectral envelope information Spec_env(1) corresponding to the first spectral line combination. If the number of spectral line ordinal numbers of the first spectral line combination B is 1, the first squared spectral coefficient corresponding to the unique spectral line ordinal number is logarithmically scaling to obtain the first fused spectral envelope information corresponding to the first spectral line combination. If there are multiple first spectral line combinations, the low-frequency spectral envelope information is constructed using the first fused spectral envelope information Spec_env(i) of the multiple first spectral line combinations, or if there is only one first spectral line combination, the first fused spectral envelope information Spec_env(i) of that first spectral line combination is used as the low-frequency spectral envelope information.
[0092] In this embodiment, when performing spectral envelope information fusion processing based on first fusion configuration data for the spectrum of a low-frequency signal, the first fusion configuration data is used to indicate which spectral lines need to be fused, and this is obtained by statistically measuring experiments. When the coding of the low-frequency signal is AI ultra-wideband speech coding, AI ultra-wideband speech coding has relatively strong modeling capabilities for speech and possesses noise reduction capabilities, so it is necessary to introduce a variable to quantify and predict the noise reduction result, and the energy envelope of the low-frequency portion can be used as a predictor variable. Statistical measurements on a large dataset have shown that when the energy envelope selection range for the low-frequency portion is the first fusion configuration data, relatively good coding rates and speech frequency quality can be achieved.
[0093] In a series of embodiments, obtaining spectral flatness information of a high-frequency spectrum by performing spectral flatness extraction processing on the high-frequency spectrum in step 104 can be achieved by the following technical proposal, which involves obtaining third fusion configuration data of the high-frequency spectrum, the third fusion configuration data including the spectral line ordinal number of each third spectral line combination, obtaining the geometric mean value of the third spectral line combination and the arithmetic mean value of the third spectral line combination for each third spectral line combination, using the comparison value between the geometric mean value and the arithmetic mean value of the third spectral line combination as spectral flatness information of the third spectral line combination, and generating spectral flatness information of the high-frequency spectrum based on the spectral flatness information of at least one third spectral line combination.
[0094] As an example, the third fusion configuration data can be stored locally on the terminal or on a server in the form of a data table, and can be read directly from the terminal or retrieved from the server by the first terminal device. An example of the third fusion configuration data is shown in Table 3, from which it can be seen that the third fusion configuration data includes two third spectral line combinations, with the spectral line ordinal numbers of the first third spectral line combination being 0 to 39 and the spectral line ordinal numbers of the second third spectral line combination being 40 to 80, and each spectral line having its own spectral coefficient.
[0095] [Table 3]
[0096] As an example, the spectral coefficients of the high-frequency spectrum are merged based on Table 3, and the spectral flatness information of each third spectral line combination in the high-frequency spectrum is extracted as shown in formula (3).
[0097]
number
[0098] In the formula, nume(i) and demo(i) represent the geometric mean and arithmetic mean of the i-th third spectral line combination in the high-frequency spectrum, respectively, while Flatness(i) represents a comparison between the geometric mean and arithmetic mean of the i-th third spectral line combination, where i represents the ordinal number of the third spectral line combination.
[0099] In the following, we will explain in detail using the third spectral line combination with ordinal number 1 as an example. From Table 3, we can see that the quantity of the third spectral line combination is 2. For the first third spectral line combination A shown in Table 3, we extract the spectral coefficients for each spectral line ordinal number corresponding to the third spectral line combination A from the high-frequency spectrum, that is, from the spectral coefficient m0 of the spectrum with spectral line ordinal number 0 to the spectral coefficient m of the spectrum with spectral line ordinal number 39. 39 Extracting up to this point, the spectral coefficients m0 of spectra with a spectral ordinal number of 0 to the spectral coefficients m of spectra with a spectral ordinal number of 39. 39 Based on this, the arithmetic mean and geometric mean of the third spectral line combination A are determined, and the comparison value between the geometric mean and the arithmetic mean is used as the spectral flatness information for the third spectral line combination A. If there are multiple third spectral line combinations, the spectral flatness information for the high-frequency spectrum is constructed using the spectral flatness information Flatness(i) of the multiple third spectral line combinations, or if there is only one third spectral line combination, the spectral flatness information Flatness(i) of that third spectral line combination is used as the spectral flatness information for the high-frequency spectrum.
[0100] Table 3, which shows the spectral flatness fusion in the high-frequency portion, is based on the critical band in the psychoacoustic model as its theoretical foundation, and in specific experiments, it is obtained by comprehensively considering the quality and code rate of the BWE. The critical band is a result obtained from experiments based on psychoacoustics, and specifically reflects the conversion between physical-mechanical stimulation and neuroelectric stimulation in the cochlea of the human inner ear. The neuroelectric stimulation converted by the human ear is consistent for a specific frequency and other frequencies within a specific range near it, as well as for pure tone speech frequency signals. In other words, it is not necessary to achieve excessively high frequency domain resolution by using an excessively high code rate. Statistical analysis and measurements of large datasets have shown that when the spectral flatness fusion selection range in the high-frequency portion is the result of the third fusion configuration, relatively good code rates and speech frequency quality can be achieved.
[0101] In a series of embodiments, obtaining the geometric mean value of the above-mentioned third spectral line combination can be achieved by the following technical method: extract the spectral coefficients corresponding to each spectral line ordinal number of the third spectral line combination from the high-frequency spectrum, square the spectral coefficients of each spectral line ordinal number to obtain the third squared spectral coefficient of each spectral line ordinal number, and if there are multiple spectral line ordinal numbers in the third spectral line combination, multiply the third squared spectral coefficients of the multiple spectral line ordinal numbers to obtain the first multiplication result, and then, based on the number of spectral line ordinal numbers, perform a root operation on the first multiplication result to obtain the geometric mean value corresponding to the third spectral line combination.
[0102] For example, the calculation process for the geometric mean can be seen in formula (4).
[0103]
number
[0104] Based on the above example, for the third spectral line combination A (i = 1) with ordinal number 1, spectral coefficients corresponding to each spectral line ordinal number of the third spectral line combination A are extracted from the high-frequency spectrum, that is, from the spectral coefficient m0 of the spectrum with spectral line ordinal number 0 to the spectral coefficient m of the spectrum with spectral line ordinal number 39. 39 Extract the spectral coefficients for each spectral ordinal number, and then square the spectral coefficients m for each spectral ordinal number. k 2 For example, m0 2 , m1 2 If the number of spectral ordinal numbers in the third spectral line combination is multiple, the third squared spectral coefficients of the multiple spectral ordinal numbers are multiplied to obtain the first multiplication result.
[0105]
number
[0106] Obtaining this is equivalent to raising 40 third spectral squarening coefficients to a power, and then, based on the number of spectral line ordinal numbers, the first multiplication result is subjected to a square root operation (i.e., raised to the power of 40) to obtain the geometric mean value nume(1) corresponding to the third spectral line combination.
[0107] In a series of embodiments, obtaining the arithmetic mean of the above-mentioned third spectral line combinations can be achieved by the following technical method: extracting spectral coefficients corresponding to each spectral line ordinal number of the third spectral line combination from the high-frequency spectrum, squaring the spectral coefficients of each spectral line ordinal number to obtain the third squared spectral coefficient of each spectral line ordinal number, and if there are multiple spectral line ordinal numbers in the third spectral line combination, adding the third squared spectral coefficients of the multiple spectral line ordinal numbers to obtain the third addition result, and averaging the third addition result based on the number of spectral line ordinal numbers to obtain the arithmetic mean corresponding to the third spectral line combination.
[0108] For example, the calculation process for the geometric mean can be seen in formula (5).
[0109]
number
[0110] Based on the above example, for the third spectral line combination with ordinal number 1 (i = 1), the spectral coefficients corresponding to each spectral line ordinal number of the third spectral line combination A are extracted from the high-frequency spectrum, that is, from the spectral coefficient m0 of the spectrum with spectral line ordinal number 0 to the spectral coefficient m of the spectrum with spectral line ordinal number 39. 39 Extract the spectral coefficients for each spectral ordinal number, and then square the spectral coefficients m for each spectral ordinal number. k 2 For example, m0 2 , m1 2 If the number of spectral ordinal numbers in the third spectral line combination is multiple, the third squared spectral coefficients of the multiple spectral ordinal numbers are added together to obtain the third summation result.
[0111]
number
[0112] Obtaining this is equivalent to adding 40 third spectral squarening coefficients, and then averaging is performed on the third addition result based on the number of spectral line ordinal numbers (i.e., dividing by 40) to obtain the arithmetic mean demo(1) corresponding to the third spectral line combination.
[0113] In step 105, the spectral flatness information of the high-frequency spectrum and the spectral envelope information of the audio frequency signal are subjected to quantization coding to obtain a bandwidth-extended code stream of the audio frequency signal, and a coding code stream of the audio frequency signal is obtained by combining the bandwidth-extended code stream and the code stream of the low-frequency signal.
[0114] In a series of embodiments, obtaining a bandwidth-extended code stream for the audio frequency signal by quantizing and coding the spectral flatness information of the high-frequency spectrum and the spectral envelope information of the audio frequency signal in step 105 can be achieved by the following technical proposal, which involves obtaining a quantization table for spectral flatness information and a quantization table for spectral envelope information, quantizing the spectral flatness information of the high-frequency spectrum according to the quantization table for spectral flatness information to obtain the quantization result for spectral flatness, quantizing the spectral envelope information of the audio frequency signal according to the quantization table for spectral envelope information to obtain the quantization result for spectral envelope, and constructing a bandwidth-extended code stream for the audio frequency signal using the quantization results for spectral flatness and spectral envelope.
[0115] In a series of embodiments, obtaining the quantization table of spectral flatness information and the quantization table of spectral envelope information described above is the following technical proposal, namely, obtaining multiple audio sample signals and filtering each audio sample signal to obtain low-frequency and high-frequency sample signals of the audio sample signal, wherein the frequency of the low-frequency sample signal is lower than the frequency of the high-frequency sample signal, and then performing frequency domain conversion on the low-frequency sample signal to obtain the low-frequency sample spectrum, and performing frequency domain conversion on the high-frequency sample signal to obtain the high-frequency sample spectrum, and finally, the low-frequency and high-frequency sample spectra are processed This can be achieved by a technical proposal that involves performing a spectral envelope extraction process to obtain spectral envelope information of an audio sample signal, a spectral flatness extraction process on the high-frequency spectrum to obtain spectral flatness information of an audio sample signal, a clustering process of the spectral flatness information of multiple audio sample signals to obtain multiple spectral flatness clustering centers, and a process to construct a quantization table of spectral flatness information based on the multiple spectral flatness clustering centers, and a clustering process of the spectral envelope information of multiple audio sample signals to obtain multiple spectral envelope clustering centers, and a process to construct a quantization table of spectral envelope information based on the multiple spectral envelope clustering centers.
[0116] As an example, the process from acquiring multiple audio sample signals to filtering the audio sample signals to obtaining spectral envelope information and spectral flatness information of the audio sample signals can be seen by referring to a specific embodiment of step 104.
[0117] As an example, the quantization table used for quantizing spectral flatness information is shown in Table 4. Table 4 embodies the clustering centers for each spectral flatness, corresponding to the four clustering centers obtained through clustering. In the subsequent quantization process, spectral flatness information A is quantized so that it becomes the clustering center with the smallest numerical difference from A among the four clustering centers.
[0118] [Table 4]
[0119] As an example, the quantization table used for the spectral envelope information of the high-frequency portion is shown in Table 5. Table 5 embodies the clustering centers of the spectral envelope obtained by clustering based on the first and second subbands of the high-frequency portion of the sample data, and corresponds to 31 clustering centers obtained through clustering. In the subsequent quantization process, the spectral envelope information A of the first and second subbands of the high-frequency portion is quantized so that it becomes the clustering center with the smallest numerical difference from A among the 31 clustering centers.
[0120] [Table 5]
[0121] As an example, the quantization table used for the spectral envelope information in the high-frequency portion is shown in Table 6. Table 6 embodies the clustering centers of the spectral envelope obtained by clustering based on the third and fourth subbands of the high-frequency portion of the sample data, and corresponds to the eight clustering centers obtained by the clustering process. In the subsequent quantization process, the spectral envelope information A of the third and fourth subbands of the high-frequency portion is quantized so that it becomes the clustering center with the smallest numerical difference from A among the eight clustering centers.
[0122] [Table 6]
[0123] As an example, the spectral envelope quantization table for the low-frequency portion is shown in Table 7. Table 7 embodies the clustering centers of the spectral envelope obtained by clustering based on the low-frequency portion of the sample data, corresponding to the eight clustering centers obtained through the clustering process. In the subsequent quantization process, the spectral envelope information A of the low-frequency portion is quantized so that it becomes the clustering center with the smallest numerical difference from A among the eight clustering centers.
[0124] [Table 7]
[0125] The generation process for Tables 4-7 was obtained through experimental statistics. By clustering a large number of audio frequency files based on the above flow, a statistical distribution based on a large number of audio frequency distributions was ultimately obtained. Considering the code rate and audio frequency quality comprehensively, clustering quantization was performed on this statistical distribution to finally generate Tables 4-7.
[0126] The quantization coding method effectively compresses and displays spectral flatness information and spectral envelope information, reducing the amount of data for these information, avoiding excessive use of communication resources, and effectively improving communication efficiency.
[0127] The audio frequency processing method provided in this embodiment enables effective coding by combining spectral envelope information and spectral flatness information for high-frequency portions without requiring a higher code rate compared to related technologies. This allows for effective characterization of high-frequency portions with low complexity, enabling the recovery of a more truthful and natural audio frequency signal in the subsequent decoding process. In particular, when the encoder is based on a neural network model, low-frequency signals are modeled via the neural network, and the errors associated with the results obtained after code transmission differ significantly from the errors in results obtained by non-AI audio codecs. The errors in the results obtained by decoding are even greater, meaning that errors due to inaccurate coding and decoding of high-frequency signals become even more pronounced. The high-frequency portions reconstructed at the decoding terminal using the bandwidth extension method of related technologies have significant noise. However, by applying the audio frequency processing method provided in this embodiment, it is possible to reconstruct accurate high-frequency portions and recover a more truthful and natural audio frequency signal.
[0128] Referring to Figure 3B, which is a schematic flowchart of the audio frequency processing method provided by the embodiment of the present invention, the steps shown in Figure 3B will be explained in combination.
[0129] In step 201, the coding code stream is decomposed to obtain a bandwidth-extended code stream and the code stream of the low-frequency signal.
[0130] As an example, referring to Figure 8, which is a schematic diagram of decoding in the audio frequency processing method provided by the embodiment of the present invention, the decoding terminal decomposes the received coding code stream into a BWE code stream and the code stream of the low frequency signal. The code stream of the low frequency signal is recovered by an AI ultra-wideband audio decoder, and the low frequency signal and the BWE code stream are recovered into a high frequency code stream by a BWE decoder provided by the embodiment of the present invention, the time domain of the high frequency code stream is converted into a high frequency signal, and an ultra-wideband signal is generated from the high frequency signal and the low frequency signal by a group of combining filters.
[0131] In step 202, the code stream of the low-frequency signal is decoded to obtain a low-frequency signal, and the low-frequency signal is subjected to frequency domain conversion to obtain the low-frequency spectrum of the low-frequency signal.
[0132] As an example, referring to Figure 5, Figure 5 is a schematic diagram of bandwidth expansion decoding in the audio frequency processing method provided by the embodiment of the present invention. The low-frequency signal in Figure 5 is the low-frequency signal obtained by decoding. The low-frequency signal obtained by decoding is subjected to frequency domain conversion processing to obtain the low-frequency spectrum of the low-frequency signal. The frequency domain conversion processing can be MDCT processing or DCT processing.
[0133] In step 203, the bandwidth-extended code stream is inversely quantized to obtain spectral flatness information and spectral envelope information.
[0134] Since the bandwidth-extended code stream is obtained by quantizing and coding spectral flatness information and spectral envelope information, spectral flatness information and spectral envelope information can be obtained by inverse quantization decoding.
[0135] In step 204, a high-frequency spectrum reconstruction process is performed based on the spectral flatness information, spectral envelope information, and low-frequency spectrum to obtain the high-frequency spectrum.
[0136] Referring to Figure 3C in a series of embodiments, Figure 3C is a schematic flowchart of the audio frequency processing method provided by the embodiment of the present invention. In step 204, a high-frequency spectrum reconstruction process is performed based on spectral flatness information, spectral envelope information, and low-frequency spectrum to obtain a high-frequency spectrum. This can be achieved by steps 2041 to 2044 shown in Figure 3C.
[0137] In step 2041, the low-frequency spectrum is subjected to spectral flatness extraction processing to obtain low-frequency spectral flatness information of the low-frequency spectrum, and subband spectral flatness information of each low-frequency subband in the low-frequency spectrum is extracted from the low-frequency spectral flatness information.
[0138] In a series of embodiments, obtaining low-frequency spectral flatness information of a low-frequency spectrum by performing a spectral flatness extraction process on the low-frequency spectrum in step 2041 can be achieved by the following technical proposal, which involves obtaining fourth fusion configuration data of the low-frequency spectrum, the fourth fusion configuration data including the spectral line ordinal number of each fourth spectral line combination, obtaining the geometric mean value of the fourth spectral line combination and the arithmetic mean value of the fourth spectral line combination for each fourth spectral line combination, using the comparison value between the geometric mean value of the low-frequency spectrum and the arithmetic mean value of the low-frequency spectrum as spectral flatness information of the fourth spectral line combination, and generating spectral flatness information of the low-frequency spectrum based on the spectral flatness information of at least one fourth spectral line combination.
[0139] In a series of embodiments, obtaining the geometric mean value of the above-mentioned fourth spectral line combination can be achieved by the following technical method: extract the spectral coefficients corresponding to each spectral line ordinal number of the fourth spectral line combination from the low-frequency spectrum, square the spectral coefficients of each spectral line ordinal number to obtain the fourth squared spectral coefficient of each spectral line ordinal number, and if there are multiple spectral line ordinal numbers in the fourth spectral line combination, multiply the fourth squared spectral coefficients of the multiple spectral line ordinal numbers to obtain a second multiplication result, and then, based on the number of spectral line ordinal numbers, perform a root operation on the second multiplication result to obtain the geometric mean value corresponding to the fourth spectral line combination.
[0140] In a series of embodiments, obtaining the arithmetic mean of the above-mentioned fourth spectral line combinations can be achieved by the following technical method: extracting spectral coefficients corresponding to each spectral line ordinal number of the fourth spectral line combination from the low-frequency spectrum, squaring the spectral coefficients of each spectral line ordinal number to obtain the fourth squared spectral coefficient of each spectral line ordinal number, and if there are multiple spectral line ordinal numbers in the fourth spectral line combination, adding the fourth squared spectral coefficients of the multiple spectral line ordinal numbers to obtain the fourth summation result, and averaging the fourth summation result based on the number of spectral line ordinal numbers to obtain the arithmetic mean corresponding to the fourth spectral line combination.
[0141] An embodiment for determining the low-frequency spectral flatness information of the low-frequency spectrum in step 2041 can refer to the real-time method for extracting the spectral flatness information of the high-frequency spectrum in step 104, the only difference being that the processing target has been changed from the high-frequency spectrum to the low-frequency spectrum, and therefore the fourth spectral line combination used is also different from the third spectral line combination.
[0142] In step 2042, subband spectral flatness information corresponding to each high-frequency subband of the high-frequency spectrum is extracted from the spectral flatness information, and subband spectral envelope information corresponding to each high-frequency subband of the high-frequency spectrum is extracted from the spectral envelope information.
[0143] In step 2043, for each high-frequency subband of the high-frequency spectrum, the difference in spectral flatness values between the subband flatness information of each low-frequency subband in the low-frequency spectrum and the subband spectral flatness information of the high-frequency subband is determined, and the low-frequency subband with the smallest difference in spectral flatness values is determined as the target spectrum.
[0144] In step 2044, amplitude value adjustment processing is performed on the target spectrum corresponding to each high-frequency subband, based on the subband spectral envelope information corresponding to each high-frequency subband of the high-frequency spectrum and the numerical difference in spectral flatness corresponding to each high-frequency subband. Simultaneously, the adjustment results corresponding to multiple high-frequency subbands are spliced into the high-frequency spectrum.
[0145] In a series of embodiments, performing amplitude value adjustment processing on the target spectrum corresponding to each high-frequency subband in step 2044, in accordance with the subband spectral envelope information corresponding to each high-frequency subband of the high-frequency spectrum and the numerical difference in spectral flatness corresponding to each high-frequency subband, can be realized by the following technical proposal, which involves: determining white noise suitable for the numerical difference in spectral flatness of the high-frequency subbands and adding white noise suitable for the target spectrum to obtain a composite target spectrum; determining the spectral envelope information of the composite target spectrum and determining the numerical difference in spectral envelopes between the spectral envelope information of the composite target spectrum and the spectral envelope information of the high-frequency subbands; and performing adjustments to the amplitude value of the composite target spectrum based on the numerical difference in spectral envelopes.
[0146] As an example, the specific recovery process can be as shown in Figure 6. First, the low-frequency spectrum is subjected to spectral flatness analysis and calculation to obtain the spectral flatness of the low-frequency portion. The calculation process can be found by referring to formulas (7) to (9). Then, the low-frequency portion closest to each high-frequency subband is selected as the target spectrum according to the spectral flatness information of the high-frequency portion. Next, the target spectrum is energy-fine-tuned according to the difference in spectral flatness information and spectral envelope information. Finally, multiple subbands in the high-frequency portion are spliced to obtain a perfect high-frequency spectrum, and then adjusted with a gradient filter to obtain a perfect high-frequency spectrum. An inverse time-frequency transform of the MDCT is performed on the high-frequency spectrum to obtain a high-frequency signal. The recovered high-frequency signal and the low-frequency signal obtained by decoding with a decoder are input into a group of orthogonal mirror image mixed filters and combined and filtered to obtain an ultra-wideband audio signal.
[0147] In the audio frequency processing method provided in this embodiment, the high-frequency spectrum is reconstructed by performing combined processing according to the spectrum, spectral envelope information, and spectral flatness information of the high-frequency portion of the low-frequency signal recovered by the decoding terminal, and the decoding terminal controls the error, preventing the audio encoder (especially an ultra-low code-rate audio encoder based on NN modeling) from amplifying the coding error in the low-frequency portion in the high-frequency portion. As a result, the sound quality of the decoding is greatly improved.
[0148] In step 205, the high-frequency spectrum is subjected to time-domain transformation to obtain a high-frequency signal, and the low-frequency signal and the high-frequency signal are combined to obtain an audio frequency signal corresponding to the coding code stream.
[0149] As an example, an inverse time-frequency transform of an MDCT is performed on a high-frequency spectrum to obtain a high-frequency signal. The recovered high-frequency signal and the low-frequency signal obtained by decoding with a decoder are input to a group of orthogonal mirror image mixed filters and combined and filtered to obtain an audio frequency signal.
[0150] Referring to Figure 3D, which is a schematic flowchart of the audio frequency processing method provided by the embodiment of the present invention, Figure 3D shows a complete coding and decoding process.
[0151] In step 301, the coding terminal filters the audio frequency signal to obtain low-frequency and high-frequency signals.
[0152] In step 302, the coding terminal processes the low-frequency signal to obtain a code stream of the low-frequency signal.
[0153] In step 303, the coding terminal performs frequency domain conversion on the low-frequency signal to obtain a low-frequency spectrum, and also performs frequency domain conversion on the high-frequency signal to obtain a high-frequency spectrum.
[0154] In step 304, the coding terminal performs spectral envelope extraction processing on the low-frequency spectrum and the high-frequency spectrum to obtain spectral envelope information of the audio frequency signal, and also performs spectral flatness extraction processing on the high-frequency spectrum to obtain spectral flatness information of the high-frequency spectrum.
[0155] In step 305, the coding terminal performs quantization coding on the spectral flatness information of the high-frequency spectrum and the spectral envelope information of the audio frequency signal to obtain a bandwidth-extended code stream of the audio frequency signal, and constructs a coding code stream of the audio frequency signal using the bandwidth-extended code stream and the code stream of the low-frequency signal.
[0156] In step 306, the coding terminal sends the coding code stream to the decoding terminal.
[0157] In step 307, the decoding terminal decomposes the coding code stream to obtain a bandwidth-extended code stream and a code stream of the low-frequency signal.
[0158] In step 308, the decoding terminal decodes the code stream of the low-frequency signal to obtain a low-frequency signal, and also performs frequency domain conversion on the low-frequency signal to obtain the low-frequency spectrum of the low-frequency signal.
[0159] In step 309, the decoding terminal performs inverse quantization on the bandwidth-extended code stream to obtain spectral flatness information and spectral envelope information.
[0160] In step 310, the decoding terminal performs a high-frequency spectrum reconstruction process based on the spectral flatness information, spectral envelope information, and low-frequency spectrum to obtain the high-frequency spectrum.
[0161] In step 311, the decoding terminal performs a time-domain transformation on the high-frequency spectrum to obtain a high-frequency signal, and also performs a synthesis process with the low-frequency signal and the high-frequency signal to obtain an audio frequency signal corresponding to the coding code stream.
[0162] By filtering the audio frequency signal to obtain low-frequency signals and high-frequency signals, coding the low-frequency signals to obtain a code stream of the low-frequency signals, extracting spectral envelope information of the audio frequency signal and spectral flatness information of the high-frequency signal from the low-frequency spectrum of the low-frequency signal and the high-frequency spectrum of the high-frequency signal, quantizing and coding the spectral flatness information and spectral envelope information to obtain a bandwidth-extended code stream of the audio frequency signal, and combining this with the code stream of the low-frequency signal to form a coding code stream of the audio frequency signal, effective coding of the high-frequency signal can be achieved using the spectral envelope information and spectral flatness information, improving the perfection of coding in the high-frequency portion, and by combining the spectral information of the low-frequency signal, spectral envelope information, and spectral flatness information of the high-frequency portion recovered at the decoding terminal to reconstruct the high-frequency spectrum, the quality of the audio frequencies obtained in the subsequent decoding process can be improved.
[0163] The following describes an exemplary application of one actual application scenario of the embodiment of this invention.
[0164] In a series of embodiments, a client is operated on the first terminal, and the client can be of various types, such as an instant messaging client, a network conferencing client, a live streaming client, a browser, etc. In response to a voice frequency collection command issued by the sender (e.g., the initiator of a network conference, the main caster, the speaker of a voice call, etc.), the client activates the microphone equipped on the first terminal to collect a voice frequency signal, and filters the collected voice frequency signal to obtain low-frequency and high-frequency signals. The frequency of the low-frequency signal is lower than the frequency of the high-frequency signal. The low-frequency signal is coded to obtain a code stream of the low-frequency signal. The low-frequency signal is then subjected to frequency domain conversion to obtain a low-frequency spectrum, and the high-frequency signal is subjected to frequency domain conversion to obtain a high-frequency spectrum. The low-frequency spectrum and the high-frequency spectrum are subjected to spectral envelope extraction to obtain spectral envelope information of the audio frequency signal. The high-frequency spectrum is subjected to spectral flatness extraction to obtain spectral flatness information of the high-frequency spectrum. The spectral flatness information of the high-frequency spectrum and the spectral envelope information of the audio frequency signal are subjected to quantization coding to obtain a bandwidth-expanded code stream of the audio frequency signal. The bandwidth-expanded code stream and the code stream of the low-frequency signal constitute the coded code stream of the audio frequency signal. Next, the client can send the coded code stream to the server via the network, and the server can send the code stream to a second terminal related to the receiving side (e.g., participants in a network conference, an audience, or a receiver of a voice call).After a client (for example, an instant messaging client, a network conferencing client, a live streaming client, or a browser) receives a coding code stream sent by a server, it decomposes the coding code stream to obtain a bandwidth-extended code stream and a code stream of the low-frequency signal, decodes the code stream of the low-frequency signal to obtain a low-frequency signal, and performs frequency-domain conversion on the low-frequency signal to obtain the low-frequency spectrum of the low-frequency signal, inversely quantizes the bandwidth-extended code stream to obtain spectral flatness information and spectral envelope information, and performs high-frequency spectrum reconstruction based on the spectral flatness information, spectral envelope information and the low-frequency spectrum to obtain a high-frequency spectrum, the frequencies of the high-frequency spectrum being higher than the frequencies of the low-frequency spectrum, and performs time-domain conversion on the high-frequency spectrum to obtain a high-frequency signal, and then synthesizes the low-frequency signal and the high-frequency signal to obtain an audio frequency signal corresponding to the coding code stream.
[0165] In the voice frequency processing method provided in the embodiment of the present invention, the coding terminal compresses the low-frequency portion of the voice frequency signal to obtain a code stream of the low-frequency signal, and at the same time executes a bandwidth expansion plan based on spectral flatness information to realize code transmission of ultraband voice at an extremely low code rate.
[0166] Referring to Figure 4, Figure 4 is a schematic diagram of the bandwidth expansion coding in the audio frequency processing method provided by the embodiment of the present invention. The input signal is an ultra-wideband signal with a sampling rate of 32 kilohertz (kHz). A 32kHz sampling rate means sampling 32,000 times per second to obtain 32,000 sampling points. The input signal is divided into frames according to a frame length of 640 points, with 640 sampling points forming one frame. Each frame has a frame length of 640 points and a frame duration of 0.02 seconds. The signal is processed through a group of QMF band division filters to obtain a high-frequency portion with a frame length of 320 points and a low-frequency portion with a frame length of 320 points. In the following text, these will be referred to as the high-frequency signal and the low-frequency signal, respectively.
[0167] Based on a sampling point with a frame length of 320 and a sampling point with a frame shift of 160, the high-frequency signal and low-frequency signal are subjected to MDCT time-frequency transformation, respectively, to obtain corresponding high-frequency spectra and low-frequency spectra. The sampling point with a frame shift of 160 refers to a time interval corresponding to a sampling point where the time difference between the start positions of two adjacent frames is 160.
[0168] Low-frequency and high-frequency spectra are fused according to their respective spectral envelope fusion tables to extract spectral envelope information. The formula used for extracting spectral envelope information is as shown in formula (6).
[0169]
number
[0170] In the formula, m k represents the kth spectral coefficient of the MDCT transformation result, and i represents the ordinal number of the spectral envelope. For example, if i is 1, then m0, m1, ..., m 19 Each term is squared, and the squared results are then added together.
[0171] The spectral envelope fusion tables for the high-frequency and low-frequency portions used in the embodiments of this invention are shown in Tables 8 and 9, respectively.
[0172] First, to explain by combining Table 8, Table 8 shows the fusion process for the four sets of MDCT conversion results. The spectral coefficients from the 0th to the 19th of the MDCT conversion results are fused based on formula (6), which corresponds to fusion of the spectral lines from the 0th to the 19th, the spectral coefficients from the 20th to the 54th of the MDCT conversion results are fused based on formula (6), the spectral coefficients from the 55th to the 89th of the MDCT conversion results are fused based on formula (6), and the spectral coefficients from the 90th to the 130th of the MDCT conversion results are fused based on formula (6).
[0173] [Table 8]
[0174] Table 8, which shows the spectral envelope fusion in the high-frequency region, is based on the critical band in the psychoacoustic model and is obtained by comprehensively considering BWE quality and code rate in specific experiments. The critical band is a result obtained from psychoacoustic experiments and specifically reflects the conversion between physical-mechanical stimulation and neuroelectric stimulation in the cochlea of the human inner ear. The neuroelectric stimulation converted by the human ear is consistent for a specific frequency and other frequencies within a specific range near it, as well as for pure tone speech frequency signals. In other words, it is not necessary to achieve excessively high frequency domain resolution by using an excessively high code rate. Based on multiple experimental measurements, and using code rate and BWE quality as evaluation indicators of the experimental results, the data shown in Table 8 was obtained.
[0175] Furthermore, combining this with Table 9, Table 9 shows the fusion process for one set of MDCT conversion results, where spectral coefficients 80 through 150 of the MDCT conversion results are fused based on formula (6).
[0176] [Table 9]
[0177] Table 9 of the spectral envelope fusion for the low-frequency portion was also obtained through statistical and measured experiments. When the encoder used in the low-frequency portion is an AI ultra-wideband speech encoder, the AI ultra-wideband speech encoder has relatively strong modeling capabilities for speech and possesses noise reduction capabilities. Therefore, it is necessary to introduce a variable to quantify and predict its noise reduction effect. The energy envelope of the low-frequency portion can be used as a predictor variable, and statistical and measured data on a large dataset revealed that when the energy envelope selection range for the low-frequency portion is the data shown in Table 9, relatively accurate and stable predictive values can be obtained, and the complexity and code rate are also acceptable. Considering factors such as calculation accuracy, stability, complexity, and code rate comprehensively, the data shown in Table 9 was selected as the spectral envelope fusion table for the low-frequency portion.
[0178] High-frequency spectra were merged according to a corresponding spectral flatness fusion table, and spectral flatness information was extracted. Formulas (7) to (9) can be used as a reference for the calculations to extract spectral flatness information.
[0179]
number
[0180] In the formula, m knume(i) represents the kth spectral coefficient of the MDCT conversion result, nume(i) and demo(i) represent the geometric mean and arithmetic mean of each spectral line in the MDCT conversion result, respectively, and the spectral flatness information Flatness(i) is a comparison value between the geometric mean and the arithmetic mean. The spectral flatness information reflects whether the audio frequency corresponding to the spectrum in question is closer due to white noise or a pure tone signal closer due to a single frequency. i represents the ordinal number of the spectral flatness information; for example, if i is 1, then m0, m1, ..., m 39 Each of these values is squared, and the spectral flatness information is determined based on the squared results.
[0181] The high-frequency spectrum was analyzed by extracting spectral flatness information based on a spectral flatness fusion table for the high-frequency portion, and the results are shown in Table 10.
[0182] [Table 10]
[0183] Table 10, which shows the spectral flatness fusion in the high-frequency region, is based on the critical band in the psychoacoustic model as its theoretical foundation, and is obtained by comprehensively considering the quality of the BWE and the code rate in specific experiments. The critical band is a result obtained from experiments based on psychoacoustics, and specifically reflects the conversion between physical-mechanical stimulation and neuroelectric stimulation in the cochlea of the human inner ear. The neuroelectric stimulation converted by the human ear is consistent for a certain frequency and other frequencies within a specific range near it, as well as for pure tone speech frequency signals. In other words, it is not necessary to achieve excessively high frequency domain resolution by using an excessively high code rate. Based on multiple experimental measurements, and using the code rate and BWE quality as evaluation indicators of the experimental results, the data shown in Table 10 was obtained.
[0184] The spectral envelope information and spectral flatness information were quantized and coded according to their respective quantization tables to form a BWE code stream. The quantization table used for quantizing the spectral flatness information is shown in Table 11.
[0185] [Table 11]
[0186] The generation process for Table 11 was obtained by statistically analyzing experiments. The spectral flatness was calculated for a large number of audio frequency files based on the above flow, ultimately yielding a statistical distribution based on a large number of audio frequency distributions. Considering the code rate and audio frequency quality comprehensively, this statistical distribution was clustered and quantized to finally generate Table 11. The generation methods for the spectral envelope quantization tables 12 (high-frequency portion, first and second subbands), 13 (high-frequency portion, third and fourth subbands), and 14 (low-frequency portion) are similar to those for Table 11. The specific results for all quantization tables correlate with experimental statistics, and the dimensions of the quantization tables can be flexibly adjusted according to specific application scenarios.
[0187] The spectral envelope quantization tables for the first and second subbands in the high-frequency region are shown in Table 12.
[0188] [Table 12]
[0189] The spectral envelope quantization tables for the third and fourth subbands in the high-frequency region are shown in Table 13.
[0190] [Table 13]
[0191] The spectral envelope quantization table for the low-frequency portion is shown in Table 14.
[0192] [Table 14]
[0193] Figure 5 is a schematic diagram of bandwidth expansion decoding in the audio frequency processing method provided by the embodiment of the present invention, where the decoding terminal recovers an ultra-wideband audio frequency signal after receiving a BWE code stream and a low-frequency signal. After receiving the BWE code stream, the decoding terminal recovers spectral envelope information and spectral flatness information via a decoding and inverse quantization module. A low-frequency spectrum is obtained by performing an MDCT time-frequency transformation on the low-frequency time-domain signal. The high-frequency spectrum is recovered according to the low-frequency spectrum, high-frequency spectral envelope information, and high-frequency spectral flatness information. The specific recovery process is shown in Figure 6. The recovery process is as follows: First, the low-frequency spectrum is analyzed for flatness to obtain the spectral flatness of the low-frequency portion. The calculation process can be found by referring to formulas (7) to (9). Then, according to the spectral flatness information of the high-frequency portion, the low-frequency portion closest to each high-frequency subband is selected as the target spectrum. Next, the target spectrum is energy-fine-tuned according to the difference in spectral flatness information and spectral envelope information. Finally, multiple subbands in the high-frequency portion are spliced to obtain a perfect high-frequency spectrum, and then adjusted with a gradient filter to obtain a perfect high-frequency spectrum. The time-frequency inverse transform of the MDCT is performed on the high-frequency spectrum to obtain a high-frequency signal. The recovered high-frequency signal and the low-frequency signal obtained by decoding with a decoder are input into a group of quadrature mirror image mixed filters and combined and filtered to obtain an ultra-wideband audio signal.
[0194] The audio frequency processing method provided in this embodiment reconstructs the high-frequency spectrum by performing combined determination and adjustment according to the spectrum of the low-frequency signal recovered by the decoding terminal, the original spectral envelope information contained in the BWE side information, and the high-frequency spectral flatness information. This minimizes the amplification of coding errors in the low-frequency portion by the high-frequency portion in ultra-low code rate audio encoders (especially ultra-low code rate audio encoders based on NN modeling), thus significantly improving the sound quality of the decoding.
[0195] Referring to Figure 7, which is a schematic diagram of coding in the audio frequency processing method provided by the embodiment of the present invention, in the coding terminal, an ultra-wideband signal passes through a group of analysis filters to obtain a high-frequency portion and a low-frequency portion, and an AI ultra-wideband audio encoder codes and compresses the low-frequency portion to obtain a code stream of the low-frequency signal. The low-frequency portion and the high-frequency portion simultaneously become inputs to the BWE encoder submitted by the embodiment of the present invention, and a BWE code stream is generated, and together with the code stream of the low-frequency signal, the final code stream is implemented.
[0196] Referring to Figure 8, Figure 8 is a schematic diagram of decoding in the audio frequency processing method provided by the embodiment of the present invention, in which a decoding terminal decomposes the received coding code stream into a BWE code stream and the code stream of the low frequency signal. The code stream of the low frequency signal is recovered as a low frequency signal via an AI ultra-wideband audio decoder, and the low frequency signal and the BWE code stream are recovered as a high frequency code stream by passing through the BWE decoder submitted by the embodiment of the present invention, the high frequency code stream is converted in the time domain to a high frequency signal, and the high frequency signal and the low frequency signal are passed through a group of combining filters to generate an ultra-wideband signal.
[0197] The bandwidth expansion technology in the audio frequency processing method provided in this embodiment can achieve ultra-wideband audio coding at extremely low code rates when combined with an AI ultra-wideband audio encoder. In this embodiment, by adding spectral flatness information and spectral envelope information as BWE side information, the broadband signal of the AI ultra-wideband audio decoder is expanded to an ultra-wideband signal under extremely low complexity. Furthermore, by adding error control to the decoder side for the low-frequency files generated by this NN modeling-based audio encoder, the influence of low-frequency quantization noise on the reconstructed high-frequency signal during bandwidth expansion is reduced compared to other bandwidth expansion methods in related technologies.
[0198] In the following, an exemplary structure in which the audio frequency processing device 455 provided in this embodiment is implemented as a software module will be described. In a series of embodiments, as shown in Figure 2A, the software module in the audio frequency processing device 455 stored in the memory 450 is a band division module 4551 arranged to filter the audio frequency signal to obtain a low frequency signal and a high frequency signal, wherein the frequency of the low frequency signal is lower than the frequency of the high frequency signal; a coding module 4552 arranged to code the low frequency signal to obtain a code stream of the low frequency signal; and a frequency domain conversion module arranged to perform frequency domain conversion on the low frequency signal to obtain a low frequency spectrum and also perform frequency domain conversion on the high frequency signal to obtain a high frequency spectrum. The system may include: a module 4553; an extraction module 4554 arranged to perform spectral envelope extraction processing on low-frequency and high-frequency spectra to obtain spectral envelope information of the audio frequency signal, and spectral flatness extraction processing on the high-frequency spectrum to obtain spectral flatness information of the high-frequency spectrum; and a quantization module 4555 arranged to perform quantization coding processing on the spectral flatness information of the high-frequency spectrum and the spectral envelope information of the audio frequency signal to obtain a bandwidth-expanded code stream of the audio frequency signal, and to construct a coding code stream of the audio frequency signal using the bandwidth-expanded code stream and the code stream of the low-frequency signal.
[0199] In a series of embodiments, the extraction module 4554 is further configured to perform spectral envelope extraction on low-frequency spectra to obtain low-frequency spectral envelope information of low-frequency spectra, and to perform spectral envelope extraction on high-frequency spectra to obtain high-frequency spectral envelope information of high-frequency spectra, thereby constituting spectral envelope information of an audio frequency signal using the low-frequency spectral envelope information and the high-frequency spectral envelope information.
[0200] In a series of embodiments, the extraction module 4554 is further configured to acquire first fusion configuration data of low-frequency spectra, the first fusion configuration data includes the spectral ordinal number of each first spectral line combination, and for each first spectral line combination, it is configured to perform the following processes: extract spectral coefficients corresponding to each spectral ordinal number of the first spectral line combination from the low-frequency spectrum; square the spectral coefficients of each spectral ordinal number to obtain the first squared spectral coefficient of each spectral ordinal number; if there are multiple spectral ordinal numbers of the first spectral line combination, add the first squared spectral coefficients of the multiple spectral ordinal numbers to obtain a first addition result; logarithmize the first addition result to obtain first fusion spectral envelope information corresponding to the first spectral line combination; and generate low-frequency spectral envelope information based on the first fusion spectral envelope information of at least one first spectral line combination.
[0201] In a series of embodiments, the extraction module 4554 is further configured to acquire second fusion configuration data of high-frequency spectra, the second fusion configuration data includes the spectral ordinal number of each second spectral line combination, and for each second spectral line combination, it is configured to perform the following processes: extract spectral coefficients corresponding to each spectral ordinal number of the second spectral line combination from the high-frequency spectrum; square the spectral coefficients of each spectral ordinal number to obtain the second squared spectral coefficient of each spectral ordinal number; if there are multiple spectral ordinal numbers of the second spectral line combination, add the second squared spectral coefficients of the multiple spectral ordinal numbers to obtain a second addition result; logarithmize the second addition result to obtain second fusion spectral envelope information corresponding to the second spectral line combination; and generate high-frequency spectral envelope information based on the second fusion spectral envelope information of at least one second spectral line combination.
[0202] In a series of embodiments, the extraction module 4554 is further configured to acquire third fusion configuration data of high-frequency spectra, the third fusion configuration data includes the spectral ordinal number of each third spectral line combination, and for each third spectral line combination, it is configured to perform the following processes: acquiring the geometric mean of the third spectral line combination and acquiring the arithmetic mean of the third spectral line combination; comparing the geometric mean of the third spectral line combination with the arithmetic mean of the third spectral line combination to obtain spectral flatness information for the third spectral line combination; and generating spectral flatness information of high-frequency spectra based on the spectral flatness information of at least one third spectral line combination.
[0203] In a series of embodiments, the extraction module 4554 is configured to further acquire third fusion configuration data of high-frequency spectra, the third fusion configuration data includes the spectral ordinal number of each third spectral line combination, and for each third spectral line combination, it performs the following processes: extracting spectral coefficients corresponding to each spectral ordinal number of the third spectral line combination from the high-frequency spectrum; squaring the spectral coefficients of each spectral ordinal number to obtain the third squared spectral coefficient of each spectral ordinal number; if there are multiple spectral ordinal numbers of the third spectral line combination, multiplying the third squared spectral coefficients of the multiple spectral ordinal numbers to obtain a first multiplication result; setting the first multiplication result as a square root based on the number of spectral ordinal numbers to obtain a geometric mean corresponding to the third spectral line combination; and constructing the geometric mean of the third spectral line combination using the geometric mean of multiple third spectral line combinations.
[0204] In a series of embodiments, the extraction module 4554 is configured to further acquire third fusion configuration data of high-frequency spectra, the third fusion configuration data includes the spectral ordinal number of each third spectral line combination, and for each third spectral line combination, it performs the following processes: extracting spectral coefficients corresponding to each spectral ordinal number of the third spectral line combination from the high-frequency spectrum; squaring the spectral coefficients of each spectral ordinal number to obtain the third squared spectral coefficient of each spectral ordinal number; if there are multiple spectral ordinal numbers of the third spectral line combination, adding the third squared spectral coefficients of the multiple spectral ordinal numbers to obtain a third addition result; averaging the third addition results based on the number of spectral ordinal numbers to obtain an arithmetic mean corresponding to the third spectral line combination; and constructing the arithmetic mean of the third spectral line combination using the arithmetic mean of multiple third spectral line combinations.
[0205] In a series of embodiments, the quantization module 4555 is further configured to acquire a quantization table for spectral flatness information and a quantization table for spectral envelope information, to quantize the spectral flatness information of the high-frequency spectrum according to the quantization table for spectral flatness information to obtain a quantized result for spectral flatness, to quantize the spectral envelope information of the audio frequency signal according to the quantization table for spectral envelope information to obtain a quantized result for the spectral envelope, and to construct a bandwidth-extended code stream of the audio frequency signal using the quantized results for spectral flatness and the quantized result for spectral envelope.
[0206] In a series of embodiments, the quantization module 4555 further acquires a plurality of audio sample signals and processes each audio sample signal to obtain low-frequency and high-frequency sample signals, wherein the frequency of the low-frequency sample signal is lower than the frequency of the high-frequency sample signal. The low-frequency sample signal is subjected to frequency domain conversion to obtain a low-frequency sample spectrum, and the high-frequency sample signal is subjected to frequency domain conversion to obtain a high-frequency sample spectrum. The low-frequency and high-frequency sample spectra are subjected to spectral envelope extraction to obtain spectral envelope information of the audio sample signal, and the high-frequency spectrum is subjected to spectral flatness extraction to obtain spectral flatness information of the audio sample signal. The system is configured to perform the following steps: obtaining degree information; clustering the spectral flatness information of multiple audio sample signals to obtain multiple spectral flatness clustering centers and spectral flatness corresponding to each spectral flatness clustering center, and constructing a quantization table of spectral flatness information based on the multiple spectral flatness clustering centers and the spectral flatness information of each spectral flatness clustering center; and clustering the spectral envelope information of multiple audio sample signals to obtain multiple spectral envelope clustering centers and spectral envelope information corresponding to each spectral envelope clustering center, and constructing a quantization table of spectral envelope information based on the multiple spectral envelope clustering centers and the spectral envelope information corresponding to each spectral envelope clustering center.
[0207] In a series of embodiments, the coding module 4552 is configured to further filter the audio frequency signal to obtain low-frequency and high-frequency signals of the audio frequency signal, wherein the frequency of the low-frequency signal is lower than the frequency of the high-frequency signal, feature extraction is performed on the low-frequency signal to obtain a first feature of the low-frequency signal, high-frequency analysis is performed on the high-frequency signal to obtain a second feature of the high-frequency signal, the feature dimension of the second feature is lower than the feature dimension of the first feature, and the first and second features are quantized and coded to obtain a code stream of the low-frequency signal of the audio frequency signal.
[0208] In the following, an exemplary structure in which the audio frequency processing device 555 provided in this embodiment is implemented as a software module will be described. In a series of embodiments, as shown in Figure 2B, the software module in the audio frequency processing device 555 stored in the memory 550 includes a decomposition module 5551 arranged to decompose a coding code stream to obtain a bandwidth-extended code stream and the code stream of the low-frequency signal, a decoding module 5552 arranged to decode the code stream of the low-frequency signal to obtain a low-frequency signal and to perform frequency-domain conversion on the low-frequency signal to obtain the low-frequency spectrum of the low-frequency signal, and an inverse quantization process of the bandwidth-extended code stream to obtain spectral flatness information and spectral pulsate The system may include: an inverse quantization module 5553 arranged to obtain entanglement information; a reconstruction module 5554 arranged to obtain a high-frequency spectrum by performing a reconstruction process of the high-frequency spectrum based on spectral flatness information, spectral envelope information, and the low-frequency spectrum, wherein the frequency of the high-frequency spectrum is higher than the frequency of the low-frequency spectrum; and a time-domain conversion module 5555 arranged to obtain a high-frequency signal by performing a time-domain conversion process on the high-frequency spectrum, and to obtain an audio frequency signal corresponding to the coding code stream by combining the low-frequency signal and the high-frequency signal.
[0209] In a series of embodiments, the reconstruction module 5554 further performs spectral flatness extraction processing on the low-frequency spectrum to obtain low-frequency spectral flatness information of the low-frequency spectrum, extracts subband spectral flatness information corresponding to each high-frequency subband of the high-frequency spectrum from the spectral flatness information, extracts subband spectral envelope information corresponding to each high-frequency subband of the high-frequency spectrum from the spectral envelope information, determines the difference in spectral flatness values between the subband spectral flatness information of each low-frequency subband in the low-frequency spectrum and the subband spectral flatness information of the high-frequency subband for each high-frequency subband of the high-frequency spectrum, determines the low-frequency subband with the smallest difference in spectral flatness values as the target spectrum, performs amplitude value adjustment processing on the target spectrum corresponding to each high-frequency subband according to the subband envelope information corresponding to each high-frequency subband of the high-frequency spectrum and the difference in spectral flatness values corresponding to each high-frequency subband, and arranges to splice the adjustment results corresponding to multiple high-frequency subbands into the high-frequency spectrum.
[0210] In a series of embodiments, the reconstruction module 5554 is further configured to perform the following processes for each high-frequency subband: determining white noise suitable for the numerical difference in spectral flatness of the high-frequency subbands and adding suitable white noise to the target spectrum to obtain a composite target spectrum; determining the spectral envelope information of the composite target spectrum and determining the numerical difference in spectral envelopes between the spectral envelope information of the composite target spectrum and the spectral envelope information of the high-frequency subbands; and adjusting the amplitude value of the composite target spectrum based on the numerical difference in spectral envelopes.
[0211] In a series of embodiments, the reconstruction module 5554 is further configured to acquire the geometric mean of the low-frequency spectrum and the arithmetic mean of the low-frequency spectrum, and to use the comparison value between the geometric mean of the low-frequency spectrum and the arithmetic mean of the low-frequency spectrum as low-frequency spectral flatness information.
[0212] In a series of embodiments, the reconstruction module 5554 is further configured to acquire fourth fusion configuration data of the low-frequency spectrum, the fourth fusion configuration data includes the spectral ordinal number of each fourth spectral line combination, and for each fourth spectral line combination, it is configured to perform the following processes: extracting spectral coefficients corresponding to each spectral ordinal number of the fourth spectral line combination from the low-frequency spectrum; squaring the spectral coefficients of each spectral ordinal number to obtain the fourth squared spectral coefficient of each spectral ordinal number; if there are multiple spectral ordinal numbers of the fourth spectral line combination, multiplying the fourth squared spectral coefficients of the multiple spectral ordinal numbers to obtain a second multiplication result; setting the second multiplication result as a square root based on the number of spectral ordinal numbers to obtain a geometric mean corresponding to the fourth spectral line combination; and constructing the geometric mean of the low-frequency spectrum using the geometric mean of the multiple fourth spectral line combinations.
[0213] In a series of embodiments, the reconstruction module 5554 is further configured to acquire fourth fusion configuration data of the low-frequency spectrum, the fourth fusion configuration data includes the spectral ordinal number of each fourth spectral line combination, and for each fourth spectral line combination, it is configured to perform the following processes: extracting spectral coefficients corresponding to each spectral ordinal number of the fourth spectral line combination from the low-frequency spectrum; squaring the spectral coefficients of each spectral ordinal number to obtain the fourth squared spectral coefficient of each spectral ordinal number; if there are multiple spectral ordinal numbers of the fourth spectral line combination, adding the fourth squared spectral coefficients of the multiple spectral ordinal numbers to obtain a fourth summation result; averaging the fourth summation results based on the number of spectral ordinal numbers to obtain an arithmetic mean corresponding to the fourth spectral line combination; and constructing an arithmetic mean of the low-frequency spectrum using the arithmetic mean of the multiple fourth spectral line combinations.
[0214] In this embodiment, a computer program product is provided, which includes computer-executable instructions, and these computer-executable instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions to cause the electronic device to perform the audio frequency processing method of this embodiment.
[0215] In this embodiment, a computer-readable storage medium is provided which stores executable commands. When a computer-readable command is stored in the storage medium and executed by a processor, the processor is guided to execute an audio frequency processing method provided in this embodiment, for example, the audio frequency processing method shown in Figures 3A to 3D.
[0216] In a series of embodiments, the computer-readable storage medium may be memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage device, optical disk, or CD-ROM, and may also be various devices containing one or any combination of the above memories.
[0217] In a series of embodiments, computer-executable instructions can take the form of programs, software, software modules, scripts, or code, and can be written in any form of programming language (compiled or interpreted language, or declarative or procedural language), and can be deployed in any form, and may include independently deployed programs, or other units deployed as modules, components, subprograms, or other units suitable for use within a computing environment.
[0218] For example, computer-executable instructions are not necessarily stored in files corresponding to the file system; they may also be stored in files of other programs or data, for instance, within one or more scripts of a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program being discussed, or within multiple common files (e.g., in files for one or more modules, subprograms, or code sections).
[0219] For example, a computer-executable command may be configured to run on a single electronic device, or it may run on multiple electronic devices located at a single location, or it may run on multiple electronic devices distributed across multiple locations and interconnected by a communication network.
[0220] In summary, according to the embodiment of the present invention, by filtering the audio frequency signal to obtain a low-frequency signal and a high-frequency signal, coding the low-frequency signal to obtain a code stream of the low-frequency signal, extracting spectral envelope information of the audio frequency signal and spectral flatness information of the high-frequency signal from the low-frequency spectrum of the low-frequency signal and the high-frequency spectrum of the high-frequency signal, quantizing and coding the spectral flatness information and spectral envelope information to obtain a bandwidth-extended code stream of the audio frequency signal, and combining this with the code stream of the low-frequency signal to form a coding code stream of the audio frequency signal, effective coding of the high-frequency signal can be achieved via the spectral envelope information and spectral flatness information, the perfection of coding the high-frequency portion can be increased, and the quality of the audio frequency obtained in the subsequent decoding process can be improved.
[0221] The above description is merely an example of the present application and does not limit the scope of protection of the present application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of the present application shall all be included within the scope of protection of the present application.
Claims
1. A method for processing audio frequencies performed by an electronic device, The coding code stream is decomposed to obtain a bandwidth-extended code stream and a code stream of low-frequency signals. The code stream of the low-frequency signal is decoded to obtain a low-frequency signal, and the low-frequency signal is subjected to frequency domain conversion to obtain the low-frequency spectrum of the low-frequency signal. The aforementioned bandwidth-extended code stream is subjected to inverse quantization to obtain spectral flatness information and spectral envelope information. Based on the spectral flatness information, the spectral envelope information, and the low-frequency spectrum, a high-frequency spectrum reconstruction process is performed to obtain a high-frequency spectrum. This includes processing the high-frequency spectrum in the time domain to obtain a high-frequency signal, and combining the low-frequency signal and the high-frequency signal to obtain an audio frequency signal corresponding to the coding code stream. Based on the aforementioned spectral flatness information, spectral envelope information, and low-frequency spectrum, a high-frequency spectrum reconstruction process is performed to obtain a high-frequency spectrum. The low-frequency spectrum is subjected to a spectral flatness extraction process to obtain low-frequency spectral flatness information of the low-frequency spectrum, and subband spectral flatness information of each low-frequency subband in the low-frequency spectrum is extracted from the low-frequency spectral flatness information. The process involves extracting subband spectral flatness information corresponding to each high-frequency subband of the high-frequency spectrum from the spectral flatness information, and extracting subband spectral envelope information corresponding to each high-frequency subband of the high-frequency spectrum from the spectral envelope information. For each high-frequency subband of the high-frequency spectrum, the difference in spectral flatness values between the subband spectral flatness information of each low-frequency subband in the low-frequency spectrum and the subband spectral flatness information of the high-frequency subband is determined, and the low-frequency subband with the smallest difference in spectral flatness values is determined as the target spectrum. A speech frequency processing method comprising: performing amplitude value adjustment processing on a target spectrum corresponding to each high-frequency subband according to subband spectral envelope information corresponding to each high-frequency subband of the high-frequency spectrum and the numerical difference in spectral flatness corresponding to each high-frequency subband; and splicing the adjustment results corresponding to a plurality of high-frequency subbands into the high-frequency spectrum.
2. Performing amplitude value adjustment processing on the target spectrum corresponding to each high-frequency subband, in accordance with the subband spectral envelope information corresponding to each high-frequency subband of the aforementioned high-frequency spectrum and the numerical difference in spectral flatness corresponding to each high-frequency subband, is: For the target spectrum corresponding to each of the aforementioned high-frequency subbands, A process to determine a white noise suitable for the numerical difference in spectral flatness of the high-frequency subband, and to add the suitable white noise to the target spectrum to obtain a composite target spectrum, The process involves determining the spectral envelope information of the composite target spectrum and determining the difference in spectral envelope values between the spectral envelope information of the composite target spectrum and the spectral envelope information of the high-frequency subband. The method according to claim 1, comprising: performing a process to adjust the amplitude value of the composite target spectrum based on the spectral envelope numerical difference.
3. Memory for storing executable commands for the computer, An electronic device comprising: a processor for realizing the audio frequency processing method according to claim 1 or 2, which executes a computer-executable command stored in the memory;
4. A computer program that causes a computer to implement the audio frequency processing method described in either claim 1 or 2.