An audio encoding method, apparatus, device, and storage medium

By performing frame-by-frame processing and type differentiation of audio signals in high-concurrency and weak-network environments, and using different encoding methods to process audio frames, the encoding efficiency and quality issues of the Audio Vivid standard in low-bitrate and narrow-band voice communication scenarios are solved, achieving higher encoding quality and user experience.

CN121171235BActive Publication Date: 2026-02-06MALANSHAN AUDIO & VIDEO LABORATORY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511709101.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2026-02-06
Estimated Expiration
2045-11-20

AI Technical Summary

Technical Problem

The existing Audio Vivid audio coding standard is difficult to meet the real-time voice transmission requirements of high concurrency and weak network environments in low bit rate and narrow band voice communication scenarios, resulting in the inability to achieve optimal coding efficiency and quality. In particular, it has application limitations in voice transmission in unstable networks such as large-scale online meetings, smart home multi-device collaboration, high-speed rail and remote areas, as well as in emerging scenarios such as VR and metaverse.

Method used

By performing frame-based processing on audio signals under high concurrency and weak network conditions, the temporal and frequency domain characteristics of audio frames are determined, audio frame types are distinguished, and different encoding methods are used for processing: a low-rate speech mode encoder is used for frames containing human language, an Audio Vivid encoder is used for frames containing music, and a weighted fusion encoding method is used for transition frames to generate the target audio signal.

Benefits of technology

It improves audio encoding quality and enhances the user experience, especially in real-time voice transmission under high concurrency and weak network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121171235B_ABST
    Figure CN121171235B_ABST
Patent Text Reader

Abstract

The application discloses an audio encoding method, device and equipment and a storage medium, and relates to the field of audio processing. The method comprises the following steps: obtaining an initial audio signal in a target environment, performing frame processing on the initial audio signal to obtain corresponding audio frames; determining the type of each audio frame based on the time domain characteristics and frequency domain characteristics of each audio frame; if the audio frame is a first type of frame, encoding the first type of frame based on a preset speech encoder to obtain a corresponding first audio bitstream; if the audio frame is a second type of frame, encoding the second type of frame based on an Audio Vivid encoder to obtain a corresponding second audio bitstream; if the audio frame is a third type of frame, encoding the third type of frame by using a preset weighted fusion encoding mode to obtain a corresponding third audio bitstream; and determining a target audio signal based on the first audio bitstream, the second audio bitstream and the third audio bitstream. Therefore, the application can improve the encoding quality of audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing, and in particular to an audio encoding method, apparatus, device, and storage medium. Background Technology

[0002] Audio Vivid, as the world's first AI-based audio codec standard, is forward-looking in the field of immersive audio, but it has shortcomings in native speech encoding capabilities.

[0003] Audio Vivid exhibits significant functional gaps in scenarios such as low-bitrate real-time communication and narrowband voice adaptation. It does not differentiate between audio signal types during encoding, employing a uniform encoding method. While this simplifies the encoder processing flow, it may result in suboptimal encoding efficiency and quality for certain specific voice types. Audio Vivid struggles to meet the demands of real-time voice transmission in high-concurrency and weak network environments, such as large-scale online conferences, multi-device collaboration in smart homes, voice transmission in unstable networks like high-speed rail and remote areas, and bandwidth allocation during multi-user real-time interaction in emerging scenarios like VR (Virtual Reality) and the metaverse. Therefore, the existing Audio Vivid standard is ill-suited for low-bitrate and narrowband voice communication scenarios, failing to provide efficient and high-quality voice encoding support for these scenarios, resulting in significant application limitations.

[0004] Therefore, improving the encoding quality of audio is a pressing technical problem that needs to be solved. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide an audio encoding method, apparatus, device, and storage medium that can improve the encoding quality of audio. The specific solution is as follows:

[0006] Firstly, this application provides an audio encoding method, including:

[0007] In a target environment where the concurrency meets the preset high concurrency conditions and the network quality meets the preset weak network conditions, the initial audio signal is acquired, and the initial audio signal is processed into frames to obtain the corresponding audio frames.

[0008] The short-time energy, zero-crossing rate, and pitch period of each audio frame are determined to determine the time-domain characteristics of each audio frame. The spectral flatness and spectral centroid of each audio frame are determined to determine the frequency-domain characteristics of each audio frame. Based on the time-domain characteristics and the frequency-domain characteristics, the type of each audio frame is determined.

[0009] If the audio frame is a first type of frame, then the first type of frame is encoded based on a preset speech encoder configured with a preset low-rate speech mode to obtain the corresponding first audio bitstream; the first type of frame is an audio frame that contains only human language.

[0010] If the audio frame is a second type of frame, then the second type of frame is encoded based on the Audio Vivid encoder to obtain the corresponding second audio bitstream; the second type of frame is an audio frame containing music but not human language.

[0011] If the audio frame is a third type of frame, then the third type of frame is encoded using a preset weighted fusion coding method to obtain the corresponding third audio bitstream; the third type of frame is an audio frame that transitions between the first type of frame and the second type of frame;

[0012] The target audio signal is determined based on the first audio stream, the second audio stream, and the third audio stream, so as to transmit the target audio signal in the target environment.

[0013] Optionally, the step of performing frame segmentation on the initial audio signal to obtain corresponding audio frames includes:

[0014] The initial audio signal is segmented into frames based on preset sampling points and the Hanning window function to determine the corresponding audio frames.

[0015] Optionally, determining the short-time energy, zero-crossing rate, and pitch period of each audio frame to determine the temporal characteristics of each audio frame includes:

[0016] Determine the square value of each audio frame, and determine the short-time energy of each audio frame based on the square value;

[0017] Based on a preset sampling order, the first audio frame is selected as the first target audio frame in each of the audio frames, and the next audio frame after the first target audio frame is determined as the second target audio frame based on the preset sampling order.

[0018] A first function value is determined based on the first target audio frame and a preset symbolic function, and a second function value is determined based on the second target audio frame and the preset symbolic function, so as to determine the difference between the first function value and the second function value;

[0019] Jump to the step of selecting the first audio frame as the first target audio frame in each audio frame based on the preset sampling order, until all audio frames have been selected, and determine the zero-crossing rate of each audio frame based on the obtained differences and the number of preset sampling points;

[0020] The pitch period of each audio frame is determined based on a preset short-time autocorrelation function and each audio frame.

[0021] The temporal characteristics of each audio frame are determined based on the short-time energy, the zero-crossing rate, and the pitch period.

[0022] Optionally, determining the spectral flatness and spectral centroid of each audio frame to determine the frequency domain characteristics of each audio frame includes:

[0023] Perform a Fast Fourier Transform on each of the audio frames to obtain the corresponding spectrum;

[0024] Determine the spectral flatness of the spectrum within a preset frequency range, and determine the spectral centroid of the spectrum;

[0025] The frequency domain characteristics of each audio frame are determined based on the spectral flatness and the spectral centroid.

[0026] Optionally, determining the type of each audio frame based on the time-domain features and the frequency-domain features includes:

[0027] The frame type of an audio frame whose time-domain features indicate that the pitch period is within a preset period, the short-time energy is greater than a preset energy threshold, and the frequency-domain features indicate that the spectral flatness is less than a first preset flatness threshold is determined as a first type of frame;

[0028] The frame type of an audio frame whose time-domain features indicate that the pitch period is not within a preset period and whose frequency-domain features indicate that the spectral flatness is greater than a second preset flatness threshold and the spectral centroid is greater than a preset centroid threshold is determined as a second type of frame.

[0029] If adjacent audio frames have different frame types, then the frame type of the adjacent audio frames is determined to be a third type of frame.

[0030] Optionally, the encoding of the first type of frames based on a preset speech encoder configured with a preset low-rate speech mode to obtain the corresponding first audio bitstream includes:

[0031] Determine the linear prediction coefficients, pitch period, and quantization residual data of the first type of frame;

[0032] A filter model is constructed based on the linear prediction coefficients, and a voiced excitation signal and / or an unvoiced excitation signal is generated based on the pitch period.

[0033] The voiced excitation signal and / or the unvoiced excitation signal, along with the quantization residual data, are input into the filter model to obtain a time-domain signal corresponding to the first type of frame. The time-domain signal is then subjected to a preset post-filtering process based on a preset speech encoder configured with a preset low-rate speech mode to obtain the corresponding first audio bitstream.

[0034] Optionally, the step of encoding the second type of frames based on the Audio Vivid encoder to obtain the corresponding second audio bitstream includes:

[0035] Based on the Audio Vivid encoder, the Audio Vivid bitstream corresponding to the second type of frame is determined, and the improved discrete cosine transform coefficients are extracted from the Audio Vivid bitstream;

[0036] The improved discrete cosine transform coefficients are inversely quantized, and the resulting processed coefficients are inversely transformed to obtain the second audio bitstream.

[0037] Secondly, this application provides an audio encoding apparatus, comprising:

[0038] The audio frame segmentation module is used to acquire the initial audio signal and perform frame segmentation processing on the initial audio signal to obtain the corresponding audio frames in a target environment where the concurrency meets the preset high concurrency conditions and the network quality meets the preset weak network conditions.

[0039] The type determination module is used to determine the short-time energy, zero-crossing rate, and pitch period of each audio frame to determine the time-domain characteristics of each audio frame, determine the spectral flatness and spectral centroid of each audio frame to determine the frequency-domain characteristics of each audio frame, and determine the type of each audio frame based on the time-domain characteristics and the frequency-domain characteristics.

[0040] The first audio frame encoding module is used to encode the first type of frame based on a preset speech encoder configured with a preset low-rate speech mode if the audio frame is a first type of frame, to obtain the corresponding first audio bitstream; the first type of frame is an audio frame that contains only human language.

[0041] The second audio frame encoding module is used to encode the second type of frame based on the Audio Vivid encoder to obtain the corresponding second audio bitstream if the audio frame is a second type of frame; the second type of frame is an audio frame containing music but not human language.

[0042] The third audio frame encoding module is used to encode the third type of frame using a preset weighted fusion encoding method if the audio frame is a third type of frame, to obtain the corresponding third audio bitstream; the third type of frame is an audio frame that transitions between the first type of frame and the second type of frame.

[0043] An audio determination module is used to determine a target audio signal based on the first audio stream, the second audio stream, and the third audio stream, so as to transmit the target audio signal in the target environment.

[0044] Thirdly, this application provides an electronic device, comprising:

[0045] Memory, used to store computer programs;

[0046] A processor for executing the computer program to implement the aforementioned audio encoding method.

[0047] Fourthly, this application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned audio encoding method.

[0048] In this application, under a target environment where the concurrency meets a preset high-concurrency condition and the network quality meets a preset weak network condition, an initial audio signal is acquired, and the initial audio signal is processed into frames to obtain corresponding audio frames. The short-time energy, zero-crossing rate, and pitch period of each audio frame are determined to determine the time-domain characteristics of each audio frame. The spectral flatness and spectral centroid of each audio frame are determined to determine the frequency-domain characteristics of each audio frame. Based on the time-domain characteristics and the frequency-domain characteristics, the type of each audio frame is determined. If the audio frame is a first-type frame, the first-type frame is encoded based on a preset speech encoder configured with a preset low-rate speech mode to obtain the corresponding first audio code. The audio stream is as follows: the first type of frame is an audio frame containing only human language; if the audio frame is a second type of frame, it is encoded using an AudioVivid encoder to obtain a corresponding second audio stream; the second type of frame is an audio frame containing music but not human language; if the audio frame is a third type of frame, it is encoded using a preset weighted fusion coding method to obtain a corresponding third audio stream; the third type of frame is an audio frame that transitions between the first type of frame and the second type of frame; a target audio signal is determined based on the first audio stream, the second audio stream, and the third audio stream so as to transmit the target audio signal in the target environment. As can be seen above, when the concurrency meets the preset high concurrency conditions and the network quality meets the preset weak network conditions, the initial audio signal is first acquired, and then frame-segmented processing is performed on the initial audio signal to obtain the corresponding audio frames. Next, the time-domain characteristics of each audio frame are determined by determining the short-time energy, zero-crossing rate, and pitch period of each audio frame. At the same time, the spectral flatness and spectral centroid of each audio frame are determined to clarify the frequency-domain characteristics of each audio frame. Then, the type of each audio frame is determined based on the time-domain and frequency-domain characteristics. If an audio frame belongs to the first type of frame, a preset speech encoder configured with a preset low-rate speech mode is used to encode the first type of frame to obtain the corresponding first audio bitstream. If an audio frame belongs to the second type of frame, the audio is encoded using an audio encoder configured with a preset low-rate speech mode. The Vivid encoder encodes the second type of frame to obtain the corresponding second audio bitstream. If an audio frame belongs to the third type of frame, it is encoded using a preset weighted fusion coding method to obtain the corresponding third audio bitstream. Finally, the target audio signal is determined based on the first, second, and third audio bitstreams for transmission in the target environment. In this way, this application can improve the audio encoding quality, thereby enhancing the user experience to some extent. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0050] Figure 1 This is a flowchart of an audio encoding method disclosed in this application;

[0051] Figure 2 This is a flowchart of a specific audio encoding method disclosed in this application;

[0052] Figure 3 This is a flowchart of an audio frame decoding process disclosed in this application;

[0053] Figure 4 This is a schematic diagram of the structure of an audio encoding device disclosed in this application;

[0054] Figure 5 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0055] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0056] Currently, Audio Vivid exhibits significant functional gaps in scenarios such as low-bitrate real-time communication and narrowband voice adaptation. It does not differentiate between audio signal types during encoding, employing a uniform encoding method. While this simplifies the encoder processing flow, it may result in suboptimal encoding efficiency and quality for certain specific voice types. Especially in high-concurrency and weak network environments for real-time voice transmission, such as large-scale online conferences, multi-device collaboration in smart homes, voice transmission in unstable networks like high-speed rail and remote areas, and bandwidth allocation for multi-user real-time interaction in emerging scenarios like VR and metaverse, Audio Vivid struggles to meet the demands. Therefore, the existing Audio Vivid standard is ill-suited for low-bitrate and narrowband voice communication scenarios, failing to provide efficient and high-quality voice encoding support for these scenarios, resulting in significant application limitations. To address this, this application provides an audio encoding method, apparatus, device, and storage medium that can improve audio encoding quality.

[0057] See Figure 1 , Figure 2 as well as Figure 3 As shown, an embodiment of the present invention discloses an audio encoding method, including:

[0058] Step S11: Under the target environment where the concurrency meets the preset high concurrency condition and the network quality meets the preset weak network condition, the initial audio signal is acquired, and the initial audio signal is processed into frames to obtain the corresponding audio frames.

[0059] First of all, it should be noted that Figure 2 and Figure 3 The “speech” frames appearing in the frame are of type 1, which are audio frames containing only human language; the “audio” frames are of type 2, which are audio frames containing music and non-human language; and the “switch frame” frames are of type 3, which are audio frames that transition between type 1 and type 2 frames.

[0060] In this embodiment, under the target environment where the concurrency meets the preset high concurrency conditions and the network quality meets the preset weak network conditions, the initial audio signal is acquired. The initial audio signal may come from locally stored audio files, audio data collected in real time by external devices, etc.

[0061] Subsequently, the initial audio signal is framed based on preset sampling points and a Hanning window function to determine the corresponding audio frames. Before framing, the frame length is set to match the length of the Audio Vivid encoder. A frame length of 1024 sampling points can be set during framing; if the signal sampling rate is 16kHz, then this frame length corresponds to a frame duration of 64ms. The frame shift can be set to 512 sampling points to reduce inter-frame spectral leakage. A Hanning window needs to be applied to the initial audio signal during framing. The formula for the Hanning window is as follows:

[0062] ;

[0063] In the formula, This is the index of the sampling point, ranging from 0 to 1023.

[0064] Furthermore, the function for applying the Hanning window is as follows:

[0065] ;

[0066] In the formula, This is the original input signal, i.e., the initial audio signal; For frame index.

[0067] Step S12: Determine the short-time energy, zero-crossing rate, and pitch period of each audio frame to determine the time-domain characteristics of each audio frame; determine the spectral flatness and spectral centroid of each audio frame to determine the frequency-domain characteristics of each audio frame; and determine the type of each audio frame based on the time-domain characteristics and the frequency-domain characteristics.

[0068] In this embodiment, after obtaining each audio frame, multi-dimensional feature extraction is required to determine the frame type.

[0069] First, considering the time-domain characteristics, the short-time energy, zero-crossing rate, and pitch period of each audio frame reflecting signal strength need to be calculated sequentially. Specifically, for the short-time energy of each audio frame, the square value of each audio frame is determined, and the short-time energy of each audio frame is determined based on the square value. Furthermore, the energy of the speech signal is usually higher than the background noise; in other words, the short-time energy of the speech signal (i.e., the first type of frame) is generally... The formula for calculating short-time energy can be shown below:

[0070] ;

[0071] The zero-crossing rate of each audio frame reflects the high-frequency characteristics of the signal. Based on a preset sampling order, the first audio frame is selected as the first target audio frame, and the next audio frame after the first target audio frame is determined as the second target audio frame. A first function value is determined based on the first target audio frame and a preset sign function, and a second function value is determined based on the second target audio frame and the preset sign function, to determine the difference between the first and second function values. The process jumps to the step of selecting the first audio frame as the first target audio frame based on the preset sampling order, until all audio frames have been selected, and the zero-crossing rate of each audio frame is determined based on the obtained differences and the number of preset sampling points. The formula for calculating the zero-crossing rate is as follows:

[0072] ;

[0073] In the formula, This represents the sign function. Furthermore, regarding the zero-crossing rate, generally, it applies to voiceless speech. Voiced sounds and audio signals .

[0074] For the pitch period, the pitch period of each audio frame is determined based on a preset short-time autocorrelation function and each audio frame. That is, it can be determined using the short-time autocorrelation function. The formula for calculating the autocorrelation function to detect the periodicity of the pitch can be as follows:

[0075] ;

[0076] Furthermore, speech signals typically occur at the pitch period (which can be the interval from the 20th to the 200th sample point). Peak values: The audio signal (i.e., the second type of frame) has no significant peak values.

[0077] Finally, the temporal characteristics of each audio frame are determined based on the short-time energy, the zero-crossing rate, and the pitch period.

[0078] For frequency domain feature extraction, a Fast Fourier Transform (FFT) is first performed on each audio frame to obtain the corresponding spectrum. That is, a 2048-point FFT is performed on the framed signal to obtain the spectrum. Padding to 2048 points improves resolution. The formula for calculating the spectrum is as follows:

[0079] ;

[0080] In the formula, This is the zero-padded frame signal; in the spectrum, the first 1024 points are... The last 1024 points were 0.

[0081] The spectral flatness of the spectrum within a preset frequency range is determined. Spectral flatness can distinguish between harmonics and broadband signals. The formula for calculating spectral flatness is as follows:

[0082] ;

[0083] In the formula, It can cover the main frequency bands. Furthermore, it supports voice signals. audio signal .

[0084] Next, the spectral centroid of the spectrum is determined. The spectral centroid reflects the frequency of energy concentration. The formula for calculating the spectral centroid is shown below:

[0085] ;

[0086] Finally, the frequency domain characteristics of each audio frame are determined based on the spectral flatness and the spectral centroid.

[0087] When classifying frames based on the above characteristics, the frame type of audio frames whose temporal characteristics indicate that the pitch period is within a preset period, whose short-time energy is greater than a preset energy threshold, and whose frequency domain characteristics indicate that the spectral flatness is less than a first preset flatness threshold are determined as the first type of frame. Here, the preset period can be the interval from the 20th sampling point to the 200th sampling point; the first preset flatness threshold can be 0.2; and the preset energy threshold can be... .

[0088] Audio frames whose time-domain characteristics indicate that the pitch period is not within a preset period and whose frequency-domain characteristics indicate that the spectral flatness is greater than a second preset flatness threshold and the spectral centroid is greater than a preset centroid threshold are classified as second-type frames. The second preset flatness threshold can be 0.4; the preset centroid threshold can be 800Hz.

[0089] In each of the aforementioned audio frames, if adjacent audio frames have different frame types, then the frame type of the adjacent audio frames is determined to be a third type of frame. For example, a third type of frame can be a transition frame from a first type of frame to a second type of frame, or a transition frame from a second type of frame to a first type of frame. Furthermore, the transition direction of the third type of frame can be marked with the symbol 'd', where 'd' can be 0 or 1.

[0090] Step S13: If the audio frame is a first type of frame, then the first type of frame is encoded based on a preset speech encoder configured with a preset low-rate speech mode to obtain the corresponding first audio bitstream; the first type of frame is an audio frame that only contains human language.

[0091] In this embodiment, when an audio frame is determined to be a first-type frame, a preset speech encoder configured with a preset low-rate speech mode is used for encoding. The rate of the preset low-rate speech mode can be 4.8 kbps. First, the linear prediction coefficients, pitch period, and quantization residual data of the first-type frame are determined to complete the decoding operation of the first-type frame. A filter model is constructed based on the linear prediction coefficients, and a voiced excitation signal such as a pulse sequence and / or a deaf excitation signal such as white noise is generated based on the pitch period.

[0092] The voiced excitation signal and / or the unvoiced excitation signal, along with the quantization residual data, are input into the filter model to obtain a time-domain signal corresponding to the first type of frame. Then, based on a preset speech encoder configured with a preset low-bitrate speech mode, the time-domain signal undergoes a preset post-filtering process, such as formant enhancement, to obtain the corresponding first audio bitstream. The preset low-bitrate speech mode effectively reduces the bitstream rate while ensuring speech intelligibility, making it suitable for transmission requirements in the target environment.

[0093] Step S14: If the audio frame is a second type of frame, then the second type of frame is encoded based on the Audio Vivid encoder to obtain the corresponding second audio bitstream; the second type of frame is an audio frame containing music and non-human language.

[0094] In this embodiment, when an audio frame is determined to be a second-type frame, an Audio Vivid encoder is used for encoding. Specifically, the Audio Vivid bitstream corresponding to the second-type frame is determined based on the Audio Vivid encoder, and Modified Discrete Cosine Transform (MDCT) coefficients are extracted from the Audio Vivid bitstream to complete the decoding operation of the second-type frame. The MDCT coefficients are then inversely quantized, and the resulting processed coefficients are inversely transformed to obtain the second audio bitstream. Furthermore, the information extracted from the Audio Vivid bitstream may also include spatial metadata such as the three-dimensional coordinates of the sound object, motion trajectory parameters, and environmental acoustic characteristics, as well as encoding control information. This extracted information can contribute to the generation of the second audio bitstream. This process fully utilizes the advantages of the Audio Vivid encoder in encoding non-speech audio such as music, ensuring that the audio quality of the second-type frame is well preserved during the encoding process.

[0095] Step S15: If the audio frame is a third type of frame, then the third type of frame is encoded using a preset weighted fusion coding method to obtain the corresponding third audio bitstream; the third type of frame is an audio frame that transitions between the first type of frame and the second type of frame.

[0096] In this embodiment, a preset weighted fusion coding method is used to process the third type of frames. This method combines the coding logic of the first and second types of frames, and achieves smooth coding of transition frames by dynamically allocating weights. Specifically, the weight ratio of the preset speech encoder and the Audio Vivid encoder, i.e., the transition weight, is determined, and the output results of the two encoders are weighted and fused. The formula for fusing the output results is as follows: In the formula, The output is a voice encoding (i.e., the first audio stream); Encode the output for Audio Vivid (second audio stream). Additionally, transition weights... The calculation formula is as follows:

[0097] ;

[0098] In the formula, t is the index of the intra-frame sampling point, and Speech is classified as the first type of frame, and audio is classified as the second type of frame.

[0099] Furthermore, decoding is required before encoding the third type of frame; the corresponding decoding fusion formula can be: In the formula, This represents the information obtained by decoding the first type of frame; This represents the information obtained by decoding the second type of frame.

[0100] Step S16: Determine the target audio signal based on the first audio stream, the second audio stream, and the third audio stream, so as to transmit the target audio signal in the target environment.

[0101] In this embodiment, the first audio stream, the second audio stream, and the third audio stream are spliced ​​and integrated to form a target audio signal, which can be stably transmitted in an environment that meets the requirements of high concurrency and weak network. Figure 2 The "bit multiplexer" in this context refers to the integration of encoded data from different sources or through different processing paths into a unified encoded bitstream output, which is the target audio signal.

[0102] As can be seen above, when the concurrency meets the preset high concurrency conditions and the network quality meets the preset weak network conditions, the initial audio signal is first acquired, and then frame-segmented processing is performed on the initial audio signal to obtain the corresponding audio frames. Next, the time-domain characteristics of each audio frame are determined by determining the short-time energy, zero-crossing rate, and pitch period of each audio frame. At the same time, the spectral flatness and spectral centroid of each audio frame are determined to clarify the frequency-domain characteristics of each audio frame. Then, the type of each audio frame is determined based on the time-domain and frequency-domain characteristics. If an audio frame belongs to the first type of frame, a preset speech encoder configured with a preset low-rate speech mode is used to encode the first type of frame to obtain the corresponding first audio bitstream. If an audio frame belongs to the second type of frame, the audio is encoded using an audio encoder configured with a preset low-rate speech mode. The Vivid encoder encodes the second type of frame to obtain the corresponding second audio bitstream. If an audio frame belongs to the third type of frame, it is encoded using a preset weighted fusion coding method to obtain the corresponding third audio bitstream. Finally, the target audio signal is determined based on the first, second, and third audio bitstreams for transmission in the target environment. In this way, this application can improve the audio encoding quality, thereby enhancing the user experience to some extent.

[0103] Accordingly, see Figure 4 As shown, this application provides an audio encoding apparatus, including:

[0104] The audio framing module 11 is used to acquire an initial audio signal and perform framing processing on the initial audio signal to obtain corresponding audio frames in a target environment where the concurrency meets a preset high concurrency condition and the network quality meets a preset weak network condition.

[0105] The type determination module 12 is used to determine the short-time energy, zero-crossing rate and pitch period of each audio frame to determine the time-domain characteristics of each audio frame, determine the spectral flatness and spectral centroid of each audio frame to determine the frequency-domain characteristics of each audio frame, and determine the type of each audio frame based on the time-domain characteristics and the frequency-domain characteristics.

[0106] The first audio frame encoding module 13 is used to encode the first type of frame based on a preset speech encoder configured with a preset low-rate speech mode if the audio frame is a first type of frame, to obtain the corresponding first audio bitstream; the first type of frame is an audio frame that only contains human language.

[0107] The second audio frame encoding module 14 is used to encode the second type of frame based on the Audio Vivid encoder to obtain the corresponding second audio bitstream if the audio frame is a second type of frame; the second type of frame is an audio frame containing music but not human language.

[0108] The third audio frame encoding module 15 is used to encode the third type of frame using a preset weighted fusion encoding method if the audio frame is a third type of frame, to obtain the corresponding third audio bitstream; the third type of frame is an audio frame that transitions between the first type of frame and the second type of frame.

[0109] The audio determination module 16 is used to determine a target audio signal based on the first audio bitstream, the second audio bitstream, and the third audio bitstream, so as to transmit the target audio signal in the target environment.

[0110] In some specific embodiments, the audio framing module 11 specifically includes:

[0111] The audio framing unit is used to perform framing processing on the initial audio signal based on preset sampling points and Hanning window function to determine the corresponding audio frames.

[0112] In some specific embodiments, the type determination module 12 specifically includes:

[0113] An energy determination unit is used to determine the square value of each audio frame and to determine the short-time energy of each audio frame based on the square value.

[0114] An audio frame determination unit is used to select the first audio frame as the first target audio frame from among the audio frames based on a preset sampling order, and to determine the next audio frame of the first target audio frame as the second target audio frame based on the preset sampling order.

[0115] The difference determination unit is used to determine a first function value based on the first target audio frame and a preset symbolic function, and to determine a second function value based on the second target audio frame and the preset symbolic function, so as to determine the difference between the first function value and the second function value;

[0116] The zero-crossing rate determination unit is used to jump to the step of selecting the first audio frame as the first target audio frame in each audio frame based on the preset sampling order, until all audio frames have been selected, so as to determine the zero-crossing rate of each audio frame based on the obtained differences and the number of preset sampling points.

[0117] The period determination unit is used to determine the pitch period of each audio frame based on a preset short-time autocorrelation function and each audio frame;

[0118] The temporal feature determination unit is used to determine the temporal features of each audio frame based on the short-time energy, the zero-crossing rate, and the pitch period.

[0119] In some specific embodiments, the type determination module 12 specifically includes:

[0120] The spectrum determination unit is used to perform a fast Fourier transform on each of the audio frames to obtain the corresponding spectrum;

[0121] A centroid determination unit is used to determine the spectral flatness of the spectrum within a preset frequency range and to determine the spectral centroid of the spectrum.

[0122] A frequency domain feature determination unit is used to determine the frequency domain features of each audio frame based on the spectral flatness and the spectral centroid.

[0123] In some specific embodiments, the type determination module 12 specifically includes:

[0124] The first type determination unit is used to determine the frame type of an audio frame whose time-domain features indicate that the pitch period is within a preset period, the short-time energy is greater than a preset energy threshold, and the frequency-domain features indicate that the spectral flatness is less than a first preset flatness threshold as a first type frame.

[0125] The second type determination unit is used to determine the frame type of an audio frame that is a second type of frame, where the time domain features indicate that the pitch period is not located within a preset period and the frequency domain features indicate that the spectral flatness is greater than a second preset flatness threshold and the spectral centroid is greater than a preset centroid threshold.

[0126] The third type determination unit is used to determine the frame type of adjacent audio frames as a third type frame if the frame types of adjacent audio frames are different.

[0127] In some specific embodiments, the first audio frame encoding module 13 specifically includes:

[0128] The data determination unit is used to determine the linear prediction coefficients, pitch period, and quantization residual data of the first type of frame;

[0129] The signal generation unit is used to construct a filter model based on the linear prediction coefficients and generate a voiced excitation signal and / or an unvoiced excitation signal based on the pitch period.

[0130] The first bitstream determination unit is used to input the voiced excitation signal and / or the unvoiced excitation signal and the quantization residual data into the filter model to obtain a time-domain signal corresponding to the first type of frame, and to perform a preset post-filtering process on the time-domain signal based on a preset speech encoder configured as a preset low-rate speech mode to obtain the corresponding first audio bitstream.

[0131] In some specific embodiments, the second audio frame encoding module 14 specifically includes:

[0132] The coefficient extraction unit is used to determine the AudioVivid bitstream corresponding to the second type of frame based on the AudioVivid encoder, and extract the improved discrete cosine transform coefficients from the AudioVivid bitstream.

[0133] The second bitstream determination unit is used to perform inverse quantization on the improved discrete cosine transform coefficients and perform inverse transform on the obtained processed coefficients to obtain the second audio bitstream.

[0134] Furthermore, embodiments of this application also disclose an electronic device, Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the audio encoding method disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0135] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0136] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0137] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the audio encoding method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.

[0138] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned disclosed audio encoding method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0139] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0140] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0141] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0142] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0143] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. An audio encoding method, characterized by, The method comprises the steps of: In a target environment where the concurrency meets a preset high concurrency condition and the network quality meets a preset weak network condition, an initial audio signal is acquired, and frame processing is performed on the initial audio signal to obtain corresponding audio frames; The short-time energy, zero-crossing rate and pitch period of each audio frame are determined to determine the time domain features of each audio frame, the spectral flatness and spectral centroid of each audio frame are determined to determine the frequency domain features of each audio frame, and the type of each audio frame is determined based on the time domain features and the frequency domain features; If the audio frame is a first type frame, the first type frame is encoded based on a preset speech encoder configured in a preset low-rate speech mode to obtain a corresponding first audio bitstream; The first type frame is an audio frame containing only human language, and the first type frame is an audio frame whose time domain features indicate that the pitch period is within a preset period, the short-time energy is greater than a preset energy threshold, the frequency domain features indicate that the spectral flatness is less than a first preset flatness threshold; If the audio frame is a second type frame, the second type frame is encoded based on an Audio Vivid encoder to obtain a corresponding second audio bitstream; the second type frame is an audio frame containing music and non-human language, and the second type frame is an audio frame whose time domain features indicate that the pitch period is not within the preset period and whose frequency domain features indicate that the spectral flatness is greater than a second preset flatness threshold and the spectral centroid is greater than a preset centroid threshold; If the audio frame is a third type frame, the third type frame is encoded using a preset weighted fusion encoding mode to obtain a corresponding third audio bitstream; the third type frame is an audio frame that transitions between the first type frame and the second type frame, and the third type frame is an adjacent audio frame with different frame types among the audio frames; A target audio signal is determined based on the first audio bitstream, the second audio bitstream and the third audio bitstream, so as to transmit the target audio signal in the target environment.

2. The audio encoding method of claim 1, wherein, The frame processing on the initial audio signal to obtain corresponding audio frames comprises: The initial audio signal is frame-processed based on a preset sampling point and a Hanning window function to determine corresponding audio frames.

3. The audio encoding method of claim 2, wherein, The determination of the short-time energy, zero-crossing rate and pitch period of each audio frame to determine the time domain features of each audio frame comprises: The square value of each audio frame is determined, and the short-time energy of each audio frame is determined based on the square value; A first target audio frame is selected as the first target audio frame in each audio frame based on a preset sampling sequence, and the next audio frame of the first target audio frame is determined as a second target audio frame based on the preset sampling sequence; A first function value is determined based on the first target audio frame and a preset sign function, and a second function value is determined based on the second target audio frame and the preset sign function, to determine the difference between the first function value and the second function value; jumping to the step of selecting a first audio frame as a first target audio frame in each of the audio frames based on a preset sampling sequence until each of the audio frames is selected, to determine a zero-crossing rate of each of the audio frames based on the obtained difference values and the number of preset sampling points; determining a pitch period of each of the audio frames based on a preset short-term autocorrelation function and each of the audio frames; determining a time-domain feature of each of the audio frames based on the short-term energy, the zero-crossing rate and the pitch period.

4. The audio encoding method of claim 1, wherein, The determination of the spectral flatness and the spectral centroid of each of the audio frames to determine a frequency-domain feature of each of the audio frames comprises: performing fast Fourier transform on each of the audio frames to obtain a corresponding spectrum; determining the spectral flatness of the spectrum within a preset frequency range and determining the spectral centroid of the spectrum; determining the frequency-domain feature of each of the audio frames based on the spectral flatness and the spectral centroid.

5. The audio encoding method of claim 1, wherein, The encoding of the first type of frame based on a preset speech encoder configured in a preset low-rate speech mode to obtain a corresponding first audio bitstream comprises: determining linear prediction coefficients, a pitch period and quantized residual data of the first type of frame; constructing a filter model based on the linear prediction coefficients and generating a voiced excitation signal and / or an unvoiced excitation signal based on the pitch period; inputting the voiced excitation signal and / or the unvoiced excitation signal and the quantized residual data into the filter model to obtain a time-domain signal corresponding to the first type of frame, and performing a preset post-filtering process on the time-domain signal based on a preset speech encoder configured in a preset low-rate speech mode to obtain a corresponding first audio bitstream.

6. The audio encoding method of any of claims 1 to 5, wherein, The encoding of the second type of frame based on an Audio Vivid encoder to obtain a corresponding second audio bitstream comprises: determining an Audio Vivid bitstream corresponding to the second type of frame based on an Audio Vivid encoder and extracting improved discrete cosine transform coefficients from the Audio Vivid bitstream; performing inverse quantization processing on the improved discrete cosine transform coefficients and performing inverse transform on the obtained processed coefficients to obtain a second audio bitstream.

7. An audio encoding apparatus characterized by comprising: It comprises: an audio framing module configured to obtain an initial audio signal and perform framing processing on the initial audio signal to obtain corresponding audio frames in a target environment where the concurrency degree meets a preset high-concurrency condition and the network quality meets a preset weak-network condition; a type determination module configured to determine a short-term energy, a zero-crossing rate and a pitch period of each of the audio frames to determine a time-domain feature of each of the audio frames, determine a spectral flatness and a spectral centroid of each of the audio frames to determine a frequency-domain feature of each of the audio frames, and determine a type of each of the audio frames based on the time-domain feature and the frequency-domain feature; a first audio frame encoding module configured to encode a first type of frame based on a preset speech encoder configured in a preset low-rate speech mode to obtain a corresponding first audio bitstream if the audio frame is a first type of frame. The first type of frame is an audio frame containing only human language, and the first type of frame is an audio frame whose time domain feature indicates that a pitch period is located in a preset period, a short-time energy is greater than a preset energy threshold, and a frequency domain feature indicates that a spectral flatness is less than a first preset flatness threshold; The second audio frame encoding module is configured to, if the audio frame is a second type of frame, encode the second type of frame based on an Audio Vivid encoder to obtain a corresponding second audio bitstream; the second type of frame is an audio frame containing music and non-human language, and the second type of frame is an audio frame whose time domain feature indicates that a pitch period is not located in the preset period and a frequency domain feature indicates that a spectral flatness is greater than a second preset flatness threshold and a spectral centroid is greater than a preset centroid threshold; The third audio frame encoding module is configured to, if the audio frame is a third type of frame, encode the third type of frame using a preset weighted fusion encoding mode to obtain a corresponding third audio bitstream; the third type of frame is an audio frame that is between the first type of frame and the second type of frame, and the third type of frame is adjacent audio frames of different frame types in each of the audio frames; The audio determination module is configured to determine a target audio signal based on the first audio bitstream, the second audio bitstream, and the third audio bitstream, so as to transmit the target audio signal in the target environment.

8. An electronic device, comprising: Comprise: A memory for saving a computer program; A processor for executing the computer program to implement the audio encoding method of any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, For saving a computer program; wherein the computer program is executed by the processor to implement the audio encoding method of any one of claims 1 to 6. For saving a computer program; wherein the computer program is executed by the processor to implement the audio encoding method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Encoder selection

    CN107408383A

  • Audio coding method and device, computer readable medium and electronic equipment

    CN118038882A