Multi-channel audio signal hybrid coding method and device, and storage medium
By employing dynamic mixing and compression processing, and using techniques such as anti-aliasing filtering, loudness-weighted averaging, frequency domain mixing, and auditory masking curve compression, the problems of signal distortion and low efficiency in the mixing of multiple audio signals are solved, achieving efficient and high-quality multi-channel audio signal encoding.
Patent Information
- Application Number
- CN202511343287.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2026-01-09
AI Technical Summary
In existing technologies, mixing multiple audio signals can easily lead to signal distortion and insufficient dynamic range. Traditional coding methods are inefficient and cannot effectively utilize the characteristics of the mixed signal.
By employing dynamic mixing and compression processing, and using methods such as anti-aliasing filtering, loudness-weighted averaging, frequency domain mixing, auditory masking curve compression, and lossless coding, combined with DCT transform and encapsulation technology, efficient mixing and coding of multiple audio signals can be achieved.
It effectively solves the signal distortion problem, avoids level overload, improves coding efficiency, adapts to the characteristics of human hearing, ensures signal quality and dynamic range, and is suitable for real-time processing of multiple audio signals.
Smart Images

Figure CN121306149A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio signal processing technology, and in particular to a method, device and storage medium for multi-channel audio signal mixing and encoding. Background Technology
[0002] In the field of audio processing technology, with the continuous development of multimedia applications, the processing and transmission of audio signals have become increasingly important. From the early simple mono audio to the widely used stereo, surround sound and other multi-channel audio, the progress of audio technology has greatly improved the user's listening experience. In various audio scenarios, such as conference systems, radio stations, and audio production, it is often necessary to process and transmit multiple audio signals.
[0003] In existing technologies, multi-channel audio mixing typically employs a simple linear superposition method, which easily leads to signal distortion and insufficient dynamic range. At the same time, traditional audio coding methods are inefficient after multi-channel signal mixing and cannot effectively utilize the characteristics of the mixed signal. Summary of the Invention
[0004] (a) Purpose of the invention
[0005] To address the technical problems existing in the background art, this invention proposes a multi-channel audio signal mixing encoding method, device, and storage medium. Through dynamic mixing and compression processing, it effectively solves the signal distortion problem caused by the traditional linear superposition method. Furthermore, dynamic mixing ensures the reasonable superposition of each signal and avoids level overload caused by simple addition.
[0006] (II) Technical Solution
[0007] This invention provides a method for mixing and encoding multiple audio signals, characterized by the following steps:
[0008] Digital audio signals are acquired through multiple audio input channels. The acquired digital audio signals are subjected to anti-aliasing filtering to obtain an audio signal in 48kHz / 16bit format. The 48kHz / 16bit audio signal is processed using a loudness-weighted averaging method. The loudness-weighted averaged audio signal is then processed using frequency domain mixing technology to obtain a mixed signal. The mixed signal is dynamically compressed based on a calculated auditory masking curve to obtain a compressed signal. The compressed signal is then mixed and encoded to obtain a compressed data stream. Finally, the compressed data stream is encapsulated and output.
[0009] Preferably, the threshold calculation model for the masking curve is as follows: ,in This represents the final masking threshold for the Zth critical frequency band, while The absolute hearing threshold under the installation environment. The energy of the Z-band is calculated using the following formula: ,in and These are the upper and lower limits of the critical frequency band. This is the signal power spectral density, and the threshold calculation model for the masking curve is... This is the frequency band offset, which is 15dB for low frequencies and 10dB for high frequencies. This is the masking and expansion effect.
[0010] Preferably, the specific steps for mixing and encoding the compressed signal to obtain a compressed data stream are as follows: the compressed signal is converted from a spatial domain signal to a frequency domain representation using DCT transform technology; the frequency domain coefficients of the compressed signal in the frequency domain representation are quantized to remove details that are not sensitive to the human eye / ear; and the quantized compressed signal is subjected to lossless compression encoding to obtain a compressed data stream.
[0011] Preferably, the specific steps for encapsulating and outputting the compressed data stream are as follows: adding necessary header information, metadata, and synchronization tags to the compressed data stream according to the encapsulation format.
[0012] Preferably, the metadata includes sampling rate and number of channels information.
[0013] Preferably, the acquired multiple audio signals are subjected to echo cancellation and dynamic mixing.
[0014] Preferably, the specific steps for echo cancellation of the acquired audio signals include: performing adaptive filter estimation on the acquired audio signals and subtracting the echo path signal, and updating the filter coefficients using the gradient descent method.
[0015] A multi-channel audio signal mixing and encoding device includes a memory and at least one processor, wherein the memory stores computer-readable instructions; the at least one processor invokes the computer-readable instructions in the memory to execute various steps of the multi-channel audio signal mixing and encoding method.
[0016] A computer-readable storage medium storing computer-readable instructions, which, when executed by a processor, implement various steps of a multi-channel audio signal mixing encoding method.
[0017] Compared with the prior art, the above-mentioned technical solution of the present invention has the following beneficial technical effects:
[0018] In this invention, the technical solution effectively solves the signal distortion problem caused by the traditional linear superposition method through dynamic mixing and compression processing. Furthermore, dynamic mixing ensures the reasonable superposition of each signal and avoids level overload caused by simple addition. Attached Figure Description
[0019] Figure 1 This is a flowchart of a multi-channel audio signal mixing and encoding method proposed in this invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments and the accompanying drawings. It should be understood that these descriptions are merely exemplary and not intended to limit the scope of the invention. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.
[0021] In the description of the invention, it should be noted that the terms "upper," "lower," "inner," "outer," "front end," "rear end," "both ends," "one end," and "the other end," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0022] In the description of the invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installed," "equipped with," and "connected," etc., should be interpreted broadly. For example, "connected" can be a fixed connection, such as welding, riveting, or bonding; it can also be a detachable connection, such as threaded connection, keyed connection, or pin connection; or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; or it can be a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0023] Example 1:
[0024] like Figure 1 As shown, the present invention proposes a multi-channel audio signal mixing and encoding method, which specifically includes the following steps: acquiring multiple audio signals, dynamically mixing the acquired multiple audio signals to obtain a mixed signal, dynamically compressing the obtained mixed signal to obtain a compressed signal, mixing and encoding the compressed signal to obtain a compressed data stream, and encapsulating and outputting the compressed data stream.
[0025] This technical solution effectively solves the signal distortion problem caused by traditional linear superposition methods through dynamic mixing and compression. Dynamic mixing ensures reasonable superposition of signals from various channels, avoiding level overload caused by simple addition. Dynamic compression optimizes the dynamic range of the signal, ensuring that the final output signal maintains an appropriate loudness level. Compared with existing technologies, this method improves coding efficiency while maintaining signal quality, making it particularly suitable for real-time processing of multiple audio signals. Through a phased signal processing flow, it achieves a complete solution from signal acquisition to encapsulated output.
[0026] Furthermore, the threshold calculation model for the masking curve is as follows: ,in This represents the final masking threshold for the Zth critical frequency band, while The absolute hearing threshold under the installation environment. The energy of the Z-band is calculated using the following formula: ,in and These are the upper and lower limits of the critical frequency band. This is the signal power spectral density, and the threshold calculation model for the masking curve is... This is the frequency band offset, which is 15dB for low frequencies and 10dB for high frequencies. This is the masking extension effect, and the formula for calculating the masking extension effect is: , For nearby shielding source energy, and The attenuation coefficient is... This represents the number of effective masking sources, with specific values based on the ISO 389-7 standard. This allows for the effective calculation of the hybrid type, thus better adapting to the characteristics of human hearing.
[0027] Furthermore, this application also proposes specific steps for acquiring multiple audio signals, including: acquiring digital audio signals through multiple audio input channels, and performing anti-aliasing filtering on the acquired digital audio signals to obtain audio signals in 48kHz / 16bit format.
[0028] Therefore, this technical solution effectively avoids aliasing and distortion problems in multiple audio signals after acquisition through precise sampling and filtering. Furthermore, the standardized 48kHz / 16bit format provides a unified input reference for subsequent signal processing, resolving quality issues caused by inconsistent sampling of multiple signals.
[0029] Specifically, the audio signal acquisition process includes acquiring digital audio signals through multiple audio input channels, and performing anti-aliasing filtering on the acquired digital audio signals to obtain a 48kHz / 16bit audio signal. Audio input channels can be microphone arrays, sound card interfaces, etc. Different audio input channels can adapt to different audio acquisition scenarios. For example, microphone arrays are suitable for audio acquisition in conference scenarios, while sound card interfaces are suitable for audio acquisition on computers. The anti-aliasing filter can be a low-pass filter, whose function is to prevent high-frequency signals from aliasing into low-frequency signals, thus ensuring the quality of the audio signal.
[0030] Furthermore, this application proposes to process a 48kHz / 16bit audio signal using a loudness-weighted averaging method, and to process the loudness-weighted averaged audio signal using frequency domain mixing technology to obtain a mixed signal.
[0031] Specifically, the loudness-weighted average method is a signal processing method based on the characteristics of human hearing. Signals in different frequency bands are weighted according to their perceived loudness. A-weighted or C-weighted curves can be used to preprocess the signals, and low-frequency and high-frequency components are appropriately attenuated according to equal loudness curves.
[0032] Therefore, this technical solution effectively solves the signal distortion problem in the multi-channel audio mixing process by combining weighted processing based on auditory perception characteristics and frequency domain mixing technology. Compared with the simple linear superposition method, this method can better maintain the dynamic range of the signal, while improving mixing efficiency through frequency domain processing. Specifically, loudness weighting avoids excessive amplification of signals in frequency bands that are not sensitive to the human ear, while frequency domain mixing enables more precise control of signal components, thereby obtaining a higher quality mixed signal.
[0033] Furthermore, this application proposes the following specific steps for dynamically compressing the obtained mixed signal: calculating the auditory masking curve based on the ISO 389-7 standard to perform dynamic range compression on the mixed signal to obtain a compressed signal.
[0034] This technical solution achieves efficient dynamic compression of mixed signals by utilizing the auditory masking effect. Dynamic compression based on the auditory masking curve can better adapt to the characteristics of human hearing and improve compression efficiency while ensuring sound quality.
[0035] Furthermore, this application proposes specific steps for hybrid encoding of compressed signals to obtain compressed data streams. Specifically, the compressed signal is first converted from a spatial domain signal to a frequency domain representation using DCT transform technology. DCT transform can effectively concentrate signal energy, facilitating subsequent processing. The frequency domain coefficients of the compressed signal in the frequency domain representation are then quantized. During the quantization process, details that are not sensitive to the human ear are removed according to the psychoacoustic model, thereby reducing the amount of data. Finally, the quantized compressed signal is subjected to lossless compression encoding using Huffman coding or arithmetic coding to obtain the final compressed data stream.
[0036] As a preferred implementation, the DCT transform can be implemented using an improved fast algorithm to improve computational efficiency. The quantization process can be based on the critical frequency band division, and different quantization step sizes can be used for the coefficients of different frequency bands.
[0037] This technical solution effectively solves the problem of low coding efficiency after mixing multiple audio signals by using frequency domain transformation and selective quantization. This solution can make full use of the frequency domain characteristics of the mixed signal and significantly improve the compression efficiency while ensuring sound quality. The combination of frequency domain transformation and quantization allows the coding process to be optimized for human hearing characteristics, avoiding the coding redundancy problem in traditional methods.
[0038] Furthermore, this application also proposes a specific implementation method for encapsulating and outputting compressed data streams. During the encapsulation and output process, necessary header information, metadata, and synchronization markers are added to the compressed data stream according to a preset encapsulation format. The header information is used to identify the starting position and basic attributes of the data stream, the metadata includes audio parameter information such as sampling rate and number of channels, and the synchronization markers are used to realize the synchronization and error recovery of the data stream.
[0039] After adding header information, metadata, and synchronization flags, it can be: 0x47 0x1F 0xFF 0x10 0x00 0xB00x0D 0x00 0x01 0xC1 0x00 0x00 0x00 0x01 0xF0 0x01 0x02 0xB0 0x17 0x00 0x01 0xC1 0x00 0x00 0x00 0x00 0x01 0xE0 0x00 0x00 0x80 0x80 0x00 0x00 0x00 0x01 0x67 / / SPS 0x52 0x01 0x48 0x44 0x54 0x56 / / The "HDTV" identifier, as a preferred implementation, allows for the selection of common audio container formats such as WAV or MP4. The WAV format uses a RIFF block structure, while the MP4 format uses a box-based hierarchical structure. The header information can be...
[0040] Therefore, this technical solution solves the problems of integrity and identifiability of multi-channel mixed audio data during transmission and storage through standardized encapsulation. The structured encapsulation method ensures compatibility of compressed data streams across different systems and platforms, while the addition of synchronization markers improves the reliability of data transmission.
[0041] Furthermore, this application also proposes a scheme for echo cancellation and dynamic mixing of multiple acquired audio signals.
[0042] Specifically, the echo cancellation step includes: estimating the echo path signal using an adaptive filter and subtracting the echo path signal from the original audio signal, wherein the filter coefficients are dynamically updated using the gradient descent method. The dynamic mixing step processes the 48kHz / 16bit format audio signal using a loudness-weighted averaging method and combines it with frequency domain mixing techniques to generate a mixed signal.
[0043] This application also proposes specific steps for echo cancellation of multiple acquired audio signals, including: performing adaptive filter estimation on the obtained audio signals and subtracting the echo path signal, and updating the filter coefficients using the gradient descent method.
[0044] As a preferred implementation, echo cancellation can be achieved through adaptive filtering using the LMS (Least Mean Square) algorithm, where the step size parameter is dynamically adjusted according to the signal characteristics to balance the convergence speed and steady-state error.
[0045] Therefore, this scheme effectively suppresses acoustic feedback interference in multi-channel audio mixing through pre-echo cancellation. The dynamic update of the adaptive filter ensures the tracking capability of echo path changes. On this basis, dynamic mixing optimizes the spectral distribution of multi-channel signals through perceptual weighting, thereby improving the subsequent coding efficiency during the mixing stage. This technical scheme significantly reduces the nonlinear distortion of the mixed signal while retaining a more complete dynamic range characteristic.
[0046] Example 2:
[0047] A multi-channel audio signal mixing and encoding device includes a memory and at least one processor. The memory stores computer-readable instructions. The at least one processor invokes the computer-readable instructions in the memory to execute various steps of the multi-channel audio signal mixing and encoding method.
[0048] Example 3:
[0049] A computer-readable storage medium storing computer-readable instructions, which, when executed by a processor, implement the various steps of a multi-channel audio signal mixing encoding method.
[0050] It should be understood that the specific embodiments described above are merely illustrative or explanatory of the principles of the invention and do not constitute a limitation thereof. Therefore, any modifications, equivalent substitutions, improvements, etc., made without departing from the spirit and scope of the invention should be included within the protection scope of the invention. Furthermore, the appended claims are intended to cover all variations and modifications falling within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.
Claims
1. A method for mixing and encoding multi-channel audio signals, characterized in that, Specifically, the following steps are included: Digital audio signals are acquired through multiple audio input channels. The acquired digital audio signals are subjected to anti-aliasing filtering to obtain an audio signal in 48kHz / 16bit format. The 48kHz / 16bit audio signal is processed using a loudness-weighted averaging method. The loudness-weighted averaged audio signal is then processed using frequency domain mixing technology to obtain a mixed signal. The mixed signal is dynamically compressed based on a calculated auditory masking curve to obtain a compressed signal. The compressed signal is then mixed and encoded to obtain a compressed data stream. Finally, the compressed data stream is encapsulated and output.
2. The multi-channel audio signal mixing and encoding method according to claim 1, characterized in that, The threshold calculation model for the masking curve is as follows: ,in This represents the final masking threshold for the Zth critical frequency band, while The absolute hearing threshold under the installation environment. The energy of the Z-band is calculated using the following formula: ,in and These are the upper and lower limits of the critical frequency band. This is the signal power spectral density, and the threshold calculation model for the masking curve is... This is the frequency band offset, which is 15dB for low frequencies and 10dB for high frequencies. This is the masking and expansion effect.
3. The multi-channel audio signal mixing and encoding method according to claim 1, characterized in that, The specific steps for mixing and encoding the compressed signal to obtain a compressed data stream are as follows: the compressed signal is converted from a spatial domain signal to a frequency domain representation using DCT transform technology; the frequency domain coefficients of the compressed signal in the frequency domain representation are quantized to remove details that are not sensitive to the human eye / ear; and the quantized compressed signal is subjected to lossless compression encoding to obtain a compressed data stream.
4. The multi-channel audio signal mixing and encoding method according to claim 3, characterized in that, The specific steps for encapsulating and outputting the compressed data stream are as follows: add necessary header information, metadata, and synchronization tags to the compressed data stream according to the encapsulation format.
5. The multi-channel audio signal mixing and encoding method according to claim 4, characterized in that, The metadata includes sampling rate and number of channels.
6. The multi-channel audio signal mixing and encoding method according to claim 1, characterized in that, The acquired audio signals are subjected to echo cancellation and dynamic mixing.
7. The multi-channel audio signal mixing and encoding method according to claim 6, characterized in that, The specific steps for echo cancellation of the acquired audio signals include: performing adaptive filter estimation on the acquired audio signals and subtracting the echo path signal, and updating the filter coefficients using the gradient descent method.
8. A multi-channel audio signal mixing and encoding device, characterized in that, The method includes a memory and at least one processor, wherein the memory stores computer-readable instructions; the at least one processor invokes the computer-readable instructions in the memory to perform various steps of the multi-channel audio signal mixing encoding method as described in any one of claims 1-7.
9. A computer-readable storage medium storing computer-readable instructions thereon, characterized in that, When the computer-readable instructions are executed by a processor, they implement the steps of the multi-channel audio signal mixing encoding method as described in any one of claims 1-7.
Citation Information
Cited By
Multi-audio mixed playing method for white noise equipment
CN122224196A