Multi-channel audio synchronous processing method, system, equipment and medium
By combining an audio phase-locked loop and a hardware timer, a multi-channel audio processing method was developed to solve the problems of phase inconsistency and delay asynchrony in multi-channel audio synchronization devices, achieving high-precision audio synchronization and improving the output audio quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN MINRRAY IND CORP LTD
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-15
AI Technical Summary
Existing multi-channel audio synchronization devices suffer from phase inconsistency and delay asynchrony during the mixing process, resulting in output audio distortion and hollow sound, which cannot meet the needs of high-quality multi-channel audio pickup devices.
A multi-channel audio processing method is adopted, which adjusts the main control clock of the receiving module through an audio phase-locked loop, uses a hardware timer to timestamp, and combines a circular buffer and a dynamic resampling factor to perform Lagrange interpolation and frequency domain transformation to ensure that the sampling rates of each signal are aligned synchronously, and corrects the residual delay through sliding cross-correlation calculation.
It achieves precise alignment of audio frames in each channel, eliminates frequency and delay deviations, improves the quality and synchronization of output audio, ensures high precision in subsequent processing stages, and eliminates minute phase and delay deviations.
Smart Images

Figure CN122053019A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of sound pickup technology, and in particular to a multi-channel audio synchronization processing method, system, device, and medium. Background Technology
[0002] With the rapid rise of content creation formats such as short videos, live streaming, and vlogs, the market demand for high-quality, multi-channel synchronous audio pickup is becoming increasingly prominent. Diverse scenarios such as multi-person interviews, group meetings, and live performances all require stable and synchronous multi-channel audio acquisition to ensure sound pickup quality.
[0003] Existing similar products have significant shortcomings in multi-channel audio synchronization performance. In particular, during multi-channel audio mixing, problems such as phase inconsistency and delay asynchrony are prone to occur, resulting in distortion, hollow sound, and frequency cancellation in the output audio. This seriously affects the sound pickup quality and the professionalism of content creation, and fails to meet the market's core demand for high-quality multi-channel sound pickup equipment. Summary of the Invention
[0004] This invention provides a multi-channel audio synchronization processing method, system, device, and medium to solve the problem of poor audio quality output by multi-channel audio pickup devices.
[0005] This invention discloses a multi-channel audio processing method applied to a multi-channel sound pickup system, the multi-channel sound pickup system comprising:
[0006] There are N receiving modules, where N ≥ 2; There are 2N transmitting modules, and each of the receiving modules is connected to 2 transmitting modules; The main control module is connected to all the receiving modules. The main control module has N independent frequency division isolation channels, and each frequency division isolation channel corresponds to one of the receiving modules. The multi-channel audio processing method includes the following steps: Each of the aforementioned transmitting modules is driven to synchronously acquire sound signals, and the received 2N sound signals are converted into corresponding time-domain signals; Using the main control module's own clock as a reference, the frame synchronization signal of each receiving module is read. Based on the phase difference and / or frequency difference between the frame synchronization signal and its own clock, the main control clock output by each receiving module to the transmitting module is adjusted through an audio phase-locked loop. The arrival delay of each time domain signal is obtained by timestamping the audio frames in each time domain signal using a hardware timer and counting the timestamps of the audio frames with the same frame number in all time domain signals. Based on the arrival delay, one path is selected as the target signal from all time-domain signals, and the processing time or data transmission method of the remaining time-domain signals is adjusted based on the target signal. Once all time-domain signals are in the processing area, one time-domain signal is selected from all the time-domain signals as a reference signal, and the remaining time-domain signals are compared with the reference signal to obtain the delay deviation of each time-domain signal. Based on the delay deviation, a dynamic resampling factor α is assigned to each of the time-domain signals, with the value of α ranging from 0.999 to 1.001. Based on the dynamic resampling factor, a third-order Lagrange interpolation is performed on each of the time-domain signals to obtain a synchronization sampling point. The interpolated time-domain signal is transformed into the frequency domain to obtain the frequency domain signal, and the signal-to-noise ratio and energy index of each frequency domain signal are obtained. Based on the signal-to-noise ratio and the energy index, the optimal path frequency domain signal is selected from all the frequency domain signals, and the optimal path frequency domain signal is restored to the time domain signal for output.
[0007] Optionally, the step of adjusting the processing time or data transmission method of the remaining time-domain signals based on the target signal includes: A circular buffer and a processing area are set in the main control module. If the remaining time-domain signals arrive earlier than the target signal, the remaining time-domain signals are buffered in the circular buffer until the target signal arrives. Then, the remaining time-domain signals and the target signal are sent to the processing area. If the remaining time-domain signals arrive later than the target signal, the remaining time-domain signals are transmitted to the processing area through the I2S DMA data transmission method.
[0008] Optionally, the step of comparing the time-domain signal with a preset reference signal to obtain the delay deviation of each time-domain signal includes: The time-domain signal is subjected to sliding cross-correlation calculation with a preset reference signal to obtain the position of maximum cross-correlation value between each time-domain signal and the reference signal, and the delay deviation is calculated based on the position of maximum cross-correlation value.
[0009] Optionally, the step of performing third-order Lagrange interpolation on each of the time-domain signals based on the dynamic resampling factor includes: If the sampling point of the remaining time-domain signal is earlier than that of the reference signal, then α<1 is set, and the remaining time-domain signal is compressed. If the sampling point of the remaining time-domain signal is later than that of the reference signal, then α>1 is set, and the remaining time-domain signal is stretched.
[0010] Optionally, after the step of performing third-order Lagrange interpolation on each of the time-domain signals based on the dynamic resampling factor, the method includes: Perform a short-time Fourier transform on each of the time-domain signals to obtain the frequency-domain phase spectrum; Using the frequency domain phase spectrum of the reference signal as the reference phase spectrum, the coherence between the remaining frequency domain phase spectra and the reference phase spectrum is calculated. If the coherence exceeds a preset coherence threshold and the phase difference exceeds a preset phase threshold, then the signal segment is inverted.
[0011] Optionally, before the step of restoring the optimal path frequency domain signal to a time domain signal output, the following steps are included: Analyze the spectral characteristics and energy distribution of the frequency domain signal to identify the current application scenario, select a corresponding signal optimization algorithm based on the current application scenario, and optimize the optimal path frequency domain signal based on the signal optimization algorithm; or Calculate the signal energy of each frequency domain signal, obtain the median of all signal energies as the reference energy, and perform gain compensation for the remaining frequency domain signals based on the reference energy.
[0012] Optionally, before the step of restoring the optimal path frequency domain signal to a time domain signal output, the method further includes: The optimal path frequency domain signal and the remaining frequency domain signals are subjected to sliding cross-correlation calculation to obtain the relative residual delay; based on the relative residual delay, the remaining frequency domain signals are subjected to Lagrange interpolation calculation or sinc resampling processing. The phase coherence of the remaining frequency domain signals after processing is obtained, the phase reversal point is obtained based on the phase coherence, and the phase reversal point is subjected to frequency domain phase rotation.
[0013] The present invention also discloses a multi-channel sound pickup system, comprising: There are N receiving modules, where N ≥ 2; There are 2N transmitting modules, and each of the receiving modules is connected to 2 transmitting modules; The main control module is connected to all the receiving modules. The main control module has N independent frequency division isolation channels, and each frequency division isolation channel corresponds to one of the transmitting modules. The main control module is used to implement the steps of the method described above.
[0014] The present invention also discloses a storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.
[0015] The present invention also discloses a sound pickup device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.
[0016] The beneficial effects of the multi-channel audio synchronization processing method, system, device, and medium provided in this invention are as follows: By dynamically adjusting the master control clock output to each receiving module through an audio phase-locked loop, the sampling rates of each signal are forced to be precisely aligned, eliminating the cumulative effect of frequency deviation from the source. A nanosecond-level hardware timestamp is applied to each channel's audio frame. Using the median delay channel as a reference, a buffer is used to achieve early-arriving channels waiting and late-arriving channels reading quickly, ensuring that audio frames with the same frame number in each time-domain signal have the same start time, solving the problem of random inter-frame delay caused by wireless transmission. Subsampling-level residual delay is detected through sliding cross-correlation calculation. Time axis fine-tuning is completed through Lagrange interpolation, and high-frequency phase inversion is corrected in the frequency domain, completely eliminating tiny phase and delay deviations of less than one sampling point. This lays a high-precision synchronization foundation for subsequent frequency domain transformation, channel selection, mixing processing, and other stages, improving the audio quality of subsequent output. Attached Figure Description
[0017] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a flowchart illustrating an embodiment of the multi-channel audio processing method provided by the present invention; Figure 2 This is a schematic diagram of an embodiment of the multi-channel pickup system provided by the present invention; Figure 3 This is a schematic diagram of the structure of an embodiment of the receiving module provided by the present invention; Figure 4 This is a schematic diagram of the structure of an embodiment of the transmitting module provided by the present invention; Figure 5 This is a schematic diagram of the structure of an embodiment of the main control module provided by the present invention; Figure 6 This is a schematic diagram of the internal structure of a pickup device in one embodiment of the present invention.
[0018] The labels for the attached figures are as follows: 10. Multi-channel pickup system; 11. Transmitter module; 12. Receiver module; 13. Main control module. Detailed Implementation
[0019] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0020] Please refer to the following: Figure 1 and Figure 2 , Figure 1 This is a flowchart illustrating an embodiment of the multi-channel audio processing method provided by the present invention. Figure 2 This is a schematic diagram of an embodiment of the multi-channel audio pickup system provided by the present invention. The multi-channel audio pickup system 10 includes N receiving modules 12, where N≥2, and 2N transmitting modules 11. Each receiving module 12 is connected to two transmitting modules 11, and a main control module 13 is connected to all receiving modules 12. The main control module 13 is built using an RV1106B main control processing chip. Wireless audio transmission between the transmitting modules 11 and the receiving modules 12 is achieved through a 2.4GHz ISM radio frequency link, and physical connection, clock synchronization, command interaction, and analog audio transmission are achieved through standardized cold shoe contacts. Both the transmitting modules 11 and the receiving modules 12 adopt a miniaturized, low-power hardware design with built-in independent lithium batteries. The main control module 13 balances portability and processing power, making it suitable for mobile audio pickup scenarios.
[0021] Figure 2 This describes the case where N=3. In other implementation scenarios, N can also be 2, 4, 5, 6, or other values. Please refer to the relevant documentation. Figure 3 and Figure 4 , Figure 3 This is a schematic diagram of a structure of an embodiment of the receiving module 12 provided by the present invention. Figure 4 This is a schematic diagram of an embodiment of the transmitting module 11 provided by the present invention. In this embodiment, each of the three receiving modules 12 (RX) supports time-division / frequency-division access to two transmitting modules 11 (TX), forming three independent 2TX+1RX wireless pickup units. The three wireless pickup units are uniformly managed by the main control module 13, realizing centralized acquisition and processing of 6-channel audio. The receiving modules 12 and the main control module 13 are connected by a cold shoe contact structure, supporting hot-swapping without power interruption, enabling rapid deployment and replacement of modules. By increasing or decreasing the number of transmitting modules 11 and receiving modules 12, configurations of one-to-two (1RX+2TX), one-to-four (2RX+4TX), and one-to-six (3RX+6TX) can be freely switched without changing the hardware structure.
[0022] The transmitting module 11 is equipped with an ATS3031 radio frequency chip, which includes a radio frequency transmitting circuit and a 2.4GHz antenna (built-in ceramic antenna). It supports independent frequency division isolation channel communication with the corresponding receiving module 12 to realize wireless transmission of audio signals.
[0023] The receiving module 12 integrates an ATS3031 RF chip, which includes an RF receiving circuit, a 2.4GHz antenna (external patch antenna), and an RF demodulation circuit. It supports time-division / frequency-division mechanism to simultaneously receive RF signals from two transmitting modules 11, and realizes demodulation and decoding of two audio signals. The ATS3031 uses an independent channel, which is frequency-division isolated from the channels of the other two receiving modules 12, so that they do not interfere with each other.
[0024] Please refer to the following: Figure 5 , Figure 5 This is a schematic diagram of an embodiment of the main control module 13 provided by the present invention. The main control module 13 is equipped with an RV1106B main control chip (with built-in NPU neural network processing unit, RISC-VVector hardware acceleration unit, and APLL audio phase-locked loop module), which is the core computing unit of the system; it has a built-in high-precision crystal oscillator (±1ppm) as the sole time reference for the entire system; it is responsible for running multi-channel audio synchronization, noise reduction, scene adaptation and other algorithms, and at the same time realizes unified management of modules and command issuance.
[0025] The main control module 13 is also equipped with two cascaded ES7210 ADC analog-to-digital converter chips, which are the core of the 6-channel analog audio acquisition. Each ES7210 supports 3 channels of analog audio acquisition. The two cascaded chips realize 6 channels, 48kHz sampling rate, and 24-bit depth high-precision analog-to-digital conversion, converting the 6 channels of analog audio signals input from the RX module into digital PCM signals, which are then transmitted to the RV1106B for subsequent processing.
[0026] The main control module 13 is equipped with three standardized cold shoe interface sockets (matching the cold shoe protrusions of the receiving module 12). Each interface socket integrates corresponding metal contacts to realize physical connection with the three receiving modules 12, analog audio reception, MCLK clock output, I2C / UART communication, and power supply output. It supports hot-swapping of the receiving modules 12, and automatically detects the module access status during plugging and unplugging, and synchronously updates the number of channels and processing links.
[0027] The main control module 13 has N independent frequency division isolation channels, each corresponding to one receiver module 12. Specifically, the main control module 13, based on the system's core RF chip ATS3031, configures three independent frequency division isolation channels (denoted as channel 1, channel 2, and channel 3) in the 2.4GHz ISM band, setting the center frequency and bandwidth (non-overlapping) of each channel to ensure that the isolation between channels meets the requirements for interference-free wireless transmission. The wireless transmission of the three receiver modules 12 is carried out in three non-overlapping channels, ensuring that even in the same spatial environment, the three sets of 2TX+1RX RF signals will not interfere with each other, guaranteeing the stability of the wireless transmission of six audio signals.
[0028] The multi-channel audio processing method provided by this invention includes the following steps: S101: Drives each transmitting module to synchronously acquire sound signals and converts the received 2N sound signals into corresponding time-domain signals.
[0029] In a specific implementation scenario, the main control module issues a pickup drive command to each receiving module, and each receiving module then forwards the pickup command to the two connected transmitting modules. Upon receiving the pickup command, each transmitting module simultaneously activates its respective pickup circuit to acquire an analog audio signal. This signal is then sent to the receiving module, which forwards it to the main control module. After receiving 2N channels of analog audio signals, the main control module converts the audio signals into time-domain signals using its built-in analog-to-digital converter (ADC).
[0030] S102: Using the main control module's own clock as a reference, read the frame synchronization signal of each receiving module, and adjust the main control clock output by each receiving module to the transmitting module through an audio phase-locked loop based on the phase difference and / or frequency difference between the frame synchronization signal and its own clock.
[0031] In a specific implementation scenario, the transmission path of each time-domain signal is used as its corresponding channel (transmitter module - receiver module - main control module). The high-precision crystal oscillator clock built into the RV1106B (accuracy within ±1ppm, far superior to the ±20ppm of the receiver and transmitter modules) serves as the sole time reference for the entire multi-channel audio pickup system, eliminating frequency deviation at its source. The clock period and frame interval corresponding to a nominal sampling rate of 48kHz and a frame length of 20ms are defined. Each receiver module returns FSYNC (Frame Synchronization), which is the start marker pulse for each audio frame (one pulse every 20ms, strictly bound to the audio frame).
[0032] The main control module collects and compares the phase difference and frequency difference between the FSYNC pulses fed back by each receiving module and its own clock. The phase difference refers to the position offset of the FSYNC pulse from the theoretical frame start pulse of its own clock at the same time point (reflecting instantaneous delay, such as FSYNC being 5 clock cycles late). The frequency difference refers to the ratio of the actual interval of the FSYNC pulses to the 20ms frame interval of the reference clock (reflecting sampling rate deviation, such as an FSYNC interval of 20.004ms, indicating that the actual sampling rate of the receiving module is slightly lower than 48kHz). The phase difference and frequency difference are used as input parameters of APLL (Audio Phase-Locked Loop). The APLL dynamically adjusts the main control clock MCLK clock frequency output to each receiving module according to the deviation value. Each receiving module has an independent APLL adjustment channel, supporting individual calibration of each receiving module.
[0033] The system continuously collects FSYNC feedback from each receiving module, calculates the deviation, and adjusts the MCLK to form a real-time closed-loop control until the deviation corresponding to the FSYNC of each receiving module meets preset conditions, such as phase difference < 1 clock cycle and frequency difference < ±0.1ppm. The frequency of the master control clock (MCLK) output to each receiving module is dynamically adjusted through an audio phase-locked loop, ensuring that the MCLK of all channels is completely synchronized with the master clock. This ultimately achieves forced alignment of the sampling rate of each channel to 48kHz, eliminating the cumulative effect of crystal oscillator error at its source. The FSYNC pulses of each receiving module are completely synchronized with the clock of the master control module, ensuring consistent frame generation rhythm.
[0034] The actual sampling rate of the transmitting and receiving modules is determined by the master clock MCLK. By dynamically adjusting the MCLK frequency of each receiving module, its actual sampling rate can be forced to be fully aligned with the 48kHz reference of the master control module. The APLL audio phase-locked loop is a clock calibration module specifically optimized for audio scenarios. It adopts linear frequency modulation tuning, with each frequency adjustment amount <0.001%. Compared with ordinary PLLs, the adjustment process has no clock glitches and no frequency jumps, avoiding audio distortion caused by clock fluctuations.
[0035] S103: Timestamp the audio frames in each time domain signal using a hardware timer, count the timestamps of audio frames with the same frame number in all time domain signals, and obtain the arrival delay of each time domain signal.
[0036] Even if the sampling rates are perfectly aligned, the audio data of the same frame number will still arrive at the main control module at different times after the signals sent by each receiving module are transmitted wirelessly (due to different distances and channel interference) and acquired by hardware (for example, frame 100 of channel 0 arrives at 100.000ms, while frame 100 of channel 1 arrives at 100.003ms). This frame-level random delay will cause inconsistencies in the starting points of subsequent framing and FFT, resulting in phase confusion in frequency domain processing.
[0037] In a specific implementation scenario, the main control module uses a built-in hardware timer (driven by its own clock) to stamp each audio frame arriving at each channel with a nanosecond-level timestamp (accuracy ±1ns). The timestamp is directly written into the header of the audio frame and bound to the time domain data without software intervention (avoiding delay errors caused by software timing).
[0038] S104: Select one path as the target signal from all time-domain signals based on the arrival delay, and adjust the processing time or data transmission method of the remaining time-domain signals based on the target signal.
[0039] In a specific implementation scenario, the timestamps of audio frames with the same frame number in each time-domain signal are statistically analyzed. The arrival delay of each channel's frame data (arrival time - theoretical frame start time of the reference clock) is calculated. All arrival delays are sorted, and the time-domain signal with the median delay is selected as the target channel. The median reference can minimize the overall waiting time. For example, if the delays of each time-domain signal are 1μs, 2μs, 3μs, 4μs, 5μs, and 6μs, 3μs is selected as the reference. Only the first two channels need to wait 2μs and 1μs, and the last three channels need to catch up to 3μs, 2μs, and 1μs, balancing synchronization accuracy and processing delay, and avoiding an increase in end-to-end delay due to excessive waiting.
[0040] The main control module allocates an independent circular buffer (sized to fit 2-3 frames of data to avoid overflow) for each time-domain signal, and controls the buffer based on the delay difference between the remaining time-domain signals and the target signal. Specifically, if the time-domain signal is received earlier than the target signal, the audio frame is written into the circular buffer and buffered according to the theoretical processing time of the target signal until the unified frame processing start time is reached. If the time-domain signal is received later than the target signal, the frame data is quickly read through I2S DMA (Inter-IC Sound Direct Memory Access, integrated circuit built-in audio bus direct memory access) high-speed data transmission (hardware-level transmission, no CPU intervention, transmission rate much higher than audio data rate), skipping non-critical software checks, ensuring that the data is sent to the circular buffer before the unified processing time.
[0041] When all audio frames with the same frame number in the time domain signals of all channels are buffered into the circular buffer and reach the unified processing time of the target signal, the main control module triggers a frame processing interrupt, sending all audio frames into the subsequent processing area simultaneously to ensure that the start time of each audio frame is completely consistent.
[0042] A circular buffer is used instead of a regular buffer to avoid empty buffer waits and full buffer overflows for frame data, adapting to the continuous transmission characteristics of real-time audio. The target latency difference for frame-level alignment is <1μs (far less than 20.8μs per sampling period), ensuring minimal pressure on subsequent sample-point-level fine-tuning.
[0043] The actual arrival time of each frame of data is marked by a nanosecond-level hardware timestamp. The median delay channel is used as a reference (balancing synchronization accuracy and processing latency). A circular buffer is used to cache early frames and catch up with late frames by high-speed reading through DMA (Direct Memory Access). Ultimately, audio data with the same frame number transmitted from each channel enters the subsequent processing module at the same time, achieving strict alignment of the frame start time.
[0044] S105: After all the time-domain signals of all channels are in the processing area, select one channel from all the time-domain signals as the reference signal, and compare the remaining time-domain signals with the reference signal to obtain the delay deviation of each channel's time-domain signal.
[0045] Frame-level alignment can ensure that the start time of audio frames is consistent, but there may still be a deviation of less than one sampling period in the sampling points within the same frame (e.g., 0.2~0.8 sampling points, corresponding to 4.16~16.64μs). Although this deviation is small, the human ear is extremely sensitive to phase differences in the high-frequency band (>4kHz). During multi-channel mixing, comb filtering effects (hollow, muffled sound, high-frequency loss) may occur, and even local frequency sound cancellation may occur. Subsampling-level fine-tuning is necessary.
[0046] In a specific implementation scenario, the quality of each time-domain signal is evaluated, and based on the evaluation results, one time-domain signal is selected as a reference signal from all the time-domain signals. A sliding cross-correlation calculation is performed on the remaining time-domain signals and the reference signal to find the position where the cross-correlation value of each remaining time-domain signal is the maximum. The position with the maximum cross-correlation value is the optimal alignment position between the two signals. Using this optimal alignment position, the subsampling delay deviation can be accurately calculated (with an accuracy of 1 / 8 of a sampling point, corresponding to 2.6 μs, far exceeding the human ear's perception threshold).
[0047] S106: Assign a dynamic resampling factor α to each time-domain signal based on the delay deviation, and perform third-order Lagrange interpolation on each time-domain signal based on the dynamic resampling factor to obtain the synchronization sampling point.
[0048] In a specific implementation scenario, a dynamic resampling factor α (strictly limited to 0.999~1.001) is assigned to each remaining time-domain signal based on the detected deviation. If the sampling points of an audio frame of a remaining time-domain signal are 0.1~0.8 sampling points earlier than the sampling points of the audio frame of the reference signal, α is set < 1 (e.g., 0.999, 0.9995), performing a slight compression on the signal. For example, a frame with 960 sampling points generates 959 or 959.5 new sampling points, effectively slowing down the signal to align it with the reference signal. If the sampling points of an audio frame of a remaining time-domain signal are 0.1~0.8 sampling points later than the sampling points of the audio frame of the reference signal, α is set > 1 (e.g., 1.0005, 1.001), performing a slight stretch on the signal. For example, a frame with 960 sampling points generates 961 or 960.5 new sampling points, effectively "speeding up" the signal to catch up with the reference channel. The range of α from 0.999 to 1.001 completely covers the residual deviation after hardware and frame-level synchronization, and the rate of change is less than 0.1%, so the human ear does not perceive any distortion.
[0049] Perform third-order Lagrange interpolation on each time-domain signal (balancing interpolation accuracy and computational load; RV1106B single-core can process in real time), generate new sampling points based on the resampling factor α, realize micro-calibration of the time axis of the sampling points, and eliminate subsampling delay bias.
[0050] In other implementation scenarios, frequency domain phase consistency verification is performed on the aligned time-domain signals. A short-time Fourier transform (STFT) is performed on the time-domain signals to extract the frequency domain phase spectrum of each signal. Using the phase spectrum of the reference signal as a benchmark, the phase coherence of other time-domain signals with respect to the reference signal is calculated. Phase coherence reflects the phase consistency of the two signals. Phase correction is performed on the high-frequency (>4kHz) time-domain signals. If the phase coherence of a channel in the high-frequency range is less than a preset threshold (e.g., 0.8), it is determined to be a phase inversion. Frequency domain phase rotation is used to correct it to be consistent with the reference channel, completely avoiding the comb filtering effect. Correction is only performed on the >4kHz high-frequency range because the human ear is not sensitive to phase differences in the low-frequency range (<4kHz), and phase deviations in the low-frequency range do not cause significant comb filtering, reducing computational load.
[0051] The refined 6-channel time-domain signals are standardized and framed according to a frame length of 20ms (960 sampling points) and a frame shift of 10ms (480 sampling points), and the output is a sampling point-level aligned 6-channel time-domain PCM (Pulse Code Modulation) signal.
[0052] Each time-domain signal achieves strict alignment at the sampling point level, with residual delay deviation <1 / 8 sampling point (approximately 2.6μs) and phase difference <±5° (high frequency band); it completely eliminates comb filtering and sound cancellation problems caused by subsampling level deviation, and the output synchronous time-domain PCM signal fully meets the synchronization requirements of subsequent frequency domain transformation, channel selection, mixing and other processing.
[0053] S107: Perform frequency domain transformation on the interpolated time domain signal to obtain the frequency domain signal, and obtain the signal-to-noise ratio and energy index of each frequency domain signal.
[0054] In a specific implementation scenario, the time-domain signal is converted into frequency-domain spectral features (which can accurately analyze each frequency component, such as high-frequency attenuation and noise bands). At the same time, preprocessing is used to eliminate signal distortion and ensure channel synchronization, providing analyzable frequency-domain data for subsequent channel selection, scene recognition, etc.
[0055] Specifically, to compensate for the natural attenuation of high frequencies (3~8kHz) in human voices, each time-domain signal is pre-emphasized and filtered. A transfer function is used... (Current sampling point - 0.97 × previous sampling point), implemented using Q15 dot conversion.
[0056] Since audio is a time-varying signal, it needs to be segmented into short frames (20ms) for analysis (to ensure that the signal is approximately stable within a single frame). However, when FFT (Fast Fourier Transform) processes aperiodic signals, frequency components leak into adjacent frequency bands. Therefore, a Hanning window is added to suppress spectral leakage. The Hanning window formula is multiplied frame by frame onto the time-domain signal to suppress spectral leakage.
[0057] Short-time Fourier transform (SFT) converts each frame's time-domain signal into a complex frequency-domain spectrum (containing the amplitude and phase at each frequency point), which forms the basis for all subsequent frequency-domain processing. Zero-padding is performed on the sampling points of each audio frame, increasing the number of zeros from 960 to 1024. Zero-padding does not add information but ensures the FFT point count is a power of 2. The main control module's RVV (RISC-V Vector Extension) instruction set... Point FFT has hardware acceleration, increasing computing power by more than 50%. It calls the RVV instruction set to execute FFT, with input in Q15 format (pre-emphasized fixed-point signals) and output upgraded to Q31 format (to prevent computational overflow, Q31 has a larger dynamic range).
[0058] S108: Select the optimal path frequency domain signal from all frequency domain signals based on signal-to-noise ratio and energy index, and restore the optimal path frequency domain signal to the time domain signal output.
[0059] In a specific implementation scenario, the best and most stable channel with the best sound quality is selected from all the channel frequency domain signals as the optimal channel frequency domain signal (the reference channel for subsequent subsampling refinement). This avoids inferior channels with high noise and unstable energy from lowering the overall sound quality. At the same time, strict switching rules prevent sudden changes in sound quality caused by frequent channel switching.
[0060] The energy of each audio frame is calculated and smoothed to eliminate energy fluctuations caused by sudden noise / silence, resulting in more stable energy values. Specifically: The instantaneous energy is:
[0061] After first-order smoothing, the energy of the audio frame is:
[0062] in, The energy value of the nth sampling point. Let be the instantaneous energy of the m-th audio frame. =0.9~0.98, The closer it is to 1, the stronger the smoothing effect.
[0063] The signal-to-noise ratio (SNR) for each audio frame can be calculated using VAD (Voice Activity Detection) to identify speech / silence segments (based on energy and spectral characteristics; for example, if the energy is above a threshold and has high-frequency speech features, it is considered a speech segment). Noise power is updated only in silence segments (to avoid including speech signals as noise). Specifically, SNR can be calculated using the following companies:
[0064] in, The signal energy of the entire audio frame. The noise energy, ε=1e-6, is used to prevent division by zero.
[0065] Furthermore, the SNR value is standardized to the range of 0 to 1 to facilitate the scoring comparison between different signals (the absolute values of SNR of different channels vary greatly, and normalization makes it easier to determine).
[0066] Based on the SNR and energy values of the frequency domain signals of different paths, each frequency domain signal is scored. For example, the score = SNR_norm × weight 1 + energy normalization value × weight 2, and the highest score is taken as the optimal frequency domain signal.
[0067] It should be noted that the scores of each signal will fluctuate depending on the actual situation. Therefore, scores can be obtained according to a preset period, and the current optimal path frequency domain signal can be switched based on the score. When switching is required, the following three conditions must be met simultaneously: ① Candidate path score - current main path score > Δ (Δ = 0.1, to avoid switching due to minor differences); ② Duration of the current path as the optimal path > T_hold (T_hold > 200ms, to ensure stability); ③ Interval between two switching operations > T_min (T_min > 500ms, to avoid repeated switching).
[0068] Since the optimal channel frequency domain signal is switchable, in order to ensure that the audio quality does not change when switching at any time, the algorithm strategy is dynamically adjusted according to the actual sound pickup scenario (maximizing the adaptation to scenario requirements), while balancing the volume of all channel signals to avoid any signal being too loud / too soft, ensuring that the output volume is stable and distortion-free.
[0069] It can extract the spectral features of all road frequency domain signals (such as low-frequency energy ratio >60% in wind noise environment, high-frequency noise in noisy street) and energy distribution (stable energy in quiet indoor environment). Based on the rule base / AI model, it calculates the confidence of 4 types of scenarios (quiet indoor / noisy street / large venue / wind noise environment). The confidence of new scenarios is >0.8 and lasts for 3~5 seconds (to avoid misidentification). It can smoothly switch the currently used signal optimization algorithm combination (such as using AI deep noise reduction + beamforming in noisy environment, and only light noise reduction in quiet environment).
[0070] The RMS (Root Mean Square) energy of each frequency domain signal is calculated in real time. Based on the median RMS, ±12dB gain compensation is dynamically applied to the remaining frequency domain signals (increase gain if it is too light, decrease gain if it is too heavy). It should be noted that the gain is only adjusted in the speech segment determined by VAD (to avoid amplifying noise in the silent segment), and the peak level is monitored (the gain is reduced if it exceeds -3dBFS) to prevent clipping distortion.
[0071] In other implementation scenarios, the subsampling-level residual delay between signals can be eliminated by correcting the correlation coefficient. The optimal frequency domain signal is used as the master signal, and the high-frequency phase deviation of the remaining frequency domain signals is corrected to prevent comb filtering effects during multi-channel mixing. Using the master signal as a reference, the remaining frequency domain signals are slid within a ±1 sampling point range with a 1 / 8 sampling point step size to calculate the normalized cross-correlation value. The peak position is the residual delay (accuracy 2.6 μs, far exceeding the human ear's perception threshold). The residual delay is then fine-tuned using Lagrange interpolation / sinc resampling to eliminate it.
[0072] The phase coherence of the remaining frequency domain signals and the main signal is calculated in the frequency domain. For the high frequency band >4kHz, if the coherence is <0.8 and the phase difference is ≈180°, the frequency domain signal in this band is inverted and subjected to anti-comb filtering.
[0073] In other implementation scenarios, a 4-6 band dynamic EQ (63Hz / 250Hz / 1kHz / 2kHz / 4kHz / 8kHz) is used. The gain of each frequency band is adjusted specifically according to the actual sound pickup scenario. For example, the core vocal frequency band of 1-2kHz is boosted to enhance speech intelligibility. A dual-loop control system of slow AGC and fast limiter is employed. The slow AGC stabilizes the output volume within the professional audio ideal loudness range of -24dBFS to -18dBFS with a response time of several hundred milliseconds. The fast limiter has a startup time of less than 1ms, effectively suppressing sudden peak signals and preventing overflow distortion. Simultaneously, the AGC and EQ are optimized in tandem. When high gain is applied by AGC, the low-frequency EQ gain is automatically reduced to prevent low-frequency noise from being amplified synchronously, thus balancing volume stability and sound quality purity. This improves vocal clarity, stabilizes the output volume to avoid sudden fluctuations, and effectively prevents noise from being amplified synchronously.
[0074] The processed optimal path frequency domain signal is restored to the time domain signal and then output. First, de-emphasis filtering is performed to cancel the effect of the preceding pre-emphasis filtering, through the inverse transfer function of the pre-emphasis filtering. The Q15 fixed-point method is used to synchronously execute all frequency domain signals to restore the natural spectral characteristics of the speech and avoid excessive brightness in the high frequencies of the output audio. Then, ISTFT (Inverse Short-Time Fourier Transform) is performed, calling the RVV instruction set of the main control module to execute 1024-point IFFT for hardware acceleration. The inverse short-time fourier transform is completed by using a 50% overlap rate with Hanning window overlap and addition to restore the frequency domain spectrum to the time domain frame sequence, canceling frame distortion and ensuring signal continuity. Subsequently, signal splicing and output are carried out. Discrete time domain frames are spliced in a 10ms frame shift order, so that the last 480 samples of the previous frame overlap and add with the first 480 samples of the next frame to form a continuous signal. Then, a 48kHz / 24bit single-channel PCM (Pulse Code Modulation) signal is output and sent to external devices such as power amplifiers, sound cards, and Bluetooth modules.
[0075] Finally, hot-swap adaptation is completed. By detecting the level changes of the cold shoe contacts in real time, the connection and disconnection actions of the transmitting and receiving modules are perceived. The software automatically updates the number of channels and adjusts the audio processing link synchronously. Relying on the interrupt mechanism of the main control module, a fast response is achieved. It supports module plugging and unplugging without power interruption and without audio interruption during the process, and flexibly adapts to dynamic changes in the number of channels.
[0076] As described above, in this embodiment, the main control clock output to each receiving module is dynamically adjusted by an audio phase-locked loop, forcing precise alignment of the sampling rates of each signal, eliminating the cumulative effect of frequency deviation from the source, and assigning nanosecond-level hardware timestamps to each channel's audio frames. Using the median delay channel as a reference, a buffer is used to achieve early arrival channel waiting and late arrival channel fast reading, ensuring that audio frames with the same frame number in each time domain signal have the same start time, solving the problem of random inter-frame delay caused by wireless transmission. Subsampling-level residual delay is detected by sliding cross-correlation calculation, and time axis fine-tuning is completed by Lagrange interpolation. High-frequency phase reversal is corrected in the frequency domain, completely eliminating tiny phase and delay deviations of less than one sampling point. This lays a high-precision synchronization foundation for subsequent frequency domain transformation, channel selection, mixing processing, and other stages, improving the audio quality of subsequent output.
[0077] Figure 6 A schematic diagram of the internal structure of a pickup device in one embodiment is shown. For example... Figure 6As shown, the microphone includes a processor, a memory, and a network interface connected via a system bus. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and may also store a computer program. When executed by the processor, this computer program enables the processor to implement a multi-channel audio processing method. The internal memory may also store a computer program, which, when executed by the processor, enables the processor to execute methods applied to the multi-channel audio processing. Those skilled in the art will understand that… Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the intelligent interactive device to which the present application is applied. The specific sound pickup device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0078] In one embodiment, a pickup device is provided, including a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps described above.
[0079] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, causes the processor to perform the steps described above.
[0080] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0081] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0082] It should be understood that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Those skilled in the art can modify the technical solutions described in the above embodiments, or make equivalent substitutions for some of the technical features; and all such modifications and substitutions should fall within the protection scope of the appended claims of the present invention.
[0083] It should be understood that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Those skilled in the art can modify the technical solutions described in the above embodiments, or make equivalent substitutions for some of the technical features; and all such modifications and substitutions should fall within the protection scope of the appended claims of the present invention.
Claims
1. A multi-channel audio processing method, characterized in that, This technology is applied to a multi-channel sound pickup system, which includes: There are N receiving modules, where N ≥ 2; There are 2N transmitting modules, and each of the receiving modules is connected to 2 transmitting modules; The main control module is connected to all the receiving modules. The main control module has N independent frequency division isolation channels, and each frequency division isolation channel corresponds to one of the receiving modules. The multi-channel audio processing method includes the following steps: Each of the aforementioned transmitting modules is driven to synchronously acquire sound signals, and the received 2N sound signals are converted into corresponding time-domain signals; Using the main control module's own clock as a reference, the frame synchronization signal of each receiving module is read. Based on the phase difference and / or frequency difference between the frame synchronization signal and its own clock, the main control clock output by each receiving module to the transmitting module is adjusted through an audio phase-locked loop. The arrival delay of each time domain signal is obtained by timestamping the audio frames in each time domain signal using a hardware timer and counting the timestamps of the audio frames with the same frame number in all time domain signals. Based on the arrival delay, one path is selected as the target signal from all time-domain signals, and the processing time or data transmission method of the remaining time-domain signals is adjusted based on the target signal. Once all time-domain signals are in the processing area, one time-domain signal is selected from all the time-domain signals as a reference signal, and the remaining time-domain signals are compared with the reference signal to obtain the delay deviation of each time-domain signal. Based on the delay deviation, a dynamic resampling factor α is assigned to each of the time-domain signals, with the value of α ranging from 0.999 to 1.
001. Based on the dynamic resampling factor, a third-order Lagrange interpolation is performed on each of the time-domain signals to obtain a synchronization sampling point. The interpolated time-domain signal is transformed into the frequency domain to obtain the frequency domain signal, and the signal-to-noise ratio and energy index of each frequency domain signal are obtained. Based on the signal-to-noise ratio and the energy index, the optimal path frequency domain signal is selected from all the frequency domain signals, and the optimal path frequency domain signal is restored to the time domain signal for output.
2. The multi-channel audio processing method according to claim 1, characterized in that, The step of adjusting the processing time or data transmission method of the remaining time-domain signals based on the target signal includes: A circular buffer and a processing area are set in the main control module. If the remaining time-domain signals arrive earlier than the target signal, the remaining time-domain signals are buffered in the circular buffer until the target signal arrives. Then, the remaining time-domain signals and the target signal are sent to the processing area. If the remaining time-domain signals arrive later than the target signal, the remaining time-domain signals are transmitted to the processing area through the I2S DMA data transmission method.
3. The multi-channel audio processing method according to claim 1, characterized in that, The step of comparing the time-domain signal with a preset reference signal to obtain the delay deviation of each channel of the time-domain signal includes: The time-domain signal is subjected to sliding cross-correlation calculation with a preset reference signal to obtain the position of maximum cross-correlation value between each time-domain signal and the reference signal, and the delay deviation is calculated based on the position of maximum cross-correlation value.
4. The multi-channel audio processing method according to claim 3, characterized in that, The step of performing third-order Lagrange interpolation on each of the time-domain signals based on the dynamic resampling factor includes: If the sampling point of the remaining time-domain signal is earlier than that of the reference signal, then α<1 is set, and the remaining time-domain signal is compressed. If the sampling point of the remaining time-domain signal is later than that of the reference signal, then α>1 is set, and the remaining time-domain signal is stretched.
5. The multi-channel audio processing method according to claim 1, characterized in that, After the step of performing third-order Lagrange interpolation on each of the time-domain signals based on the dynamic resampling factor, the following steps are included: Perform a short-time Fourier transform on each of the time-domain signals to obtain the frequency-domain phase spectrum; Using the frequency domain phase spectrum of the reference signal as the reference phase spectrum, the coherence between the remaining frequency domain phase spectra and the reference phase spectrum is calculated. If the coherence exceeds a preset coherence threshold and the phase difference exceeds a preset phase threshold, then the signal segment is inverted.
6. The multi-channel audio processing method according to claim 1, characterized in that, Before the step of restoring the optimal path frequency domain signal to a time domain signal output, the following steps are included: Analyze the spectral characteristics and energy distribution of the frequency domain signal to identify the current application scenario, select a corresponding signal optimization algorithm based on the current application scenario, and optimize the optimal path frequency domain signal based on the signal optimization algorithm; or Calculate the signal energy of each frequency domain signal, obtain the median of all signal energies as the reference energy, and perform gain compensation for the remaining frequency domain signals based on the reference energy.
7. The multi-channel audio processing method according to claim 6, characterized in that, Before the step of restoring the optimal path frequency domain signal to a time domain signal output, the method further includes: The optimal path frequency domain signal and the remaining frequency domain signals are subjected to sliding cross-correlation calculation to obtain the relative residual delay; based on the relative residual delay, the remaining frequency domain signals are subjected to Lagrange interpolation calculation or sinc resampling processing. The phase coherence of the remaining frequency domain signals after processing is obtained, the phase reversal point is obtained based on the phase coherence, and the phase reversal point is subjected to frequency domain phase rotation.
8. A multi-channel sound pickup system, characterized in that, include: There are N receiving modules, where N ≥ 2; There are 2N transmitting modules, and each of the receiving modules is connected to 2 transmitting modules; The main control module is connected to all the receiving modules. The main control module has N independent frequency division isolation channels, and each frequency division isolation channel corresponds to one of the transmitting modules. The main control module is used to implement the steps of the method as described in any one of claims 1-7.
9. A storage medium, characterized in that, The device stores a computer program that, when executed by a processor, causes the processor to perform the steps of the method as described in any one of claims 1 to 7.
10. A sound pickup device, characterized in that, It includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method as described in any one of claims 1 to 7.