A multi-mode audio and video processing device and control method with fully isolated acquisition and HDMI audio closed-loop processing
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-19
- Publication Date
- 2026-08-14
AI Technical Summary
1. 麦克风采集前端易受USB、HDMI、电源等数字系统干扰,地环路与共模噪声导致底噪偏高,难以实现高保真采集;
1、模拟前端全隔离采集架构,彻底消除地环路与共模噪声,底噪低、保真度高、阻抗可调,满足多种音频设备的高保真接入需求,接口性能好,灵活性高;相比采用音频变压器隔离的传统方案,本发明的全隔离采集架构可将模拟输入的THD从传统音频变压器的0.01%等级降低至话放芯片的0.0001%等级以下,且输入阻抗可调,兼容动圈麦克风、电容麦克风、高阻乐器等各类信号源。
Smart Images

Figure CN122578792A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio and video processing technology, specifically to an audio and video processing device and control method with fully isolated microphone acquisition, HDMI audio closed-loop processing, and multi-mode dynamic link switching. Background Technology
[0002] Existing audio and video processing equipment generally suffers from the following technical defects: 1. The microphone acquisition front end is susceptible to interference from digital systems such as USB, HDMI, and power supply. Ground loops and common-mode noise result in high background noise, making it difficult to achieve high-fidelity acquisition. 2. Traditional HDMI audio processing typically supports extraction output. The extracted audio, after external processing, can usually only be played out through an amplifier or speakers. It is impossible to reinject the processed audio into the HDMI output, let alone achieve a complete digital replacement of the original audio. 3. Voice cancellation, mixing, and audio effects processing are mostly integrated into the same chip. Due to the chip's power consumption and computing power, the chip selection flexibility is poor, and the effect is limited. 4. Recording, karaoke, live streaming and other scenarios reuse the same audio link, and it is impossible to achieve performance differences between different modes through dynamic switching at the hardware level. Summary of the Invention
[0003] To address the problems mentioned in the background section, the technical solution of this invention is as follows: A multi-mode audio and video processing device includes an analog front-end acquisition module, a digital isolation module, a main processing unit, a first digital signal processing unit, a second digital signal processing unit, a control unit, and an HDMI interface unit. The analog front-end acquisition module is used for impedance matching, analog amplification, and analog-to-digital conversion of microphone signals; The digital isolation module includes an I2S high-speed digital isolation unit, a control signal optocoupler isolation unit, and an isolation power supply unit, which enables complete electrical isolation between the analog front-end acquisition module and the device's digital system. The HDMI interface unit is used to receive and output HDMI audio and video signals; The main processing unit is configured to extract audio data from the input HDMI signal and send the extracted audio data to the first digital signal processing unit. The first digital signal processing unit is used to perform human voice cancellation processing on the audio data. It performs a 180-degree phase reversal on the data of one channel of the accompaniment audio and superimposes it with the data of the other channel, and performs gain smoothing by connecting a mid-frequency bandpass filter from 200Hz to 4kHz in series on the superposition path. The second digital signal processing unit is used to mix and process the audio data after human voice removal; The main processing unit uses digital mixing to configure the original HDMI audio percentage to 0% and the processed audio percentage to 100%, thus completely replacing the audio and injecting it back into the output HDMI signal. The device supports at least recording mode, karaoke mode, and live streaming mode. The control unit dynamically switches the audio path of each mode by configuring the internal link of the chip and the external analog switch. The audio links, functional modules, sampling rates and processing strategies of the three working modes are different. The control unit configures the internal register of the chip to enable or disable the functional module through the I2C interface according to the mode command, and simultaneously controls the external analog switch to switch the physical path of the I2S data stream through the GPIO level.
[0004] Preferably, the main processing unit is an HDMI processing chip with HDMI audio extraction, digital mixing and audio re-injection functions, and the main processing unit is selected from MS2131S, MS2133 or an HDMI processing chip with equivalent functions.
[0005] Preferably, the first digital signal processing unit is a PTN1118; the second digital signal processing unit is selected from a PTN1011 or a functionally equivalent audio DSP chip.
[0006] Preferably, the analog front-end acquisition module includes an ADC chip, which uses independent analog power supply and independent grounding, so that the noise floor of the acquired signal is lower than -110dB and the total harmonic distortion is lower than 0.0001%, achieving high-fidelity acquisition and adjustable impedance.
[0007] Preferably, the ADC chip is a PCM1863 or a high-fidelity ADC chip with equivalent functionality.
[0008] Preferably, the high-speed digital isolation unit in the digital isolation module is a digital isolation chip, and the control signal optocoupler isolation unit is an optocoupler device. The digital isolation chip and optocoupler device can be replaced with functionally equivalent isolation elements.
[0009] Preferably, in karaoke mode, the following complete closed-loop processing is performed: HDMI audio extraction, vocal removal, audio mixing, digital replacement, and HDMI re-injection output.
[0010] Preferably, in recording mode, the audio processing module is disabled, and dry audio acquisition and automatic mixing of multiple accompaniments are performed to maintain a high sampling rate and high-fidelity output.
[0011] Preferably, in live streaming mode, human voice cancellation processing is disabled, multi-channel audio mixing is enabled, and an independent monitoring channel is configured. The main processing unit establishes an isochronous transmission channel conforming to the UAC standard protocol via the USB bus, converts the parsed PC-side downlink audio data packets into I2S format streams, and routes them to the monitoring channel.
[0012] A multi-mode audio and video processing control method includes the following steps: S1. After the microphone signal is processed by the analog front end, it achieves complete electrical isolation from the digital system through I2S high-speed digital isolation, control signal optocoupler isolation and isolated power supply; S2. The main processing unit extracts audio data from the input HDMI signal; S3. The extracted audio data is sent to the first digital signal processing unit to perform human voice cancellation; S4. The audio after removing human voices is sent to the second digital signal processing unit to complete mixing and sound effects processing; S5. The main processing unit completely replaces the original audio signal in the output HDMI signal with the processed audio signal through digital mixing. That is, the proportion of the original audio signal in the output is zero, and the proportion of the processed audio signal is 100%. S6. The control unit configures the internal links of the chip and the external analog switch according to the selected working mode, and dynamically switches the audio path.
[0013] Compared with the prior art, the beneficial effects of the present invention are: 1. The fully isolated acquisition architecture of the analog front end completely eliminates ground loops and common-mode noise, resulting in low noise floor, high fidelity, and adjustable impedance, meeting the high-fidelity access requirements of various audio devices. It also offers good interface performance and high flexibility. Compared to the traditional solution that uses audio transformer isolation, the fully isolated acquisition architecture of this invention can reduce the THD of analog input from the 0.01% level of traditional audio transformers to below the 0.0001% level of microphone preamp chips. Furthermore, the input impedance is adjustable, making it compatible with various signal sources such as dynamic microphones, condenser microphones, and high-impedance musical instruments.
[0014] 2. The HDMI audio closed-loop processing architecture is not limited to a specific main processing chip model and supports the replacement of equivalent chips such as MS2131S and MS2133. The HDMI audio closed-loop replacement solution not only broadens the sources of audio for karaoke scenarios, but also transmits audio and video synchronously through HDMI after re-injection. Compared with the traditional solution where video goes through the HDMI TV and audio goes through the analog power amplifier and speakers, its audio and video synchronization accuracy is higher, reducing the audio and video offset from the usual 100~200ms to less than 50ms.
[0015] 3. Three-mode hardware-level dynamic link switching meets the differentiated needs of high-fidelity recording, strong karaoke effects, and multi-channel live streaming. The recording mode can achieve 192K high sampling and high dynamic dry audio recording. The karaoke mode can achieve a maximum of -30dB of voice cancellation and various microphone sound effects. The live streaming mode can achieve various microphone sound effects and loop playback of the acquisition port to meet the needs of live PK. The aforementioned maximum -30dB voice cancellation depth performance is derived by measuring the steady-state attenuation difference of the central sound image energy at the first DSP output pin and performing logarithmic conversion under laboratory calibration conditions with a standard human voice frequency band constant amplitude stereo test excitation signal input from 1kHz to 4kHz. The audio and video offset parameter of less than 50ms is strictly calculated based on the mathematical accumulation of the hardware drive clock rate of the I2S port inside the main processing chip and the total instruction cycle consumption time when the dual DSP is fully loaded with algorithm firmware. From the physical operation limit, it is defined that it must be lower than the industry-recognized critical standard of 50ms.
[0016] 4. High integration, strong compatibility, and wide applicability; core technology performance is maintained even after component replacement. Compared to traditional video capture cards, this invention integrates the audio processing links (voice removal, multi-channel mixing, and audio effects processing) required for karaoke and live streaming modes, achieving complete audio and video synchronization processing without the need for an external sound card. Compared to traditional live streaming sound cards, this invention has built-in HDMI video capture and loop-out functions, supporting real-time fusion output of 4K video signals and processed audio, avoiding synchronization delays and compatibility issues caused by external capture cards. Attached Figure Description
[0017] Figure 1 This is a block diagram of the overall hardware architecture of the system of the present invention; Figure 2 This is a block diagram of the fully isolated configurable analog acquisition architecture of the present invention; Figure 3 This is a schematic diagram of the HDMI audio closed-loop processing flow and audio re-injection replacement of the present invention; Figure 4 This is a block diagram of the signal link in the karaoke mode of the present invention; Figure 5 This is a block diagram of the signal link in the recording mode of the present invention; Figure 6 This is a block diagram of the signal link in the live streaming mode of this invention; Figure 7 This is a circuit connection diagram showing the implementation of the DSP2 input accompaniment signal to the Loop signal when switching the karaoke live streaming mode in this invention.
[0018] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0020] The purpose of this invention is to provide a multi-mode audio and video processing device and control method with fully isolated acquisition and HDMI audio closed-loop processing, which realizes high anti-interference of microphone acquisition, HDMI audio extraction-processing-digital replacement-reinjection closed loop, and multi-mode hardware-level dynamic link switching.
[0021] I. Overall Technical Solution A multi-mode audio and video processing device includes an analog front-end acquisition module, a digital isolation module, a main processing unit, a first digital signal processing unit, a second digital signal processing unit, a control unit, and an HDMI interface unit. The analog front-end acquisition module is responsible for impedance matching, analog amplification, and analog-to-digital conversion of the microphone signal.
[0022] The digital isolation module includes an I2S high-speed digital isolation unit, a control signal optocoupler isolation unit, and an isolation power supply unit, which enables complete electrical isolation between the analog front end and the whole digital system. They do not share a ground, power supply, or control path, ensuring signal fidelity, impedance stability, and configurable parameters.
[0023] The HDMI interface unit enables audio and video signal input and output.
[0024] The main processing unit extracts audio from the input HDMI and sends it to the first digital signal processing unit to perform voice cancellation; the audio after voice cancellation enters the second digital signal processing unit to complete mixing and sound effects processing.
[0025] The main processing unit sets the original HDMI audio percentage to 0% through digital mixing, and sets the processed audio percentage to 100%, achieving complete replacement and reinjection to the HDMI output, while maintaining direct video signal transmission.
[0026] The device supports three working modes: recording, karaoke, and live streaming. The control unit dynamically switches the audio paths, functional modules, and processing strategies for each mode by configuring the internal registers of the chip and external analog switches.
[0027] Preferably, the main processing unit uses a chip with HDMI audio extraction and re-injection functions. Different models of equivalent chips can be selected according to the actual application scenario to adapt to different performance and cost requirements.
[0028] Preferably, the first digital signal processing unit and the second digital signal processing unit can be replaced with different models of DSP chips that have the same processing functions.
[0029] Preferably, the isolation components in the digital isolation module can be selected from different models of digital isolation chips or optocouplers according to the transmission rate and anti-interference requirements, without affecting the implementation of the overall isolation architecture.
[0030] See the overall hardware architecture diagram of the system. Figure 1 .
[0031] II. Fully Isolated Configurable Analog Acquisition Architecture The analog front end uses an independent analog area, including impedance matching circuit, amplifier circuit and ADC conversion circuit.
[0032] The I2S digital audio signal is transmitted using a high-speed digital isolator. The control signal is transmitted in isolation using an optocoupler isolation unit and an I2C bidirectional isolation unit. The power supply uses an independent isolated power supply module, which completely decouples the analog acquisition part from the back-end digital system, completely blocks interference, and maintains audio fidelity and adjustable impedance.
[0033] Preferably, the high-speed digital isolator can be a different type of digital isolation chip, such as capacitive or magnetically coupled, and the optocoupler isolation transmission can be a high-speed optocoupler or an ordinary optocoupler, as long as it can achieve electrical isolation of the control signal.
[0034] Preferably, the ADC conversion circuit uses a high-fidelity ADC chip, such as PCM1863 or other equivalent chips, depending on the sampling rate requirements. The core of this design is to maintain an independent analog power supply and independent grounding architecture.
[0035] See the block diagram of the fully isolated configurable analog acquisition architecture. Figure 2 .
[0036] III. HDMI Audio Closed-Loop Processing Architecture The input HDMI audio and video signals enter the main processing unit; The audio data is extracted and sent to the first DSP to perform voice cancellation; After processing, the audio enters the second DSP to complete mixing, sound effects, and volume adjustment; The main processing unit performs a complete replacement using digital mixing: 0% original audio, 100% new audio; The replaced audio is re-embedded in the HDMI output, enabling karaoke, mixing, and live streaming output via the HDMI port.
[0037] Preferably, the main processing unit can be selected from chips such as MS2131S and MS2133 that have HDMI audio extraction and re-injection functions. Its core is to realize the closed-loop link of audio extraction-processing-re-injection, rather than being limited to a specific chip model.
[0038] See the HDMI audio closed-loop processing flow and audio re-injection replacement diagram. Figure 3 .
[0039] IV. Three-mode dynamic audio link switching Recording mode Turn off the audio effects of DSP1 and DSP2, perform high-fidelity dry audio acquisition, automatically mix multiple accompaniment tracks, and output high-sampling, high-quality, low-latency recording and monitoring signals.
[0040] Karaoke Mode Enable the HDMI audio extraction and processing loop, perform voice removal, mixing, and audio effects processing, and complete the HDMI audio replacement output.
[0041] Live streaming mode Turn off and bypass DSP1 (voice cancellation), enable multi-input mixing (including the Loop channel), configure an independent monitoring channel, and adapt to live streaming and multi-source audio input.
[0042] The device uses the I2S clock of the main processor unit MS2131S as the system's I2S bus clock. Other components, including DSP1, DSP2, ADC, and DAC, are connected to this bus as I2S slaves. This ensures clock synchronization of the audio data streams across all modules. Different modes are dynamically switched at the hardware level through internal chip configuration and external analog switches. Each mode has independent audio paths and performance indicators: the control unit configures internal chip registers via the I2C interface to enable or disable each functional module, and simultaneously controls external analog switches via GPIO to switch the direction of the I2S data stream. The audio links in the three operating modes are simply appropriate trimmings of the complete operating link of the device. Therefore, users can dynamically switch operating modes according to their needs.
[0043] Recording mode: DSP audio module is off, high sampling rate ADC is enabled, and the audio link is "analog front end → digital isolation → main processing unit → mixing → output"; See the signal link diagram for recording mode: Figure 5 Karaoke Mode: Enables dual-DSP full-link audio, with the audio link being "HDMI audio → dual-DSP processing → digital replacement → output"; See the signal link diagram for karaoke mode: Figure 4 Live streaming mode: Turn off the first DSP and enable the second DSP mixing function. The audio link is "multi-source audio → second DSP mixing → independent monitoring → output". See the signal link diagram for live streaming mode: Figure 6 Preferably, the control logic for mode switching can be implemented through MCU programming, and different models of external analog switches can be selected according to the number of channels and bandwidth requirements. The core is to achieve functional differentiation of different modes through hardware link switching. V. Diagram of HDMI audio replacement in karaoke mode In karaoke mode, the digital mixing ratio of the original HDMI audio and the processed audio is configured as "0% original audio, 100% processed audio" to achieve a complete replacement. The original HDMI audio signal is shielded within the main processing unit and is not output. The audio processed by dual DSP (accompaniment with vocals removed + microphone sound effect signal) is embedded into the HDMI video stream to form a new HDMI audio and video signal output; The video signal is passed through without any processing, ensuring image quality and transmission rate. Specifically, the main processing unit integrates a TMDS decoding and re-encoding core. The so-called video signal pass-through means that after the TMDS signal is unpacked at the receiving end, the extracted video pixel data is not sent to any image scaling, color space conversion, or frame rate conversion engine for processing. Instead, it is directly routed to the TMDS re-encoding end on the internal bus. Within this extremely short unpacking and reassembly cycle, the audio processing module of the main processing unit repackages the received processed new audio data into audio data islands according to the specifications and precisely overwrites and embeds them into the blanking period position of the re-encoded output video stream. Thus, at the physical link of the underlying communication protocol, data packet erasure and complete audio replacement are truly achieved without interfering with the original image quality matrix. The diagram illustrating the HDMI audio closed-loop processing flow and audio re-injection replacement in karaoke mode can be found here. Figure 3 . After the device is powered on, the control unit MCU (AT32F403) is responsible for initializing the system and configuring the digital isolation and analog front-end parameters. Then, according to preset or user instructions, the control unit configures the internal links of the chip and the external analog switches, switching to the corresponding operating mode.
[0044] (The control unit MCU is connected via I2C interface and GPIO respectively) Figure 4 , attached Figure 5 , attached Figure 6 (The signal link switching main processing unit, DSP1, DSP2, and the working status of the two analog switches) Example 1: The working process of the fully isolated microphone acquisition circuit is as follows: (See the block diagram of the isolated configurable simulation acquisition architecture:) Figure 2 ) The microphone signal is input via a 3-in-1 XLR connector, filtered, and then enters a preamplifier (preferably OPA1612) to complete impedance matching and gain amplification. The specific implementation details of the impedance matching circuit are as follows: the non-inverting input of the preamplifier OPA1612 is connected in parallel with multiple sets of precision feedback resistors through an analog switch array. The MCU dynamically switches the input resistance value according to the identified access device type (dynamic microphone, condenser microphone, or high-impedance musical instrument) through an optocoupler isolation unit, so that its input impedance can be mapped in real time within the range of 600Ω to 1MΩ, thereby ensuring that the original signals with different physical properties achieve optimal power transmission and spectral flatness during the pickup stage; then it is converted into an I2S digital signal by a high-fidelity ADC (PCM1863).
[0045] The I2S signal is transmitted to the system's I2S bus through a multi-channel high-speed digital isolation chip (preferably CA-IS3663LB). Considering that the high-speed digital isolation chip will generate picosecond-level edge jitter during pulse transmission, in order to support low total harmonic distortion (THD) performance below 0.0001%, the logic bridge of this solution is as follows: the system MasterClock generated by the main processing unit MS2131S is directly used as a reference and is returned to the analog front end through the independent channel of the isolation chip, which is forced to be used as the hardware driving master clock of the ADC chip PCM1863. Because the sampling logic of the ADC is locked on this high-order clock edge, it induces its conversion period to achieve strict phase synchronization with the system master clock, thereby neutralizing the random timing bias introduced by the isolation device on the physical link and ensuring the sampling accuracy and high fidelity of the sampled data. The control signals and port status signals of the MCU are implemented through optocouplers (preferably ELQ3H7) and I2C bidirectional isolators (ISO1541) to achieve bidirectional isolated communication control.
[0046] The entire analog front-end is powered by an independent isolated power supply module, not sharing a power ground with USB or HDMI. Specifically, this isolated power supply module integrates a 48V high-frequency DC-DC regulator boost circuit to generate standard phantom power. Because this phantom power is completely decoupled from digital ground through the secondary coil of the isolation transformer, it is injected into the XLR connector through a bias resistor at the balanced input, thereby inducing stable charge migration in the polarization head of the condenser microphone. This achieves compatible power supply for the condenser microphone under a fully isolated architecture, eliminating common-mode noise interference introduced by external phantom power. This achieves complete separation between analog and digital circuits, ensuring that the analog input is unaffected by noise interference from digital circuits, interfaces, and even external devices (such as PCs, HDMI players, and TVs). It maintains a device noise floor of -110dB even in extremely harsh environments, fully leveraging the performance of the ADC.
[0047] Example 2: Audio workflow in 3 working modes: (See the signal link diagram for karaoke mode:) Figure 4 ) The main processing unit (preferably MS2131S or MS2133) extracts I2S format audio data from the HDMI stream and transmits it to the first DSP (preferably PTN1118) through the chip's I2S_1 port.
[0048] Based on the user-preset vocal cancellation level (parameters obtained from the MCU via the I2C bus), vocal attenuation is performed on the accompaniment audio. Furthermore, the essential technical approach to this vocal cancellation level lies in the MCU extracting the corresponding Q value and center frequency gain of the mid-frequency bandpass dynamic attenuation filter from a pre-set coefficient mapping table within the DSP, according to the user-selected level instruction. As the level increases, the algorithm logic automatically adjusts the suppression bandwidth of the wide-pass filter and increases the coherent cancellation weight in the core frequency band from 200Hz to 4kHz. Through this adaptive parameter mapping, the device can accurately and deeply strip the central vocal spectrum while maintaining the fullness of the background accompaniment, based on the specific complexity of the vocal reverb in the original audio source. The processed audio is sent to the I2S_3 port of the second DSP (PTN1011) via the I2S_1 bus. This vocal attenuation processing is specifically executed by the phase cancellation algorithm operator inside the first DSP. After receiving the accompaniment audio, the processing module first separates its stereo left and right channel PCM data. Utilizing the acoustic characteristic that in commercial music mixing, the lead vocals are usually evenly distributed in the center of the sound field (i.e., both left and right channels contain the same amplitude and phase signal), the processing module attenuates one of the channels... After a strict 180-degree phase reversal calculation, the signal is superimposed and mixed with the other channel, resulting in coherent cancellation of the vocal waveform in the center image. Simultaneously, a 200Hz to 4kHz mid-frequency bandpass dynamic attenuation filter is connected in series on the superposition path to smooth the frequency band gain, thereby accurately stripping the vocal spectrum in the accompaniment according to the preset depth level. At the same time, the microphone input signal is converted into an I2S signal by a fully isolated acquisition circuit and sent to the I2S_2 port of the PTN1011. After the microphone signal sent to the PTN1011 undergoes reverb, EQ, pitch shifting and other sound effects processing, it is mixed with the accompaniment sound from the PTN1118 and the final audio is sent back to the MS2131S via the I2S_2 bus.
[0049] The MS2131S utilizes the digital mixer at its on-chip HDMI output port, configuring the original HDMI audio percentage to 0% and the newly received audio (processed by the DSP) percentage to 100%, thus completely replacing the original HDMI audio with the input audio. This completes the closed-loop processing of HDMI audio. In this closed-loop physical link, the extracted audio data flows along the high-speed synchronous clock of the I2S bus. The total instruction cycle time for the external dual DSP chips to execute algorithms such as human voice phase cancellation and frequency domain reverberation is strictly limited to the order of 2 to 5 milliseconds in the hardware architecture. This tiny audio system processing latency is directly accommodated by the main processing unit. Because it is far below the 50-millisecond lower limit threshold for human visual and auditory cortex perception of audio-visual misalignment, no additional latency alignment compensation is required for video frame buffering. The key stitching logic lies in the fact that the main processing unit MS2131S has a fixed hardware pipeline micro-delay register configured in its internal TMDS recoding engine. The delay value of this register is precisely calibrated to 3.5ms (taking the midpoint between 2ms and 5ms for dual DSP audio processing). This microsecond-level video dwell action induces the through-video pixel stream at the output endpoint to achieve phase alignment with the audio data packet that has completed the algorithm calculation and is returned via the I2S bus. This eliminates the time difference caused by the computing power overhead of the audio DSP at the underlying protocol level, so that the final output HDMI signal achieves high-fidelity synchronization with audio and video offset of less than 50ms without the need for large-capacity frame buffer compensation. The processed audio stream is directly injected back into the real-time rendering stream, which naturally achieves high-precision synchronization at the physical and sensory levels.
[0050] See the HDMI audio closed-loop processing flowchart and the HDMI audio replacement diagram in karaoke mode: Figure 3 (See the signal link diagram for recording mode:) Figure 5 ) In recording mode, the control unit configures the audio routing register of the main processing unit (MS2131S) via I2C, directly sending the I2S microphone signal from fully isolated acquisition and the accompaniment source (Bluetooth / OTG) signal to the internal mixer of the main processing unit, bypassing the first and second DSPs. Simultaneously, the control unit controls an external analog switch via GPIO to disconnect the I2S data input of the two DSPs, putting them into low-power bypass mode. The main processing unit outputs the mixed dry signal to the USB recording port (supporting 192kHz / 24bit) on one side and to the monitoring DAC on the other, achieving zero-latency monitoring. All audio processing modules are disabled in this mode to ensure the original recording quality.
[0051] Preferably, the ADC chip used in recording mode is the PCM1863, whose independent analog power supply and grounding design ensures high fidelity of dry audio acquisition. Other high-fidelity ADC chips with the same power supply architecture can also be selected.
[0052] (See the signal link diagram for live streaming mode:) Figure 6 ) In live streaming mode, the control unit configures the main processing unit via I2C to stop sending HDMI-extracted audio data to the first DSP, and controls the analog switch to directly send the loopback audio (i.e., the accompaniment played on the PC or the other party's audio during a live chat) from the USB port to the I2S3 input port of the second DSP. Regarding the specific acquisition path of the loopback audio, the main processing unit internally enables an audio endpoint controller conforming to the USB Audio Class (UAC) standard protocol. This controller, acting as a high-speed USB slave device, establishes an isochronous transfer channel with the PC master controller via a USB bus handshake, receiving... It parses the downlink audio data packets from the PC operating system, unpacks and buffers them, and physically converts them into a continuous I2S format serial digital audio stream. In practice, the UAC controller has a dynamic buffer threshold monitoring logic. When the downlink data packet rate fluctuates on the PC side, the controller adjusts the sampling frequency parameter of the UAC feedback endpoint to achieve adaptive alignment of the input flow. This dynamic optimization ensures that the accompaniment or loopback audio from the PC side will not produce pops or drops in sound during the conversion into a serial digital signal, thus providing a deterministic phase reference for the multi-channel mixing of the second DSP.
[0053] Ultimately, the signal is routed to the input of the analog switch via the corresponding functional pins of the chip. Simultaneously, the fully isolated microphone signal is converted by an ADC and sent to the I2S2 port of the second DSP. The second DSP mixes the signal processed by microphone effects (reverb, noise reduction, etc.) with the Loopback audio, outputting one path to the monitoring DAC for real-time monitoring by the broadcaster, and sending the other path back to the main processing unit. The main processing unit replaces the original HDMI audio with this mixed signal and outputs it to the HDMI OUT and USB live streaming ports. The first DSP is completely bypassed and put into sleep mode, reducing latency and power consumption.
[0054] Preferably, the second DSP in live streaming mode can be replaced with an equivalent chip with multi-channel mixing capabilities without affecting the configuration and use of the independent monitoring channel.
[0055] Example 3: Implementation of switching live streaming mode for karaoke, with DSP2 inputting accompaniment signals to Loop signals.
[0056] The main processing unit MS2131S acts as an I2S master device, outputting I2S data streams to other connected slave devices. This includes outputting the LRCK clock signal through pin 69, the SCLK clock signal through pin 70, and the OUT data signal through pin 72. Both the LRCK and SCLK clock signals are simultaneously connected to pins 13 and 14 (the chip's I2S3 clock input ports) of DSP1 (PTN1118) and DSP2 (PTN1011). The OUT data signal is simultaneously connected to pin 5 (the chip's I2S3 data input port) of PTN1118 and pin 1 (A1 terminal) of the analog switch (BL1551B). Pin 6 (the chip's I2S3 data output port) of PTN1118 is connected to pin 3 (A2 terminal) of the analog switch (BL1551B), and pin 4 (common terminal) of the analog switch (BL1551B) is connected to pin 5 (the chip's I2S3 data input port) of PTN1011.
[0057] When the device selects the karaoke mode, the main control unit controls the MS2131S to extract the background accompaniment audio from the HDMI port and send it to the output port of the I2S via the I2C interface. The audio data stream is directly transmitted through the I2S port to the I2S3 port of the PTN1118 connected to this port. After removing the human voice from the audio data, the PTN1118 transmits the signal data to the A2 pin (pin 3) of the analog switch through the output pin (pin 8) of the I2S3 port. The main control unit selects to pull the control pin (pin 6) of the analog switch low to 0, and selects the connection from A2 to the common terminal (connecting pins 3 and 4 of the chip). The audio signal after removing the human voice from the PTN1118 is transmitted through the analog switch to the I2S3 signal input port (pin 5) of the PTN1011. This realizes the link of HDMI extraction of accompaniment, removal of human voice, and then transmission to the mixing DSP for mixing in karaoke mode.
[0058] When the user switches the device to live streaming mode, the main control unit controls the MS2131S to stop extracting HDMI audio for accompaniment and instead obtain the audio played from the USB port as a loop audio source to transmit to the I2S port. It also pulls the control pin (pin 6) of the analog switch BL1551B high to enable the connection from A1 to the common terminal (connecting pins 1 and 4 of the chip). The original loop audio data stream output by the MS2131S is transmitted to the I2S3 signal input port (pin 5) of the PTN1011 through the analog switch. This achieves the link where the loop signal is directly transmitted to the mixing DSP for mixing in live streaming mode. The loop audio data stream bypasses the PTN1118, resulting in shorter live audio latency. Because the audio stream bypasses the PTN1118, the PTN1118 will be set to sleep mode in live streaming mode, disabling its function.
[0059] The audio data stream switching process of the monitoring port DAC (PCM5102A) is similar to that described above.
[0060] The circuit connection relationship for switching from karaoke to live streaming mode and from the DSP2 input accompaniment signal to the Loop signal is shown in [link to karaoke live streaming mode]. Figure 7 .
[0061] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A multi-mode audio and video processing device, characterized in that, It includes an analog front-end acquisition module, a digital isolation module, a main processing unit, a first digital signal processing unit, a second digital signal processing unit, a control unit, and an HDMI interface unit. The analog front-end acquisition module is used for impedance matching, analog amplification, and analog-to-digital conversion of microphone signals; The digital isolation module includes an I2S high-speed digital isolation unit, a control signal optocoupler isolation unit, and an isolation power supply unit, which enables complete electrical isolation between the analog front-end acquisition module and the device's digital system. The HDMI interface unit is used to receive and output HDMI audio and video signals; The main processing unit is configured to extract audio data from the input HDMI signal and send the extracted audio data to the first digital signal processing unit. The first digital signal processing unit is used to perform human voice cancellation processing on the audio data. It performs a 180-degree phase reversal on the data of one channel of the accompaniment audio and superimposes it with the data of the other channel, and performs gain smoothing by connecting a mid-frequency bandpass filter from 200Hz to 4kHz in series on the superposition path. The second digital signal processing unit is used to mix and process the audio data after human voice removal; The main processing unit uses digital mixing to configure the original HDMI audio percentage to 0% and the processed audio percentage to 100%. The device supports at least recording mode, karaoke mode, and live streaming mode. The control unit dynamically switches the audio path of each mode by configuring the internal link of the chip and the external analog switch. The audio links, functional modules, sampling rates and processing strategies of the three working modes are different. The control unit configures the internal register of the chip to enable or disable the functional module through the I2C interface according to the mode command, and simultaneously controls the external analog switch to switch the physical path of the I2S data stream through the GPIO level.
2. The multi-mode audio and video processing device according to claim 1, characterized in that, The main processing unit is an HDMI processing chip with HDMI audio extraction, digital mixing and audio re-injection functions, and the main processing unit is selected from MS2131S, MS2133 or an HDMI processing chip with equivalent functions.
3. The multi-mode audio and video processing device according to claim 1, characterized in that, The first digital signal processing unit is specifically PTN1118; the second digital signal processing unit is selected from PTN1011 or an equivalent audio DSP chip.
4. The multi-mode audio and video processing device according to claim 1, characterized in that, The analog front-end acquisition module includes an ADC chip, which uses independent analog power supply and independent grounding, so that the noise floor of the acquired signal is lower than -110dB and the total harmonic distortion is lower than 0.0001%.
5. The multi-mode audio and video processing device according to claim 4, characterized in that, The ADC chip is a PCM1863 or a high-fidelity ADC chip with equivalent functionality.
6. The multi-mode audio and video processing device according to claim 1, characterized in that, The high-speed digital isolation unit in the digital isolation module is a digital isolation chip, and the control signal optocoupler isolation unit is an optocoupler device. The digital isolation chip and optocoupler device can be replaced with functionally equivalent isolation elements.
7. The multi-mode audio and video processing device according to claim 1, characterized in that, In karaoke mode, the following complete closed-loop processing is performed: HDMI audio extraction, vocal removal, audio mixing, digital replacement, and HDMI re-injection output.
8. The multi-mode audio and video processing device according to claim 1, characterized in that, In recording mode, the audio processing module is disabled, and dry audio acquisition and automatic mixing of multiple accompaniments are performed to maintain a high sampling rate and high-fidelity output.
9. A multi-mode audio and video processing device according to claim 1, characterized in that, In live streaming mode, human voice cancellation is disabled, multi-channel audio mixing is enabled, and an independent monitoring channel is configured. The main processing unit establishes an isochronous transmission channel conforming to the UAC standard protocol via the USB bus, converts the parsed PC-side downlink audio data packets into I2S format streams, and routes them to the monitoring channel.
10. A multi-mode audio and video processing control method, applied to the apparatus of any one of claims 1-9, characterized in that, Includes the following steps: S1. After the microphone signal is processed by the analog front end, it achieves complete electrical isolation from the digital system through I2S high-speed digital isolation, control signal optocoupler isolation and isolated power supply; S2. The main processing unit extracts audio data from the input HDMI signal; S3. The extracted audio data is sent to the first digital signal processing unit to perform human voice cancellation; S4. The audio after removing human voices is sent to the second digital signal processing unit to complete mixing and sound effects processing; S5. The main processing unit completely replaces the original audio signal in the output HDMI signal with the processed audio signal through digital mixing. That is, the proportion of the original audio signal in the output is zero, and the proportion of the processed audio signal is 100%. S6. The control unit configures the internal links of the chip and the external analog switch according to the selected working mode, and dynamically switches the audio path.