Display screen sound amplification control method and device
By optimizing the weighted calculation of microphone and speaker signals using quantum circuits, the problem of linkage between display and audio equipment was solved, achieving signal fusion and feedback suppression between speakers and microphones, thus improving audio quality and clarity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING MYSHER TECH
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-10
AI Technical Summary
In existing sound reinforcement solutions, the monitor and audio equipment cannot be linked, which can easily cause acoustic feedback and howling between the speakers and microphones, affecting the normal conduct of the meeting.
By constructing parameterized quantum circuits, the weighted calculation of microphone and speaker signals is optimized. Combined with inverse filtering and feedback suppression functions, dynamic fusion and feedback suppression of speaker and microphone signals are achieved. Quantum computing is used to assist in optimizing the weighting coefficients to improve audio signal quality.
It effectively suppresses acoustic feedback howling, achieves signal fusion between speakers and microphones, improves audio quality and clarity in meetings, and adapts to complex acoustic environments.
Smart Images

Figure CN121842584A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of audio processing, artificial intelligence, and quantum computing, and in particular to a method and apparatus for controlling the amplification of a display screen. Background Technology
[0002] With the development of information technology and intelligentization in conferences, large display screens have become one of the core devices in various educational and conference scenarios. Audio amplification, as a crucial link in the transmission of conference information, directly affects the efficiency of conference communication. Existing amplification solutions mostly consist of microphones used in environments such as external speakers, central control units, and power amplifiers.
[0003] These solutions typically involve external devices that occupy a large area and require considerable space, which may not be sufficient for small meeting rooms. Furthermore, the large screen and amplification equipment cannot work in tandem, as external speakers and microphones cannot be directly connected to the conference or educational screens. Often, the screen and speakers play content independently, without coordinated control or playback. This can easily lead to feedback between the speakers and microphones, causing howling and disrupting the normal proceedings of the meeting.
[0004] Therefore, developing a display screen amplification control method and device is key to improving the acoustic feedback effect between speakers and microphones in meetings. Summary of the Invention
[0005] The present invention aims to solve these problems by providing a display screen amplification control method and apparatus to meet the acoustic feedback effect between speakers and microphones in a conference.
[0006] This invention provides a method for controlling the sound amplification of a display screen, comprising the following steps:
[0007] Step 1: Acquire N raw audio signals from microphones using the microphone acquisition circuit, where N is an integer greater than 0 and not exceeding 8; acquire M raw audio data from speakers using the speaker acquisition circuit, where M is an integer greater than 0 and not exceeding 2.
[0008] Step 2: Preprocess the original audio signals from the N microphones and the original audio data from the M speakers to obtain N+M preprocessed signals;
[0009] Step 3: Perform weighted calculations on the N+M preprocessed signals to obtain the target audio signal;
[0010] Step four: According to the control instructions of the large screen control circuit, determine whether it is allowed to release the target audio signal; if it is allowed to release, output the target audio signal to the power amplifier circuit for amplification; and play the amplified audio signal through the speaker circuit.
[0011] Step two includes:
[0012] The original audio signals from the N microphones are subjected to format conversion and gain adjustment.
[0013] The original audio data of the M-channel loudspeakers is subjected to voltage division processing, reducing its voltage from 17-28V to 0-5V;
[0014] The N+M channels of signals are subjected to denoising, noise cancellation, and noise suppression processing.
[0015] The preprocessed signal is used as the input for the weighted calculation.
[0016] The weighted calculation is performed using the following formula:
[0017]
[0018] in:
[0019] Sout represents the weighted fused target audio signal;
[0020] Smici represents the pre-processed audio signal from the i-th microphone;
[0021] Sspkj represents the pre-processed audio signal acquired by the j-th speaker;
[0022] αi represents the weighting coefficient of the i-th microphone signal, which is dynamically adjusted according to the meeting scenario;
[0023] βj represents the weighting coefficient of the j-th speaker signal, used for feedback suppression and signal fusion;
[0024] The weighting coefficients satisfy
[0025] The step of performing denoising, noise cancellation, and noise suppression processing on the N+M channels includes, in which, the denoising process involves constructing an inverse filtering model with the following filtering function:
[0026]
[0027] in:
[0028] H(f) represents the filter response at frequency f;
[0029] G represents the noise reduction gain coefficient, with a value ranging from 10dB to 35dB;
[0030] Pnoise(f) represents the noise power spectral density at frequency f;
[0031] Psignal(f) represents the signal power spectral density at frequency f;
[0032] ∈ is a very small positive number used to prevent the denominator from being zero;
[0033] The filtered signal is used as the input for the weighted calculation.
[0034] The step four is followed by an audio feedback suppression step, wherein the audio feedback suppression step uses the following howling suppression function:
[0035] F suppress (t)=γ·∫0 t e -λ(t-τ) ·|S fb (τ)|·sgn(S fb (τ))dτ where:
[0036] Fsuppress(t) represents the amount of feedback suppression at time t;
[0037] Sfb(τ) represents the feedback signal at time τ, which is obtained by the speaker acquisition circuit;
[0038] γ represents the inhibition strength coefficient;
[0039] λ represents the attenuation factor, which controls the attenuation rate of the historical feedback signal;
[0040] sgn(·) is the sign function, used to determine the phase of the feedback signal;
[0041] Used to achieve closed-loop howling suppression.
[0042] In step three, the N+M preprocessed signals are weighted and calculated to obtain the target audio signal;
[0043] The optimization process for the weight coefficients (αi, βj) includes the following steps:
[0044] A parameterized quantum circuit is constructed, mapping the weighted coefficient optimization problem to a ground state search problem for qubits; wherein the parameters of the quantum circuit include the angle θk of the quantum rotation gate and the coupling strength φkl of the entanglement gate;
[0045] The time-domain or frequency-domain characteristics of the preprocessed N+M audio signals are converted into the initial input state |ψin> of the quantum circuit through a classical coding layer;
[0046] Running the parameterized quantum circuit on the quantum processor yields the output state |ψout>.
[0047] Measure the expected value of the output state <ψout|H|ψout>, where H is the Hamiltonian constructed based on the audio signal fusion quality index;
[0048] Using a classical optimizer, the parameters θk and φkl of the quantum circuit are iteratively updated with the goal of minimizing the desired value;
[0049] When the expected value converges to below the threshold, the optimal output state is reached. Decode the corresponding optimal set of weight coefficients
[0050] The target audio signal Sout is obtained by weighting the data using the set of optimal weighting coefficients.
[0051] According to another aspect of the present invention, the present invention also provides a display screen amplification control device, comprising:
[0052] The acquisition module is configured to acquire N raw audio signals from microphones via a microphone acquisition circuit, where N is an integer greater than 0 and not exceeding 8; and to acquire M raw audio data from speakers via a speaker acquisition circuit, where M is an integer greater than 0 and not exceeding 2.
[0053] The preprocessing module is configured to preprocess the original audio signals from the N microphones and the original audio data from the M speakers to obtain N+M preprocessed signals.
[0054] The weighted calculation module is configured to perform weighted calculations on the N+M preprocessed signals to obtain the target audio signal;
[0055] The judgment and processing module is configured to determine whether the release of the target audio signal is allowed according to the control command of the large screen control circuit; if the release is allowed, the target audio signal is output to the power amplifier circuit for amplification; and the amplified audio signal is played through the speaker circuit.
[0056] The preprocessing module includes:
[0057] The original audio signals from the N microphones are subjected to format conversion and gain adjustment.
[0058] The original audio data of the M-channel loudspeakers is subjected to voltage division processing, reducing its voltage from 17-28V to 0-5V;
[0059] The N+M channels of signals are subjected to denoising, noise cancellation, and noise suppression processing.
[0060] The preprocessed signal is used as the input for the weighted calculation.
[0061] The weighted calculation uses the following formula:
[0062]
[0063] in:
[0064] Sout represents the weighted fused target audio signal;
[0065] Smici represents the pre-processed audio signal from the i-th microphone;
[0066] Sspkj represents the pre-processed audio signal acquired by the j-th speaker;
[0067] αi represents the weighting coefficient of the i-th microphone signal, which is dynamically adjusted according to the meeting scenario;
[0068] βj represents the weighting coefficient of the j-th speaker signal, used for feedback suppression and signal fusion;
[0069] The weighting coefficients satisfy
[0070] The step of performing denoising, noise cancellation, and noise suppression processing on the N+M channels of signals includes, in which the denoising process involves constructing an inverse filtering model, the filtering function of which is:
[0071]
[0072] in:
[0073] H(f) represents the filter response at frequency f;
[0074] G represents the noise reduction gain coefficient, with a value ranging from 10dB to 35dB;
[0075] Pnoise(f) represents the noise power spectral density at frequency f;
[0076] Psignal(f) represents the signal power spectral density at frequency f;
[0077] ∈ is a very small positive number used to prevent the denominator from being zero;
[0078] The filtered signal is used as the input for the weighted calculation.
[0079] It also includes an audio feedback suppression module, which is configured to include an audio feedback suppression step, wherein the audio feedback suppression step is controlled by the following howling suppression function:
[0080] F suppress (t)=γ·∫0 t e -λ(t-τ) ·||S fb (τ)|·sgn(S fb (τ))dτ where:
[0081] Fsuppress(t) represents the amount of feedback suppression at time t;
[0082] Sfb(τ) represents the feedback signal at time τ, which is obtained by the speaker acquisition circuit;
[0083] γ represents the inhibition strength coefficient;
[0084] λ represents the attenuation factor, which controls the attenuation rate of the historical feedback signal;
[0085] sgn(·) is the sign function, used to determine the phase of the feedback signal;
[0086] Used to achieve closed-loop howling suppression.
[0087] The weighted calculation module includes a quantum processing unit, which includes:
[0088] A parameterized quantum circuit is constructed, mapping the weighted coefficient optimization problem to a ground state search problem for qubits; wherein the parameters of the quantum circuit include the angle θk of the quantum rotation gate and the coupling strength φkl of the entanglement gate;
[0089] The time-domain or frequency-domain characteristics of the preprocessed N+M audio signals are converted into the initial input state |ψin> of the quantum circuit through a classical coding layer;
[0090] Running the parameterized quantum circuit on the quantum processor yields the output state |ψout>.
[0091] Measure the expected value of the output state <ψout|H|ψout>, where H is the Hamiltonian constructed based on the audio signal fusion quality index;
[0092] Using a classical optimizer, the parameters θk and φkl of the quantum circuit are iteratively updated with the goal of minimizing the desired value;
[0093] When the expected value converges to below the threshold, the optimal output state is reached. Decode the corresponding optimal set of weight coefficients
[0094] The target audio signal Sout is obtained by weighting the data using the set of optimal weighting coefficients.
[0095] The purpose of this invention is to provide a display screen amplification control method, comprising the following steps:
[0096] Step 1: Acquire N raw audio signals from microphones via a microphone acquisition circuit, where N is an integer greater than 0 and not exceeding 8; acquire M raw audio data from speakers via a speaker acquisition circuit, where M is an integer greater than 0 and not exceeding 2. Step 2: Preprocess the N raw audio signals from microphones and the M raw audio data from speakers to obtain N+M preprocessed signals. Step 3: Perform weighted calculations on the N+M preprocessed signals to obtain the target audio signal. Step 4: Based on the control command from the large screen control circuit, determine whether to allow the release of the target audio signal; if release is allowed, output the target audio signal to the power amplifier circuit for amplification; play the amplified audio signal through the speaker circuit.
[0097] Noise reduction is achieved by suppressing noise in the target acquisition signal; anti-feedback function is achieved by processing the target acquisition signal.
[0098] The above description is merely an overview of this solution. In order to better understand the technical means of this invention and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this invention more apparent and understandable, specific embodiments of this invention are described below. Attached Figure Description
[0099] To more clearly illustrate this technical solution, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this technology. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0100] Figure 1 This is a schematic flowchart illustrating a first embodiment of a display screen amplification control method of the present invention;
[0101] Figure 2 This invention schematically illustrates a circuit principle block diagram of a display screen amplification control method;
[0102] Figure 3 This is a schematic flowchart illustrating Embodiment 2 of the display screen amplification control method of the present invention;
[0103] Figure 4 This invention schematically illustrates a structural block diagram of a display screen amplification control device;
[0104] Figure 5 This invention schematically illustrates the structure of another display screen amplification control device. Detailed Implementation
[0105] The embodiments of the present invention will be described in detail below, but the present invention can be implemented in many different ways as defined and covered by the claims.
[0106] To make the technical problems, technical solutions, and beneficial effects to be solved by this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this application.
[0107] It should be noted that when a component is referred to as being "fixed to" or "set on" another component, it can be directly on or indirectly on that other component. When a component is referred to as being "connected to" another component, it can be directly connected to or indirectly connected to that other component.
[0108] It should be understood that the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.
[0109] like Figures 1-2 As shown in Embodiment 1, a method for controlling the sound amplification of a display screen includes the following steps:
[0110] Step 1: Acquire N raw audio signals from microphones using the microphone acquisition circuit, where N is an integer greater than 0 and not exceeding 8; acquire M raw audio data from speakers using the speaker acquisition circuit, where M is an integer greater than 0 and not exceeding 2.
[0111] Step 2: Preprocess the original audio signals from the N microphones and the original audio data from the M speakers to obtain N+M preprocessed signals;
[0112] Step 3: Perform weighted calculations on the N+M preprocessed signals to obtain the target audio signal;
[0113] Step four: According to the control instructions of the large screen control circuit, determine whether it is allowed to release the target audio signal; if it is allowed to release, output the target audio signal to the power amplifier circuit for amplification; and play the amplified audio signal through the speaker circuit.
[0114] This embodiment provides a basic flow of a display screen amplification control method. For example... Figure 1 and Figure 2As shown, this method first acquires N channels of raw audio signals, such as N=4 or 8, through a microphone array arranged at the bottom of the large screen, for example, a 2×4 or 1×4 arrangement. Simultaneously, it acquires M channels of raw audio data from the output of the built-in speaker of the large screen through a speaker acquisition circuit, for example, M=1 or 2, which are used for subsequent feedback suppression.
[0115] Next, the audio processing circuit preprocesses the acquired N+M channels of signals. The preprocessed N+M channels of signals are then fed into the weighted algorithm unit for fusion calculation to obtain the target audio signal.
[0116] Subsequently, the large-screen control circuit, typically the main control chip of the large screen, sends control commands to the audio processing circuit via a USB or UART interface. These commands determine whether to allow the release of the target audio signal for amplification; for example, if the system's mute mode is detected, release is prohibited.
[0117] Once permission is granted, the audio processing circuit outputs the target audio signal to the power amplifier circuit via the lineout line. The power amplifier circuit amplifies the signal, which is then played back by the speaker circuit integrated with the large screen, achieving a low-latency, integrated sound amplification effect.
[0118] Step two includes:
[0119] The original audio signals from the N microphones are subjected to format conversion and gain adjustment.
[0120] The original audio data of the M-channel loudspeakers is subjected to voltage division processing, reducing its voltage from 17-28V to 0-5V;
[0121] The N+M channels of signals are subjected to denoising, noise cancellation, and noise suppression processing.
[0122] The preprocessed signal is used as the input for the weighted calculation.
[0123] Specifically, for the N raw audio signals acquired by the microphones, preprocessing includes format conversion and gain adjustment. Specifically, if the microphone is a digital microphone, the audio format conversion circuit converts its PDM (Pulse Density Modulation) signal into an I2S or TDM signal supported by the audio main control chip; if it is an analog microphone, the signal preprocessing unit first differentially amplifies it before performing analog-to-digital conversion. Simultaneously, the system automatically adjusts the gain based on the ambient sound pressure level to ensure appropriate signal strength.
[0124] For the M-channel signals acquired by the speaker, the core of its preprocessing is voltage division. Since the signal voltage directly acquired from the back end of the speaker amplifier is relatively high, such as 17-28V, while the operating voltage of the audio processing circuit is usually 3.3V or 5V, it is necessary to use a voltage divider circuit, such as a resistor divider network, to reduce it to a safe range of 0-5V for subsequent processing.
[0125] Finally, the preprocessing stage also includes unified denoising, noise cancellation, and noise suppression. These processed N+M signals, as clean input sources, are fed into the next stage of weighted calculation unit.
[0126] For a detailed circuit block diagram, please refer to [link / reference]. Figure 1 and Figure 2 As shown, it includes a microphone acquisition circuit 11, an audio processing circuit 12, a large screen control circuit 13, a speaker acquisition circuit 14, a power amplifier circuit 15, and a speaker circuit 16.
[0127] The microphone acquisition circuit 11 is characterized by acquiring N (0 < N ≥ 8) audio signals to be processed; the audio processing circuit 12 is connected to the microphone acquisition circuit 11, the large screen control circuit 13, and the speaker acquisition circuit 14, and performs preprocessing and weighted calculation on the N audio signals to be processed acquired by the microphone acquisition circuit and the M audio signals played by the speakers, and then outputs the processed audio.
[0128] The large screen control circuit 13 is connected to the audio processing circuit 12, and controls the audio processing circuit to determine whether to amplify the sound and the actions of other audio processors.
[0129] The speaker acquisition circuit 14 is connected to the audio processing circuit 12 to acquire the audio signals played by M (0 < M ≥ 2) speakers.
[0130] The power amplifier circuit 15 is connected to the audio processing circuit 12 and the large screen audio processing circuit to amplify the target audio released by the audio processing circuit; the speaker circuit 16 is connected to the power amplifier circuit 15 to play the audio signal amplified by the power amplifier circuit.
[0131] The audio processing circuit 12 is further configured to perform noise suppression and frequency shifting on the audio signal acquired by the microphone acquisition circuit 11; the power amplifier circuit 15 is further configured to amplify the target acquisition signal after noise suppression and frequency shifting.
[0132] Noise reduction is achieved by suppressing noise in the target acquisition signal; anti-feedback function is achieved by shifting the frequency of the target acquisition signal.
[0133] The weighted calculation uses the following formula:
[0134]
[0135] in:
[0136] Sout represents the weighted fused target audio signal;
[0137] Smici represents the pre-processed audio signal from the i-th microphone;
[0138] Sspkj represents the pre-processed audio signal acquired by the j-th speaker;
[0139] αi represents the weighting coefficient of the i-th microphone signal, which is dynamically adjusted according to the meeting scenario;
[0140] βj represents the weighting coefficient of the j-th speaker signal, used for feedback suppression and signal fusion;
[0141] The weighting coefficients satisfy This is to ensure the stability of the output signal amplitude and avoid signal overload or distortion caused by weighted calculation.
[0142] Specifically,
[0143] The system has multiple pre-stored scene modes, and the main control chip switches modes based on instructions from the large screen control circuit or automatic detection:
[0144] Scenario 1, Large-scale speech mode: The system identifies the direction of the sound source and increases the weight of the microphone channel pointing towards the speaker's position, for example, increasing its αi from 0.125 to 0.3, while decreasing the weight of microphones in other directions to achieve directional sound pickup and enhance speech clarity. βj remains at its initial negative value.
[0145] Scenario 2, Remote Video Conferencing Mode: To ensure clear transmission of both local participants' voices and remote audio, the system employs a balanced αi weight. Simultaneously, to more effectively prevent echoes caused by remote audio being picked up by the microphone and transmitted back to the remote location after being played through the local speaker, the system appropriately increases the negative value of βj, for example, adjusting it from -0.1 to -0.15, to enhance echo cancellation.
[0146] Scenario 3, Small Discussion Mode: The system uses equal weights for αi to encourage free speech from all parties. In this mode, the negative value of βj can be slightly reduced, for example, adjusted to -0.05, because the speaker volume is usually low in this mode, the risk of feedback is low, and excessive elimination may affect the naturalness of the speech.
[0147] A concrete example could be a system with 8 microphones (N=8) and 2 speaker acquisition channels (M=2), where initially all αi = 0.125 and βj = -0.1. When the system switches to a large-scale presentation mode and detects the speaker in front of microphones 1 and 2, α1 and α2 are adjusted to 0.3, the remaining αi are adjusted to 0.05, while β1 and β2 remain at -0.1. Calculation verification: 0.3 + 0.3 + 0.05 * 6 + (-0.1) * 2 = 1, which satisfies the normalization condition.
[0148] The step of performing denoising, noise cancellation, and noise suppression processing on the N+M channels includes, in which, the denoising process involves constructing an inverse filtering model with the following filtering function:
[0149]
[0150] in:
[0151] H(f) represents the filter response at frequency f;
[0152] G represents the noise reduction gain coefficient, with a value ranging from 10dB to 35dB;
[0153] Pnoise(f) represents the noise power spectral density at frequency f;
[0154] Psignal(f) represents the signal power spectral density at frequency f;
[0155] ∈ is a very small positive number used to prevent the denominator from being zero;
[0156] The filtered signal is used as the input for the weighted calculation.
[0157] Specifically,
[0158] Specific application process:
[0159] The audio processing circuit divides the N+M channels of the time domain signal into frames, for example, each frame is 20ms, and performs a Fast Fourier Transform (FFT) to convert them to the frequency domain, obtaining S(f).
[0160] The VAD module determines whether the current frame is a silent frame. If so, it updates Pnoise(f) according to the recursive averaging method described above.
[0161] The noise reduction gain G is determined based on the overall SNR of the current frame.
[0162] Substitute Pnoise(f), Psignal(f), G, and ∈ into the formula to calculate the filter response H(f) of the current frame.
[0163] Multiplying the original spectrum S(f) by H(f) yields the denoised spectrum S′(f).
[0164] The inverse fast Fourier transform (IFFT) is performed on S′(f), and the signal is restored to the denoised time domain signal using the overlay and addition method for subsequent weighted calculation.
[0165] Furthermore, according to another aspect of the invention, such as Figure 3 As shown in Embodiment 2 of the present invention, after step four, the present invention further includes step five, an audio feedback suppression step, wherein the audio feedback suppression step uses the following howling suppression function:
[0166] F suppress (t)=γ·∫0 t e -λ(t-τ) ·|S fb (τ)|·sgn(S fb (τ))dτ
[0167] in:
[0168] Fsuppress(t) represents the amount of feedback suppression at time t;
[0169] Sfb(τ) represents the feedback signal at time τ, which is obtained by the speaker acquisition circuit;
[0170] γ represents the inhibition strength coefficient;
[0171] λ represents the attenuation factor, which controls the attenuation rate of the historical feedback signal;
[0172] sgn(·) is the sign function, used to determine the phase of the feedback signal;
[0173] Used to achieve closed-loop howling suppression.
[0174] Specifically,
[0175] The system analyzes the spectrum of Sfb(τ) in real time. If the peak energy of a specific frequency f0 is found to increase by more than 20dB continuously within 5 consecutive audio frames (approximately 100ms), then f0 is determined to be a howling frequency.
[0176] For the howling frequency f0, extract the data of its signal Sfb(τ) within the time window, substitute it into the above integral formula, and calculate the instantaneous suppression amount Fsuppress(t) for f0.
[0177] The suppression factor Fsuppress(t) is converted in real time into a depth parameter of a narrowband notch filter for frequency f0. This notch filter is dynamically inserted into the output path of the audio processing circuit, producing an attenuation of depth Fsuppress(t) in dB at f0.
[0178] After suppression is applied, the system continues to monitor the energy at f0. If the energy decreases, γ and λ are maintained; if the energy continues to rise, γ is increased to strengthen the suppression. When the energy at f0 remains below the safety threshold for a relatively long time, such as 2 seconds, the suppression is gradually reduced, allowing the notch filter depth to recover slowly until it is turned off, in order to avoid unnecessary sound quality damage.
[0179] This closed-loop processing ensures that the system can quickly, accurately, and adaptively eliminate feedback while maintaining the quality of the original audio to the greatest extent possible.
[0180] According to another aspect of the present invention, in embodiment four, in step three above, the N+M preprocessed signals are weighted to obtain the target audio signal; the optimization process of the weight coefficients (αi, βj) includes the following steps:
[0181] A parameterized quantum circuit is constructed, mapping the weighted coefficient optimization problem to a ground state search problem for qubits; wherein the parameters of the quantum circuit include the angle θk of the quantum rotation gate and the coupling strength φkl of the entanglement gate;
[0182] The time-domain or frequency-domain characteristics of the preprocessed N+M audio signals are converted into the initial input state |ψin> of the quantum circuit through a classical coding layer;
[0183] Running the parameterized quantum circuit on the quantum processor yields the output state |ψout>.
[0184] Measure the expected value of the output state <ψout|H|ψout>, where H is the Hamiltonian constructed based on the audio signal fusion quality index;
[0185] Using a classical optimizer, the parameters θk and φkl of the quantum circuit are iteratively updated with the goal of minimizing the desired value;
[0186] When the expected value converges to below the threshold, the optimal output state is reached. Decode the corresponding optimal set of weight coefficients
[0187] The target audio signal Sout is obtained by weighting the data using the set of optimal weighting coefficients.
[0188] The purpose of Example 4 is to implement a weight coefficient optimization method for weighted fusion of audio signals using a hybrid quantum-classical computing model. The core of this method is to transform the weight allocation, which relies on empirical formulas or fixed scene patterns in traditional methods, into a dynamic and adaptive optimization problem solved within a quantum computing framework, aiming to achieve better speech intelligibility and fusion results in complex acoustic environments.
[0189] This system adds a communication interface to the existing audio processing circuit, which can be either real quantum hardware or a high-performance simulator. The main control chip of the audio processing circuit, acting as the classical computing component, is responsible for signal preprocessing and final weighted calculation, while the optimization of the weighting coefficients is assisted by the quantum coprocessor.
[0190] The optimization objective is defined as finding a set of weight coefficients {αi, βj} such that the fused target signal Sout simultaneously satisfies the following conditions: maximizing speech intelligibility (evaluated using a pre-trained deep learning model); minimizing signal distortion (measured by calculating the mean square error with the original microphone signal set); and maximizing feedback suppression (evaluated by assessing the residual energy of the signal in the output obtained from the loudspeaker).
[0191] The weighted sum of these three objectives is constructed into a cost function C({αi,βj}). Optimizing the weight coefficients is transformed into finding the {αi,βj} combination that minimizes C. The specific steps are as follows: executed by the audio main control chip, the main control chip extracts specific features of the N-channel microphone signals Smici and the M-channel speaker acquisition signals Sspkj of the current frame, such as the short-time energy, zero-crossing rate, and the first few coefficients of the Mel-frequency cepstral coefficients (MFCC) of each signal, forming a feature vector.
[0192] The feature vector is embedded through a classic embedding layer, such as a small neural network or a transformation matrix. Mapped to a set of angles This set of angles is used to initialize the initial quantum state |ψin> of the parameterized quantum circuit (VQC), typically achieved through a series of Ry rotation gates for each qubit. Where Q is the number of qubits.
[0193] Running on a quantum coprocessor, a parameterized quantum circuit (VQC) is executed: a VQC consists of a series of gates with adjustable parameters. Its structure typically includes an encoding layer: as described above, an initial state is prepared using input features. The function of rotating gates with adjustable parameters, such as Ry(θk), Rz(θk+1), and fixed entanglement gates, such as the controlled-NOT gate CNOT, can be viewed as introducing coupling φkl alternately to form a hypothetical state, i.e., Ansatz.
[0194] The circuit operates in the initial state |ψin> and outputs the final quantum state. It depends on all adjustable parameters and fixed entanglement structure
[0195] We need to measure the output state |ψout>. We need to define a Hamiltonian with the expected value...<h>=<ψout|H|ψout> should be associated with the classical cost function C. A simple mapping approach is to design H as a diagonal operator whose eigenvalues correspond to "penalties" for different combinations of weight coefficients.
[0196] In fact, a more feasible approach is to measure the expected value of each qubit on the Z-axis from |ψout>. <zq>(Values range from [-1, +1]). These Q measurements are then passed through a classic decoding layer, such as a linear transformation, mapping them back to a real vector. This vector, after Softmax normalization, is interpreted as the optimal set of weight coefficients we are seeking. Candidate solutions.
[0197] The audio control chip uses this set of candidate weight coefficients to actually calculate the Sout of a frame, and calculates the classic cost function value C based on the three objectives mentioned above: sharpness, distortion, and feedback suppression.
[0198] Classical optimizers, such as gradient descent, ADAM, or gradient-free optimizers specifically designed for VQC, aim to minimize C and generate a new set of VQC parameters. Send this new set of parameters to the quantum coprocessor and repeat the above steps. After several iterations, such as 50-100 times, the iteration stops when the value of the cost function C no longer decreases significantly or reaches a preset threshold. The resulting set of weight coefficients is then obtained. This is the approximate optimal solution for the current audio frame under the current acoustic environment.
[0199] In practical applications, to balance computational latency and performance, not every audio frame undergoes a complete quantum optimization. The system continuously monitors the rate of change of the acoustic environment. When the environment is stable, the weights optimized in the previous frame are used; when a significant change in the acoustic environment is detected, such as movement, opening a door, or the appearance of a new noise source, a complete quantum optimization process is triggered to quickly and adaptively find new optimal weights.
[0200] In this second embodiment, by introducing a quantum hybrid model, this method can find the global or near-global optimal solution more efficiently in a huge, non-convex weight coefficient search space. Compared with traditional fixed strategies or classical optimization algorithms, it can significantly improve the clarity and naturalness of speech amplification in complex and dynamically changing conference room environments, while also having stronger robustness.
[0201] like Figure 4 and Figure 5 As shown, according to another aspect of the present invention, the present invention also provides a display screen amplification control device, comprising:
[0202] The acquisition module is configured to acquire N raw audio signals from microphones via a microphone acquisition circuit, where N is an integer greater than 0 and not exceeding 8; and to acquire M raw audio data from speakers via a speaker acquisition circuit, where M is an integer greater than 0 and not exceeding 2.
[0203] The preprocessing module is configured to preprocess the original audio signals from the N microphones and the original audio data from the M speakers to obtain N+M preprocessed signals.
[0204] The weighted calculation module is configured to perform weighted calculations on the N+M preprocessed signals to obtain the target audio signal;
[0205] The judgment and processing module is configured to determine whether the release of the target audio signal is allowed according to the control command of the large screen control circuit; if the release is allowed, the target audio signal is output to the power amplifier circuit for amplification; and the amplified audio signal is played through the speaker circuit.
[0206] The preprocessing module includes:
[0207] The original audio signals from the N microphones are subjected to format conversion and gain adjustment.
[0208] The original audio data of the M-channel loudspeakers is subjected to voltage division processing, reducing its voltage from 17-28V to 0-5V;
[0209] The N+M channels of signals are subjected to denoising, noise cancellation, and noise suppression processing.
[0210] The preprocessed signal is used as the input for the weighted calculation.
[0211] The weighted calculation uses the following formula:
[0212]
[0213] in:
[0214] Sout represents the weighted fused target audio signal;
[0215] Smici represents the pre-processed audio signal from the i-th microphone;
[0216] Sspkj represents the pre-processed audio signal acquired by the j-th speaker;
[0217] αi represents the weighting coefficient of the i-th microphone signal, which is dynamically adjusted according to the meeting scenario;
[0218] βj represents the weighting coefficient of the j-th speaker signal, used for feedback suppression and signal fusion;
[0219] The weighting coefficients satisfy
[0220] The step of performing denoising, noise cancellation, and noise suppression processing on the N+M channels of signals includes, in which the denoising process involves constructing an inverse filtering model, the filtering function of which is:
[0221]
[0222] in:
[0223] H(f) represents the filter response at frequency f;
[0224] G represents the noise reduction gain coefficient, with a value ranging from 10dB to 35dB;
[0225] Pnoise(f) represents the noise power spectral density at frequency f;
[0226] Psignal(f) represents the signal power spectral density at frequency f;
[0227] ∈ is a very small positive number used to prevent the denominator from being zero;
[0228] The filtered signal is used as the input for the weighted calculation.
[0229] Among them, such as Figure 5 As shown, it also includes an audio feedback suppression module, which is configured as an audio feedback suppression step, wherein the audio feedback suppression step is controlled by the following howling suppression function:
[0230] F suppress (t)=γ·∫0 t e -λ(t-τ) ·|S fb (τ)|·sgn(S fb (τ))dτ
[0231] in:
[0232] Fsuppress(t) represents the amount of feedback suppression at time t;
[0233] Sfb(τ) represents the feedback signal at time τ, which is obtained by the speaker acquisition circuit;
[0234] γ represents the inhibition strength coefficient;
[0235] λ represents the attenuation factor, which controls the attenuation rate of the historical feedback signal;
[0236] sgn(·) is the sign function, used to determine the phase of the feedback signal;
[0237] Used to achieve closed-loop howling suppression.
[0238] The weighted calculation module includes a quantum processing unit, which includes:
[0239] A parameterized quantum circuit is constructed, mapping the weighted coefficient optimization problem to a ground state search problem for qubits; wherein the parameters of the quantum circuit include the angle θk of the quantum rotation gate and the coupling strength φkl of the entanglement gate;
[0240] The time-domain or frequency-domain characteristics of the preprocessed N+M audio signals are converted into the initial input state |ψin> of the quantum circuit through a classical coding layer;
[0241] Running the parameterized quantum circuit on the quantum processor yields the output state |ψout>.
[0242] Measure the expected value of the output state <ψout|H|ψout>, where H is the Hamiltonian constructed based on the audio signal fusion quality index;
[0243] Using a classical optimizer, the parameters θk and φkl of the quantum circuit are iteratively updated with the goal of minimizing the desired value;
[0244] When the expected value converges to below the threshold, the optimal output state is reached. Decode the corresponding optimal set of weight coefficients
[0245] The target audio signal Sout is obtained by weighting the data using the set of optimal weighting coefficients.
[0246] The specific embodiments of the apparatus claims can be referred to the above-described method embodiments, and will not be repeated here to avoid repetition.
[0247] The above are merely preferred embodiments of the present invention and are not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.< / zq> < / h>
Claims
1. A display screen amplification control method, characterized by, The method comprises the following steps: Step 1: collecting N-channel microphone original audio signals through a microphone acquisition circuit, wherein N is an integer greater than 0 and not more than 8; collecting M-channel loudspeaker original audio data through a loudspeaker acquisition circuit, wherein M is an integer greater than 0 and not more than 2; Step 2: preprocessing the N-channel microphone original audio signals and the M-channel loudspeaker original audio data to obtain N+M-channel preprocessed signals; Step 3: performing weighted calculation on the N+M-channel preprocessed signals to obtain a target audio signal; Step 4: judging whether to release the target audio signal according to a control instruction of a large-screen control circuit; if the release is allowed, outputting the target audio signal to a power amplifier circuit for amplification; and playing the amplified audio signal through a loudspeaker circuit.
2. The display screen amplification control method of claim 1, wherein, The step 2 comprises: format conversion and gain adjustment on the N-channel microphone original audio signals; voltage division processing on the M-channel loudspeaker original audio data, dividing the voltage from 17-28V to 0-5V; noise removal, noise elimination and noise suppression processing on the N+M-channel signals; The preprocessed signals are used as inputs for the weighted calculation.
3. The display screen amplification control method of claim 2, wherein, The weighted calculation adopts the following formula: wherein: Sout represents the target audio signal after weighted fusion; Smici represents the preprocessed audio signal of the i-th microphone; Sspkj represents the preprocessed audio signal of the j-th loudspeaker; αi represents the weight coefficient of the i-th microphone signal, which is dynamically adjusted according to the conference scene; βj represents the weight coefficient of the j-th loudspeaker signal, which is used for feedback suppression and signal fusion; The weight coefficients satisfy 4. The display screen amplification control method of claim 3, wherein, The noise removal processing in the step of processing the N+M-channel signals to remove noise, eliminate noise and suppress noise comprises constructing an inverse filter model, and the filter function is: wherein: H(f) represents the filter response at frequency f; G represents the noise reduction gain coefficient, and the value range is 10dB-35dB; Pnoise(f) represents the noise power spectral density at frequency f; Psignal(f) represents the signal power spectral density at frequency f; ∈ is a very small positive number, which is used to prevent the denominator from being zero; The filtered signals are used as inputs for the weighted calculation.
5. The display screen amplification control method of claim 4, wherein, After the step 4, an audio feedback suppression step is further included, wherein the control audio feedback suppression step adopts the following howling suppression function: F suppress (t) = γ · ∫0 t e -λ(t-τ) ·||S fb (τ)|·sgn(S fb (τ)) dτ where: Fsuppress(t) represents the feedback suppression amount at time t; Sfb(τ) represents the feedback signal at time τ, which is obtained by the loudspeaker acquisition circuit; γ represents the suppression intensity coefficient; λ represents the attenuation factor, which controls the decay rate of the historical feedback signal; sgn(·) is a sign function, which is used to determine the phase of the feedback signal; which is used to realize closed-loop howling suppression.
6. The display screen amplification control method of claim 5, wherein, The step 3 performs weighted calculation on the N+M-channel preprocessed signals to obtain a target audio signal; In the step, the optimization process of the weight coefficients (αi, βj) comprises the following steps: constructing a parameterized quantum circuit to map the weighted coefficient optimization problem to a ground state search problem of quantum bits; wherein the parameters of the quantum circuit include the angle θk of the quantum rotation gate and the coupling strength φkl of the entanglement gate; Converting time domain or frequency domain features of the preprocessed N+M audio signals into initial input states of a quantum circuit through a classical encoding layer; Running the parameterized quantum circuit on a quantum processor to obtain an output state; Measuring an expected value of the output state, wherein H is a Hamiltonian constructed according to an audio signal fusion quality index; Using a classical optimizer to iteratively update parameters θk and φkl of the quantum circuit to minimize the expected value; when the expected value converges below the threshold value, the corresponding optimal weight coefficient set is decoded from the optimal output state Using the optimal set of weight coefficients to perform weighted calculation to obtain a target audio signal Sout.
7. A display screen amplification control device, characterized by, It comprises: A collection module configured to collect N-channel microphone raw audio signals through a microphone acquisition circuit, wherein N is an integer greater than 0 and not more than 8; and collect M-channel loudspeaker raw audio data through a loudspeaker acquisition circuit, wherein M is an integer greater than 0 and not more than 2; A preprocessing module configured to preprocess the N-channel microphone raw audio signals and M-channel loudspeaker raw audio data to obtain N+M preprocessed signals; A weighted calculation module configured to perform weighted calculation on the N+M preprocessed signals to obtain a target audio signal; A judgment processing module configured to determine whether to release the target audio signal according to a control instruction of a large screen control circuit; if release is allowed, the target audio signal is output to a power amplifier circuit for amplification; and the amplified audio signal is played through a loudspeaker circuit.
8. The display screen amplification control device of claim 7, wherein, The preprocessing module comprises: Performing format conversion and gain adjustment on the N-channel microphone raw audio signals; Performing voltage division processing on the M-channel loudspeaker raw audio data to divide the voltage from 17-28V to 0-5V; Performing noise removal, noise elimination and noise suppression processing on the N+M signals; The preprocessed signals are used as inputs for weighted calculation; The weighted calculation uses the following formula: Wherein: Sout represents the target audio signal after weighted fusion; Smici represents the preprocessed audio signal of the i th microphone; Sspkj represents the preprocessed audio signal of the j th loudspeaker; αi represents the weight coefficient of the i th microphone signal, which is dynamically adjusted according to the conference scene; βj represents the weight coefficient of the j th loudspeaker signal, which is used for feedback suppression and signal fusion; The weight coefficients satisfy In the noise removal, noise elimination and noise suppression processing of the N+M signals, the noise removal processing includes constructing an inverse filter model with a filter function as follows: Wherein: H(f) represents the filter response at frequency f; G represents the noise reduction gain coefficient, which ranges from 10dB to 35dB; Pnoise(f) represents the noise power spectral density at frequency f; Psignal(f) represents the signal power spectral density at frequency f; ∈ is a very small positive number to prevent the denominator from being zero; The filtered signals are used as inputs for weighted calculation.
9. The display screen amplification control device of claim 8, wherein, It also comprises an audio feedback suppression module configured to perform an audio feedback suppression step, wherein the control audio feedback suppression step uses the following howling suppression function: F suppress (t) = γ · ∫0 t e -λ(t-τ) ·||S fb (τ)|·sgn(S fb (τ))dτ wherein: Fsuppress(t) represents the feedback suppression amount at time t; Sfb(τ) represents the feedback signal at time τ, obtained by a speaker acquisition circuit; γ represents a suppression intensity coefficient; λ represents a decay factor, controlling the decay speed of the historical feedback signal; sgn(·) is a sign function, used to determine the phase of the feedback signal; to realize closed-loop howling suppression.
10. The display screen amplification control device of claim 9, wherein, The weighted calculation module includes a quantum processing unit, which includes: Construct a parameterized quantum circuit to map the weighted coefficient optimization problem to a ground state search problem of quantum bits; wherein the parameters of the quantum circuit include the angle θk of the quantum rotation gate and the coupling strength φkl of the entanglement gate; Convert the time domain or frequency domain features of the preprocessed N+M audio signals into the initial input state ∣ψin> of the quantum circuit through a classical encoding layer; Run the parameterized quantum circuit on a quantum processor to obtain an output state ∣ψout>; Measure the expected value <ψout∣H∣ψout> of the output state, wherein H is a Hamiltonian constructed according to the audio signal fusion quality index; Use a classical optimizer to iteratively update the parameters θk and φkl of the quantum circuit with the goal of minimizing the expected value; when the expected value converges below the threshold value, the corresponding optimal weight coefficient set is decoded from the optimal output state Use the optimal set of weight coefficients to perform weighted calculation to obtain the target audio signal Sout.