An ai-driven dynamic audio fence control method

CN122290622APending Publication Date: 2026-06-26JIANGSU COLLEGE OF INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU COLLEGE OF INFORMATION TECH
Filing Date
2026-04-12
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve precise control and adaptive processing of dynamic noise sources in large open spaces, leading to noise pollution that distracts attention, reduces work and communication efficiency, and is complex to deploy and unable to adapt to complex and ever-changing acoustic environments.

Method used

A beamforming microphone array combined with a DSP processor and an AI neural network is used for signal acquisition, feature extraction, speech noise classification, and dynamic beamforming. Combined with 3D audio rendering and head tracking technology, it achieves precise focusing of the target speech and suppression of background noise.

Benefits of technology

It achieves distortion-free focusing of target speech and precise suppression of background noise in dynamic scenes, adapts to various scenario requirements, reduces deployment costs, improves communication efficiency and user comfort, and is suitable for open offices, video conferencing and public spaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290622A_ABST
    Figure CN122290622A_ABST
Patent Text Reader

Abstract

This invention discloses an AI-driven dynamic audio fence control method, specifically including the following steps: Step 1: Signal acquisition and preprocessing using a beamforming microphone array and a DSP processor; Step 2: AI feature extraction and speech noise classification of standardized multi-channel digital audio signals; Step 3: MVDR dynamic beamforming processing based on speech noise classification labels and noise feature masks; Step 4: 3D audio rendering of the beamformed single-channel target speech signal; Step 5: Effect monitoring and parameter iteration based on the target speech signal and feedback signals from the audio output module; Step 6: Pure target speech output playback based on the 3D rendered time-domain target speech signal. The advantages of this invention are: ensuring a continuous improvement in the target speech signal-to-noise ratio (SNR), ultimately transmitting pure target speech to the speaker, forming a personal sound bubble-like virtual audio fence around the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an audio fence control method, specifically an AI-driven dynamic audio fence control method. Background Technology

[0002] Noise can severely distract people in modern work and life settings. Studies have shown that noise interference can increase the degree of distraction and significantly increase the error rate in work and communication, leading to decreased work focus and low meeting communication efficiency. Therefore, reducing noise has become a common pursuit for people seeking a high-quality environment.

[0003] Currently, noise control methods mainly rely on passive noise reduction technologies, including using noise-isolating headphones, laying sound-absorbing materials in the environment, and setting up physical sound barriers. These methods achieve noise reduction by absorbing or reflecting sound waves, but they have significant technical limitations. Their coverage is limited, only applicable to personal devices or small, fixed areas, and cannot handle dynamically distributed noise sources in large open spaces. Furthermore, their noise filtering accuracy is low and easily affected by dynamic environmental changes, such as speaker movement or the simultaneous presence of multiple noise sources, leading to a significant decrease in the clarity of the target speech. Additionally, deployment and maintenance are difficult, requiring dedicated hardware and complex algorithm optimizations, and they cannot adapt to complex and changing acoustic environments in real time.

[0004] With the development of Active Noise Control (ANC) and microphone array technologies, the technology of creating virtual audio isolation zones by combining beamforming arrays with digital signal processing (DSP) has become a research hotspot. Existing related systems capture sound signals through microphone arrays and use phase adjustment to focus speech in a specific direction, thereby suppressing interference noise from other directions. However, early technologies were limited by computational costs and hardware performance, making it difficult to handle dynamic scenarios involving speaker movement and changes in noise source location. In recent years, the integration of artificial intelligence (AI) and audio processing technologies has improved signal processing capabilities, such as using neural networks to achieve functions like feedback suppression and anomaly detection. However, existing related patent technologies are mostly limited to static audio processing scenarios or only applicable to specific applications such as anomaly detection, lacking the ability to adaptively control dynamic noise in large open spaces. Although some spatial audio patents introduce Time-varying Head-Related Transfer Functions (HRTFs), they do not fully integrate AI algorithms to achieve real-time parameter adjustment of audio fences. These technical problems seriously restrict the large-scale promotion and application of audio fence technology in practical scenarios, and there is an urgent need for a dynamic audio fence solution with intelligent adaptive capabilities. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide an AI-driven dynamic audio fence control method that can significantly reduce the distracting effect of noise pollution on people's attention, improve user comfort, reduce the negative impact of noise pollution on the human auditory and nervous systems, and promote barrier-free communication.

[0006] To address the aforementioned technical problems, the AI-driven dynamic audio fence control method of the present invention includes the following steps:

[0007] Step 1: Use a beamforming microphone array and a DSP processor for signal acquisition and preprocessing;

[0008] Step 2: Perform AI feature extraction and speech noise classification on the standardized multi-channel digital audio signal;

[0009] Step 3: Perform MVDR dynamic beamforming processing based on the speech noise classification labels and noise feature masks;

[0010] Step 4: Perform 3D audio rendering on the single-channel target speech signal after beamforming;

[0011] Step 5: Monitor the effect and iterate parameters based on the target speech signal and the feedback signal from the audio output module;

[0012] Step 6: Output and play the clean target speech based on the 3D rendered temporal target speech signal.

[0013] Furthermore, step 1 specifically includes:

[0014] A beamforming microphone array synchronously acquires multi-channel acoustic signals from the target area and its surrounding environment. After converting the acoustic signals into electrical signals, they are transmitted to a DSP processor for preprocessing to obtain standardized multi-channel digital audio signals. ,in The sequence number of the signal sampling points of the multi-channel digital audio signal. The channel number for the beamforming microphone array. =1,2,3,..., , The total number of channels in a beamforming microphone array.

[0015] Furthermore, in step 1, the preprocessing includes sequentially executed signal amplification, anti-aliasing filtering, and analog-to-digital converter (ADC). The signal amplification method employs a differential programmable gain amplifier (PGA); the anti-aliasing filtering method employs a second-order or higher active low-pass anti-aliasing filter; and the ADC method employs a synchronous multi-channel ADC, which outputs a standardized multi-channel digital audio signal. .

[0016] Furthermore, step 2 specifically includes: processing the preprocessed multi-channel digital audio signal... The input is fed into a CNN-RNN hybrid neural network, where the CNN network first extracts multi-channel digital audio signals. The local spectral features, also known as the spectral feature map. Then, the spectral feature map is captured through an RNN network. The temporal correlation features are used to obtain an audio feature vector that fuses the spectrum and temporal sequence. ; to feature vector The input is fed into a transformer-based classification model, which combines the environmental noise modeling results from an encoder-decoder neural network to classify and identify the feature vectors, outputting classification labels for the speech noise. , =1 indicates the target speech. =0 indicates background noise, and a noise feature mask is generated simultaneously. .

[0017] Furthermore, step 3 specifically includes: constructing a dynamic beamformer based on the minimum variance distortionless response (MVDR) algorithm, and classifying the speech noise according to its classification labels. With noise feature mask And determine the direction vector of the target speech. The optimal weight vector for the dynamic beamformer is obtained by solving for the inverse of the signal covariance matrix. , The mathematical expression is:

[0018] ;

[0019] In this mathematical expression, Multi-channel digital audio signal The covariance matrix, , Represents the mathematical expectation. Represents a multi-channel digital audio signal as a matrix The conjugate transpose of; For array manifold vectors, , For the first The distance of each microphone relative to the array reference point The signal wavelength of the channel signal. The imaginary unit, Represents the transpose of a matrix. This represents the number of microphones.

[0020] Furthermore, step 3 specifically includes: weighting and summing the multi-channel digital audio signals using the optimal weight vector of the dynamic beamformer to obtain the beamformed single-channel target speech signal. .

[0021] Furthermore, in step 3, the direction vector of the target speech The acquisition method specifically includes: multi-channel digital signals Two-channel signals and Perform short-time Fourier transform to obtain Phase value at each frequency point and The phase value at each frequency point will and Subtracting the phase values ​​at the same frequency point yields the phase difference. Combined acquisition of two-channel signals and The spacing between the two microphones and the signal wavelength of the channel signal You can get .

[0022] Furthermore, step 4 specifically includes introducing a time-varying head-related transfer function. By combining the head movement data of the target object collected by the head movement tracking sensor, the horizontal azimuth angle of the head is obtained. With vertical pitch angle For the target speech signal The mathematical expression for 3D spatial audio rendering is as follows:

[0023] ;

[0024] In this mathematical expression, For target speech signal The frequency domain signal obtained by Fast Fourier Transform (FFT) The frequency of the audio signal; For time-varying head-related transfer functions; The frequency domain target speech signal, after 3D rendering, is converted into a time domain target speech signal using Inverse Fast Fourier Transform (IFFT). .

[0025] Furthermore, step 5 specifically includes:

[0026] The feedback signal from the audio output module is synchronously acquired in real time using a beamforming microphone array. Calculate the signal-to-noise ratio of the feedback signal. With the speech transmission index as an indicator of speech clarity Signal-to-noise ratio The mathematical expression is:

[0027] ;

[0028] In this mathematical expression, For feedback signal The effective power of the target speech signal is determined by With the target speech signal The cross-correlation was calculated to obtain the result; For feedback signal Effective power of background noise, , For feedback signal Total effective power;

[0029] Will With preset threshold Compare: If ≥ If so, then keep the current system parameters unchanged; < Then and The monitoring indicators are iteratively optimized using the backpropagation algorithm to improve the network parameters of the CNN-RNN hybrid neural network and the transformer-based classification model, while simultaneously adjusting the direction vector of the dynamic beamformer. Resolve for the optimal weight vector To achieve adaptive optimization of system parameters and ensure Maintain at the preset threshold above.

[0030] Furthermore, step 6 specifically includes processing the 3D-rendered temporal target speech signal obtained in step 4. After being converted from digital to analog (DAC) and amplified by power, the audio is transmitted to the audio output device to complete the output and playback of the pure target voice.

[0031] The advantages of this invention are:

[0032] (1) Based on the AI-enhanced audio signal processing framework, multi-channel audio signals are collected through beamforming microphone array, signal features are extracted by AI neural network and speech and noise are classified and identified to generate dynamic adaptive filter; beam direction is adjusted by combining minimum variance distortionless response (MVDR) algorithm, and time-varying head correlation transfer function (HRTFs) is introduced to realize real-time rendering of 3D audio isolation zone; the output audio effect is monitored in real time by microphone, and the system parameters are iteratively optimized and adjusted by AI model to ensure that the target speech signal-to-noise ratio (SNR) continues to improve, and finally the pure target speech is transmitted to the speaker to form a personal sound bubble-like virtual audio fence around the target object.

[0033] (2) By deeply integrating AI algorithms with beamforming, 3D audio rendering and other technologies, the system solves the problems of poor adaptability, low filtering accuracy and complex deployment of traditional audio noise reduction and isolation technologies in dynamic scenes. It achieves significant technical effects in acoustic performance, scene adaptation, deployment and application, health and communication. Specifically, through the collaborative work of MVDR dynamic beamforming algorithm and AI speech / noise classification model, it achieves distortion-free focusing of target speech and accurate suppression of background noise. At the same time, combined with the environmental noise modeling of encoder-decoder neural network, it can effectively distinguish various common background noises such as keyboard typing, air conditioner operation and people talking, and achieve accurate isolation between target speech and background noise. The system analyzes audio signals in real time through AI model and dynamically adjusts beam direction and audio fence boundary, which can accurately track the position movement of the target speaker and adapt to the dynamic acoustic environment of multiple noise sources. Combined with time-varying HRTFs and head movement tracking technology, it realizes dynamic rendering of audio fence in 3D space and adapts to the target speech. The elephant's head movement can meet the needs of various scenarios such as open offices, video conferencing, and even public spaces, overcoming the limitations of traditional technologies that are only applicable to static scenarios. It employs a highly integrated DSP processor and optical MEMS microphone array to achieve end-to-end low-latency processing of audio signals. The system features a modular design, with each hardware module combined through standardized interfaces, eliminating the need for complex wiring. The arrangement and number of microphone arrays can be flexibly adjusted according to scenario requirements, reducing deployment, installation, and maintenance costs in different scenarios. Simultaneously, the lightweight software algorithm can run on embedded hardware without relying on high-performance servers, further reducing usage costs. A dynamic audio fence constructed through 3D audio rendering technology forms individual sound bubbles around the target object, accurately focusing the target speech in 3D space while shielding external noise interference, providing users with an immersive audio experience. In video conferencing scenarios, it can build independent audio fences for multiple speakers, achieving clear separation of multiple target voices and improving the communication experience and participation in meetings.

[0034] (3) By effectively isolating background noise, the distracting effect of noise pollution on people's attention is significantly reduced, which can improve work focus and meeting communication efficiency, and reduce work errors caused by noise interference; at the same time, it avoids the ear discomfort caused by wearing traditional noise-isolating headphones for a long time. It achieves noise reduction through non-contact virtual audio fence, which improves the comfort of long-term use and reduces the negative impact of noise pollution on the human auditory system and nervous system; the system's high-precision speech filtering and dynamic tracking capabilities can be adapted to the auxiliary communication scenarios of hearing-impaired people. By amplifying and noise-reducing the target speech and transmitting it to the hearing aid device, it can improve the hearing-impaired people's ability to recognize the target speech; at the same time, in multi-person communication scenarios, it can achieve clear speech separation, provide support for auxiliary communication for people with language communication disorders, and promote the realization of barrier-free communication. Attached Figure Description

[0035] Figure 1 This is a flowchart of the AI-driven dynamic audio fence control method of the present invention. Detailed Implementation

[0036] The AI-driven dynamic audio fence control method of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0037] To achieve dynamic audio fence control, the following AI-driven dynamic audio fence system is constructed. This system includes an audio acquisition module and a connected DSP processor. The DSP processor is also connected to the terminal playback device. The DSP processor runs an AI feature extraction and classification module, a dynamic beamforming module, a 3D audio rendering module, an effect monitoring and parameter iteration module, and an audio output module. The modules work together to achieve full-process processing from audio signal acquisition to pure target voice output. The system architecture is modular and supports flexible deployment and expansion.

[0038] The audio acquisition module uses a beamforming microphone array arranged linearly or circularly. The microphone array can be an optical MEMS microphone array to complete the synchronous acquisition of multi-channel acoustic signals, capture the audio signals of the target area and the surrounding environment, convert the acoustic signals into electrical signals and transmit them to the subsequent processing module to achieve low-loss and high-synchronization signal acquisition.

[0039] The AI ​​feature extraction and classification module is equipped with a CNN-RNN hybrid neural network and a transformer-based architecture model. After digitally processing the acquired multi-channel audio electrical signals, it extracts key features such as the spectral features and temporal patterns of the signals. The model completes the accurate classification and recognition of target speech and background noise, providing feature basis for the generation of dynamic filters. Background noise includes keyboard sounds, air conditioner sounds, and irrelevant conversation sounds.

[0040] The dynamic beamforming module constructs a beamformer based on the minimum variance distortionless response (MVDR) algorithm. According to the output of the AI ​​classification module, it dynamically adjusts the beam direction and coverage to achieve precise focusing on the speech of the target speaker, while suppressing noise signals in non-target directions and generating a dynamic filter that adapts to the current acoustic environment.

[0041] The 3D audio rendering module introduces Time-varying Head-Related Transfer Functions (HRTFs), which, combined with head motion tracking data of the target object, perform 3D spatial audio rendering on the target speech signal after beamforming. The boundary and range of the audio fence are dynamically adjusted in 3D space to realize the spatial construction of personal sound bubbles.

[0042] The effect monitoring and parameter iteration module monitors the signal effect of the audio output module in real time through the microphone array, extracts key indicators such as the signal-to-noise ratio (SNR) and speech clarity of the output signal, and feeds the monitoring data back to the AI ​​feature extraction and classification module and the dynamic beamforming module. The AI ​​model iteratively optimizes the network parameters and beam adjustment parameters of the MVDR algorithm based on the monitoring data to achieve closed-loop adaptive control of the system.

[0043] The audio output module transmits the purified target speech signal, which has undergone 3D rendering and noise reduction, to terminal playback devices such as speakers to complete the final audio output.

[0044] This AI-driven dynamic audio fence system overcomes the limitations of traditional beamforming technology's fixed direction by introducing a CNN-RNN hybrid neural network and a transformer-based architecture. It analyzes multi-channel audio input signals in real time, dynamically adjusting the direction and coverage of the audio beam to achieve real-time tracking of speaker movement and noise source changes, enabling adaptive expansion and contraction of the virtual audio fence. It integrates speech feature extraction (spectral and temporal patterns) with environmental noise modeling techniques, employing an encoder-decoder neural network for end-to-end audio signal processing. This accurately distinguishes target speech from various background noise interferences, improving speech recognition and filtering accuracy in complex acoustic environments. Furthermore, it incorporates the time-varying head-related transfer function H... The audio fence is deeply integrated with AI dynamic control algorithms and head tracking technology, enabling dynamic rendering in 3D space to adapt to the head movements of the target object, thus improving the spatial adaptability and immersive experience of the audio fence. It adopts a highly integrated DSP processor and optical MEMS microphone array to achieve low-latency processing of audio signals (end-to-end processing latency meets the real-time requirements of actual scenarios). The system is modularly designed, eliminating the need for complex wiring and reducing deployment and maintenance costs in different scenarios. A closed-loop feedback control system based on signal-to-noise ratio (SNR) is constructed. By monitoring the output audio effect in real time, iterative optimization of AI model parameters and beamforming parameters is achieved, ensuring that the system maintains a high SNR and speech intelligibility in dynamic acoustic environments.

[0045] like Figure 1 As shown, the AI-driven dynamic audio fence control method of the present invention includes the following specific steps:

[0046] Step 1: Use a beamforming microphone array and a DSP processor for signal acquisition and preprocessing;

[0047] In a preferred but non-limiting embodiment of the present invention, step 1 specifically includes:

[0048] A beamforming microphone array synchronously acquires multi-channel acoustic signals from the target area and its surrounding environment. These acoustic signals are converted into electrical signals and then transmitted to a DSP processor for preprocessing. The analog electrical signals are then converted into 16-bit / 48kHz digital audio signals, eliminating DC offset and high-frequency aliasing interference during signal acquisition, resulting in standardized multi-channel digital audio signals. ,in The sequence number of the signal sampling points of the multi-channel digital audio signal. The channel number for the beamforming microphone array. =1,2,3,..., , The total number of channels in a beamforming microphone array.

[0049] In a preferred but non-limiting embodiment of the present invention, step 1 includes a preprocessing procedure consisting of sequentially executed signal amplification, anti-aliasing filtering, and analog-to-digital converter (ADC). The signal amplification method employs a differential programmable gain amplifier (PGA), which amplifies the weak differential analog signal output from the beamforming microphone array with low noise, suppresses environmental electromagnetic interference with common-mode rejection ratio (CMRR), maintains signal fidelity, and automatically adjusts the gain according to the input signal amplitude to avoid saturation and clipping. The anti-aliasing filtering method uses a second-order or higher active low-pass anti-aliasing filter. The cutoff frequency of the active low-pass anti-aliasing filter is set to 1.2 times the highest frequency of the target speech, such as 3.6kHz–4.2kHz, filtering out high-frequency components above the Nyquist frequency to prevent frequency aliasing during ADC sampling. A linear phase design is used to avoid introducing speech waveform distortion. The ADC method employs a synchronous multi-channel high-precision ADC with a sampling rate of up to 48kHz and a bit depth of 16bit / 24bit. Multi-channel synchronous sampling ensures time alignment of each channel in the microphone array, and its output is a standardized multi-channel digital audio signal. .

[0050] Step 2: Perform AI feature extraction and speech noise classification on the standardized multi-channel digital audio signal;

[0051] In a preferred but non-limiting embodiment of the present invention, step 2 specifically includes:

[0052] Preprocessed multi-channel digital audio signal The input is fed into a CNN-RNN hybrid neural network, where the CNN network first extracts multi-channel digital audio signals. The local spectral features, also known as the spectral feature map. Then, the spectral feature map is captured through an RNN network. The temporal correlation features are used to obtain an audio feature vector that fuses the spectrum and temporal sequence. ; to feature vector The input is fed into a transformer-based classification model, which combines the environmental noise modeling results from an encoder-decoder neural network to classify and identify the feature vectors, outputting classification labels for the speech noise. , =1 indicates the target speech. =0 indicates background noise, and a noise feature mask is generated simultaneously. This provides a basis for noise suppression in subsequent beamforming.

[0053] It should be noted that the methods for extracting local spectral features using CNN networks include: processing multi-channel digital audio signals. Each digital audio signal is framed and windowed. After frame-by-frame windowing, a Short-Time Fourier Transform (STFT) is performed to obtain a spectrogram. The spectrogram is then input into a Convolutional Neural Network (CNN). After multiple convolutions and pooling layers, Mel-frequency cepstral coefficients (MFCC), spectral entropy, spectral flatness, and local texture features of the spectrogram are obtained. Multiple convolutions scan local regions of the audio spectrogram, automatically capturing acoustic features such as texture, edges, peaks, valleys, and frequency band variations to form the convolution result. Pooling layers downsample the convolution result, preserving key features, compressing data, and improving robustness to interference, thus obtaining a spectral feature map. To ensure the model remains stable under noise, the local texture features of the spectrogram refer to the local stripes, bands, dots, edges, connectivity and other image structure features automatically extracted by the convolutional layers of the CNN in the time-frequency two-dimensional plane of the audio spectrogram. These features are used to characterize the morphological differences of speech, steady-state noise and impulse noise in the time-frequency domain.

[0054] Furthermore, spectral feature maps are captured using RNN networks. The temporal correlation features are used to obtain an audio feature vector that fuses the spectrum and temporal sequence. The methods specifically include:

[0055] Spectral feature map The data is fed into a bidirectional recurrent neural network (RNN) to learn the temporal structure of speech, which includes initial consonant → final vowel, noise persistence patterns, and transient impact (keyboard) features, thereby obtaining temporal features. ,Will and By splicing and merging, we obtain .

[0056] Furthermore, the feature vector The input is fed into a transformer-based classification model, which combines the environmental noise modeling results from an encoder-decoder neural network to classify and identify the feature vectors, outputting classification labels for the speech noise. , =1 indicates the target speech. =0 indicates background noise, and a noise feature mask is generated simultaneously. The methods specifically include:

[0057] The spectrum and temporal audio feature vectors fed into the encoder of the Transformer A multi-head self-attention mechanism is used to globally correlate features across all time frames within the sequence and output global temporal enhancement features. This automatically captures the dependencies between speech and noise over long time axes, improving feature stability and classification robustness in low signal-to-noise ratio and aliasing environments. The global temporal enhancement features output by the Transformer encoder are fed into a fully connected layer and mapped to binary classification logistic values. These logistic values ​​are then normalized using Softmax to generate posterior probabilities for speech and noise. The classification labels for speech and noise are output based on the maximum probability. , =1 indicates the target speech. =0 indicates background noise; and Input to an encoder-decoder neural network, encoder-decoder neural network encoder pairs For noise with a value of 0, three types of inherent noise characteristics are extracted: noise power spectrum, noise spatial correlation, and noise type (steady-state / transient / impulse). These characteristics form a noise feature code. The decoder then applies this noise feature code to each sampling point. Each microphone channel Perform point-by-point, channel-by-channel prediction and output the probability of noise presence at the current position; the mask value is directly determined by the noise presence probability, without the need for custom weights or coefficients, i.e., if the channel... Sampling points For noise, output the effective suppression value for subsequent beamforming noise suppression; for target speech, output the pass-through value to maintain the target speech without distortion, and finally output the channel-by-channel, sample-by-sample noise feature mask. .

[0058] For example, multi-channel digital audio signals Local spectral features such as Mel frequency cepstral coefficients (MFCC) or spectrogram features.

[0059] Step 3: Perform MVDR dynamic beamforming processing based on the speech noise classification labels and noise feature masks;

[0060] In a preferred but non-limiting embodiment of the present invention, step 3 specifically includes:

[0061] A dynamic beamformer is constructed based on the minimum variance distortionless response (MVDR) algorithm, and the speech noise classification labels output in step 2 are used as the basis for the beamformer. With noise feature mask And determine the direction vector of the target speech. The optimal weight vector for the dynamic beamformer is obtained by solving for the inverse of the signal covariance matrix. This achieves distortion-free amplification of speech from the target direction and minimum variance suppression of noise from non-target directions. The mathematical expression is:

[0062] ;

[0063] In this mathematical expression, Multi-channel digital audio signal The covariance matrix, , Represents the mathematical expectation. Represents a multi-channel digital audio signal as a matrix The conjugate transpose of the matrix, a single multi-channel digital audio signal. That is the first one of the matrix Line 1 Column elements, It reflects the spatial correlation of multi-channel signals and is calculated from the real-time acquired audio signals; The array manifold vector is formed by the geometry of the microphone array and the direction vector of the target speech. Decide, , For the first The distance of each microphone relative to the array reference point The signal wavelength of the channel signal. The imaginary unit, Represents the transpose of a matrix. Number of microphones; This is the optimal weight vector for the MVDR beamformer, whose value changes dynamically to achieve adaptive adjustment of beam direction and gain.

[0064] In a preferred but non-limiting embodiment of the present invention, step 3 further includes:

[0065] The multi-channel digital audio signals are weighted and summed using the optimal weight vector of the dynamic beamformer to obtain the single-channel target speech signal after beamforming. ,in The target speech signal, after beamforming processing, suppresses background noise and retains the original characteristics of the speech in the target direction without distortion.

[0066] It should be noted that the array reference point can be the geometric center of the beamforming microphone array.

[0067] In a preferred but non-limiting embodiment of the present invention, in step 3, the direction vector of the target speech The beam pointing angle varies over time and is determined by the position of the target speaker; the direction vector of the target speech. The methods for obtaining it specifically include:

[0068] For multi-channel digital signals Two-channel signals and Perform short-time Fourier transform to obtain Phase value at each frequency point and The phase value at each frequency point will and Subtracting the phase values ​​at the same frequency point yields the phase difference. Combined acquisition of two-channel signals and The spacing between the two microphones and the signal wavelength of the channel signal You can get .

[0069] Step 4: Perform 3D audio rendering on the single-channel target speech signal after beamforming;

[0070] In a preferred but non-limiting embodiment of the present invention, step 4 specifically includes:

[0071] Introducing a time-varying head-related transfer function By combining the head movement data of the target object collected by the head movement tracking sensor, the horizontal azimuth angle of the head is obtained. With vertical pitch angle (These are all dynamic parameters that change over time), for the target speech signal output in step 3 The mathematical expression for 3D spatial audio rendering is as follows:

[0072] ;

[0073] In this mathematical expression, For target speech signal The frequency domain signal obtained by Fast Fourier Transform (FFT) The frequency of the audio signal; The head-related transfer function is a time-varying function determined by the head movement angle of the target object and the signal frequency. It reflects the transmission characteristics of sound from different spatial directions by the human ear and is dynamically updated with the head movement angle. The frequency domain target speech signal, after 3D rendering, is converted into a time domain target speech signal using Inverse Fast Fourier Transform (IFFT). This enables the rendering of audio fences in 3D space, allowing the target speech to be focused spatially around the target object, forming a personal sound bubble.

[0074] Step 5: Monitor the effect and iterate parameters based on the target speech signal and the feedback signal from the audio output module;

[0075] In a preferred but non-limiting embodiment of the present invention, step 5 specifically includes:

[0076] The feedback signal from the audio output module is synchronously acquired in real time using a beamforming microphone array. Calculate the signal-to-noise ratio of the feedback signal. With the speech transmission index as an indicator of speech clarity Signal-to-noise ratio The mathematical expression is:

[0077] ;

[0078] In this mathematical expression, For feedback signal The effective power of the target speech signal is determined by With the target speech signal The cross-correlation was calculated to obtain the result; For feedback signal Effective power of background noise, , For feedback signal Total effective power;

[0079] Will With preset threshold Compare: If ≥ If so, then keep the current system parameters unchanged; < Then and The monitoring indicators are iteratively optimized using the backpropagation algorithm to optimize the network parameters of the CNN-RNN hybrid neural network and the transformer-based classification model, while the direction vector of the dynamic beamformer is adjusted. Resolve for the optimal weight vector To achieve adaptive optimization of system parameters and ensure Maintain at the preset threshold above.

[0080] It should be noted that the preset threshold Configure according to the actual needs of the scenario, such as an office scenario. =15dB.

[0081] Furthermore, real-time calculations will be performed. , As a quality evaluation metric for system adaptive optimization, an unsupervised loss function is constructed. The parameters of the CNN-RNN hybrid neural network and the Transformer classification network are iteratively updated through the backpropagation algorithm, so that the model output converges towards higher signal-to-noise ratio and higher speech intelligibility. The specific implementation method is as follows:

[0082] Using real-time monitored voice quality metrics as the optimization target, the optimization objective is to maximize... and maximizing Construct the real-time loss function Loss(n) = coefficient 1 × (set SNR target value − ) + coefficient 2 × (set STI target value − ),when and If the values ​​fall below the set SNR and STI target values ​​respectively, Loss(n) will increase, triggering optimization. and Once the set SNR and STI target values ​​are reached, Loss(n) will approach 0, at which point optimization stops. Loss(n) is then propagated backward along the network's forward path, updating the Transformer classification model parameters, CNN convolutional layer parameters, RNN temporal modeling parameters, and Encoder-Decoder noise modeling parameters in sequence. The update rule uses gradient descent to gradually reduce Loss(n), thus making the output speech clearer and the noise less.

[0083] Step 6: Output and play the clean target speech based on the 3D rendered temporal target speech signal.

[0084] In a preferred but non-limiting embodiment of the present invention, step 6 specifically includes:

[0085] The temporal target speech signal obtained in step 4 after 3D rendering After being converted from digital to analog (DAC) and amplified by power, the audio is transmitted to audio output devices such as speakers to complete the output and playback of pure target speech, forming a dynamic audio fence around the target object and effectively isolating the target speech from background noise.

[0086] Examples of application scenarios for this invention are as follows:

[0087] In an open-plan office setting, a linearly arranged optical MEMS microphone array is deployed at the edge of the office desk. The array is 0.5m to 1m long and is equipped with a head tracking sensor and a desktop DSP processing module. The audio output is connected to a speaker. The system is set to a single-person focus mode. The 3D spatial range of the audio fence is controlled within a spherical area with a radius of 0.5m centered on the user's head. The system tracks the user's head and body movements in real time, dynamically adjusts the beam direction, and suppresses noise such as conversations and keyboard typing from nearby colleagues.

[0088] In video conferencing scenarios, a circular microphone array is deployed in the center of the conference table, with 8 to 16 microphones and an array radius of 0.2m to 0.3m. It is equipped with a panoramic head tracking device and a wall-mounted DSP processing module, and the audio output is connected to the conference speakers. The system has a multi-person conference mode, which can simultaneously identify and track the position of multiple target speakers, build an independent audio fence for each speaker, achieve simultaneous focusing of multi-target speech, and suppress ambient noise outside the conference room.

[0089] In public space scenarios, it employs a distributed deployment of multiple microphone arrays, coupled with a cloud-based DSP computing power expansion module, with the audio output connected to directional speakers in the public area. The system is set up with a regional isolation mode, dividing the public space into multiple independent audio isolation zones. The audio fence boundary of each isolation zone can be dynamically adjusted according to the flow of people, achieving audio signal isolation between different areas and avoiding mutual interference.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention without departing from the spirit and scope of the present invention. Any modifications or equivalent substitutions should be covered within the scope of protection of the claims of the present invention.

Claims

1. An AI-driven dynamic audio fence control method, characterized in that, Includes the following steps: Step 1: Use a beamforming microphone array and a DSP processor for signal acquisition and preprocessing; Step 2: Perform AI feature extraction and speech noise classification on the standardized multi-channel digital audio signal; Step 3: Perform MVDR dynamic beamforming processing based on the speech noise classification labels and noise feature masks; Step 4: Perform 3D audio rendering on the single-channel target speech signal after beamforming; Step 5: Monitor the effect and iterate parameters based on the target speech signal and the feedback signal from the audio output module; Step 6: Output and play clean target speech based on the 3D rendered temporal target speech signal.

2. The AI-driven dynamic audio fence control method according to claim 1, characterized in that: Step 1 specifically includes A beamforming microphone array synchronously acquires multi-channel acoustic signals from the target area and its surrounding environment. After converting the acoustic signals into electrical signals, they are transmitted to a DSP processor for preprocessing to obtain standardized multi-channel digital audio signals. ,in The sequence number of the signal sampling points of the multi-channel digital audio signal. The channel number for the beamforming microphone array. =1,2,3,..., , The total number of channels in a beamforming microphone array.

3. The AI-driven dynamic audio fence control method according to claim 2, characterized in that: In step 1, the preprocessing process includes sequentially executed signal amplification, anti-aliasing filtering, and analog-to-digital converter (ADC). The signal amplification method uses a differential programmable gain amplifier (PGA); the anti-aliasing filtering method uses an active low-pass filter of second order or higher. The analog-to-digital converter (ADC) method uses a synchronous multi-channel ADC, which outputs a standardized multi-channel digital audio signal. .

4. The AI-driven dynamic audio fence control method according to claim 3, characterized in that: Step 2 specifically includes Preprocessed multi-channel digital audio signal The input is fed into a CNN-RNN hybrid neural network, where the CNN network first extracts multi-channel digital audio signals. The local spectral features, also known as the spectral feature map. Then, the spectral feature map is captured through an RNN network. The temporal correlation features are used to obtain an audio feature vector that fuses the spectrum and temporal sequence. ; eigenvectors The input is fed into a transformer-based classification model, which combines the environmental noise modeling results from an encoder-decoder neural network to classify and identify the feature vectors, outputting classification labels for the speech noise. , =1 indicates the target speech. =0 indicates background noise, and a noise feature mask is generated simultaneously. .

5. The AI-driven dynamic audio fence control method according to claim 4, characterized in that: Step 3 specifically includes A dynamic beamformer is constructed based on the minimum variance distortionless response (MVDR) algorithm, and the beamformer is based on the classification labels of speech noise. With noise feature mask And determine the direction vector of the target speech. The optimal weight vector for the dynamic beamformer is obtained by solving for the inverse of the signal covariance matrix. , The mathematical expression is: ; In this mathematical expression, Multi-channel digital audio signal The covariance matrix, , Represents the mathematical expectation. Represents a multi-channel digital audio signal as a matrix The conjugate transpose of; For array manifold vectors, , For the first The distance of each microphone relative to the array reference point The signal wavelength of the channel signal. The imaginary unit, Represents the transpose of a matrix. This represents the number of microphones.

6. The AI-driven dynamic audio fence control method according to claim 5, characterized in that: Step 3 specifically also includes The multi-channel digital audio signals are weighted and summed using the optimal weight vector of the dynamic beamformer to obtain the single-channel target speech signal after beamforming. .

7. The AI-driven dynamic audio fence control method according to claim 6, characterized in that: In step 3, the direction vector of the target speech The methods for obtaining it include, specifically: For multi-channel digital signals Two-channel signals and Perform short-time Fourier transform to obtain Phase value at each frequency point and The phase value at each frequency point will and Subtracting the phase values ​​at the same frequency point yields the phase difference. Combined acquisition of two channel signals and The spacing between the two microphones and the signal wavelength of the channel signal You can get .

8. The AI-driven dynamic audio fence control method according to claim 7, characterized in that: Step 4 specifically includes Introducing a time-varying head-related transfer function By combining the head movement data of the target object collected by the head movement tracking sensor, the horizontal azimuth angle of the head is obtained. With vertical pitch angle For the target speech signal The mathematical expression for 3D spatial audio rendering is as follows: ; In this mathematical expression, For target speech signal The frequency domain signal obtained by Fast Fourier Transform (FFT) The frequency of the audio signal; For time-varying head-related transfer functions; The frequency domain target speech signal, after 3D rendering, is converted into a time domain target speech signal using Inverse Fast Fourier Transform (IFFT). .

9. The AI-driven dynamic audio fence control method according to claim 8, characterized in that: Step 5 specifically includes The feedback signal from the audio output module is synchronously acquired in real time using a beamforming microphone array. Calculate the signal-to-noise ratio of the feedback signal. With the speech transmission index as an indicator of speech clarity Signal-to-noise ratio The mathematical expression is: ; In this mathematical expression, For feedback signal The effective power of the target speech signal is determined by With the target speech signal The cross-correlation was calculated to obtain the result; For feedback signal Effective power of background noise, , For feedback signal Total effective power; Will With preset threshold Compare: If ≥ If so, then keep the current system parameters unchanged; if < Then and The monitoring indicators are iteratively optimized using the backpropagation algorithm to improve the network parameters of the CNN-RNN hybrid neural network and the transformer-based classification model, while simultaneously adjusting the direction vector of the dynamic beamformer. Resolve for the optimal weight vector To achieve adaptive optimization of system parameters and ensure Maintain at the preset threshold above.

10. The AI-driven dynamic audio fence control method according to claim 9, characterized in that: Step 6 specifically includes processing the temporal target speech signal obtained in step 4 after 3D rendering. After being converted from digital to analog (DAC) and amplified by power, the audio is transmitted to the audio output device to complete the output and playback of the pure target voice.