Wireless audio transmission optimization system based on adaptive beam forming
By using adaptive beamforming technology and deep learning and reinforcement learning methods in wireless audio transmission systems, we dynamically identify and suppress complex interference sources and optimize user experience, and solve the problems of insufficient ability to identify complex interference sources and poor adaptability in the prior art, and achieve high-quality wireless audio transmission.
Patent Information
- Application Number
- CN202510494821.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The existing anti-interference technology of wireless audio transmission has problems such as insufficient ability to identify complex interference sources, insufficient consideration of users' subjective feelings, and poor adaptability, and cannot meet the high-quality needs of wireless audio transmission in complex electromagnetic environments.
Adopting a wireless audio transmission optimization system based on adaptive beamforming, including a multi-channel signal acquisition and preprocessing module, a layered attention interference source feature extraction module, a user discomfort perception model building module, a beam zero-point dynamic adjustment module and a reinforcement learning optimization module. Through deep learning and reinforcement learning technology, interference sources can be dynamically identified and suppressed, and the user experience is optimized.
It improves the anti-interference performance and audio quality of wireless audio transmission, enhances the system's adaptability to unknown interference environments, optimizes user experience, and reduces the use of computing resources.
Smart Images

Figure CN120017192A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wireless communication, and more particularly to a wireless audio transmission optimization system based on adaptive beamforming. Background Art
[0002] Wireless audio devices have been widely used in consumer electronics, conference systems, smart homes and other fields in recent years due to their portability and wireless connection advantages. However, in actual use environments, wireless audio transmission often faces complex and changeable electromagnetic interference problems, resulting in reduced audio transmission quality and impaired user experience.
[0003] Existing wireless audio transmission anti-interference technologies mainly include spectrum expansion, adaptive frequency hopping and traditional beamforming. Spectrum expansion technology disperses signal energy into a wider frequency band, making the interference less impactful, but the technology has limited resistance to transient strong interference; adaptive frequency hopping technology can avoid interference at fixed frequencies, but it is not effective when the interference source covers multiple frequency bands; traditional beamforming technology can suppress interference from a specific direction through spatial filtering, but it has the following major defects: traditional beamforming technology usually relies on a preset interference model, and it is difficult to effectively deal with unknown or complex and changeable interference sources in the environment; the existing interference source feature extraction algorithm is inefficient and cannot quickly and accurately distinguish between major and minor interference sources, resulting in unreasonable allocation of computing resources; the existing system lacks a direct feedback mechanism for the user's actual perceived experience; the existing technology usually adopts a fixed optimization strategy and cannot be adaptively adjusted for different environments and usage scenarios, resulting in insufficient adaptability of the system when facing new interference sources.
[0004] In summary, the existing wireless audio transmission anti-interference technology has problems such as insufficient ability to identify complex interference sources, insufficient consideration of user subjective feelings, and weak adaptability, which cannot meet the high-quality requirements of wireless audio transmission in complex electromagnetic environments. Therefore, there is an urgent need for a wireless audio transmission optimization method that can accurately identify various interference sources, dynamically adjust the interference suppression strategy according to user personalized needs, and has continuous learning capabilities. Summary of the invention
[0005] The present invention provides a wireless audio transmission optimization system based on adaptive beamforming, which solves the technical problems in the related art of insufficient ability to identify complex interference sources, insufficient consideration of user subjective feelings and weak adaptability.
[0006] The present invention provides a wireless audio transmission optimization system based on adaptive beamforming, comprising: Multi-channel signal acquisition and preprocessing module, which acquires multi-channel acoustic signals and performs space-time domain preprocessing to generate feature matrix; Hierarchical attention distractor feature extraction module, which processes the feature matrix and identifies the type and location of distractors; User discomfort perception model building module, which builds a user discomfort perception model based on user physiological response data, facial expression recognition and behavior analysis, and generates a personalized interference sensitivity map; The beam null dynamic adjustment module dynamically adjusts the null position and depth of the beamforming algorithm based on the interference source type, location and personalized interference sensitivity map; The reinforcement learning optimization module receives the execution results of the beam zero point dynamic adjustment module and user feedback data, and uses the reinforcement learning method to continuously optimize the interference suppression strategy to improve the system's adaptability to unknown interference environments.
[0007] Furthermore, the multi-channel signal acquisition and preprocessing module includes: The microphone array collects multi-channel acoustic signals. The number of microphones ; Applying a spatial filter to the collected signal to perform spatial filtering processing on the multi-channel acoustic signal; Use the exponentially weighted moving average algorithm to smooth the spatially filtered signal; Applying an adaptive noise suppression algorithm to the smoothed signal to reduce signal noise; Use short-time Fourier transform to transform the denoised signal into time-frequency domain; Generate interference source feature extraction input and build a feature matrix based on time-frequency domain signals as the input of interference source feature extraction.
[0008] Furthermore, the hierarchical attention interference source feature extraction module includes: Construct a multi-scale convolutional neural network, perform feature extraction, and build a network structure containing convolution kernels of different sizes to capture time-frequency features of different scales; Channel attention calculation: apply the channel attention mechanism to each group of feature maps and calculate the channel attention weight; Feature weighting, applying the channel attention weight to the original feature map to obtain a weighted feature map; Feature fusion, fusing weighted feature maps of different scales to obtain a comprehensive feature representation; Interference source identification and location, processing fusion features, realizing the identification and location of interference sources, and obtaining interference source feature representation.
[0009] Furthermore, the multi-scale convolutional neural network includes three convolution kernels of different sizes, namely 3×3, 5×5 and 7×7, which capture time-frequency features of different scales.
[0010] Furthermore, the user discomfort perception model building module includes: Physiological data collection, collecting user physiological response data; Facial expression recognition: using computer vision algorithms to analyze user facial expressions in real time and identify micro-expression changes; Behavioral data monitoring, monitoring users’ implicit behaviors; Environmental context information collection; User discomfort calculation, building a user discomfort perception model based on the physiological data, the expression characteristics, the behavior data and the environmental context information, and calculating user discomfort; Personalized interference sensitivity map generation, based on historical data and user feedback, generates personalized interference sensitivity maps for different interference types and spatial locations.
[0011] Furthermore, the beam zero point dynamic adjustment module includes: Interference source suppression priority calculation: based on interference source characteristics and interference sensitivity maps, calculate the suppression priority of each interference source; Priority sorting and screening: sort the identified interference sources according to the calculated suppression priority and select the one with the highest priority The interference source is used as the placement target of the beam null; Constructing the interference source covariance matrix, constructing the interference source covariance matrix; Calculate the beamforming weight vector, and use the linear constrained minimum variance algorithm to calculate the beamforming weight vector; Zero point depth adaptive control: for specific types of interference sources, an adaptive zero point depth control mechanism is introduced to increase the corresponding zero point depth for interference sources with higher priority and greater sensitivity; Real-time detection of user discomfort changes. Real-time detection of user discomfort changes. When the user discomfort exceeds the preset threshold, it triggers the rapid update of the beam zero point position and depth. The beamforming is applied by applying the calculated beamforming weight vector to the input signal to obtain the beamformed output signal.
[0012] Furthermore, the reinforcement learning optimization module includes: State space and action space definition, including the interference source characteristics of the current environment User inappropriateness Beamforming Configuration Parameters The state space of historical interaction records defines the action space containing the adjustment operations of the beamforming parameters; Reward function design: Design a reward function that comprehensively considers user discomfort changes, objective sound quality evaluation indicators, and system resource consumption; Deep Q network construction, as the core algorithm of reinforcement learning, builds value network; The experience replay mechanism is implemented to store experience samples of the interaction between the system and the environment and improve learning efficiency; Double Q-learning is applied to reduce the overestimation problem; ϵ-greedy strategy application, using ϵ-greedy strategy for action selection, reducing the exploration rate over time; Periodic evaluation mechanism application: set up a periodic evaluation mechanism to evaluate the performance of the current strategy and save the high-performance strategy as the benchmark strategy; Hierarchical knowledge base construction,Construct a hierarchical knowledge base, classify and store successful interference suppression strategies according to interference types and environmental characteristics to form a strategy library.
[0013] Furthermore, the reward function is: ; in Indicates time The reward value, Respectively indicate time and of users are inappropriate, Separately for the moment Perceptual evaluation of speech quality and short-term target intelligibility index to evaluate the output audio quality. represents the output signal after beamforming; For Action The computing resource consumption, Indicates time The action vector of They are the weight coefficients of user discomfort improvement, perceptual evaluation of speech quality, short-term target intelligibility index, and computing resource consumption.
[0014] Weight coefficient, balance the importance of each indicator.
[0015] Furthermore, the calculation formula for forming the strategy library is: ; in Represents the policy library, Respectively represent the first, second, and strategy, Indicates the total number of strategies in the strategy library; For new interference scenarios, we first retrieve strategies for similar scenarios from the strategy library as the initial strategies, and then perform targeted optimization to accelerate the convergence process.
[0016] A computer-readable storage medium is used to store computer-readable instructions. When the computer-readable instructions are read by a computer, the above-mentioned wireless audio transmission optimization system based on adaptive beamforming can be executed.
[0017] The beneficial effects of the present invention are: by innovatively combining deep learning, user experience perception and reinforcement learning technology, the problem of interference with wireless audio transmission is solved, the audio quality is improved while the user experience is optimized, and the invention is practical; Through the hierarchical attention interference source feature extraction network, the present invention can efficiently capture the time-frequency features of different scales, and improve the interference source recognition accuracy and recognition speed compared with traditional methods; By establishing a user discomfort perception model, the present invention can generate a personalized interference sensitivity map based on the user's physiological response, facial expression and behavioral data, thereby achieving targeted optimization of the user experience and improving the satisfaction reported by the user. By dynamically adjusting the beam zero point position and depth, the present invention can accurately control the beam forming effect, improve the average suppression strength of interference signals, and enhance the anti-interference performance of the system; By using reinforcement learning to continuously optimize the interference suppression strategy, the present invention can quickly adjust and adapt to new interference sources, shortening the adaptation time from the traditional several minutes to several seconds, thereby improving the system's ability to adapt to unknown interference environments. By prioritizing interference sources and selectively placing zero points, optimal interference suppression can be achieved with limited computing resources, reducing computing resource usage while extending device battery life. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A module diagram of a wireless audio transmission optimization system based on adaptive beamforming in the present invention; Figure 2 It is a flow chart of the multi-channel signal acquisition and preprocessing module of the present invention; Figure 3 It is a flowchart of the hierarchical attention distractor feature extraction module of the present invention; Figure 4 A flowchart of a user discomfort perception model construction module of the present invention; Figure 5 It is a flow chart of the beam zero point dynamic adjustment module of the present invention; Figure 6 It is a flowchart of the reinforcement learning optimization module of the present invention. DETAILED DESCRIPTION
[0019] The subject matter described herein will now be discussed with reference to example implementations. It should be understood that the discussion of these implementations is only to enable those skilled in the art to better understand and implement the subject matter described herein, and the functions and arrangements of the elements discussed may be changed without departing from the scope of protection of the present specification. Various examples may omit, replace, or add various processes or components as needed. In addition, the features described in some examples may also be combined in other examples.
[0020] At least one embodiment of the present invention discloses a wireless audio transmission optimization system based on adaptive beamforming, such as Figures 1 to 6 As shown, including: Multi-channel signal acquisition and preprocessing module, which acquires multi-channel acoustic signals and performs space-time domain preprocessing to generate feature matrix; Step 1.1, the microphone array collects multi-channel acoustic signals; By including An array of microphones (of which ) collects multi-channel acoustic signals. The signals collected by the array can be expressed as a vector: ; in Indicates time The multi-channel signal vector, Respectively indicate time The first, second, and The signal of a microphone, Indicates the number of microphones, Represents the transpose of a vector.
[0021] Step 1.2, applying a spatial filter to the acquired signal; Perform spatial filtering on the collected multi-channel signals, using spatial filters Process the signal to obtain the signal after spatial filtering: ; in Indicates time The spatially filtered signal vector, is the spatial filter weight matrix, represents the conjugate transpose, For the moment A multi-channel signal vector.
[0022] Step 1.3, use the exponentially weighted moving average algorithm; The exponentially weighted moving average algorithm is used to perform time domain smoothing on the spatially filtered signal. The calculation formula is: ; in Respectively indicate time and time The smoothed signal vector is is the smoothing factor, with a value range of (0, 1). The value makes the system more sensitive to new inputs, and smaller The value enhances the smoothing effect and reduces the impact of noise; the algorithm can effectively suppress transient fluctuations in the signal and improve the system's ability to identify stable interference sources.
[0023] Step 1.4, applying an adaptive noise suppression algorithm; Apply the adaptive noise suppression algorithm to the smoothed signal to reduce the noise of the signal and obtain the reduced noise signal .
[0024] Step 1.5, use short-time Fourier transform; The denoised signal is transformed into the time-frequency domain using short-time Fourier transform (STFT) to obtain the frequency domain signal representation: ; in Indicates the frequency The time window is The frequency domain signal vector, Indicates frequency, represents the time window, Represents short-time Fourier transform.
[0025] Step 1.6, generating interference source feature extraction input; Based on the time-frequency domain signal, construct the feature matrix: ; in represents the feature matrix, Represent the frequency and time window respectively. The frequency domain signal vector, Indicates the number of frequency points, Indicates the number of time windows.
[0026] Hierarchical attention distractor feature extraction module, which processes the feature matrix and identifies the type and location of distractors; The multi-scale convolution structure and channel attention module are used to build an interference source feature extraction network, and the feature matrix is processed to achieve accurate identification and positioning of the interference source; the feature matrix generated in the multi-channel signal acquisition and preprocessing module is used As input, deep feature extraction and analysis are performed; specifically including: Step 2.1, construct a multi-scale convolutional neural network and perform feature extraction; A multi-scale convolutional neural network structure is constructed, which contains three convolution kernels of different sizes (3×3, 5×5 and 7×7) to capture time-frequency features of different scales; among them: 3×3 convolution kernel network substructure: The first layer uses 32 3×3 convolution kernels for feature extraction, the second layer uses 64 3×3 convolution kernels to further extract features, and the third layer uses 128 3×3 convolution kernels to obtain fine-grained local features; 5×5 convolution kernel network substructure: The first layer uses 32 5×5 convolution kernels for feature extraction, the second layer uses 64 5×5 convolution kernels to further extract features, and the third layer uses 128 5×5 convolution kernels to obtain medium-scale features; 7×7 convolution kernel network substructure: The first layer uses 32 7×7 convolution kernels for feature extraction, the second layer uses 64 7×7 convolution kernels to further extract features, and the third layer uses 128 7×7 convolution kernels to obtain large-scale global features; Each network substructure applies batch normalization and ReLU activation function after each convolution operation to accelerate network convergence and enhance nonlinear expression capabilities. This multi-scale and multi-level design enables the network to simultaneously capture interference pattern characteristics of different time spans and frequency bandwidths.
[0027] The characteristic matrix obtained by the multi-channel signal acquisition and preprocessing module Three sets of feature maps are obtained by processing these three convolution kernels respectively. The calculation formula is: ; in Represents three feature maps of different sizes respectively; Indicates the use of Convolution operation with large and small convolution kernels, Represents the size of the convolution kernel; It is the feature matrix obtained by the multi-channel signal acquisition and preprocessing module.
[0028] Step 2.2, channel attention calculation; Apply the channel attention mechanism to each group of feature maps and calculate the channel attention weight; first perform global average pooling (GAP) on each group of feature maps to obtain the channel descriptor: ; in Indicates Channel descriptors of group feature maps, represents the global average pooling operation; For the Group feature maps; Then the channel attention weight is generated through a two-layer fully connected network, and the calculation formula is: ; in Indicates The channel attention weights of the group feature map, are the weight matrices of the first and second layers of the fully connected network, is the ReLU activation function, is the Sigmoid activation function.
[0029] Step 2.3, feature weighting; Apply the channel attention weights to the original feature map to obtain the weighted feature map: ; in Indicates The weighted feature map of the group feature map, For the The channel attention weights of the group feature map, Represents element-wise multiplication over the channel dimension.
[0030] Step 2.4, feature fusion; Fusion of weighted feature maps of different scales to obtain a comprehensive feature representation: ; in represents the comprehensive feature representation, represents the concatenation operation in the channel dimension, They represent the weighted feature maps of the first, second and third groups respectively.
[0031] Step 2.5, interference source identification and location; The interference source classifier processes the fusion features to identify and locate the interference source; the interference source classifier consists of two components: interference type classifier and interference source locator; The interference type classifier outputs the interference type probability vector: ; in represents the interference type probability vector, is the fully connected layer, is the fusion feature representation, is the normalized exponential function; The interference source locator outputs the spatial coordinates of the interference source, expressed as: ; in represents the spatial coordinates of the interference source, is the fully connected layer, It is the fusion feature representation; Interference source spatial coordinates: ; in represents the interference source coordinates, is the azimuth, is the pitch angle, For distance estimation.
[0032] The output of this module is the interference source characteristic representation: ; in Indicates the interference source characteristics, including two parts of information: Interference type probability vector Indicates various possible interference types and their probabilities; Interference source spatial coordinates The location of the interference source in three-dimensional space is accurately located.
[0033] This information will be used for subsequent beam null optimization adjustments to provide precise target positioning for interference suppression.
[0034] User discomfort perception model building module, which builds a user discomfort perception model based on user physiological response data, facial expression recognition and behavior analysis, and generates a personalized interference sensitivity map; Integrate physiological response sensing, facial expression recognition and implicit behavior analysis data to build a user discomfort perception model and generate a personalized interference sensitivity map. This module is carried out in parallel with the hierarchical attention interference source feature extraction module to provide a decision basis for the subsequent dynamic adjustment of the beam zero point; specifically, it includes: Step 3.1, physiological data collection; Collect user physiological response data, including heart rate variability (HRV), galvanic skin response (EDA), pupil size change and other physiological indicators, and obtain the physiological indicator vector: ; in Indicates time The physiological index vector, Respectively represent the first, second, and Physiological indicators at the time The value of is the total number of physiological indicators.
[0035] Step 3.2, facial expression feature recognition; Computer vision algorithms are used to analyze the user's facial expressions in real time, identify micro-expression changes, and obtain expression feature vectors: ; in Indicates time The facial expression feature vector of Respectively represent the first, second, and facial features at the moment The strength of is the total number of expression features.
[0036] Step 3.3, behavioral data monitoring; Monitor user implicit behaviors, such as repeated position adjustment, frequent content pause, volume adjustment, etc., and obtain the behavior feature vector: ; in Indicates time The behavioral feature vector, Respectively represent the first, second, and Behavioral characteristics at the moment The intensity or frequency of is the total number of behavioral features.
[0037] Step 3.4, environmental context information collection; Collect environmental context information, including audio content type, environmental noise level, user activity status, etc., to form a context feature vector: ; in Indicates time The context feature vector of Respectively represent the first, second, and context features at time The value of is the total number of context features.
[0038] Step 3.5, user inappropriate calculation; Build a user discomfort perception model and calculate user discomfort: ; in Indicates that the user is inappropriate; They represent physiological indicators, behavioral characteristics, and expression characteristics indexes respectively; The weight coefficients representing physiological indicators, behavioral characteristics, expression characteristics and contextual characteristics are optimized by gradient descent algorithm; Respectively represent Physiological indicators, behavioral characteristics and facial features at the moment The value of represents the weighting function of contextual features, Indicates the moment; Indicates the value from 1 to All The sum is calculated, where is the total number of physiological indicators; Indicates the value from 1 to All The sum is calculated, where is the total number of behavioral features; Indicates the value from 1 to All The sum is calculated, where is the total number of expression features.
[0039] Step 3.6, generating a personalized interference sensitivity map; Generate personalized interference sensitivity maps for different interference types and spatial locations based on historical data and user feedback: ; in represents the personalized interference sensitivity map, and denote the azimuth and elevation angles respectively, Indicates the type of interference, is the graph generation function, Represents historical inappropriate data, represents the interference type probability vector, Represents the coordinates of the interference source.
[0040] The output of this module includes user inappropriate and personalized interference sensitivity maps : User inappropriateness It is a real-time changing scalar that indicates the degree of discomfort currently felt by the user; Personalized interference sensitivity map It is a three-dimensional mapping function that describes the sensitivity distribution of users caused by different spatial locations and different types of interference sources.
[0041] This information will be combined with the interference source features of the hierarchical attention interference source feature extraction module to guide the precise placement and depth control of the beam zero point to achieve personalized interference suppression effect.
[0042] The beam null dynamic adjustment module dynamically adjusts the null position and depth of the beamforming algorithm based on the interference source type, location and personalized interference sensitivity map; Combined with the interference source characteristics and the output of the user discomfort perception model, the zero point position and depth of the beamforming algorithm are dynamically adjusted to achieve targeted interference suppression; this module integrates the output results of the first two modules. On the one hand, it uses the interference source characteristics identified by the hierarchical attention interference source feature extraction module Provides accurate information on the location of interference sources and generates a personalized interference sensitivity map based on the user discomfort perception model Implement user experience driven optimization adjustments; specifically including: Step 4.1, interference source suppression priority calculation; Distractor features obtained based on hierarchical attention distractor feature extraction module The disturbance sensitivity map obtained by the user discomfort perception model building module Calculate the suppression priority of each interference source: ; in Indicates The suppression priority of each interference source, are the interference source type confidence and user sensitivity weight coefficients, Respectively represent The azimuth and elevation angles of the interference source, Indicates The maximum type probability of interference sources, Indicates the most likely type of interference source.
[0043] Step 4.2, prioritization and screening; According to the calculated suppression priority, the identified interference sources are sorted and the one with the highest priority is selected. interference sources as the placement targets of beam nulls, where Depends on the computing resources and hardware capability limitations of the system.
[0044] Step 4.3, construct the interference source covariance matrix; First, construct the interference source covariance matrix: ; in is the interference source covariance matrix, For the The weight of each interference source and its priority Proportional, For the The array manifold vector corresponding to the direction of the interference source, The conjugate transpose of ; Indicates the value from 1 to All To sum, is the number of interference sources selected.
[0045] Step 4.4, calculating the beamforming weight vector; The beamforming weight vector is calculated using the linear constrained minimum variance (LCMV) algorithm: ; in represents the beamforming weight vector, represents the inverse matrix of the interference source covariance matrix; The conjugate transpose of is the array manifold matrix containing the desired signal direction and the interference source direction: ; in Respectively represent the first and The array manifold vectors in the direction of the interference source, is the array manifold vector of the desired signal direction; is a linear constraint vector, indicating that the gain in the direction of the desired signal is 1 and the gain in the direction of the interference source is 0, expressed as: ; in Indicates transpose.
[0046] Step 4.5, zero-point depth adaptive control; For specific types of interference sources, an adaptive zero-point depth control mechanism is introduced. For interference sources with higher priority and greater sensitivity, the depth of the corresponding zero point is increased to make the suppression effect more obvious. Zero-point depth adjustment is achieved by modifying the linear constraint vector: ; in For the A linear constraint vector for the direction of the interference source; For the The target gain in the direction of the interference source is usually a negative value. The larger its absolute value is, the deeper the zero point is and the stronger the suppression effect is. The calculation formula is: ; in is the proportionality coefficient, For the The priority of each interference source, For the The sensitivity of each interference source.
[0047] Step 4.6, real-time detection of user inappropriate changes; Real-time detection of user discomfort changes in the output of the user discomfort perception model building module; Exceeding the preset threshold When , the beam zero position and depth are quickly updated, and the calculation formula is: ; in Whether to trigger the rapid update of the beam zero position and depth. True means triggering, False means not triggering. Respectively represent the user inappropriateness at the current moment and the previous moment, is the time interval, are the incremental threshold and the absolute threshold, respectively. Indicates other situations.
[0048] Step 4.7, beamforming application; Apply the calculated beamforming weight vector to the multi-channel signal collected in the multi-channel signal acquisition and preprocessing module to obtain the output signal after beamforming: ; in The output signal after beamforming is a single-channel audio signal after spatial filtering. represents the beamforming weight vector The conjugate transpose of is used to perform weighted combination of multi-channel signals; Represents the multi-channel signal collected in the multi-channel signal acquisition and preprocessing module, including the original audio data from the microphone array.
[0049] The output of this module includes the beamforming weight vector and the output signal after beamforming : beamforming weight vector is a complex vector containing the weight coefficients for each microphone channel to achieve spatial filtering; Output signal after beamforming It is a single-channel audio signal that has been processed with interference suppression, with a higher signal-to-noise ratio and clearer target sound source content.
[0050] These results not only directly provide user-perceivable audio output, but also provide an evaluation basis for subsequent reinforcement learning optimization.
[0051] The reinforcement learning optimization module receives the execution results of the beam zero point dynamic adjustment module and user feedback data, and uses the reinforcement learning method to continuously optimize the interference suppression strategy to improve the system's adaptability to unknown interference environments; Adopting deep reinforcement learning method, based on user feedback and objective sound quality evaluation indicators, the interference suppression strategy is continuously optimized to improve the system's adaptability to unknown interference environments. This module forms a closed-loop feedback optimization for the above modules, and continuously improves the dynamic adjustment strategy of the beam zero point by evaluating the beam forming results in the beam zero point dynamic adjustment module. Specifically, it includes: Step 5.1, state space and action space definition; Defining the state space Contains the interference source characteristics of the current environment (from the hierarchical attention distractor feature extraction module), user inappropriateness (from the User Discomfort Perception Model Building Module), beamforming configuration parameters (from the beam zero point dynamic adjustment module) and historical interaction records; the state vector is expressed as: ; in Indicates time The state vector of Indicates time The interference source characteristics, Indicates time of users are inappropriate, represents the beamforming weight vector, For the moment Historical interaction features, including past State and action information for each time step; represents the number of historical time steps considered; Defining the action space It includes the adjustment operations of beamforming parameters, such as null position fine-tuning, null depth adjustment, priority reallocation, etc. The action vector is expressed as: ; in Indicates time The action vector of , Respectively represent the first and The adjustment amount of the azimuth, elevation and zero depth of the interference source, represents the adjustment vector for priority assignment, Indicates the number of selected interference sources.
[0052] Step 5.2, reward function design; Designing the reward function Comprehensively consider the user discomfort changes (from the user discomfort perception model building module), objective sound quality evaluation indicators (applied to the output of the beam zero point dynamic adjustment module) and system resource consumption: ; in Indicates time The reward value, Respectively indicate time of users are inappropriate, Separately for the moment Perceptual evaluation of speech quality and short-term target intelligibility index, evaluation of the audio quality output by the beam null dynamic adjustment module, represents the output signal after beamforming; For Action The computing resource consumption; They are the weight coefficients of user discomfort improvement, perceptual evaluation of speech quality, short-term target intelligibility index, and computing resource consumption.
[0053] Step 5.3, deep Q network construction; Using the Deep Q Network (DQN) as the core algorithm for reinforcement learning, we build a value network: ; in represents the value network, represents the state vector, represents the action vector, is the network parameter. The network contains 4 fully connected layers with the following structure: Input layer, state vector Dimensions: Hidden layer 1: 256 neurons, ReLU activation function; Hidden layer 2: 128 neurons, ReLU activation function; Hidden layer 3: 64 neurons, ReLU activation function; Output layer, action space Dimension.
[0054] Step 5.4, experience replay mechanism implementation; Use experience replay technology to improve learning efficiency and build an experience pool Storage transfer sample: ; in Indicates the current state. Indicates the current action. Indicates immediate reward, Indicates the next state; In each training, a mini-batch is randomly sampled from the experience pool to update the Q network, and the optimization objective is: ; in represents the loss function of the Q network, is the current Q network parameter, is the target network parameter, is a discount factor with a value range of [0, 1], balancing the importance of immediate rewards and future rewards; Indicates the experience pool Find the expectation of the state-action-reward-next-state quadruple sampled in the experiment; For all possible actions Take the maximum value; Indicates the current state. Indicates the current action. Indicates immediate reward, Indicates the next state, Indicates the next state Next action value.
[0055] Step 5.5, double Q learning application; The DoubleDQN mechanism is introduced to reduce the overestimation problem. The current network is used to select actions, and the target network is used to evaluate the value of actions: ; in represents the loss function of the Q network, is the current Q network parameter, is the target network parameter, is the discount factor; Indicates return The action to obtain the maximum value Indicates the current state. Indicates the current action. Indicates immediate reward, Indicates the next state; Indicates the experience pool Find the expectation of the state-action-reward-next-state quadruple sampled in .
[0056] Step 5.6, ϵ greedy strategy is applied; Use the ϵ greedy strategy for action selection, reducing the exploration rate ϵ over time: ; in Indicates time The action of choice; Indicates return The action to obtain the maximum value represents the current time step; Indicates time The exploration rate is calculated as: ; in are the initial exploration rate and the final exploration rate, respectively. is the decay rate, which controls the speed at which the exploration rate decreases; Represents the base of natural logarithms, approximately equal to 2.71828.
[0057] Step 5.7, application of periodic evaluation mechanism; Set up a periodic evaluation mechanism, each processing After 100 audio clips, the current strategy performance is evaluated and the high-performance strategy is saved as the baseline strategy; the evaluation indicators include average user discomfort (the output of the user discomfort perception model building module), audio quality score (the output of the beam zero point dynamic adjustment module) and system resource efficiency; is the preset evaluation cycle, which means each processing Each audio clip is evaluated once.
[0058] Step 5.8, hierarchical knowledge base construction; Build a hierarchical knowledge base to classify and store successful interference suppression strategies according to interference types and environmental characteristics to form a strategy library: ; in Represents the policy library, Respectively represent the first, second, and strategy, Indicates the total number of strategies in the strategy library; For new interference scenarios, we first retrieve the strategies of similar scenarios in the strategy library as the initial strategies, then perform targeted optimization to accelerate the convergence process, and finally generate the optimized interference suppression strategy. .
[0059] The output of this module includes the optimized interference suppression strategy and hierarchical knowledge base : Optimized interference suppression strategy Directly feed back to the beam zero point dynamic adjustment module to guide the dynamic adjustment of the beam zero point; Hierarchical Knowledge Base As the long-term memory of the system, it stores the optimal strategies in different scenarios, enabling the system to quickly adapt to new interference environments.
[0060] Through this closed-loop optimization mechanism, the system can continuously learn and improve interference suppression performance, forming an adaptive and evolving anti-interference capability.
[0061] Through the implementation of the above five modules, the wireless audio transmission optimization system based on adaptive beamforming provided in this embodiment forms a complete closed-loop system: The multi-channel signal acquisition and preprocessing module captures and preprocesses acoustic signals to provide basic data for subsequent analysis; the hierarchical attention interference source feature extraction module and the user discomfort perception model construction module analyze interference features from two dimensions: objective and subjective. One module identifies the type and location of interference sources, and the other module evaluates the actual impact of interference on users. The beam zero point dynamic adjustment module integrates the results of the first two steps to achieve targeted beam zero point adjustment to suppress interference; The reinforcement learning optimization module continuously optimizes the strategy of the entire system through reinforcement learning, forming an adaptive and evolving anti-interference capability; A close data flow and control flow relationship is formed between the modules, which can effectively solve the technical problems faced by wireless audio transmission in complex electromagnetic interference environments and improve audio transmission quality and user experience.
[0062] A computer-readable storage medium is used to store computer-readable instructions. When the computer-readable instructions are read by a computer, the above-mentioned wireless audio transmission optimization system based on adaptive beamforming can be executed.
[0063] Here, the present invention provides an implementation example: an audio transmission scenario of a wireless headset in a high-speed rail carriage; The high-speed rail carriage is a typical complex electromagnetic interference environment with multiple interference sources: Electromagnetic interference caused by high-speed rail operation; Wi-Fi and Bluetooth signals from passengers’ mobile phones, tablets and other wireless devices; Electromagnetic radiation generated by the in-car broadcasting system; Interference signals from other passengers’ wireless headphones; Communication signals transmitted by base stations along the high-speed railway; The characteristics of these interference sources are: diverse interference types, dynamically changing locations, unstable interference intensity, and users have high requirements for audio quality (hoping to hear music or calls clearly in a noisy environment).
[0064] The wireless headset configuration used in this application example is as follows: 4 omnidirectional microphones, evenly distributed on the surface of the headset; 1 high-precision acceleration sensor to detect user head movements; 1 infrared sensor to detect the wearing status of the headset; Integrated physiological sensors to monitor changes in skin conductivity; Bluetooth 5.2 chip, supporting dual-mode transmission; Low-power ARM processor, supporting neural network reasoning; 500mAh Lithium battery.
[0065] In the high-speed rail carriage environment, the system first collects acoustic signals through 4 microphones. The sampling frequency is set to 48kHz and the sampling accuracy is 24bit. Initial spatial filter Sets the beam to point in the direction of the user's face (determined by initial calibration). Smoothing factor Set it to 0.85 to effectively suppress transient noise while preserving signal details.
[0066] After time-frequency transformation, the system constructs the feature matrix It contains data from 128 frequency points and 50 time windows, forming a 128×50 feature matrix. In the high-speed rail environment, the system uses a finer frequency resolution for the low-frequency (50-300Hz) range to better capture the low-frequency interference characteristics unique to high-speed rail.
[0067] In this environment, the system identified three main types of interference sources: Low-frequency periodic electromagnetic interference generated by high-speed rail operation (confidence 92.3%, azimuth angle 15°, pitch angle -10°); The Wi-Fi signal of the neighboring passenger’s mobile phone (confidence level 89.7%, azimuth angle 75°, pitch angle 5°); In-car broadcasting system (confidence level 94.5%, azimuth 180°, elevation 30°).
[0068] Through multi-scale convolutional network processing, the system's feature extraction effects under three different convolution kernels show that the 7×7 convolution kernel is best at capturing the low-frequency interference characteristics of high-speed rail; the 3×3 convolution kernel is most sensitive to Wi-Fi signal feature extraction; and the 5×5 convolution kernel performs best when identifying interference from broadcast systems. The channel attention mechanism automatically assigns weights of 0.45, 0.35, and 0.2 to these three groups of features, reflecting their different impacts on audio quality.
[0069] Example of building a user discomfort perception model: The system detected significant differences in user sensitivity to different interference sources in the high-speed rail environment: The physiological indicators of users to the low-frequency interference of high-speed rail did not change significantly (galvanic skin response increased by 3%); The patient showed moderate discomfort due to interference from the neighbor’s Wi-Fi signal (galvanic skin response increased by 12%, micro-expression change index was 0.27); Extremely sensitive to interference from the broadcasting system (galvanic skin response increased by 28%, micro-expression change index 0.65, frequent adjustment of headphone position).
[0070] The personalized interference sensitivity map generated by the system shows that the interference sensitivity at 75° azimuth (neighboring seat position) and 180° azimuth (broadcast system direction) is 0.65 and 0.83 respectively (full score 1.0), while the sensitivity at 15° azimuth (direction of high-speed rail interference) is only 0.31; Based on the interference source characteristics and sensitivity maps, the system calculates the suppression priority of the three interference sources: Broadcast system: 0.78 (highest priority); Neighboring Wi-Fi: 0.62 (medium priority); High-speed rail interference: 0.35 (lowest priority).
[0071] The system applies the LCMV algorithm to calculate the beamforming weight vector and imposes a deeper null in the direction of the broadcast system ( ), apply a medium-depth zero point to the neighboring Wi-Fi direction ( ), applying a shallow zero point to the high-speed rail interference direction ( ).
[0072] When the user turns his head to talk with the person sitting next to him, the system detects the direction of the user's head, adjusts the direction of the main lobe of the beam in real time, and temporarily reduces the zero point depth in the direction of the person sitting next to him to ensure the clarity of the conversation.
[0073] During the one-hour high-speed rail journey, the system continuously learns and optimizes the interference suppression strategy: Initial stage (0-15 minutes): The system is mainly in the exploration stage ( ), try different zero point configurations Mid-term stage (15-40 minutes): exploration rate decreases ( ), the system learns that broadcast sounds are usually short-lived but cause severe interference, and adopts a "fast deep zero response" strategy Late stage (40-60 minutes): Exploration rate further reduced ( ), the system learns that the interference from high-speed railways will be enhanced in the tunnel section, and deepens the zero point in the corresponding direction in advance.
[0074] The system stores these learned strategies in a hierarchical knowledge base, labeled as "high-speed rail environment-user ID-date". The next time the user takes the high-speed rail, the system can directly call the strategy library as the initial strategy without re-learning.
[0075] Compared with ordinary wireless headphones that do not use the method of the present invention, the technical effects achieved by this application example in a high-speed rail environment are as follows: Interference source identification accuracy: 94.2% and 73.8% (traditional method); Signal-to-noise ratio improvement: 11.3dB vs. 5.1dB (traditional method); User satisfaction rating (average of 10 test users, full score 10): 8.7 and 6.2 (traditional method); Adaptation time to new interference: 4.3 seconds and 78 seconds (traditional method); Battery life: 5.8 hours and 4.6 hours (traditional method); This application example proves that the technical solution of the present invention has interference suppression effect and improved user experience in an actual complex electromagnetic interference environment, and fully achieves the expected technical goals.
[0076] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation mode. The above-mentioned specific implementation mode is merely illustrative and not restrictive. Under the guidance of this embodiment, ordinary technicians in this field can also make more forms of equivalent embodiments, all of which are within the protection of this embodiment.
Claims
1. A wireless audio transmission optimization system based on adaptive beamforming, characterized in that: include: Multi-channel signal acquisition and preprocessing module, which acquires multi-channel acoustic signals and performs space-time domain preprocessing to generate feature matrix; Hierarchical attention distractor feature extraction module, Process the signature matrix to identify the type and location of interference sources; User discomfort perception model building module, which builds a user discomfort perception model based on user physiological response data, facial expression recognition and behavior analysis, and generates a personalized interference sensitivity map; The beam null dynamic adjustment module dynamically adjusts the null position and depth of the beamforming algorithm based on the interference source type, location, and personalized interference sensitivity map; The reinforcement learning optimization module receives the execution results of the beam zero point dynamic adjustment module and user feedback data, and uses the reinforcement learning method to continuously optimize the interference suppression strategy to improve the system's adaptability to unknown interference environments.
2. According to claim 1, a wireless audio transmission optimization system based on adaptive beamforming is characterized in that: The multi-channel signal acquisition and preprocessing module includes: The microphone array collects multi-channel acoustic signals. The number of microphones ; Applying a spatial filter to the collected signal to perform spatial filtering processing on the multi-channel acoustic signal; Use the exponentially weighted moving average algorithm to smooth the spatially filtered signal; Applying an adaptive noise suppression algorithm to the smoothed signal to reduce signal noise; Use short-time Fourier transform to transform the denoised signal into time-frequency domain; Generate interference source feature extraction input and build a feature matrix based on time-frequency domain signals as the input of interference source feature extraction.
3. According to claim 1, a wireless audio transmission optimization system based on adaptive beamforming is characterized in that: The hierarchical attention interference source feature extraction module includes: Construct a multi-scale convolutional neural network, perform feature extraction, and build a network structure containing convolution kernels of different sizes to capture time-frequency features of different scales; Channel attention calculation: apply the channel attention mechanism to each group of feature maps and calculate the channel attention weight; Feature weighting, applying the channel attention weight to the original feature map to obtain a weighted feature map; Feature fusion, fusing weighted feature maps of different scales to obtain a comprehensive feature representation; Interference source identification and location, processing fusion features, realizing the identification and location of interference sources, and obtaining interference source feature representation.
4. The wireless audio transmission optimization system based on adaptive beamforming according to claim 3 is characterized in that: The multi-scale convolutional neural network includes three convolution kernels of different sizes, namely 3×3, 5×5 and 7×7, which capture time-frequency features of different scales.
5. The wireless audio transmission optimization system based on adaptive beamforming according to claim 1, characterized in that: The user discomfort perception model building module includes: Physiological data collection, collecting user physiological response data; Facial expression recognition: using computer vision algorithms to analyze user facial expressions in real time and identify micro-expression changes; Behavioral data monitoring, monitoring users’ implicit behaviors; Environmental context information collection; User discomfort calculation, building a user discomfort perception model based on the physiological data, the expression characteristics, the behavior data and the environmental context information, and calculating user discomfort; Personalized interference sensitivity map generation, based on historical data and user feedback, generates personalized interference sensitivity maps for different interference types and spatial locations.
6. The wireless audio transmission optimization system based on adaptive beamforming according to claim 1, characterized in that: The beam zero point dynamic adjustment module includes: Interference source suppression priority calculation: based on interference source characteristics and interference sensitivity maps, calculate the suppression priority of each interference source; Priority sorting and screening: sort the identified interference sources according to the calculated suppression priority and select the one with the highest priority The interference source is used as the placement target of the beam null; Constructing the interference source covariance matrix, constructing the interference source covariance matrix; Calculate the beamforming weight vector, and use the linear constrained minimum variance algorithm to calculate the beamforming weight vector; Zero point depth adaptive control: for specific types of interference sources, an adaptive zero point depth control mechanism is introduced to increase the corresponding zero point depth for interference sources with higher priority and greater sensitivity; Real-time detection of user discomfort changes. Real-time detection of user discomfort changes. When the user discomfort exceeds the preset threshold, it triggers the rapid update of the beam zero point position and depth. The beamforming is applied by applying the calculated beamforming weight vector to the input signal to obtain the beamformed output signal.
7. The wireless audio transmission optimization system based on adaptive beamforming according to claim 1, characterized in that: The reinforcement learning optimization module includes: State space and action space definition, including the interference source characteristics of the current environment User inappropriateness Beamforming Configuration Parameters The state space of historical interaction records defines the action space containing the adjustment operations of the beamforming parameters; Reward function design: Design a reward function that comprehensively considers user discomfort changes, objective sound quality evaluation indicators, and system resource consumption; Deep Q network construction, as the core algorithm of reinforcement learning, builds value network; The experience replay mechanism is implemented to store experience samples of the interaction between the system and the environment and improve learning efficiency; Double Q-learning is applied to reduce the overestimation problem; ϵ-greedy strategy application, using ϵ-greedy strategy for action selection, reducing the exploration rate over time; Periodic evaluation mechanism application: set up a periodic evaluation mechanism to evaluate the performance of the current strategy and save the high-performance strategy as the benchmark strategy; Hierarchical knowledge base construction,Construct a hierarchical knowledge base, classify and store successful interference suppression strategies according to interference types and environmental characteristics to form a strategy library.
8. The wireless audio transmission optimization system based on adaptive beamforming according to claim 7, characterized in that: The reward function is: ; in Indicates time The reward value, Respectively indicate time and of users are inappropriate, Separately for the moment Perceptual evaluation of speech quality and short-term target intelligibility index to evaluate the output audio quality. represents the output signal after beamforming; For Action The computing resource consumption, Indicates time The action vector of They are the weight coefficients of user discomfort improvement, perceptual evaluation of speech quality, short-term target intelligibility index, and computing resource consumption.
9. The wireless audio transmission optimization system based on adaptive beamforming according to claim 7, characterized in that: The calculation formula for forming the strategy library is: ; Respectively represent the first, second, and strategy, Indicates the total number of strategies in the strategy library; For new interference scenarios, we first retrieve strategies for similar scenarios from the strategy library as the initial strategies, and then perform targeted optimization to accelerate the convergence process.
10. A computer-readable storage medium, characterized in that: It is used to store computer-readable instructions, and when the computer-readable instructions are read by a computer, it can run a wireless audio transmission optimization system based on adaptive beamforming as described in any one of claims 1-9.
Citation Information
Patent Citations
Cognitive interference integrated beam forming method based on deep reinforcement learning
CN117614501A
Dynamically adapting sound based on environmental characterization
US10511906B1
Beam-steering antenna array audio accessory device
US20220368394A1
Cited By
Interference suppression method for low-power-consumption HRF + HPLC dual-mode chip
CN120691968A
Communication detection system and detection method based on Beidou positioning
CN120934609A
A communication detection system and method based on BeiDou positioning
CN120934609B
Audio recording interference method and system based on acoustic feature feedback and beamforming
CN122640072A