An optimized system for wireless audio transmission based on adaptive beamforming
By using adaptive beamforming technology and deep learning methods in wireless audio transmission systems, complex interference sources are identified and suppressed, and beam formation is dynamically adjusted to optimize user experience, and the problems of insufficient ability to identify complex interference sources, insufficient consideration of users' subjective feelings, and poor adaptability in the prior art are solved, and high-quality wireless audio transmission is achieved.
Patent Information
- Application Number
- CN202510494821.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The existing anti-interference technology of wireless audio transmission has problems such as insufficient ability to identify complex interference sources, insufficient consideration of users' subjective feelings, and poor adaptability, and cannot meet the high-quality needs of wireless audio transmission in complex electromagnetic environments.
The wireless audio transmission optimization system based on adaptive beamforming is adopted. The system includes a multi-channel signal acquisition and preprocessing module, a layered attention interference source feature extraction module, a user discomfort perception model construction module, a beam zero-point dynamic adjustment module and a reinforcement learning optimization module. Through deep learning, user experience perception and reinforcement learning technology, precise identification and suppression of complex interference sources are achieved, and beam formation is dynamically adjusted to optimize user experience.
It improves the anti-interference performance and audio quality of wireless audio transmission, enhances the optimization of user experience, can quickly adapt to new interference environments, reduces the use of computing resources and extends the battery life of the device.
Smart Images

Figure CN120017192B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communication technologies, and more specifically, to a wireless audio transmission optimization system based on adaptive beamforming. Background Art
[0002] Wireless audio devices have been widely used in multiple fields such as consumer electronics, conference systems, and smart homes in recent years due to their portability and wireless connection advantages. However, in the actual usage environment, wireless audio transmission often faces complex and variable electromagnetic interference problems, resulting in a decline in audio transmission quality and impaired user experience.
[0003] Existing anti-interference technologies for wireless audio transmission mainly include methods such as spectrum spreading, adaptive frequency hopping, and traditional beamforming. The spectrum spreading technology disperses the signal energy over a wider frequency band to reduce the impact of interference, but this technology has limited resistance to transient strong interference; the adaptive frequency hopping technology can avoid interference at fixed frequency points, but it is not effective when the interference source covers multiple frequency bands; the traditional beamforming technology can suppress interference from specific directions through spatial filtering, but it has the following main defects: the traditional beamforming technology usually relies on a preset interference model and is difficult to effectively deal with unknown or complex and variable interference sources in the environment; the existing interference source feature extraction algorithms are not efficient enough to quickly and accurately distinguish between main interference sources and secondary interference sources, resulting in unreasonable allocation of computing resources; the existing system lacks a direct feedback mechanism for the actual perception experience of users; the existing technologies usually adopt a fixed optimization strategy and cannot adaptively adjust according to different environments and usage scenarios, resulting in insufficient adaptability of the system when facing new types of interference sources.
[0004] In summary, the existing anti-interference technologies for wireless audio transmission have problems such as insufficient ability to identify complex interference sources, insufficient consideration of the subjective feelings of users, and lack of strong adaptability, and cannot meet the high-quality requirements of wireless audio transmission in complex electromagnetic environments. Therefore, there is an urgent need for a wireless audio transmission optimization method that can accurately identify various interference sources, dynamically adjust interference suppression strategies according to the personalized needs of users, and has the ability of continuous learning. Summary of the Invention
[0005] The present invention provides a wireless audio transmission optimization system based on adaptive beamforming to solve the technical problems of insufficient ability to identify complex interference sources, insufficient consideration of the subjective feelings of users, and lack of strong adaptability in related technologies.
[0006] The present invention provides a wireless audio transmission optimization system based on adaptive beamforming, including:
[0007] A multi-channel signal acquisition and preprocessing module that acquires multi-channel acoustic signals and performs spatio-temporal preprocessing to generate a feature matrix;
[0008] The hierarchical attention interference source feature extraction module processes the feature matrix to identify the type and location of the interference source;
[0009] The user discomfort perception model construction module constructs a user discomfort perception model based on user physiological response data, facial expression recognition, and behavior analysis, and generates a personalized interference sensitivity map;
[0010] The beam null dynamic adjustment module dynamically adjusts the null position and depth of the beamforming algorithm in combination with the interference source type, location, and personalized interference sensitivity map;
[0011] The reinforcement learning optimization module receives the execution results of the beam null dynamic adjustment module and user feedback data, and uses reinforcement learning methods to continuously optimize the interference suppression strategy to improve the system's adaptability to unknown interference environments.
[0012] Furthermore, the multi-channel signal acquisition and preprocessing module includes:
[0013] The microphone array collects multi-channel acoustic signals, and the number of microphones for collecting multi-channel acoustic signals ;
[0014] Apply a spatial filter to the collected signals to perform spatial filtering on the multi-channel acoustic signals;
[0015] Use the exponentially weighted moving average algorithm to smooth the spatially filtered signals;
[0016] Apply an adaptive noise suppression algorithm to the smoothed signals to perform signal noise reduction;
[0017] Use the short-time Fourier transform to perform time-frequency domain transformation on the noise-reduced signals;
[0018] Generate the input for interference source feature extraction, and construct a feature matrix based on the time-frequency domain signals as the input for interference source feature extraction.
[0019] Furthermore, the hierarchical attention interference source feature extraction module includes:
[0020] Construct a multi-scale convolutional neural network for feature extraction processing, and construct a network structure containing convolutional kernels of different sizes to capture time-frequency features of different scales;
[0021] Channel attention calculation, apply the channel attention mechanism to each group of feature maps to calculate the channel attention weights;
[0022] Feature weighting, apply the channel attention weights to the original feature maps to obtain weighted feature maps;
[0023] Feature fusion, fusing weighted feature maps of different scales to obtain a comprehensive feature representation;
[0024] Interference source identification and localization, processing the fused features to achieve the identification and localization of interference sources and obtain interference source feature representations.
[0025] Furthermore, the construction of the multi-scale convolutional neural network includes three types of convolutional kernels with different sizes, namely 3×3, 5×5, and 7×7, to capture time-frequency features of different scales.
[0026] Furthermore, the user discomfort perception model construction module includes:
[0027] Physiological data collection, collecting users' physiological response data;
[0028] Facial expression feature recognition, analyzing users' facial expressions in real time through computer vision algorithms to identify micro-expression changes;
[0029] Behavior data monitoring, monitoring users' implicit behaviors;
[0030] Collection of environmental context information;
[0031] Calculation of user discomfort level, constructing a user discomfort perception model based on the physiological data, the facial expression features, the behavior data, and the environmental context information, and calculating the user discomfort level;
[0032] Generation of personalized interference sensitivity maps, generating personalized interference sensitivity maps for different interference types and spatial positions based on historical data and user feedback.
[0033] Furthermore, the beam null dynamic adjustment module includes:
[0034] Calculation of interference source suppression priorities, calculating the suppression priorities of each interference source based on the interference source features and the interference sensitivity maps;
[0035] Priority ranking and screening, ranking the identified interference sources according to the calculated suppression priorities and selecting the interference source with the highest priority as the placement target of the beam null;
[0036] Construction of the interference source covariance matrix, constructing the interference source covariance matrix;
[0037] Calculation of the beamforming weight vector, calculating the beamforming weight vector using the linearly constrained minimum variance algorithm;
[0038] Adaptive null depth control, introducing an adaptive null depth control mechanism for specific types of interference sources, increasing the corresponding null depth for interference sources with higher priorities and greater sensitivities;
[0039] Real-time detect the change of user discomfort. When the user discomfort exceeds the preset threshold, trigger the rapid update of the beam null position and depth.
[0040] Beamforming application. Apply the calculated beamforming weight vector to the input signal to obtain the beamformed output signal.
[0041] Furthermore, the reinforcement learning optimization module includes:
[0042] Definition of state space and action space. Define the state space that includes the interference source characteristics of the current environment User discomfort Beamforming configuration parameters and historical interaction records, and define the action space that includes the adjustment operations of the beamforming parameters;
[0043] Design of reward function. Design a reward function that comprehensively considers the change of user discomfort, objective speech quality evaluation indicators, and system resource consumption;
[0044] Construction of deep Q network. As the core algorithm of reinforcement learning, construct a value network;
[0045] Implementation of experience replay mechanism. Store the experience samples of the interaction between the system and the environment and improve the learning efficiency;
[0046] Application of double Q learning. Reduce the overestimation problem;
[0047] Application of ϵ-greedy policy. Adopt the ϵ-greedy policy for action selection and decrease the exploration rate over time;
[0048] Application of periodic evaluation mechanism. Set up a periodic evaluation mechanism to evaluate the performance of the current policy and save the high-performance policy as the benchmark policy;
[0049] Construction of hierarchical knowledge base. Construct a hierarchical knowledge base, classify and store the successful interference suppression strategies according to the interference type and environmental characteristics to form a policy library.
[0050] Furthermore, the reward function is:
[0051] ;
[0052] Where represents the reward value at time , respectively represent the user discomfort at times and , are respectively the perceptual evaluation speech quality and short-time objective intelligibility index at time , evaluating the output audio quality, Represents the output signal after beamforming; Is the action Of the computing resource consumption, Represents the moment Of the action vector; Are the weight coefficients of user discomfort improvement, perceived evaluation of speech quality, short-term target intelligibility index, and computing resource consumption respectively.
[0053] The weight coefficients balance the importance of various indicators.
[0054] Furthermore, the calculation formula for forming the policy library is:
[0055] ;
[0056] Where Represents the policy library, Respectively represent the 1st, 2nd, and the th policies, Represents the total number of policies in the policy library;
[0057] For a new interference scenario, first retrieve the policy of the similar scenario in the policy library as the initial policy, and then perform targeted optimization to accelerate the convergence process.
[0058] A computer-readable storage medium is used to store computer-readable instructions, which can run an optimization system for wireless audio transmission based on adaptive beamforming as described above when read by a computer.
[0059] The beneficial effects of the present invention are as follows: By innovatively combining deep learning, user experience perception, and reinforcement learning technologies, the problem of interference in wireless audio transmission is solved, the audio quality is improved while the user experience is optimized, and it has practicality;
[0060] Through the hierarchical attention interference source feature extraction network, the present invention can efficiently capture time-frequency features of different scales, improve the interference source recognition accuracy compared with traditional methods, and enhance the recognition speed;
[0061] By establishing a user discomfort perception model, the present invention can generate a personalized interference sensitivity map based on the user's physiological reactions, facial expressions, and behavior data, realizing targeted optimization of the user experience and improving the satisfaction reported by users;
[0062] By dynamically adjusting the position and depth of the beam null, the present invention can precisely control the beamforming effect, improve the average suppression intensity of interference signals, and enhance the anti-interference performance of the system;
[0063] By continuously optimizing the interference suppression strategy using reinforcement learning, the present invention can quickly adjust and adapt when facing new interference sources. The adaptation time is shortened from the traditional level of several minutes to the level of several seconds, improving the system's adaptability to unknown interference environments;
[0064] Through the prioritization of interference sources and selective null placement, it is possible to achieve the optimal interference suppression effect with limited computing resources, reducing the occupancy of computing resources and extending the battery life of the device. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 It is a module diagram of an optimized wireless audio transmission system based on adaptive beamforming in the present invention;
[0066] Figure 2 It is a flowchart of the multi-channel signal acquisition and preprocessing module of the present invention;
[0067] Figure 3 It is a flowchart of the hierarchical attention interference source feature extraction module of the present invention;
[0068] Figure 4 It is a flowchart of the user discomfort perception model construction module of the present invention;
[0069] Figure 5 It is a flowchart of the beam null dynamic adjustment module of the present invention;
[0070] Figure 6 It is a flowchart of the reinforcement learning optimization module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] Now, the subject matter described herein will be discussed with reference to exemplary embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.
[0072] In at least one embodiment of the present invention, an optimized wireless audio transmission system based on adaptive beamforming is disclosed, as Figures 1 to 6 shown, including:
[0073] A multi-channel signal acquisition and preprocessing module that acquires multi-channel acoustic signals and performs spatio-temporal domain preprocessing to generate a feature matrix;
[0074] Step 1.1, the microphone array acquires multi-channel acoustic signals;
[0075] By including An array of microphones (where ) acquires multi-channel acoustic signals, and the signals acquired by the array can be represented as a vector:
[0076] ;
[0077] where represents the multi-channel signal vector at time , respectively represent the signals of the first, second, th, th microphone at time , represents the transpose of the vector.
[0078] Step 1.2, apply a spatial filter to the acquired signals;
[0079] Perform spatial filtering on the acquired multi-channel signals, and use the spatial filter to process the signals to obtain the spatially filtered signals:
[0080] ;
[0081] where represents the spatially filtered signal vector at time , is the spatial filtering weight matrix, represents the conjugate transpose, is the multi-channel signal vector at time .
[0082] Step 1.3, use the exponentially weighted moving average algorithm;
[0083] Perform time-domain smoothing on the spatially filtered signals using the exponentially weighted moving average algorithm, and the calculation formula is:
[0084] ;
[0085] where respectively represent the smoothed signal vectors at time and time , is the smoothing factor, and its value range is (0, 1). A larger value makes the system more sensitive to new inputs, and a smaller value enhances the smoothing effect and reduces the noise impact; this algorithm can effectively suppress the transient fluctuations in the signals and improve the system's ability to identify stable interference sources.
[0086] Step 1.4, apply an adaptive noise cancellation algorithm;
[0087] Apply the adaptive noise suppression algorithm to the smoothed signal for signal denoising to obtain the denoised signal 。
[0088] Step 1.5, use the short-time Fourier transform;
[0089] Perform time-frequency domain transformation on the denoised signal, and use the short-time Fourier transform (STFT) to obtain the frequency-domain signal representation:
[0090] ;
[0091] where represents the frequency-domain signal vector with frequency and time window ; represents the frequency, represents the time window, represents the short-time Fourier transform.
[0092] Step 1.6, generate the input for interference source feature extraction;
[0093] Based on the time-frequency domain signal, construct the feature matrix:
[0094] ;
[0095] where represents the feature matrix, respectively represent the frequency-domain signal vectors with frequency and time window , represents the number of frequency points, represents the number of time windows.
[0096] The hierarchical attention interference source feature extraction module processes the feature matrix to identify the type and location of the interference source;
[0097] Adopt a multi-scale convolutional structure and a channel attention module to construct an interference source feature extraction network, process the feature matrix, and achieve accurate identification and positioning of the interference source; use the feature matrix generated in the multi-channel signal acquisition and preprocessing module as the input for in-depth feature extraction and analysis; specifically including:
[0098] Step 2.1, construct a multi-scale convolutional neural network for feature extraction processing;
[0099] Construct a multi-scale convolutional neural network structure, including three different sizes of convolutional kernels (3×3, 5×5, and 7×7) to capture time-frequency features of different scales; where:
[0100] 3×3 Convolution Kernel Network Substructure: In the first layer, 32 3×3 convolution kernels are used for feature extraction. In the second layer, 64 3×3 convolution kernels are used to further extract features. In the third layer, 128 3×3 convolution kernels are used to obtain fine-grained local features;
[0101] 5×5 Convolution Kernel Network Substructure: In the first layer, 32 5×5 convolution kernels are used for feature extraction. In the second layer, 64 5×5 convolution kernels are used to further extract features. In the third layer, 128 5×5 convolution kernels are used to obtain medium-scale features;
[0102] 7×7 Convolution Kernel Network Substructure: In the first layer, 32 7×7 convolution kernels are used for feature extraction. In the second layer, 64 7×7 convolution kernels are used to further extract features. In the third layer, 128 7×7 convolution kernels are used to obtain large-scale global features;
[0103] After each layer of convolution operation in each network substructure, batch normalization and ReLU activation function are applied to accelerate network convergence and enhance the non-linear expression ability. This multi-scale and multi-level design enables the network to capture the interference pattern features of different time spans and frequency bandwidths simultaneously.
[0104] For the feature matrix obtained by the multi-channel signal acquisition and preprocessing module It is processed by these 3 types of convolution kernels respectively to obtain 3 groups of feature maps The calculation formula is:
[0105] ;
[0106] Where respectively represent 3 types of feature maps with different sizes; represents the convolution operation using size convolution kernels, represents the size of the convolution kernel; is the feature matrix obtained by the multi-channel signal acquisition and preprocessing module.
[0107] Step 2.2, Channel Attention Calculation;
[0108] Apply the channel attention mechanism to each group of feature maps to calculate the channel attention weights; First, perform global average pooling (GAP) on each group of feature maps to obtain the channel descriptor:
[0109] ;
[0110] Where represents the channel descriptor of the th group of feature maps, represents the global average pooling operation; is the th group of feature maps;
[0111] Then, the channel attention weights are generated through a two - layer fully - connected network, and the calculation formula is:
[0112] ;
[0113] where represents the channel attention weight of the th group of feature maps, are the weight matrices of the first - layer and second - layer fully - connected networks respectively, is the ReLU activation function, is the Sigmoid activation function.
[0114] Step 2.3, Feature weighting;
[0115] Apply the channel attention weights to the original feature maps to obtain weighted feature maps:
[0116] ;
[0117] where represents the weighted feature map of the th group of feature maps, is the channel attention weight of the th group of feature maps, represents the element - wise multiplication in the channel dimension.
[0118] Step 2.4, Feature fusion;
[0119] Fuse the weighted feature maps of different scales to obtain a comprehensive feature representation:
[0120] ;
[0121] where represents the comprehensive feature representation, represents the concatenation operation in the channel dimension, represent the weighted feature maps of the first group, the second group, and the third group respectively.
[0122] Step 2.5, Interference source identification and localization;
[0123] Process the fused features through an interference source classifier to achieve the identification and localization of interference sources; the interference source classifier consists of two components: an interference type classifier and an interference source locator;
[0124] The interference type classifier outputs an interference type probability vector:
[0125] ;
[0126] where Represents the interference type probability vector, is a fully connected layer, is the fused feature representation, is the normalized exponential function;
[0127] The interferer locator outputs the spatial coordinates of the interferer, expressed as:
[0128] ;
[0129] where represents the spatial coordinates of the interferer, is a fully connected layer, is the fused feature representation;
[0130] Spatial coordinates of the interferer:
[0131] ;
[0132] where represents the coordinates of the interferer, is the azimuth angle, is the elevation angle, is the distance estimate.
[0133] The output result of this module is the interferer feature representation:
[0134] ;
[0135] where represents the interferer feature representation, containing two parts of information:
[0136] Interference type probability vector represents various possible interference types and their probabilities;
[0137] Spatial coordinates of the interferer Precisely locates the position of the interferer in three-dimensional space.
[0138] This information will be used for subsequent beam null optimization and adjustment to provide accurate target positioning for interference suppression.
[0139] The user discomfort perception model construction module constructs a user discomfort perception model based on user physiological response data, facial expression recognition, and behavior analysis, and generates a personalized interference sensitivity map;
[0140] Integrate physiological response sensing, facial expression recognition, and implicit behavior analysis data to construct a user discomfort perception model and generate a personalized interference sensitivity map. This module runs in parallel with the hierarchical attention interferer feature extraction module to jointly provide a decision-making basis for subsequent beam null dynamic adjustment; specifically including:
[0141] Step 3.1, Physiological data collection;
[0142] Collect the user's physiological response data, including physiological indicators such as heart rate variability (HRV), electrodermal activity (EDA), and pupil size changes, to obtain a physiological indicator vector:
[0143] ;
[0144] where represents the physiological indicator vector at time , respectively represent the values of the 1st, 2nd, and th physiological indicators at time , is the total number of physiological indicators.
[0145] Step 3.2, Facial expression feature recognition;
[0146] Analyze the user's facial expressions in real time through computer vision algorithms to identify micro-expression changes and obtain a facial expression feature vector:
[0147] ;
[0148] where represents the facial expression feature vector at time , respectively represent the intensities of the 1st, 2nd, and rd facial expression features at time , is the total number of facial expression features.
[0149] Step 3.3, Behavioral data monitoring;
[0150] Monitor the user's implicit behaviors, such as repeated position adjustments, frequent content pauses, volume adjustments, etc., to obtain a behavioral feature vector:
[0151] ;
[0152] where represents the behavioral feature vector at time , respectively represent the intensities or frequencies of the 1st, 2nd, and th behavioral features at time , is the total number of behavioral features.
[0153] Step 3.4, Collection of environmental context information;
[0154] Collect environmental context information, including audio content type, environmental noise level, user activity status, etc., to form a context feature vector:
[0155] ;
[0156] where represents the context feature vector at a moment , and respectively represent the values of the 1st, 2nd, th context feature at the moment , and is the total number of context features.
[0157] Step 3.5, User discomfort calculation;
[0158] Construct a user discomfort perception model and calculate the user discomfort:
[0159] ;
[0160] where represents the user discomfort; respectively represent the physiological index, behavior feature, and expression feature index; respectively represent the weight coefficients of the physiological index, behavior feature, expression feature, and context feature, which are optimized by the gradient descent algorithm; respectively represent the th physiological index, the th behavior feature, and the th expression feature at the moment ;
[0161] represents the weighted function of the context feature, represents the moment; represents the sum from 1 to for all , where is the total number of physiological indexes; represents the sum from 1 to for all , where is the total number of behavior features; represents the sum from 1 to for all , where is the total number of expression features.
[0162] Step 3.6, Generation of personalized interference sensitivity map;
[0163] Based on historical data and user feedback, generate a personalized interference sensitivity map for different interference types and spatial locations:
[0164] ;
[0165] Wherein represents the personalized interference sensitivity map, and represent the azimuth angle and the elevation angle respectively, represents the interference type, is the map generation function, represents the historical discomfort data, represents the interference type probability vector, represents the coordinates of the interference source.
[0166] The output result of this module includes the user discomfort and the personalized interference sensitivity map :
[0167] User discomfort is a scalar that changes in real time, representing the degree of discomfort currently felt by the user;
[0168] Personalized interference sensitivity map is a three-dimensional mapping function that describes the sensitivity distribution caused by different spatial positions and different types of interference sources to the user.
[0169] These information will be combined with the interference source features of the hierarchical attention interference source feature extraction module to guide the precise placement and depth control of the beam nulls, and achieve personalized interference suppression effects.
[0170] The beam null dynamic adjustment module combines the interference source type, position and the personalized interference sensitivity map to dynamically adjust the null position and depth of the beamforming algorithm;
[0171] Combining the interference source features and the output of the user discomfort perception model, dynamically adjust the null position and depth of the beamforming algorithm to achieve targeted interference suppression; this module integrates the output results of the previous two modules. On the one hand, it uses the interference source features identified by the hierarchical attention interference source feature extraction module to provide accurate interference source position information, and on the other hand, combines the personalized interference sensitivity map generated by the user discomfort perception model to achieve user experience-driven optimization adjustment; specifically including:
[0172] Step 4.1, calculating the interference source suppression priority;
[0173] Based on the interference source features obtained from the hierarchical attention interference source feature extraction module and the interference sensitivity map obtained from the user discomfort perception model construction module calculate the suppression priority of each interference source:
[0174] ;
[0175] Wherein represents the suppression priority of the th interference source, are the weight coefficients of the interference source type confidence and user sensitivity respectively, respectively represent the azimuth and elevation angles of the th interference source, represents the maximum type probability of the th interference source, indicating the most likely type of this interference source.
[0176] Step 4.2, Priority sorting and screening;
[0177] According to the calculated suppression priority, sort the identified interference sources, and select the th interference source with the highest priority as the placement target of the beam null, where depends on the computational resources and hardware capabilities of the system.
[0178] Step 4.3, Construct the interference source covariance matrix;
[0179] First, construct the interference source covariance matrix:
[0180] ;
[0181] Wherein is the interference source covariance matrix, is the weight of the th interference source, which is proportional to its priority is the array manifold vector in the direction corresponding to the th interference source, the conjugate transpose of; represents the sum from 1 to for all , is the number of selected interference sources.
[0182] Step 4.4, Calculate the beamforming weight vector;
[0183] Use the linearly constrained minimum variance (LCMV) algorithm to calculate the beamforming weight vector:
[0184] ;
[0185] Wherein represents the beamforming weight vector, represents the inverse matrix of the interference source covariance matrix; the conjugate transpose of, is the array manifold matrix including the desired signal direction and the interference source direction:
[0186] ;
[0187] where respectively represent the array manifold vectors of the first and the th interference source directions, is the array manifold vector of the desired signal direction;
[0188] is the linear constraint vector, indicating that the gain of the desired signal direction is 1 and the gain of the interference source direction is 0, expressed as:
[0189] ;
[0190] where represents the transpose.
[0191] Step 4.5, zero-depth adaptive control;
[0192] For a specific type of interference source, an adaptive zero-depth control mechanism is introduced. For interference sources with higher priority and greater sensitivity, increase the depth of their corresponding zeros to make the suppression effect more obvious. The zero-depth adjustment is achieved by modifying the linear constraint vector:
[0193] ;
[0194] where is the linear constraint vector of the th interference source direction; is the target gain of the th interference source direction, usually negative, and the larger its absolute value, the deeper the zero, and the stronger the suppression effect. The calculation formula of
[0195] ;
[0196] where is the proportionality coefficient, is the priority of the th interference source, is the sensitivity of the th interference source.
[0197] Step 4.6, real-time detection of user discomfort change;
[0198] Real-time detect the change of user discomfort output by the user discomfort perception model construction module; when the user discomfort exceeds the preset threshold When it is triggered, the beam null position and depth are updated quickly, and the calculation formula is:
[0199] ;
[0200] where indicates whether to trigger the quick update of the beam null position and depth. True means trigger, and False means not to trigger. respectively represent the user discomfort at the current moment and the previous moment. is the time interval. are the incremental threshold and the absolute threshold respectively. represents other situations.
[0201] Step 4.7, beamforming application;
[0202] Apply the calculated beamforming weight vector to the multi-channel signals collected by the multi-channel signal acquisition and preprocessing module to obtain the output signal after beamforming:
[0203] ;
[0204] where represents the output signal after beamforming, which is a single-channel audio signal after spatial filtering processing; represents the beamforming weight vector The conjugate transpose of, which performs weighted combination on multi-channel signals; represents the multi-channel signals collected by the multi-channel signal acquisition and preprocessing module, which contains the original audio data from the microphone array.
[0205] The output results of this module include the beamforming weight vector and the output signal after beamforming : The beamforming weight vector is a complex vector, which contains the weight coefficients for each microphone channel and realizes spatial filtering;
[0206] The output signal after beamforming is a single-channel audio signal after interference suppression processing, with a higher signal-to-noise ratio and clearer target sound source content.
[0207] These results not only directly provide the user-perceivable audio output, but also provide an evaluation basis for subsequent reinforcement learning optimization.
[0208] The reinforcement learning optimization module receives the execution results of the beam null dynamic adjustment module and the user feedback data, and continuously optimizes the interference suppression strategy by using the reinforcement learning method to improve the system's adaptability to unknown interference environments;
[0209] Adopt a deep reinforcement learning method to continuously optimize the interference suppression strategy based on user feedback and objective sound quality evaluation metrics, improving the system's adaptability to unknown interference environments. This module forms a closed-loop feedback optimization for the aforementioned modules. By evaluating the beamforming results in the beam null dynamic adjustment module, the dynamic adjustment strategy of the beam null is continuously improved. Specifically, it includes:
[0210] Step 5.1, Define the state space and action space;
[0211] Define the state space Include the interference source characteristics of the current environment (from the hierarchical attention interference source feature extraction module), user discomfort (from the user discomfort perception model construction module), beamforming configuration parameters (from the beam null dynamic adjustment module), and historical interaction records; The state vector is represented as: ;
[0212] Where represents the state vector at time , represents the interference source characteristics at time , represents the user discomfort at time , represents the beamforming weight vector, is the historical interaction feature at time , including the state and action information of the past time steps; represents the considered number of historical time steps;
[0213] Define the action space Include operations to adjust beamforming parameters, such as fine-tuning the null position, adjusting the null depth, reallocating priorities, etc.; The action vector is represented as:
[0214] ;
[0215] Where represents the action vector at time , ,
[0216] respectively represent the adjustment amounts of the azimuth angle, elevation angle, and null depth of the first, the th interference source, represents the adjustment vector for priority allocation, represents the number of selected interference sources.
[0217] Step 5.2, Design the reward function;
[0218] Design the reward function Comprehensively consider the change in user discomfort (from the user discomfort perception model construction module), the objective speech quality evaluation index (applied to the output of the beam null dynamic adjustment module), and the system resource consumption:
[0219] ;
[0220] where represents the reward value at time , respectively represent the user discomfort at time , are respectively the perceived evaluation speech quality and the short-time objective intelligibility index at time , evaluating the audio quality of the output of the beam null dynamic adjustment module, represents the output signal after beamforming; is the action 's computational resource consumption; are respectively the weight coefficients of user discomfort improvement, perceived evaluation speech quality, short-time objective intelligibility index, and computational resource consumption.
[0221] Step 5.3, Construct the deep Q network;
[0222] Adopt the deep Q network (DQN) as the core algorithm of reinforcement learning to construct the value network:
[0223] ;
[0224] where represents the value network, represents the state vector, represents the action vector, are the network parameters. This network contains 4 fully connected layers, and the structure is:
[0225] Input layer, the dimension of the state vector :
[0226] Hidden layer 1: 256 neurons, ReLU activation function;
[0227] Hidden layer 2: 128 neurons, ReLU activation function;
[0228] Hidden layer 3: 64 neurons, ReLU activation function;
[0229] Output layer, the dimension of the action space .
[0230] Step 5.4, Implement the experience replay mechanism;
[0231] Improve learning efficiency using experience replay technology and build an experience pool Store transfer samples:
[0232] ;
[0233] where represents the current state, represents the current action, represents the immediate reward, represents the next state;
[0234] In each training, randomly sample a mini-batch from the experience pool for Q-network update, and the optimization objective is:
[0235] ;
[0236] where represents the loss function of the Q-network, is the current Q-network parameter, is the target network parameter, is the discount factor, with a value range of [0, 1], balancing the importance of immediate reward and future reward;
[0237] represents taking the expectation of the state-action-reward-next state quadruple sampled from the experience pool ;
[0238] represents taking the maximum value over all possible actions ; represents the current state, represents the current action, represents the immediate reward, represents the next state, represents the next state the action value under.
[0239] Step 5.5, Application of Double Q-learning;
[0240] Introduce the Double Deep Q-Network (DoubleDQN) mechanism to reduce overestimation problems. The current network is used to select actions, and the target network is used to evaluate action values:
[0241] ;
[0242] where represents the loss function of the Q-network, is the current Q-network parameter, is the target network parameter, is the discount factor; Indicates returning the action that achieves the maximum value Indicates the current state, Indicates the current action, Indicates the immediate reward, Indicates the next state;
[0243] Indicates taking the expectation of the state-action-reward-next state quadruple sampled from the experience pool
[0244] Step 5.6, Application of the ϵ-greedy strategy;
[0245] Use the ϵ-greedy strategy to select actions and decrease the exploration rate ϵ over time:
[0246] ;
[0247] Where Indicates the time when the action is selected; Indicates returning the action that achieves the maximum value Indicates the current time step;
[0248] Indicates the exploration rate at time , and the calculation formula is:
[0249] ;
[0250] Where are the initial exploration rate and the final exploration rate respectively, is the decay rate, which controls the speed of the exploration rate decrease; Indicates the base of the natural logarithm, approximately equal to 2.71828.
[0251] Step 5.7, Application of the periodic evaluation mechanism;
[0252] Set up a periodic evaluation mechanism. After processing audio segments, evaluate the performance of the current strategy and save the high-performance strategy as the benchmark strategy; The evaluation metrics include the average user discomfort (the output of the user discomfort perception model construction module), the audio quality score (the output applied to the beam null dynamic adjustment module), and the system resource efficiency; Where is the preset evaluation period, indicating that an evaluation is performed every time audio segments are processed.
[0253] Step 5.8, Hierarchical knowledge base construction;
[0254] Construct a hierarchical knowledge base, classify and store successful interference suppression strategies according to interference types and environmental characteristics to form a strategy library:
[0255] ;
[0256] Among them represents the strategy library, respectively represent the 1st, 2nd, and th strategies, represents the total number of strategies in the strategy library;
[0257] For a new interference scenario, first retrieve the strategies for similar scenarios in the strategy library as the initial strategies, then perform targeted optimization to accelerate the convergence process, and finally generate optimized interference suppression strategies .
[0258] The output results of this module include the optimized interference suppression strategies and the hierarchical knowledge base :
[0259] The optimized interference suppression strategies are directly fed back to the beam null dynamic adjustment module to guide the dynamic adjustment of the beam null;
[0260] The hierarchical knowledge base serves as the long-term memory of the system, storing the optimal strategies in different scenarios, enabling the system to quickly adapt to new interference environments.
[0261] Through this closed-loop optimization mechanism, the system can continuously learn and improve its interference suppression performance, forming an adaptive and evolving anti-interference ability.
[0262] Through the implementation of the above five modules, an optimized wireless audio transmission system based on adaptive beamforming provided by this embodiment forms a complete closed-loop system:
[0263] The multi-channel signal acquisition and preprocessing module captures and preprocesses acoustic signals, providing basic data for subsequent analysis; the hierarchical attention interference source feature extraction module and the user discomfort perception model construction module analyze interference characteristics from two dimensions, objective and subjective, one identifying the type and location of the interference source and the other evaluating the actual impact of the interference on the user;
[0264] The beam null dynamic adjustment module integrates the results of the first two steps to achieve targeted beam null adjustment to suppress interference;
[0265] The reinforcement learning optimization module continuously optimizes the strategies of the entire system through reinforcement learning, forming an adaptive and evolving anti-interference ability;
[0266] Tight data flow and control flow relationships are formed among the modules, which can effectively solve the technical problems faced by wireless audio transmission in complex electromagnetic interference environments and improve audio transmission quality and user experience.
[0267] A computer-readable storage medium is used to store computer-readable instructions that, when read by a computer, can run an optimization system for wireless audio transmission based on adaptive beamforming as described above.
[0268] Here, the present invention provides an implementation example: the audio transmission scenario of wireless earphones in a high-speed train carriage;
[0269] The inside of a high-speed train carriage is a typical complex electromagnetic interference environment with multiple interference sources:
[0270] Electromagnetic interference generated by the operation of the high-speed train;
[0271] Wi-Fi and Bluetooth signals emitted by wireless devices such as passengers' mobile phones and tablets;
[0272] Electromagnetic radiation generated by the in-car broadcast system;
[0273] Interference signals generated by the wireless earphones of other passengers;
[0274] Communication signals transmitted by base stations along the high-speed train line;
[0275] These interference sources are characterized by diverse interference types, dynamically changing positions, unstable interference intensities, and high user requirements for audio quality (hoping to clearly hear music or call content in a noisy environment).
[0276] The wireless earphones used in this application example are configured as follows:
[0277] 4 omnidirectional microphones, evenly distributed on the surface of the earphones;
[0278] 1 high-precision acceleration sensor to detect the user's head movements;
[0279] 1 infrared sensor to detect the wearing state of the earphones;
[0280] An integrated physiological sensor to monitor changes in skin conductivity;
[0281] A Bluetooth 5.2 chip supporting dual-mode transmission;
[0282] A low-power ARM processor supporting neural network inference;
[0283] A 500 mAh lithium battery.
[0284] In the high-speed rail carriage environment, the system first collects acoustic signals through 4 microphones. The sampling frequency is set to 48 kHz and the sampling precision is 24 bit. The initial spatial filter is set to beam towards the direction of the user's face (determined by initial calibration). The smoothing factor is set to 0.85, which can effectively suppress transient noise while retaining signal details.
[0285] After time-frequency transformation, the system constructs a feature matrix which contains data of 128 frequency points and 50 time windows, forming a 128×50 feature matrix. In the high-speed rail environment, the system additionally adopts a finer frequency resolution for the low-frequency (50 - 300 Hz) range to better capture the low-frequency interference characteristics unique to high-speed rail.
[0286] In this environment, the system has identified three main interference sources:
[0287] Low-frequency periodic electromagnetic interference generated by high-speed rail operation (confidence level 92.3%, azimuth angle 15°, elevation angle -10°);
[0288] Wi-Fi signal of the adjacent seat passenger's mobile phone (confidence level 89.7%, azimuth angle 75°, elevation angle 5°);
[0289] In-car broadcast system (confidence level 94.5%, azimuth angle 180°, elevation angle 30°).
[0290] Through the processing of the multi-scale convolutional network, the feature extraction effects of the system under three different convolutional kernels show that the 7×7 convolutional kernel has the best effect in capturing the low-frequency interference characteristics of high-speed rail; the 3×3 convolutional kernel is the most sensitive to the extraction of Wi-Fi signal features; the 5×5 convolutional kernel performs optimally in identifying the interference of the broadcast system. The channel attention mechanism automatically assigns weights of 0.45, 0.35, and 0.2 to these three groups of features, reflecting the differences in their impact on the audio quality.
[0291] Example of constructing the user discomfort perception model:
[0292] The system detects that there are obvious differences in the sensitivity of users to different interference sources in the high-speed rail environment:
[0293] The physiological index changes of users to the low-frequency interference of high-speed rail are not obvious (the galvanic skin response increases by 3%);
[0294] They show moderate discomfort to the interference of the adjacent seat Wi-Fi signal (the galvanic skin response increases by 12%, and the micro-expression change index is 0.27);
[0295] They are extremely sensitive to the interference of the broadcast system (the galvanic skin response increases by 28%, the micro-expression change index is 0.65, and they frequently adjust the headphone position).
[0296] The system-generated personalized interference sensitivity map shows that the interference sensitivities at azimuth angles of 75° (adjacent seat position) and 180° (direction of the broadcast system) are 0.65 and 0.83 (out of 1.0), respectively, while the sensitivity at the azimuth angle of 15° (direction of high-speed rail operation interference) is only 0.31;
[0297] Based on the interference source characteristics and the sensitivity map, the system calculates the suppression priorities of the three interference sources:
[0298] Broadcast system: 0.78 (highest priority);
[0299] Adjacent seat Wi-Fi: 0.62 (medium priority);
[0300] High-speed rail interference: 0.35 (lowest priority).
[0301] The system applies the LCMV algorithm to calculate the beamforming weight vector and applies deeper nulls in the direction of the broadcast system ( ), medium-depth nulls in the direction of the adjacent seat Wi-Fi ( ), and shallow nulls in the direction of high-speed rail interference ( ).
[0302] When the user turns their head to talk to the adjacent seat, the system detects the user's head turn, adjusts the main lobe direction of the beam in real time, and temporarily reduces the null depth in the direction of the adjacent seat to ensure the clarity of the conversation.
[0303] During a one-hour high-speed rail journey, the system continuously learns and optimizes the interference suppression strategy:
[0304] Initial stage (0 - 15 minutes): The system is mainly in the exploration stage ( ), trying different null configurations
[0305] Mid-stage (15 - 40 minutes): The exploration rate decreases ( ), and the system learns that the broadcast sound usually has a short duration but serious interference, and adopts the "fast deep null response" strategy
[0306] Late stage (40 - 60 minutes): The exploration rate further decreases ( ), and the system learns that the high-speed rail interference will increase in the tunnel section, and deepens the nulls in the corresponding direction in advance.
[0307] The system stores these learned strategies in a hierarchical knowledge base, marked as "high-speed rail environment - user ID - date". When the user takes the high-speed rail next time, the system can directly call this strategy library as the initial strategy without having to learn again.
[0308] Compared with ordinary wireless earphones that do not use the method of the present invention, the technical effects achieved by this application example in the high-speed rail environment are as follows:
[0309] Interference source recognition accuracy: 94.2% and 73.8% (traditional method);
[0310] Signal-to-noise ratio improvement: 11.3 dB and 5.1 dB (traditional method);
[0311] User satisfaction score (average of 10 test users, full score 10): 8.7 and 6.2 (traditional method);
[0312] New interference adaptation time: 4.3 seconds and 78 seconds (traditional method);
[0313] Battery life: 5.8 hours and 4.6 hours (traditional method);
[0314] This application example proves that the technical solution of the present invention has an interference suppression effect and an improvement in user experience in an actual complex electromagnetic interference environment, and fully achieves the expected technical goals.
[0315] The above embodiments of the present invention have been described, but these embodiments are not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of this embodiment.
Claims
1. A wireless audio transmission optimization system based on adaptive beamforming, characterized in that: include: Multi-channel signal acquisition and preprocessing module, which acquires multi-channel acoustic signals and performs space-time domain preprocessing to generate feature matrix; Hierarchical attention distractor feature extraction module, Process the signature matrix to identify the type and location of interference sources; User discomfort perception model building module, which builds a user discomfort perception model based on user physiological response data, facial expression recognition and behavior analysis, and generates a personalized interference sensitivity map; The beam null dynamic adjustment module dynamically adjusts the null position and depth of the beamforming algorithm based on the interference source type, location, and personalized interference sensitivity map, including: Interference source suppression priority calculation: based on interference source characteristics and interference sensitivity maps, calculate the suppression priority of each interference source; Priority sorting and screening: sort the identified interference sources according to the calculated suppression priority and select the one with the highest priority The interference source is used as the placement target of the beam null; Construct interference source covariance matrix; Calculate the beamforming weight vector, and use the linear constrained minimum variance algorithm to calculate the beamforming weight vector; Zero point depth adaptive control: for specific types of interference sources, an adaptive zero point depth control mechanism is introduced to increase the corresponding zero point depth for interference sources with higher priority and greater sensitivity; Real-time detection of user discomfort changes. When user discomfort exceeds a preset threshold, it triggers a rapid update of the beam zero position and depth. Applying beamforming, applying the calculated beamforming weight vector to the input signal to obtain a beamformed output signal; The reinforcement learning optimization module receives the execution results of the beam zero point dynamic adjustment module and user feedback data, and uses the reinforcement learning method to continuously optimize the interference suppression strategy to improve the system's adaptability to unknown interference environments.
2. According to claim 1, a wireless audio transmission optimization system based on adaptive beamforming is characterized in that: The multi-channel signal acquisition and preprocessing module includes: The microphone array collects multi-channel acoustic signals. The number of microphones ; Applying a spatial filter to the collected signal to perform spatial filtering processing on the multi-channel acoustic signal; Use the exponentially weighted moving average algorithm to smooth the spatially filtered signal; Applying an adaptive noise suppression algorithm to the smoothed signal to reduce signal noise; Use short-time Fourier transform to transform the denoised signal into time-frequency domain; Generate interference source feature extraction input and build a feature matrix based on time-frequency domain signals as the input of interference source feature extraction.
3. According to claim 1, a wireless audio transmission optimization system based on adaptive beamforming is characterized in that: The hierarchical attention interference source feature extraction module includes: Construct a multi-scale convolutional neural network, perform feature extraction, and build a network structure containing convolution kernels of different sizes to capture time-frequency features of different scales; Channel attention calculation: apply the channel attention mechanism to each group of feature maps and calculate the channel attention weight; Feature weighting, applying the channel attention weight to the original feature map to obtain a weighted feature map; Feature fusion, fusing weighted feature maps of different scales to obtain a comprehensive feature representation; Interference source identification and location, processing fusion features, realizing the identification and location of interference sources, and obtaining interference source feature representation.
4. The wireless audio transmission optimization system based on adaptive beamforming according to claim 3 is characterized in that: The multi-scale convolutional neural network includes three convolution kernels of different sizes, namely 3×3, 5×5 and 7×7, which capture time-frequency features of different scales.
5. The wireless audio transmission optimization system based on adaptive beamforming according to claim 1, characterized in that: The user discomfort perception model building module includes: Physiological data collection, collecting user physiological response data; Facial expression recognition: using computer vision algorithms to analyze user facial expressions in real time and identify micro-expression changes; Behavioral data monitoring, monitoring users’ implicit behaviors; Environmental context information collection; User discomfort calculation, building a user discomfort perception model based on the physiological data, the expression characteristics, the behavior data and the environmental context information, and calculating user discomfort; Personalized interference sensitivity map generation, based on historical data and user feedback, generates personalized interference sensitivity maps for different interference types and spatial locations.
6. The wireless audio transmission optimization system based on adaptive beamforming according to claim 1, characterized in that: The reinforcement learning optimization module includes: State space and action space definition, including the interference source characteristics of the current environment User inappropriateness Beamforming Configuration Parameters The state space of historical interaction records defines the action space containing the adjustment operations of the beamforming parameters; Reward function design: Design a reward function that comprehensively considers user discomfort changes, objective sound quality evaluation indicators, and system resource consumption; Deep Q network construction, as the core algorithm of reinforcement learning, builds value network; The experience replay mechanism is implemented to store experience samples of the interaction between the system and the environment and improve learning efficiency; Double Q-learning is applied to reduce the overestimation problem; ϵ-greedy strategy application, using ϵ-greedy strategy for action selection, reducing the exploration rate over time; Periodic evaluation mechanism application: set up a periodic evaluation mechanism to evaluate the performance of the current strategy and save the high-performance strategy as the benchmark strategy; Hierarchical knowledge base construction,Construct a hierarchical knowledge base, classify and store successful interference suppression strategies according to interference types and environmental characteristics to form a strategy library.
7. The wireless audio transmission optimization system based on adaptive beamforming according to claim 6, characterized in that: The reward function is: ; in Indicates time The reward value, and Respectively indicate time and of users are inappropriate, and Separately for the moment Perceptual evaluation of speech quality and short-term target intelligibility index to evaluate the output audio quality. represents the output signal after beamforming; For Action The computing resource consumption is Indicates time The action vector of and They are the weight coefficients of user discomfort improvement, perceptual evaluation of speech quality, short-term target intelligibility index, and computing resource consumption.
8. The wireless audio transmission optimization system based on adaptive beamforming according to claim 6, characterized in that: The calculation formula for forming the strategy library is: ; in Represents the policy library, Respectively represent the first, second, and strategy, Indicates the total number of strategies in the strategy library; For new interference scenarios, we first retrieve strategies for similar scenarios from the strategy library as the initial strategies, and then perform targeted optimization to accelerate the convergence process.
9. A computer-readable storage medium, characterized in that: It is used to store computer-readable instructions, and when the computer-readable instructions are read by a computer, it can run a wireless audio transmission optimization system based on adaptive beamforming as described in any one of claims 1-8.
Citation Information
Patent Citations
Cognitive interference integrated beam forming method based on deep reinforcement learning
CN117614501A
Dynamically adapting sound based on environmental characterization
US10511906B1