Beam forming method and device, electronic equipment, storage medium and product
By acquiring multimodal perception data in the vehicle environment and using a beamforming weight optimization model trained by deep reinforcement learning, the problem of mismatch between the beamforming method in the existing technology and the dynamic environment is solved, real-time adjustment and adaptability in the dynamic environment are achieved, and the performance of the in-vehicle voice interaction system is improved.
Patent Information
- Application Number
- CN202510723097.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-09
AI Technical Summary
Existing on-board beamforming methods cannot dynamically adapt to changes in the sound field, resulting in positioning deviation and interference suppression failure, and cannot adjust beamforming weights in real time in dynamic environments.
By acquiring multimodal perception data, including original speech signals and environmental auxiliary data, and using a beamforming weight optimization model trained by deep reinforcement learning, the beamforming weights are dynamically optimized and multimodal data is integrated to improve environmental adaptability and anti-interference capabilities.
It realizes real-time adjustment of beamforming weights in a dynamic vehicle environment, improves signal anti-interference capability and environmental adaptability, and enhances the user experience of the in-vehicle voice interaction system.
Smart Images

Figure CN120612955A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of signal processing technology, and in particular to a beamforming method, device, electronic device, storage medium and product. Background Art
[0002] In in-vehicle voice interaction systems, acoustic front-end calibration is a key step in ensuring the quality of voice signal acquisition. Within this calibration, beamforming technology can improve sound source localization accuracy and directional sound pickup capabilities.
[0003] The current beamforming method in vehicle scenarios relies on static calibration and fixed weight rules, and cannot dynamically adapt to changes in the sound field (such as sudden changes in the sound field caused by opening and closing windows, and shifts in the sound source position caused by passenger movement). As a result, the calibrated beamforming algorithm may experience positioning deviations or interference suppression failures in real scenarios. Summary of the Invention
[0004] The present invention provides a beamforming method, device, electronic device, storage medium and product to solve the problem of mismatch between existing beamforming methods and dynamic environments.
[0005] According to one aspect of the present invention, a beamforming method is provided, the method comprising:
[0006] Acquire multimodal perception data of the vehicle environment; the multimodal perception data includes original voice signals and environmental auxiliary data;
[0007] Determining a multidimensional state vector based on multimodal perception data;
[0008] The multidimensional state vector is input into the beamforming weight optimization model to obtain the beamforming weight corresponding to the original speech signal; the beamforming weight optimization model is obtained based on deep reinforcement learning training.
[0009] According to another aspect of the present invention, a beamforming apparatus is provided, the apparatus comprising:
[0010] A data acquisition module is used to acquire multimodal perception data of the vehicle environment; the multimodal perception data includes original voice signals and environmental auxiliary data;
[0011] A state vector determination module, configured to determine a multidimensional state vector based on multimodal sensing data;
[0012] The beamforming weight determination module is used to input the multidimensional state vector into the beamforming weight optimization model to obtain the beamforming weight corresponding to the original speech signal; the beamforming weight optimization model is obtained based on deep reinforcement learning training.
[0013] According to another aspect of the present invention, an electronic device is provided, comprising:
[0014] at least one processor; and
[0015] a memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the beamforming method according to any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the beamforming method according to any embodiment of the present invention when executed.
[0018] According to another aspect of the present invention, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the beamforming method according to any embodiment of the present invention is implemented.
[0019] The technical solution of the embodiment of the present invention obtains multimodal perception data of the vehicle environment; the multimodal perception data includes the original voice signal and environmental auxiliary data; a multidimensional state vector is determined based on the multimodal perception data; the multidimensional state vector is input into a beamforming weight optimization model to obtain the beamforming weight corresponding to the original voice signal; the beamforming weight optimization model is obtained based on deep reinforcement learning training. This technical solution solves the problem of mismatch between existing beamforming methods and dynamic environments by fusing multimodal original voice signals and environmental auxiliary data, and then using a beamforming weight optimization model constructed based on deep reinforcement learning to drive the dynamic optimization of beamforming weights. It can adjust the beamforming weights in real time in the actual vehicle dynamic environment, improve the signal anti-interference ability and environmental adaptability, and thus effectively enhance the user experience of the vehicle voice interaction system.
[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 is a flowchart of a beamforming method provided according to embodiment 1 of the present invention;
[0023] Figure 2 is a flowchart of a beamforming method provided according to embodiment 2 of the present invention;
[0024] Figure 3 1 is a flow chart of a beamforming optimization framework according to a third embodiment of the present invention;
[0025] Figure 4 is a flowchart of a training process of a beamforming weight optimization model provided in accordance with the third embodiment of the present invention;
[0026] Figure 5 2 is a schematic structural diagram of a beamforming device according to a fourth embodiment of the present invention;
[0027] Figure 6 3 is a schematic structural diagram of an electronic device for implementing the beamforming method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0028] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0029] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0030] Example 1
[0031] Figure 1This is a flow chart of a beamforming method provided in the first embodiment of the present invention. This embodiment is applicable to the case of dynamically determining the beamforming weight of a speech signal in a vehicle environment. The method can be executed by a beamforming device, which can be implemented in the form of hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the beamforming method provided in this embodiment 1 specifically includes the following steps:
[0032] S110 , obtaining multimodal perception data of the vehicle environment; the multimodal perception data includes original speech signals and environmental auxiliary data.
[0033] Among them, multimodal perception data can refer to multi-source perception data collected based on various sensors, and the multimodal perception data specifically includes original voice signals and environmental auxiliary data; the original voice signal can be a multi-channel voice signal collected by a microphone array, and the environmental auxiliary data can include but is not limited to: posture data collected by an inertial measurement unit (IMU), ranging data collected by an infrared time of flight (TOF) sensor, temperature data collected by a temperature sensor, etc.
[0034] In an embodiment of the present invention, various sensors configured on the vehicle can be used to collect multimodal perception data of the current vehicle environment in real time. The multimodal perception data specifically includes original voice signals and environmental auxiliary data. The environmental auxiliary data is used to provide necessary dynamic environmental monitoring information for the optimization of the beamforming weights of subsequent voice signals, thereby improving the anti-interference capability of the vehicle-borne beamforming.
[0035] S120. Determine a multidimensional state vector based on the multimodal perception data.
[0036] Among them, the multidimensional state vector can be understood as a vector constructed after fusing multimodal perception data, which can comprehensively characterize the dynamic changes of the vehicle environment, user behavior and system load status. For example, the multidimensional state vector can include state vectors of the following dimensions: sound source angle error, real-time signal-to-noise ratio, temperature parameters, reverberation estimation time, array topology identification, sound speed compensation factor, calculation load factor, etc.
[0037] In an embodiment of the present invention, after obtaining the multimodal perception data, it can be preprocessed and feature extracted to generate a multidimensional state vector containing dynamic sound field characteristics, user behavior characteristics and system load state characteristics. For example, the spectral energy attenuation data of the original speech signal can be extracted, and the spectral energy attenuation data can be predicted by LSTM (Long Short-Term Memory) to obtain the reverberation estimation time. The head posture change rate can also be calculated by combining IMU posture data and Euler angular rate integral, etc. The extracted multidimensional state vector will serve as the input of the subsequent beamforming weight optimization model to drive the real-time optimization of the beamforming weight.
[0038] S130. Input the multidimensional state vector into a beamforming weight optimization model to obtain beamforming weights corresponding to the original speech signal; the beamforming weight optimization model is obtained based on deep reinforcement learning training.
[0039] The beamforming weight optimization model may be a model used to determine the optimal beamforming weights for the current vehicle environment, which may be pre-trained through deep reinforcement learning. The beamforming weights may be weighting coefficients used to control the microphone's received signal (i.e., the original speech signal). By adjusting these weights, the target direction signal can be enhanced and interference direction noise can be suppressed.
[0040] In an embodiment of the present invention, the extracted multidimensional state vector can be input into a pre-trained beamforming weight optimization model, which can be used to determine the optimal beamforming weights for the original speech signal in the current vehicle environment. The beamforming weight optimization model can be pre-trained using deep reinforcement learning. Furthermore, the determined beamforming weights can be applied to downstream speech processing tasks such as speech recognition, adaptive speech enhancement, and sound source localization, thereby enabling integration with the in-vehicle voice interaction system.
[0041] The technical solution of the embodiment of the present invention obtains multimodal perception data of the vehicle environment; the multimodal perception data includes the original voice signal and environmental auxiliary data; a multidimensional state vector is determined based on the multimodal perception data; the multidimensional state vector is input into a beamforming weight optimization model to obtain the beamforming weight corresponding to the original voice signal; the beamforming weight optimization model is obtained based on deep reinforcement learning training. This technical solution solves the problem of mismatch between existing beamforming methods and dynamic environments by fusing multimodal original voice signals and environmental auxiliary data, and then using a beamforming weight optimization model constructed based on deep reinforcement learning to drive the dynamic optimization of beamforming weights. It can adjust the beamforming weights in real time in the actual vehicle dynamic environment, improve the signal anti-interference ability and environmental adaptability, and thus effectively enhance the user experience of the vehicle voice interaction system.
[0042] Example 2
[0043] Figure 2 This is a flow chart of a beamforming method provided in the second embodiment of the present invention, which is further optimized and expanded based on the above embodiment and can be combined with various optional technical solutions in the above embodiment. Figure 2 As shown, the second embodiment provides a beamforming method, which specifically includes the following steps:
[0044] S210: Acquire multimodal perception data of the vehicle environment; the multimodal perception data includes original voice signals and environmental auxiliary data.
[0045] In an embodiment of the present invention, environmental auxiliary data may include at least: posture data collected by an inertial measurement unit (IMU), which is used to track dynamic behaviors such as head movement of the driver or passenger and changes in vehicle posture; ranging data collected by an infrared time-of-flight (TOF) sensor, which is used to measure the distance to the target object (such as a sound source or a passenger); temperature data collected by a temperature sensor, which is used to monitor the real-time temperature of the vehicle environment for sound speed compensation; resource monitoring data collected by a system monitoring tool, which is used to monitor the status of embedded hardware resources, including FPGA (Field Programmable Gate Array) logic unit utilization, memory occupancy, processing delay, power consumption, etc.
[0046] S220. Determine a multidimensional state vector based on the multimodal perception data.
[0047] In an embodiment of the present invention, the multidimensional state vector may specifically include state vectors of the following dimensions: sound source angle error, real-time signal-to-noise ratio, head posture change rate, temperature parameter, reverberation estimation time, array topology identifier, sound velocity compensation factor, and computational load factor. The process of determining each state vector is as follows:
[0048] ①Determine the sound source angle error based on the ranging data and the original speech signal
[0049] Specifically, it performs a generalized cross-correlation phase transform (GCC-PHAT) on the original speech signal to calculate the time delay difference of arrival (TDOA) between each microphone pair; combines infrared TOF ranging data with TDOA, and uses a geometric acoustic model (such as a spherical wave model) to infer the estimated direction angle θ of the sound source. estimated ; Compare the theoretical target direction θ target and the actual estimated direction θ estimated , to obtain the sound source angle error θ err.
[0050] ②Determine the real-time signal-to-noise ratio based on the original speech signal
[0051] Specifically, it includes: performing a short-time Fourier transform (STFT) on the original speech signal to obtain signal spectrum data; inputting the signal spectrum data into a pre-trained LSTM network to extract spectrum features and noise estimation, and predicting the noise spectrum data of the current frame; and determining the real-time signal-to-noise ratio (SNR) based on the signal spectrum data and noise spectrum data.
[0052] ③Determine the head posture change rate based on posture data
[0053] Specifically, it includes: performing Kalman filtering on the raw angular velocity data collected by the IMU to eliminate sensor noise; based on the filtered angular velocity data, using the Euler angular rate integral calculation to determine the corresponding head posture change rate θ h .
[0054] ④Determine temperature parameters based on temperature data
[0055] Specifically, the temperature data collected by the temperature sensor is processed using Kalman filtering to output a smoothed temperature estimation value, namely the temperature parameter T.
[0056] ⑤ Determine the estimated reverberation time based on the original speech signal
[0057] Specifically, it includes: performing STFT on the original speech signal to obtain signal spectrum energy attenuation data; inputting the signal spectrum energy attenuation data into the pre-trained TSTM network to predict the reverberation estimation time RT60; where the reverberation estimation time RT60 represents the time required for the sound energy to decay to one millionth of the initial value (i.e., a decrease of 60 decibels).
[0058] ⑥ Determine array topology based on ranging data
[0059] Specifically, it includes: combining infrared TOF ranging data and array geometric constraints (such as array spacing), and obtaining array topology identification M through threshold logic classification topo ; Among them, the array topology identifier M topo Used to represent the current physical layout status of the microphone array.
[0060] ⑦Determine the sound speed compensation factor based on temperature data and attitude data
[0061] Specifically including: determining the temperature parameter T and the head posture change rate θ hThe two are inputted into a multilayer perceptron (MLP) to obtain the sound speed compensation factor P. sound ; Among them, the sound speed compensation factor P sound Used to calibrate the sound velocity estimation deviation caused by ambient temperature changes and user behavior (such as head movement), ensuring the accuracy of sound source localization and beamforming.
[0062] ⑧Determine the calculation load factor based on resource monitoring data
[0063] Specifically, it includes: normalizing each performance indicator data (such as logic unit utilization, memory occupancy, processing delay, etc.) in the resource monitoring data; performing weighted summation on each normalized performance indicator data to obtain the calculation load factor P compute ; Among them, calculate the load factor P compute A comprehensive indicator that reflects the system's computing resource usage, used to dynamically adjust algorithm complexity and optimize energy efficiency to ensure real-time computing.
[0064] S230. Input the multidimensional state vector into the strategy network of the beamforming weight optimization model, and use the strategy network to infer the target action vector corresponding to the original speech signal; wherein the target action vector includes the parameters to be adjusted corresponding to the covariance matrix influencing hyperparameters, the constraint matrix influencing hyperparameters, and the constraint response vector influencing hyperparameters.
[0065] Among them, the policy network can refer to a network used to map the input multidimensional state vector to the corresponding action vector. The target action vector can be understood as a specific action for adjusting the signal coherence (covariance matrix), beam direction (constraint matrix) and beam direction gain (constraint response vector) of the original speech signal. It can include the parameters to be adjusted corresponding to the three types of hyperparameters: the covariance matrix affecting hyperparameters, the constraint matrix affecting hyperparameters, and the constraint response vector affecting hyperparameters. For example, the target action vector can include the adjustment amount of the forgetting factor β, the main lobe direction adjustment amount Δθ target , main lobe gain f target The adjustment amount, etc.
[0066] The covariance matrix influencing hyperparameters may refer to the hyperparameters used to control the covariance matrix of the received signal (i.e., the original speech signal), which specifically include: a forgetting factor β, used to control the forgetting rate of the historical covariance matrix and improve noise suppression capability; a recursive step size μ, used to control the weight of new data to quickly track environmental changes; and a regularization coefficient φ, used to improve the reversibility of the covariance matrix and enhance robustness.
[0067] The constraint matrix influence hyperparameters may refer to the hyperparameters of the constraint matrix used to adjust the beamforming, which specifically include: the main lobe direction adjustment amount Δθ target, used to adjust the main lobe direction (desired signal direction); the null direction adjustment amount Δθ interference , used to adjust the null direction (interference direction)
[0068] The constraint response vector influence hyperparameters may refer to the hyperparameters used to adjust the constraint response vector of beamforming, which specifically include: main lobe gain f target , used to adjust the gain value in the main lobe direction; the null depth f null , used to adjust the gain value in the null direction.
[0069] In an embodiment of the present invention, the multidimensional state vector extracted from the multimodal perception data of the current vehicle environment can be input into the strategy network of a pre-trained beamforming weight optimization model, and the strategy network is used to decide the target action vector corresponding to the current multidimensional state vector. The target action vector specifically includes: parameters to be adjusted corresponding to the covariance matrix influencing hyperparameters, the constraint matrix influencing hyperparameters, and the constraint response vector influencing hyperparameters. Subsequently, the target action vector will be used to adjust the covariance matrix, constraint matrix, and constraint response vector in the beamforming, respectively, to determine the final beamforming weight.
[0070] S240 , based on each parameter to be adjusted, respectively call a preset covariance matrix formula, a preset constraint matrix formula, and a preset constraint response vector formula to determine a corresponding covariance matrix, constraint matrix, and constraint response vector.
[0071] In an embodiment of the present invention, the corresponding covariance matrix R can be determined based on the covariance matrix influence hyperparameters, constraint matrix influence hyperparameters and constraint response vector influence hyperparameters output by the aforementioned model, respectively corresponding to the parameters to be adjusted, combined with the pre-configured preset covariance matrix formula, preset constraint matrix formula and preset constraint response vector formula. xx (t), constraint matrix C(t) and constraint response vector f(t); wherein the preset covariance matrix formula can be expressed as:
[0072] R xx (t) = β·R xx (t-1)+μ·x(t)x H (t)+φ·I
[0073] Where R xx (t) represents the covariance matrix of the received signal x(t) at time t; H represents the conjugate transpose; and I represents the identity matrix.
[0074] The preset constraint matrix formula can be expressed as:
[0075] C(t)=[a(θ target +Δθ target ),...,a(θinterference +Δθ interference )]
[0076] Where θ target and θ interference They represent the pre-calibrated main lobe direction and null direction respectively; a(θ) is the steering vector, which represents the phase response of the acoustic wave incident on the array from the direction θ.
[0077] The preset constraint response vector formula can be expressed as:
[0078] f(t)=[f target ,0,...,f null ] T
[0079] Where T represents transpose.
[0080] S250 : Based on the covariance matrix, the constraint matrix, and the constraint response vector, a preset beamforming weight closed-form solution formula is called to determine the corresponding beamforming weight.
[0081] The preset closed-form solution formula of the beamforming weight may refer to a closed-form solution formula of the beamforming weight determined in advance based on an LCMV (Linearly Constrained Minimum Variance) beamformer.
[0082] In the embodiment of the present invention, the covariance matrix R determined above can be xx , constraint matrix C and constraint response vector f, and determine the final beamforming weight w by calling the preset beamforming weight closed-form solution formula opt , where the closed-form solution formula for the preset beamforming weight can be expressed as:
[0083]
[0084] Furthermore, based on the above-mentioned embodiments of the invention, the beamforming method provided in this embodiment further includes:
[0085] The beamforming weights are applied to downstream speech processing tasks, which include at least speech signal enhancement and sound source localization.
[0086] In the embodiment of the present invention, the beamforming weight w determined above can be opt Applied to some downstream speech processing tasks, for example, the determined beamforming weights w optIt acts on the original speech signal to achieve adaptive speech enhancement. It can also determine the direction of the sound source based on the position of the beamformed main lobe. By applying the determined beamforming weights to downstream speech processing tasks, it can be linked with the in-vehicle voice interaction system, thereby improving the user experience.
[0087] The technical solution of the embodiment of the present invention extracts a multidimensional state vector from the multimodal original voice signal and environmental auxiliary data, calls the strategy network of the beamforming weight optimization model to decide the target action vector corresponding to the current multidimensional state vector, and then uses the target action vector to adjust the covariance matrix, constraint matrix and constraint response vector in the beamforming respectively, and then determines the final beamforming weight, which solves the problem of mismatch between the existing beamforming method and the dynamic environment, and can dynamically generate corresponding optimal beamforming weights according to different vehicle environments, thereby improving the signal anti-interference ability and environmental adaptability, thereby effectively improving the user experience of the vehicle voice interaction system.
[0088] Example 3
[0089] Current beamforming methods have the following main shortcomings:
[0090] ① Poor adaptability to dynamic scenarios: Existing methods rely on preset static acoustic parameters (such as beam angle and dereverberation threshold). However, the in-vehicle environment is dynamic and time-varying (such as sudden changes in the sound field caused by window opening and closing, and sound source position shift caused by passenger movement). This causes the calibrated beamforming algorithm to exhibit positioning deviations or interference suppression failures in real-world scenarios.
[0091] Insufficient response to sudden interference: Existing methods suppress noise from non-target directions by using a fixed "pickup cone." However, in the face of sudden interference (such as a sudden horn honking outside the vehicle or a rear passenger interrupting), the static beam cannot quickly adjust the null direction, causing the target speech signal-to-noise ratio (SNR) to drop sharply, thus affecting subsequent in-vehicle voice interaction.
[0092] ③ Single optimization goal: Existing methods use sound source localization accuracy or directional pickup angle control as a single optimization goal, and are unable to simultaneously consider the coordinated optimization of anti-interference (such as wind noise suppression), energy efficiency (such as low-power beamforming), and dynamic sound source tracking (such as driver head rotation) in complex sound fields.
[0093] ④ High cost of environmental adaptability calibration: Acoustic environment calibration (such as reverberation time measurement and noise spectrum modeling) needs to be repeated according to the confined space characteristics of different vehicle models (such as glass reflectivity and seat sound absorption coefficient), resulting in a long development cycle and difficulty in large-scale deployment.
[0094] The real-time optimization of dynamic beamforming weights in the present invention is essentially due to the inability of traditional static beamforming algorithms to adapt to changes in the sound field caused by head rotation and passenger displacement. The adaptive compensation of environmental parameters is essentially due to the fact that temperature changes cause sound velocity drift, and the window opening / closing state changes the sound field propagation path. Traditional beamforming algorithms (such as LCMV) rely on fixed weight strategies, which can cause mainlobe shift or increased sidelobe interference in dynamic scenarios (such as passenger head rotation and position movement).
[0095] The traditional beamforming algorithm can be modeled as follows: Consider D point source signals from directions θ1, θ2..., θ D is incident on a microphone array of arbitrary geometry containing M sensors. The output y(k) of the array beamforming can be expressed as:
[0096] y(k)=w H x(k)
[0097] Where k∈[1,2,...,K] is a time index, representing a moment in the discrete time series; w is the M*1 complex beamforming weight (vector) to be estimated; x(k) represents the received signal (i.e., the original speech signal)
[0098] For the LCMV beamformer, the beamforming weights w are the solutions to the following multilinear constrained optimization problem:
[0099]
[0100] subject to C H w=f
[0101] Where J(w) represents the cost function; R xx represents the covariance matrix (M*M dimensions), and R xx =E{x(k)x H (k)}, E{·} represents the expectation operation; C represents the constraint matrix (M*D dimension), and f represents the constraint response vector (D*1 dimension).
[0102] The closed-form solution for the optimal weight of LCMV beamforming is as follows:
[0103]
[0104] It is important to understand that the antenna array layout of different vehicle models (such as sensor position, spacing, and installation angle) will change the array's steering vector a(θ), which in turn affects the composition of the constraint matrix C. The calibration requirements of different vehicle models may adjust the definition of the constraint response vector f (such as the number of desired signal directions L and the number of interference directions DL). The constraint response vector f specifies the gain requirement in each direction (such as 1 or 0). For example, vehicle model A needs to maintain gain in 3 directions (L=3), and vehicle model B needs to maintain gain in 5 directions (L=5). This causes the structure of the constraint response vector f to change, which in turn causes the final beamforming weight w to change. opt Changes have occurred.
[0105] The body structure of different models (such as metal parts distribution, body size) will change the statistical characteristics of multipath interference and noise, thereby affecting the covariance matrix R of the received signal x(k) xx , and directly affects the beamforming weight w opt The interference suppression capability of different models may be adjusted according to the calibration requirements, thereby affecting the response characteristics of beamforming. The number of sensors M or the number of interference directions D may be different for different models, resulting in a significant change in the complexity of the matrix inversion operation. The complexity is O(M 3 ),and The larger K is, the more stable the covariance matrix estimation is, but the amount of calculation will increase significantly. At the same time, for large microphone arrays (such as M = 1000), The real-time calculation of is difficult to achieve. Traditional static beamforming weight schemes that rely on accurate models are no longer able to solve the real-time update problem of LCMV beamforming.
[0106] Based on the above problems, an embodiment of the present invention proposes a dynamic beamforming solution based on deep reinforcement learning. It realizes complex nonlinear modeling by directly establishing an end-to-end mapping from multimodal inputs (temperature, posture, sound field) to beamforming weights. It balances positioning accuracy, voice quality and computational efficiency through a reward function to achieve multi-objective collaborative control, and can completely dominate the optimization of dynamic beamforming weights and environmental parameter adaptation. The core idea of reinforcement learning is that the agent learns the optimal strategy through interaction with the environment to maximize the long-term cumulative reward. Its framework includes the following key elements: State: the current observation information of the environment (such as signal parameters, interference distribution); Action: the decision made by the agent based on the state (such as adjusting the beamforming weight); Reward: the feedback of the environment on the action (such as beam gain improvement, interference suppression effect); Policy: the mapping rule from state to action (that is, how to adjust the beam according to the current information). Reinforcement learning does not require pre-labeled data, but instead dynamically optimizes strategies through trial-and-error. It is suitable for dynamic, highly uncertain, and difficult-to-model scenarios.
[0107] Figure 3 This is a flow chart of the beamforming optimization framework provided in the third embodiment of the present invention. Figure 3 As shown in Figure 1, the framework consists of a multimodal perception layer, a feature extraction layer, a reinforcement learning decision layer, and an execution layer. The multimodal perception layer is responsible for collecting multimodal perception data from the vehicle environment, including raw voice signals from the microphone array, attitude data from the inertial measurement unit (IMU), ranging data from the infrared time-of-flight (TOF) sensor, temperature data from the temperature sensor, and resource monitoring data from the system monitoring tool.
[0108] The feature extraction layer is responsible for preprocessing and feature extraction of the collected multi-source and multi-modal perception data to generate the corresponding multi-dimensional state vector, which specifically includes the state vectors of the following dimensions: sound source angle error θ err , real-time signal-to-noise ratio SNR, head posture change rate θ h , temperature parameter T, reverberation estimation time RT60, array topology identifier M topo , sound velocity compensation factor P sound and calculate the load factor P comput The process of determining the state vector of each dimension may refer to the above embodiment and will not be described in detail here.
[0109] The reinforcement learning decision layer is responsible for determining the target action vector in the current vehicle environment through the designed action space, state space, policy gradient update algorithm and reward function.
[0110] The execution layer is responsible for completing the beamforming weight update based on the target action vector output by the model and applying it to downstream speech processing tasks.
[0111] Among them, the state space of reinforcement learning can be expressed as:
[0112] s cont =[θ err ,SNR,θ h ,T,RT60,M topo ,P sound ,P compute ]
[0113] Influence covariance matrix R xx The hyperparameters of the constraint matrix C include the forgetting factor β, the recursive step size μ, and the regularization coefficient φ; the hyperparameters that affect the constraint matrix C include the main lobe direction adjustment Δθ target and the zero sink direction adjustment amount Δθ interferenc ; The hyperparameters that affect the constrained response vector f include: main lobe gain f target and the zero-sag depth f null The above hyperparameters will serve as the action vector in the action space, that is, the action space can be expressed as:
[0114] a cont =[β,μ,φ,Δθ target ,Δθ interference ,f target ,f null ]
[0115] Among them, these actions (vectors) have a covariance matrix R xx The impact is as follows:
[0116] R xx (t) = β·R xx (t-1)+μ·x(t)x H (t)+φ·I
[0117] The impact on the constraint matrix C is as follows:
[0118] C(t)=[a(θ target +Δθ target ),...,a(θ interference +Δθ interference )]
[0119] The impact on the constraint response vector f is as follows:
[0120] f(t)=[f target ,0,...,f null ] T
[0121] Next, considering the multi-objective performance of beamforming, the reward function is designed as follows:
[0122] r t =α1·ΔSINR+α2·||Δw|| -1 -α3·Nulling_Error-α4·Power_Cost
[0123] Among them, r t represents the instantaneous reward at time step t; ΔSINR represents the change in signal to interference plus noise ratio, that is, the difference in signal to interference plus noise ratio (SINR) between the current moment and the previous moment; Δw represents the change in beamforming weight; Nulling_Error represents the nulling depth suppression error, that is, the absolute error between the actual suppression depth in the interference direction and the target value; Power_Cost represents the power consumption control term, that is, the penalty for exceeding the array element transmission limit; α1, α2, α3, and α4 represent weight coefficients.
[0124] Figure 4 This is a flow chart of the training process of the beamforming weight optimization model provided in the third embodiment of the present invention. Figure 4 As shown in Figure 2, the training process of the beamforming weight optimization model includes the following steps:
[0125] S310: Obtain a multi-dimensional state vector training set.
[0126] In the embodiment of the present invention, the specific process of obtaining the multi-dimensional state vector training set may include:
[0127] ① Obtain multimodal perception data in different vehicle environments;
[0128] ②Align the timestamps of each sensor data in the multimodal perception data;
[0129] ③ Perform feature extraction on multimodal perception data to generate a multidimensional state vector training set.
[0130] S320: Initialize and build an initial beamforming weight optimization model, and obtain a reward function configured for the initial beamforming weight optimization model.
[0131] In an embodiment of the present invention, an initial beamforming weight optimization model may be initialized and constructed, that is, the model parameters of the policy network (Actor) and the value network (Critic) are initialized, and a pre-configured reward function is obtained.
[0132] S330. In the first training stage, the covariance matrix influence hyperparameter is used as the action vector to be adjusted, and the initial beamforming weight optimization model is trained based on the multi-dimensional state vector training set and the reward function. The iterative training process is repeated until the initial beamforming weight optimization model converges to obtain the intermediate beamforming weight optimization model.
[0133] In the embodiment of the present invention, the model training process includes two stages, wherein the first training stage (basic training) fixes the constraint matrix C and the constraint response vector f, and only affects the covariance matrix R xx The three hyperparameters are adjusted; the second training stage (joint fine-tuning) reuses the network parameters of the first training stage, and unlocks the adjustment of the hyperparameters corresponding to the constraint matrix C and the constraint response vector f, and performs the covariance matrix R xx , the constraint matrix C, and the constraint response vector f correspond to the joint tuning of hyperparameters. Specifically, in the first training phase, the constraint matrix C and the constraint response vector f correspond to the adjustment of hyperparameters, and only the covariance matrix affecting the hyperparameters is used as the action in the action space. The initial beamforming weight optimization model is trained using the multi-dimensional state vector training set and the reward function. The training goal is to maximize the cumulative reward J(π) of the agent in long-term interaction. Its mathematical form is:
[0134]
[0135] Where π represents the policy, i.e., the mapping rule from state to action; E[·] represents the mathematical expectation; τ represents the trajectory, i.e., the state-action-reward sequence; γ∈[0,1] represents the discount factor, which is used to balance the importance of current rewards and future rewards; r t represents the immediate reward at time step t, which is used to reflect the optimization objective of beamforming performance.
[0136] At the same time, the proximal policy optimization (PPO) is used in the training process to maximize the objective function while ensuring the stability of the training and the applicability to high-dimensional continuous action space, and to avoid divergence by limiting the amplitude of the policy update. Its loss function can be expressed as the policy loss term L CLIP , value loss item L Value and the entropy regularization term L Entropy It consists of three parts, namely:
[0137] L PPO (θ)=L CLIP +c1·L Value -c2·L Entropy
[0138] Where c1 represents the value loss weight, which is used to adjust the contribution of the value network to the total loss and prevent the value network from overfitting; c2 represents the entropy weight, which is used to control the balance between exploration and utilization. When the entropy weight is high, the model is more inclined to explore new actions.
[0139] Among them, the policy loss term L CLIP It can be expressed as:
[0140]
[0141] Where, π θ (a t |s t ) represents the current policy network π θ In state s t Next select action a t The probability of π θold (a t |s t ) represents the action selection probability of the old policy network (before updating); A t Represents the advantage function, used to evaluate the state s t Next select action a t The degree of superiority relative to the average performance can be calculated by Generalized Advantage Estimation (GAE); the hyperparameter ε represents the Clip threshold, which is used to limit the probability ratio The fluctuation range of the strategy is to prevent the strategy from updating too much and avoid unstable training.
[0142] Value loss item L Value It can be expressed as:
[0143] L Value =(V θ (s t )-R t ) 2
[0144] Where V θ (s t ) represents the value network V θ For state s t The value estimation (expected cumulative reward) is used to evaluate the overall quality of the current state and assist in strategy network optimization; t represents the target value, which can be calculated by the actual reward and the discounted next state value:
[0145] Entropy regularization term L Entropy It can be expressed as:
[0146] L Entropy =S[πθ ]s t
[0147] Where S[·] represents the entropy of the strategy distribution. The larger the entropy, the stronger the randomness of the strategy.
[0148] The iterative training process is repeated until the average reward fluctuation for a preset number of consecutive iterations is less than a preset fluctuation threshold, or the first training phase is terminated when the maximum number of training rounds is reached, and the intermediate beamforming weight optimization model after the first training phase is obtained.
[0149] S340. In the second training stage, the covariance matrix influencing hyperparameters, the constraint matrix influencing hyperparameters and the constraint response vector influencing hyperparameters are used as the action vectors to be adjusted, and the intermediate beamforming weight optimization model is trained based on the multi-dimensional state vector training set and the reward function. The iterative training process is repeated until the intermediate beamforming weight optimization model converges to obtain the beamforming weight optimization model.
[0150] In an embodiment of the present invention, unlike the first training stage, the second training stage takes the covariance matrix influencing hyperparameters, the constraint matrix influencing hyperparameters and the constraint response vector influencing hyperparameters as actions in the action space to perform multi-parameter joint optimization. The specific training process is similar to that of the first training stage and will not be repeated here.
[0151] The beamforming method provided by the embodiment of the present invention has at least the following beneficial effects:
[0152] ① Improved adaptability to real-time dynamic environments: Traditional LCMV beamforming relies on offline calculations or recursive updates of fixed parameters. In dynamic scenarios (such as user movement and interference mutations), the covariance matrix inverse needs to be frequently recalculated, resulting in high computational delay (complexity of O(M)). 3 )) and tracking lag. Based on the improvement of reinforcement learning (RL), by adjusting the covariance matrix affecting the hyperparameters, the constraint matrix affecting the hyperparameters and the constraint response vector affecting the hyperparameters in real time, the matrix inversion operation is bypassed and the computational complexity is reduced to O(M 2 ), significantly improving beam tracking accuracy and stability in dynamic scenarios.
[0153] ② Enhanced multi-objective collaborative optimization capabilities: Traditional methods require preset fixed constraints (such as mainlobe direction and null depth), making it difficult to simultaneously optimize multiple objectives such as signal-to-interference-and-noise ratio (SINR), energy consumption, and beam stability. Improved reinforcement learning (RL)-based methods, by designing a multi-dimensional reward function, encode objectives such as SINR, interference suppression, and energy consumption as reward items, driving the RL agent to autonomously explore the optimal trade-off between objectives, thereby achieving multi-objective collaborative optimization;
[0154] ③ Improved cross-scenario generalization efficiency: Traditional algorithms require parameter redesign for different scenarios (such as vehicle models and array layouts). However, this method uses a pre-trained beamforming weight optimization model and only requires a small number of new scenario samples to complete strategy fine-tuning, achieving rapid deployment across vehicle models and hardware.
[0155] Example 4
[0156] Figure 5 This is a structural diagram of a beamforming device provided by the fourth embodiment of the present invention. Figure 5 As shown, the device includes:
[0157] The data acquisition module 41 is used to acquire multimodal perception data of the vehicle environment; the multimodal perception data includes the original voice signal and environmental auxiliary data;
[0158] A state vector determination module 42, configured to determine a multidimensional state vector based on multimodal sensing data;
[0159] The beamforming weight determination module 43 is used to input the multidimensional state vector into the beamforming weight optimization model to obtain the beamforming weight corresponding to the original speech signal; the beamforming weight optimization model is obtained based on deep reinforcement learning training.
[0160] Furthermore, based on the above-mentioned embodiment of the invention, the beamforming weight determination module 43 includes:
[0161] A model inference unit is configured to input the multidimensional state vector into the policy network of the beamforming weight optimization model and use the policy network to infer a target action vector corresponding to the original speech signal; wherein the target action vector includes parameters to be adjusted corresponding to the covariance matrix influence hyperparameters, the constraint matrix influence hyperparameters, and the constraint response vector influence hyperparameters, respectively;
[0162] A parameter determination unit is used to call a preset covariance matrix formula, a preset constraint matrix formula and a preset constraint response vector formula based on each parameter to be adjusted to determine the corresponding covariance matrix, constraint matrix and constraint response vector;
[0163] The beamforming weight determination unit is used to call a preset beamforming weight closed-form solution formula to determine the corresponding beamforming weight based on the covariance matrix, the constraint matrix and the constraint response vector.
[0164] Furthermore, based on the above-mentioned embodiments of the invention, the training process of the beamforming weight optimization model includes:
[0165] Obtain a multidimensional state vector training set;
[0166] Initialize and build the initial beamforming weight optimization model and obtain the reward function configured for the initial beamforming weight optimization model;
[0167] In the first training phase, the initial beamforming weight optimization model is trained based on the multidimensional state vector training set and the reward function, using the covariance matrix influence hyperparameter as the action vector to be adjusted. The iterative training process is repeated until the initial beamforming weight optimization model converges, resulting in an intermediate beamforming weight optimization model.
[0168] In the second training stage, the covariance matrix influencing hyperparameters, the constraint matrix influencing hyperparameters and the constraint response vector influencing hyperparameters are used as the action vectors to be adjusted. The intermediate beamforming weight optimization model is trained based on the multi-dimensional state vector training set and the reward function. The iterative training process is repeated until the intermediate beamforming weight optimization model converges to obtain the beamforming weight optimization model.
[0169] Furthermore, based on the above-mentioned embodiments of the invention, the covariance matrix affects the hyperparameters including: forgetting factor, recursive step size and regularization coefficient; the constraint matrix affects the hyperparameters including: main lobe direction adjustment amount and null direction adjustment amount; the constraint response vector affects the hyperparameters including: main lobe gain and null depth.
[0170] Furthermore, based on the above-mentioned embodiment of the invention, the environmental auxiliary data at least includes:
[0171] Attitude data collected by the inertial measurement unit;
[0172] Distance measurement data collected by infrared time-of-flight sensors;
[0173] Temperature data collected by the temperature sensor;
[0174] Resource monitoring data collected by system monitoring tools.
[0175] Furthermore, based on the above-mentioned embodiment of the invention, the state vector determination module 42 includes:
[0176] a sound source angle error determining unit, configured to determine a sound source angle error based on the ranging data and the original speech signal;
[0177] A real-time signal-to-noise ratio determination unit, configured to determine a real-time signal-to-noise ratio based on an original speech signal;
[0178] a head posture change rate determining unit, configured to determine a head posture change rate based on the posture data;
[0179] a temperature parameter determining unit, configured to determine a temperature parameter based on the temperature data;
[0180] a reverberation estimation time determining unit, configured to determine a reverberation estimation time based on an original speech signal;
[0181] an array topology identifier determining unit, configured to determine an array topology identifier based on the ranging data;
[0182] a sound speed compensation factor determining unit, configured to determine the sound speed compensation factor based on the temperature data and the attitude data;
[0183] The computing load factor determining unit is configured to determine the computing load factor based on the resource monitoring data.
[0184] Furthermore, based on the above embodiments of the invention, the reward function is expressed as:
[0185] r t =α1·ΔSINR+α2·||Δw|| -1 -α3·Nulling_Error-α4·Power_Cost
[0186] Among them, r t represents the instantaneous reward at time step t; ΔSINR represents the change in signal-to-interference-and-noise ratio; Δw represents the change in beamforming weight; Nulling_Error represents the nulling depth suppression error; Power_Cost represents the power consumption control term; α1, α2, α3, and α4 represent weight coefficients.
[0187] Furthermore, based on the above embodiments of the invention, the beamforming device further includes:
[0188] The application module is used to apply the beamforming weights to downstream speech processing tasks, where the downstream speech processing tasks include at least speech signal enhancement and sound source localization.
[0189] The beamforming device provided in the embodiment of the present invention can execute the beamforming method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0190] Example 5
[0191] Figure 6 A schematic diagram of the structure of an electronic device 50 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0192] like Figure 6As shown, the electronic device 50 includes at least one processor 51 and a memory, such as a read-only memory (ROM) 52, a random access memory (RAM) 53, etc., which is communicatively connected to the at least one processor 51. The memory stores a computer program that can be executed by the at least one processor. The processor 51 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 52 or the computer program loaded from the storage unit 58 into the random access memory (RAM) 53. Various programs and data required for the operation of the electronic device 50 can also be stored in the RAM 53. The processor 51, ROM 52, and RAM 53 are connected to each other via a bus 54. An input / output (I / O) interface 55 is also connected to the bus 54.
[0193] Multiple components in the electronic device 50 are connected to the I / O interface 55, including an input unit 56, such as a keyboard, a mouse, etc.; an output unit 57, such as various types of displays, speakers, etc.; a storage unit 58, such as a magnetic disk, an optical disk, etc.; and a communication unit 59, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 59 allows the electronic device 50 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0194] The processor 51 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Examples of the processor 51 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors for running machine learning model algorithms, a digital signal processor (DSP), and any other suitable processor, controller, microcontroller, etc. The processor 51 executes the various methods and processes described above, such as the beamforming method.
[0195] In some embodiments, the beamforming method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 58. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 50 via ROM 52 and / or communication unit 59. When the computer program is loaded into RAM 53 and executed by processor 51, one or more steps of the beamforming method described above may be performed. Alternatively, in other embodiments, processor 51 may be configured to perform the beamforming method in any other suitable manner (e.g., via firmware).
[0196] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0197] In some embodiments, the beamforming method can be implemented as a computer program, which is intangibly contained in a computer program product. When executed by a processor, the computer program implements the beamforming method of the present invention. A computer program product can be understood as a software product whose solution is primarily implemented through the computer program. The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when executed by the processor, the computer program implements the functions / operations specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0198] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0199] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0200] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0201] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0202] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0203] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A beamforming method, characterized in that: The method comprises: Acquiring multimodal perception data of the vehicle environment; the multimodal perception data includes original voice signals and environmental auxiliary data; Determining a multidimensional state vector based on the multimodal perception data; The multidimensional state vector is input into a beamforming weight optimization model to obtain the beamforming weight corresponding to the original speech signal; the beamforming weight optimization model is obtained based on deep reinforcement learning training.
2. The method according to claim 1, characterized in that Inputting the multidimensional state vector into a beamforming weight optimization model to obtain a beamforming weight corresponding to the original speech signal includes: Inputting the multidimensional state vector into the strategy network of the beamforming weight optimization model, and using the strategy network to infer a target action vector corresponding to the original speech signal; wherein the target action vector includes parameters to be adjusted corresponding to the covariance matrix influence hyperparameter, the constraint matrix influence hyperparameter, and the constraint response vector influence hyperparameter, respectively; Based on each of the parameters to be adjusted, respectively calling a preset covariance matrix formula, a preset constraint matrix formula, and a preset constraint response vector formula to determine the corresponding covariance matrix, constraint matrix, and constraint response vector; Based on the covariance matrix, the constraint matrix and the constraint response vector, a preset beamforming weight closed-form solution formula is called to determine the corresponding beamforming weight.
3. The method according to claim 1, characterized in that The training process of the beamforming weight optimization model includes: Obtain a multidimensional state vector training set; Initializing and constructing an initial beamforming weight optimization model, and obtaining a reward function configured for the initial beamforming weight optimization model; In a first training phase, the initial beamforming weight optimization model is trained based on the multidimensional state vector training set and the reward function, using the covariance matrix influence hyperparameter as the action vector to be adjusted, and the iterative training process is repeated until the initial beamforming weight optimization model converges, thereby obtaining an intermediate beamforming weight optimization model; In the second training stage, the covariance matrix influencing hyperparameters, the constraint matrix influencing hyperparameters and the constraint response vector influencing hyperparameters are used as the action vectors to be adjusted, and the intermediate beamforming weight optimization model is trained based on the multidimensional state vector training set and the reward function. The iterative training process is repeated until the intermediate beamforming weight optimization model converges, thereby obtaining the beamforming weight optimization model.
4. The method according to claim 2 or 3, characterized in that The covariance matrix affects the hyperparameters including the forgetting factor, the recursive step size and the regularization coefficient; the constraint matrix affects the hyperparameters including the main lobe direction adjustment amount and the null direction adjustment amount; the constraint response vector affects the hyperparameters including the main lobe gain and the null depth.
5. The method according to claim 1, wherein The environmental auxiliary data at least includes: Attitude data collected by the inertial measurement unit; Distance measurement data collected by infrared time-of-flight sensors; Temperature data collected by the temperature sensor; Resource monitoring data collected by system monitoring tools.
6. The method according to claim 5, characterized in that The determining of a multidimensional state vector based on the multimodal sensing data includes: Determine a sound source angle error based on the ranging data and the original voice signal; determining a real-time signal-to-noise ratio based on the original speech signal; determining a head posture change rate based on the posture data; determining a temperature parameter based on the temperature data; determining a reverberation estimation time based on the original speech signal; determining an array topology identifier based on the ranging data; determining a sound speed compensation factor based on the temperature data and the posture data; A computation load factor is determined based on the resource monitoring data.
7. The method according to claim 3, characterized in that The reward function is expressed as: r t =α1·ΔSINR+α2·||Δw|| -1 -a3·Nulling_Error-a4·Power_Cost Among them, r t represents the instantaneous reward at time step t; ΔSINR represents the change in signal-to-interference-and-noise ratio; Δw represents the change in beamforming weight; Nulling_Error represents the nulling depth suppression error; Power_Cost represents the power consumption control term; α1, α2, α3, and α4 represent weight coefficients.
8. The method according to claim 1, characterized in that Also includes: The beamforming weights are applied to downstream speech processing tasks, wherein the downstream speech processing tasks include at least speech signal enhancement and sound source localization.
9. A beamforming device, characterized in that: The device comprises: A data acquisition module is used to acquire multimodal perception data of the vehicle environment; the multimodal perception data includes original voice signals and environmental auxiliary data; a state vector determination module, configured to determine a multidimensional state vector based on the multimodal sensing data; A beamforming weight determination module is used to input the multidimensional state vector into a beamforming weight optimization model to obtain the beamforming weight corresponding to the original speech signal; the beamforming weight optimization model is obtained based on deep reinforcement learning training.
10. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, where the computer program is executed by the at least one processor to enable the at least one processor to perform the beamforming method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the beamforming method according to any one of claims 1 to 8 when executed.
12. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the beamforming method according to any one of claims 1 to 8.