Video call system with far-field voice enhancement
By combining millimeter-wave radar and microphone array modules, the system can detect user movement in real time and optimize sound source localization, thus solving the problems of lag and anti-interference in voice enhancement in far-field video call systems and improving voice quality and call fluency.
Patent Information
- Application Number
- CN202511502519.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2026-01-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing video call systems, voice enhancement technology cannot detect user movement in real time in far-field mobile scenarios, resulting in delayed sound source location, voice signal distortion, insufficient gain, and inability to compensate for voice quality issues when users move quickly.
A millimeter-wave radar module is used to continuously collect user movement trajectory information. Combined with a microphone array module and a trajectory processing and feature extraction module, a "radar trajectory-voice feature" correlation model is constructed through a sound source location prediction and positioning parameter conversion module to achieve dynamic compensation mode and optimize the sound source localization process.
It achieves real-time and accurate voice enhancement during user movement, reduces voice stuttering and noise, and improves the smoothness and communication efficiency of far-field video calls, making it suitable for scenarios such as remote work and online education.
Smart Images

Figure CN121309761A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice enhancement technology, specifically to a far-field voice enhancement video call system. Background Technology
[0002] In scenarios such as remote work, online education, and home communication, the voice quality of video call systems directly affects communication efficiency. Especially in far-field environments, the dynamic changes in the location of the sound source caused by user movement become the core challenge for voice enhancement. Existing video call voice enhancement technologies mostly rely on microphone array beamforming algorithms to focus the voice signal by locating the direction of the sound source, but this has significant limitations in scenarios where the user is moving.
[0003] In traditional solutions, microphone arrays rely solely on speech signals for sound source localization. However, due to speech signal sampling delays and noise interference, the localization response lags behind the user's actual movement speed, causing beamforming direction to deviate from the actual sound source location, resulting in speech distortion and insufficient gain. Furthermore, existing technologies lack proactive perception of the user's movement trajectory and cannot predict changes in the sound source's location. When the user moves rapidly, the speech enhancement effect is prone to abrupt breaks.
[0004] Furthermore, the existing system has not established a correlation model between movement trajectory and voice features, making it difficult to dynamically adjust the voice enhancement strategy based on trajectory changes. As a result, when the user moves a large distance or moves at a high speed, problems such as voice amplitude attenuation and a decrease in the proportion of human voice frequency cannot be compensated in a timely manner, further reducing call quality.
[0005] Therefore, there is an urgent need for a technical solution that can sense user movement in real time, predict the location of sound sources in advance, and dynamically optimize voice enhancement to solve the voice quality problem of video calls in far-field mobile scenarios. Summary of the Invention
[0006] (a) Technical problems to be solved
[0007] To address the shortcomings of existing technologies, this invention provides a far-field voice-enhanced video call system, which solves the problems mentioned in the background section.
[0008] (II) Technical Solution
[0009] To achieve the above objectives, the present invention provides the following technical solution: a far-field voice enhancement video call system, comprising:
[0010] The millimeter-wave radar module is used to continuously collect user movement-related data and obtain user movement trajectory information;
[0011] Microphone array module, used for synchronous acquisition of voice signals;
[0012] The trajectory processing and feature extraction module is used to extract trajectory features from the user's movement trajectory information and to extract speech features from the speech signal;
[0013] The sound source location prediction and positioning parameter conversion module is used to predict the sound source location change trend based on user movement trajectory information and convert the sound source location change trend into look-ahead positioning parameters used by microphone array beamforming algorithm.
[0014] The association model and dynamic compensation module are used to construct a "radar trajectory-speech feature" association model. Then, based on the association model, it is determined whether to activate the dynamic compensation mode. In the dynamic compensation mode, the sound source localization process is optimized to achieve far-field speech enhancement in mobile scenarios.
[0015] As a further aspect of the present invention: the process by which the millimeter-wave radar module acquires the user's movement trajectory includes:
[0016] With a fixed time interval T radar Transmit and receive millimeter-wave signals, and receive reflected signals reflected by the user;
[0017] Based on the Doppler effect, the user's position coordinates (x, y) at different times are calculated by detecting the frequency change of the reflected signal and combining it with radar spatial layout parameters. n ,y n ,z n ), where n represents the nth acquisition time;
[0018] The sequence of position coordinates arranged in chronological order is (x1, y1, z1), (x2, y2, z2), ..., (x... k ,y k ,z k Let L0 be the user's movement trajectory, i.e., L0 = {(x1, y1, z1), (x2, y2, z2), ..., (x...} k ,y k ,z k )}.
[0019] As a further aspect of the present invention: the trajectory features extracted by the trajectory processing and feature extraction module include trajectory displacement features and trajectory velocity features, wherein:
[0020] Trajectory displacement characteristics Δs n Through formula The calculation yields, where (x) n ,y n ,z n ) and (x n+1 ,y n+1 ,z n+1 () represents the user location coordinates from two consecutive data collections;
[0021] The trajectory velocity feature v n Through formula The calculation yields T, where T radar The acquisition time interval of the millimeter-wave radar is defined as the time-varying velocity feature, which constitutes a velocity sequence V0={v1,v2,……,v...}. k}
[0022] As a further aspect of the present invention: the speech features extracted by the trajectory processing and feature extraction module include amplitude features and frequency proportion features, wherein:
[0023] Amplitude characteristic A is expressed by the formula The calculation yields, where s i Let f be the amplitude of the speech signal at the i-th sampling point, and N be the total number of sampling points in the speech segment, where N = f s ×T voice f s T is the speech sampling frequency. voice The time interval for extracting audio segments;
[0024] Frequency proportion feature R voice The proportion of energy within the target frequency range to the total energy is determined by the formula. The calculation yields, where E voice This is the sum of the squares of the amplitudes of all sampling points within the human voice frequency band, i.e. E total It is the sum of squares of the amplitudes of all sampling points within the speech sample, i.e. The target frequency range is 300Hz–3400Hz.
[0025] As a further aspect of the present invention: the process by which the sound source location prediction and localization parameter conversion module predicts the trend of sound source location change includes:
[0026] Based on the current time t and historical trajectory data, a linear prediction method is used to predict the sound source location (x) at the next time t+T (radar). t+T(radar) p ,y t+T(radar) p ,z t+T(radar) p );
[0027] Along the x-axis, through ;
[0028] Calculate the rate of change of position r along the x-axis. x :
[0029] Where, x t It is the x-axis coordinate at time t, x t−T(radar)It is the x-axis coordinate at time t−T (radar), r x Indicates the x-axis direction at T radar The change in position within, i.e., the rate of movement along the x-axis;
[0030] Similarly, through: and The y-axis moving speed r is obtained respectively. y and the speed of movement r in the z-axis direction z ;
[0031] Then through ;
[0032] Predict the x-axis coordinate at time t+T (radar). t+T(radar) p ;
[0033] Similarly, through and Predict the x-axis coordinate y at time t+T (radar) respectively. t+T(radar) p and z-axis coordinates t+T(radar) p ;
[0034] That is, the predicted value (x) of the sound source location at the next moment is obtained. t+T(radar) p ,y t+T(radar) p ,z t+T(radar) p );
[0035] By continuously calculating and rolling out the predicted location of the sound source at a series of future moments, a trend of sound source location change is formed.
[0036] As a further aspect of the present invention: the forward-looking positioning parameters include the distance d between the sound source and the center of the microphone array, the horizontal azimuth angle α, and the vertical elevation angle β;
[0037] in:
[0038] Distance d is obtained through the formula The calculation yields, where (x) s p ,y s p ,z s p ) is the predicted sound source location (x t+T(radar) p ,y t+T(radar) p ,z t+T(radar) p The relabeling of (x) m ,y m,z m () represents the coordinates of the center position of the microphone array;
[0039] The horizontal azimuth angle α is obtained through the formula The calculation shows that α is used to guide the adjustment of the microphone array beam in the horizontal direction;
[0040] Vertical elevation angle β is obtained through the formula The calculated value of β is used to guide the adjustment of the microphone array beam in the vertical direction.
[0041] As a further aspect of the present invention: the "radar trajectory-voice feature" association model includes:
[0042] Amplitude correlation formula: , where a1 is the influence coefficient of trajectory velocity feature on amplitude feature, and b1 is the basic amplitude when trajectory velocity feature is 0;
[0043] Frequency percentage correlation formula: , where a2 is the influence coefficient of trajectory displacement feature on frequency proportion feature, and b2 is the basic human voice proportion when trajectory displacement feature remains unchanged;
[0044] Composite Risk Index Formula: Where c1 and c2 are preset weighting coefficients, R risk Used to reflect the possibility of speech enhancement tomography.
[0045] As a further aspect of the present invention: the conditions for the association model and the dynamic compensation module to activate the dynamic compensation mode include: velocity condition: current trajectory velocity feature v current >0.5m / s; Risk condition: Current composite risk index R risk,current >R threshold , where R threshold This is a preset risk threshold.
[0046] As a further aspect of the present invention: the process by which the correlation model and dynamic compensation module optimize the sound source localization process in dynamic compensation mode includes:
[0047] The look-ahead positioning parameters are directly integrated into the microphone array positioning process, and the microphone array beamforming is adjusted based on the positioning parameters predicted by the radar.
[0048] The weighting coefficient w(θ) for beamforming is calculated based on the look-ahead positioning parameters.
[0049] The formula is: Where M0 is the number of microphones, d m Let θ be the distance between the m-th microphone and the center of the array, θ be the desired beam direction, and ϕ be the actual angle of arrival of the sound source. mLet be the installation angle of the m-th microphone, and λ be the wavelength of the speech signal. c is the speed of light, f voice The center frequency of the speech signal;
[0050] By adjusting the signal weights of each microphone channel, the array beam is directed toward the predicted sound source location, thus controlling the sound source localization response time to within 50 milliseconds.
[0051] As a further aspect of the present invention: the sound source location is associated with the user location, and the user location is identified as the sound source location, with the sound source changing as the user moves.
[0052] (III) Beneficial Effects
[0053] This invention provides a far-field voice enhancement video call system. Compared with the prior art, it has the following advantages:
[0054] The millimeter-wave radar module continuously collects user movement trajectory information and calculates the user's real-time location coordinates by combining the Doppler effect. This solves the problems of traditional voice positioning, which relies on a single voice signal and is susceptible to noise interference and has ambiguous positioning. By directly associating the user's location with the sound source location, the initial accuracy of sound source positioning is ensured from a physical perspective.
[0055] The sound source location prediction module uses historical trajectory data and a linear prediction algorithm to calculate the sound source location in advance, converting it into beamforming parameters such as horizontal azimuth angle α and vertical elevation angle β. This "prediction-adjustment" mechanism breaks the lag of traditional "real-time response," enabling the microphone array to be aligned with the moving sound source in advance, significantly reducing positioning delay.
[0056] The trajectory processing and feature extraction module simultaneously extracts trajectory and speech features, establishing a mapping relationship between movement state and speech characteristics through a "radar trajectory-speech feature" correlation model. For example, it accurately locates target speech by utilizing the energy proportion of the human voice frequency band, effectively filtering out environmental noise and improving the purity of the speech signal.
[0057] When the user's movement speed exceeds 5 m / s or the composite risk index exceeds the threshold, the system automatically activates the dynamic compensation mode. By directly integrating the look-ahead positioning parameters into the beamforming weight calculation, the array beam is directed towards the predicted sound source location in real time, keeping the positioning response time within 50 milliseconds. This avoids speech "discontinuity" and "distortion" problems when the user moves quickly, ensuring stable speech enhancement during movement.
[0058] The system integrates physical trajectory data from millimeter-wave radar with voice signal data from a microphone array to form a "dual-modal" positioning and enhancement scheme. Compared to traditional systems that rely solely on voice signals, it significantly improves anti-interference capabilities in far-field environments, high-noise environments, or scenarios where users move frequently, reducing positioning failures caused by weak voice signals or strong noise.
[0059] The system adapts to different usage scenarios through adjustable parameters. For example, by adjusting the parameters in the speech segment extraction interval or beamforming weight formula, it can meet the speech enhancement needs of different spatial scales such as home, office, and meeting, thereby improving the system's versatility.
[0060] Through precise sound source tracking, forward beam adjustment, and dynamic compensation, the system can maintain clear voice acquisition and enhancement effects while the user is moving, reducing issues such as voice stuttering, blurriness, and noise in video calls, and improving the smoothness and communication efficiency of far-field video calls. It is especially suitable for scenarios that require free movement, such as remote work, family video chat, and online education.
[0061] In summary, this invention effectively solves the core problems of voice enhancement in far-field mobile scenarios, such as positioning lag, weak anti-interference, and poor adaptability, through multimodal fusion of millimeter-wave radar and microphone array, look-ahead localization prediction, dynamic compensation mechanism, and multi-feature association model, and significantly improves the performance and user experience of video call systems. Attached Figure Description
[0062] Figure 1 This is a system block diagram of a far-field voice enhancement video call system according to the present invention.
[0063] Figure 2 This is a schematic diagram of the process of a far-field voice enhancement video call system according to the present invention. Detailed Implementation
[0064] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0065] Please see Figure 1 and Figure 2 As shown, the embodiments of the present invention provide the following technical solutions:
[0066] As an embodiment of the present invention:
[0067] This invention relates to a far-field voice-enhanced video call system, comprising:
[0068] The millimeter-wave radar module is used to continuously collect user movement-related data and obtain user movement trajectory information;
[0069] Microphone array module, used for synchronous acquisition of voice signals;
[0070] The trajectory processing and feature extraction module is used to extract trajectory features from the user's movement trajectory information and to extract speech features from the speech signal;
[0071] The sound source location prediction and positioning parameter conversion module is used to predict the sound source location change trend based on user movement trajectory information and convert the sound source location change trend into look-ahead positioning parameters used by microphone array beamforming algorithm.
[0072] The association model and dynamic compensation module are used to construct a "radar trajectory-speech feature" association model. Then, based on the association model, it is determined whether to activate the dynamic compensation mode. In the dynamic compensation mode, the sound source localization process is optimized to achieve far-field speech enhancement in mobile scenarios.
[0073] The implementation steps of a far-field voice enhancement video call system are as follows:
[0074] The first step involves continuously collecting user movement-related data using millimeter-wave radar to obtain user movement trajectory information; simultaneously, a microphone array synchronously collects voice signals at fixed time intervals T. voice Extract audio segments as audio samples for the corresponding trajectory data;
[0075] The second step is to predict the trend of sound source location change based on the trajectory information and convert it into look-ahead positioning parameters that can be used by the microphone array beamforming algorithm.
[0076] Thirdly, a "radar trajectory-speech feature" correlation model is constructed to determine whether to activate the dynamic compensation mode in order to achieve far-field speech enhancement in mobile scenarios.
[0077] Example 1, as the foundational solution of this invention, innovatively integrates a millimeter-wave radar module and a microphone array module to construct a complete technical framework of "trajectory acquisition - feature extraction - sound source prediction - dynamic compensation." Its core advantage lies in overcoming the limitations of traditional far-field voice enhancement technology in user mobility scenarios: it continuously acquires the user's movement trajectory using millimeter-wave radar, combines trajectory processing and feature extraction modules to achieve correlation analysis between trajectory and voice features, generates look-ahead positioning parameters through a sound source location prediction module, enabling the microphone array beamforming algorithm to adapt to the sound source's movement trend in advance; simultaneously, it dynamically determines compensation needs using a "radar trajectory - voice feature" correlation model, initiating a dynamic compensation mode to optimize the positioning process. This solution effectively solves the problems of voice signal distortion and positioning delay in mobile scenarios, providing a more stable and clearer voice enhancement effect for far-field video calls.
[0078] As a second embodiment of the present invention:
[0079] In its specific implementation, compared to Embodiment 1, the technical solution of this embodiment differs from that of Embodiment 1 only in that this embodiment also proposes a specific method for millimeter-wave radar to capture user movement trajectories:
[0080] The data acquisition principle of millimeter-wave radar is that millimeter-wave radar transmits millimeter-wave signals of a specific frequency. When the millimeter-wave signal encounters the user, it is reflected, and the radar receives the reflected signal.
[0081] In this embodiment, the user is the target object;
[0082] According to the Doppler effect, the frequency of the reflected signal changes when the user moves. The changed frequency fr is related to the user's moving speed v and the relative position of the radar and the user.
[0083] Its basic relationship is expressed as follows: ;
[0084] Among them, f r θ is the frequency of the reflected signal received by the radar, θ is the angle between the user's moving direction and the radar's transmitted signal direction, c is the speed of light, which is a constant. In this embodiment, its value is 3×10⁸ m / s, v is the user's moving speed, and f₀ is the initial frequency of the radar transmitted signal.
[0085] By continuously detecting the frequency f of the reflected signal r By combining the changes in the radar with the parameters corresponding to the radar's spatial layout, the user's position coordinates at different times can be calculated, thereby capturing the user's movement trajectory.
[0086] In actual implementation, millimeter-wave radar operates at a fixed time interval T. radarThe system transmits and receives signals, acquiring a set of reflected signal data each time. After signal processing, the current user location coordinates (x, y, y) are calculated. n ,y n ,z n ), where n represents the nth data acquisition time;
[0087] In this context, signal processing refers to the basic processing of filtering and amplifying the reflected signal data, which is an existing technology. The focus of this invention is not on the details of this signal processing, so it will not be described in detail.
[0088] As the number of data collections increases, we obtain (x1, y1, z1), (x2, y2, z2), ..., (x k ,y k ,z k The corresponding user location coordinate sequence, these location points arranged in chronological order constitute the user movement trajectory L0, that is, L0={(x1,y1,z1), (x2,y2,z2), ..., (x k ,y k ,z k )}.
[0089] Example 2, building upon Example 1, further clarifies the specific principles and methods of millimeter-wave radar capturing user movement trajectories. Its core value lies in enhancing the scientific rigor and reliability of trajectory acquisition. By introducing the Doppler effect principle, it details the quantitative relationship between radar reflected signal frequency and user movement speed and relative position, clarifying the process of calculating user position coordinates based on changes in reflected signal frequency. This refined design transforms millimeter-wave radar trajectory acquisition from a "black box operation" into a quantifiable and verifiable technical process, ensuring the accuracy of user movement trajectory data. This provides high-quality foundational data support for subsequent stages such as sound source location prediction and feature correlation analysis, enhancing the overall system's technical feasibility and stability.
[0090] As an embodiment of the present invention:
[0091] In its specific implementation, compared to Embodiment 1 and Embodiment 2, the technical solution of this embodiment combines the solutions of Embodiment 1 and Embodiment 2. The only difference between this embodiment and Embodiment 1 and Embodiment 2 is that this embodiment also proposes an implementation method for trajectory data processing and feature extraction, as detailed below:
[0092] Basic features are extracted from the user movement trajectory L0 acquired by millimeter-wave radar. These basic features include:
[0093] Step 2.1, Extraction of trajectory displacement features:
[0094] pass ;
[0095] Calculate the displacement Δs between two adjacent points. n That is, trajectory displacement characteristics;
[0096] Among them, (x) n ,y n ,z n ) and (x n+1 ,y n+1 ,z n+1 ) represents the user location coordinates between two consecutive data collections, Δs n This indicates the distance the user moves within the time interval between two consecutive data collections;
[0097] Step 2.2, Extraction of trajectory velocity features:
[0098] Combined with the data acquisition time interval T radar ,pass ;
[0099] Calculate the user's time interval T radar Average speed v within n That is, trajectory velocity characteristics;
[0100] By continuously calculating the movement speed for each time period using the above method, a sequence of user movement speed over time is obtained: V0={v1,v2,……,v k This is used to predict the trend of sound source location changes and determine the activation conditions of dynamic compensation mode.
[0101] Simultaneously, speech features are extracted from the synchronously acquired speech signals. These speech features include the amplitude features and frequency proportion features of the speech.
[0102] Step 2.3, Extraction of amplitude features:
[0103] pass Calculate the average amplitude A of each speech segment, i.e., the amplitude feature;
[0104] Among them, s i Let f be the amplitude of the speech signal at the i-th sampling point, and N be the total number of sampling points in the speech segment, where N = f s ×T voice f s This is a fixed value, referring to the speech sampling frequency;
[0105] Step 2.4, Extraction of frequency proportion features:
[0106] A basic frequency analysis of the speech samples is performed using Fast Fourier Transform (FFT) to statistically analyze the proportion of the target frequency range. In this embodiment, human voices are mainly concentrated in the 300Hz-3400Hz range, which is taken as the target frequency range. Subsequently, the proportion R of energy within the target frequency range to the total energy is calculated. voice The method is as follows:
[0107] First, the speech samples are segmented by frequency, that is, into human voice frequency band and non-human voice frequency band;
[0108] Subsequently passed Calculate the sum of squares of the amplitudes of all sampling points within the human voice frequency band;
[0109] And through Calculate the sum of squares of the amplitudes of all sampling points within the speech sample, i.e., the total energy E. total ;
[0110] After that, through Calculate the proportion R of energy within the target frequency range to the total energy. voice That is, frequency proportion characteristics;
[0111] Among them, R voice The lower the value, the more severe the noise or speech distortion.
[0112] Example 3 focuses on the details of feature extraction from trajectory data and speech signals, improving the accuracy of system analysis through quantitative feature definition. On the one hand, displacement and velocity features are extracted from the radar trajectory to form a velocity sequence, providing a quantitative basis for predicting the trend of sound source location changes and judging dynamic compensation conditions. On the other hand, amplitude and frequency proportion features are extracted from the speech signal, clarifying the calculation method of the energy proportion of the human voice frequency band, which can intuitively reflect the noise interference level of the speech signal. The accurate extraction of these features upgrades the correlation analysis of "radar trajectory-speech features" from qualitative judgment to quantitative calculation, not only laying a data foundation for the subsequent construction of correlation models, but also enabling real-time perception of speech quality fluctuations through feature changes, providing reliable feature indicators for the activation of dynamic compensation mode.
[0113] As an embodiment of the present invention:
[0114] In its specific implementation, compared to Embodiments 1, 2, and 3, the difference between this embodiment and Embodiments 1, 2, and 3 lies only in that this embodiment also proposes specific steps for predicting the trend of sound source location changes and providing look-ahead localization parameters:
[0115] Step 3.1 Prediction of the trend of sound source location change
[0116] Based on user movement trajectory and speed information obtained by millimeter-wave radar, predict the trend of sound source location change;
[0117] In this embodiment, it is assumed that the user's movement is continuous, and the location of the sound source is inferred in the near future based on current and historical trajectory data.
[0118] In this embodiment, the sound source is associated with the user's location. The user's location is considered to be the sound source location because the sound source moves with the user when the user makes a sound.
[0119] Let the current time be t. The user location data collected from t−m×T (radar) to time t is given by the coordinates (x, y). t−mT(radar), y t−mT(radar), z t−mT(radar) ) to (x t ,y t ,z t );
[0120] Where t−mT(radar) is an expression used to specify a specific time point, where t represents the current time, m is a preset value, a positive integer used to represent the "number of historical data," meaning the number of times millimeter-wave radar data was collected to analyze user movement patterns, etc., and T(radar) is the collection time interval T of the millimeter-wave radar. radar , which represents the past time corresponding to m "radar acquisition time intervals Tradar" based on the current time t;
[0121] For example, if m=2, Tradar=50 milliseconds, and the current time t is the 1000th millisecond, then t−mTradar=1000−2×50=900 milliseconds.
[0122] Subsequently, a linear prediction method is used to predict the sound source location (x) at the next time step t+T (radar). t+T(radar) p ,y t+T(radar) p ,z t+T(radar) p );
[0123] Taking the x-axis direction as an example, through ;
[0124] Calculate the rate of change of position r along the x-axis. x :
[0125] Where, x t It is the x-axis coordinate at time t, x t−T(radar) It is the x-axis coordinate at time t−T (radar), r x Indicates the x-axis direction at T radar The change in position within, i.e., the rate of movement along the x-axis;
[0126] Similarly, through: and The y-axis moving speed r is obtained respectively. y and the speed of movement r in the z-axis direction z ;
[0127] Then through ;
[0128] Predict the x-axis coordinate at time t+T (radar). t+T(radar) p ;
[0129] Similarly, through and Predict the x-axis coordinate y at time t+T (radar) respectively. t+T(radar) p and z-axis coordinates t+T(radar) p ;
[0130] That is, the predicted value (x) of the sound source location at the next moment is obtained. t+T(radar) p ,y t+T(radar) p ,z t+T(radar) p );
[0131] Then, through continuous rolling calculations, the predicted location of the sound source at a series of future moments is obtained, which serves as the location change trend and provides look-ahead positioning parameters for the microphone array beamforming algorithm.
[0132] Step 3.2: Look-ahead positioning parameter conversion and provision
[0133] The predicted sound source location change trend information is converted into the positioning parameters required by the microphone array beamforming algorithm;
[0134] Microphone array beamforming algorithms adjust the beam direction based on the angle and distance information of the sound source position relative to the microphone array to achieve target speech extraction.
[0135] The specific method is as follows:
[0136] First, the coordinates of the center position of the microphone array are marked as (x m ,y m ,z m ), predicting the location of the sound source (x t+T(radar) p ,y t+T(radar) p ,z t+T(radar) p ) Rewritten as (x s p,y s p ,z s p );
[0137] The center position of the microphone array is predetermined during system installation;
[0138] Then through ;
[0139] Calculate the distance d between the sound source and the center of the microphone array;
[0140] In the x−y plane, according to the definition of the tangent function for a right triangle, the tangent value of the azimuth angle α is expressed as: ;
[0141] Subsequently passed The azimuth angle α of the sound source relative to the microphone array on the horizontal plane is calculated.
[0142] Theoretically, α ranges from −90° to 90°, but by combining the signs of actual Δx and Δy, its specific direction within the complete horizontal azimuth range of −180° to 180° can be further determined:
[0143] If Δx>0 and Δy>0, then α is between 0° and 90°, representing the horizontal orientation of the sound source in the first quadrant of the microphone array center;
[0144] If Δx < 0 and Δy > 0, then α is between 90° and 180°. In actual calculations, the arctan result will be negative, so a 180° correction needs to be added, indicating that the sound source is in the horizontal direction of the second quadrant.
[0145] If Δx < 0 and Δy < 0, then α is between −180° and −90° (in actual calculations, the arctan result will be negative, so a 180° correction needs to be added, indicating that the sound source is in the horizontal direction of the third quadrant).
[0146] If Δx > 0 and Δy < 0, then α is between −90° and 0°, indicating that the sound source is in the horizontal direction of the fourth quadrant.
[0147] Among them, α defines the deflection angle of the sound source relative to the center of the microphone array on the horizontal plane, which is used to guide the adjustment of the microphone array beam in the horizontal direction so that the beam is directed towards the horizontal direction of the sound source.
[0148] In the vertical plane, according to the definition of the tangent function, the formula for the tangent value of the elevation angle β is expressed as:
[0149] ;
[0150] Subsequently passed The elevation angle β of the sound source relative to the microphone array on the vertical plane is calculated.
[0151] The value of β ranges from −90° to 90°:
[0152] When Δz>0, β is a positive value, indicating that the sound source is above the center of the microphone array. The larger β is, the higher the sound source is in the vertical direction.
[0153] When Δz < 0, β is negative, indicating that the sound source is below the center of the microphone array. The larger the absolute value of β, the lower the sound source is in the vertical direction.
[0154] When Δz=0, β=0°, indicating that the sound source and the center of the microphone array are in the same vertical plane at the same level.
[0155] Among them, the elevation angle β defines the deflection angle of the sound source relative to the center of the microphone array on the vertical plane. It is used to guide the adjustment of the microphone array beam in the vertical direction, so that the beam can be accurately oriented towards the sound source in the vertical dimension. This ensures that the microphone array beam can accurately cover the sound source position in three-dimensional space, thereby improving the target speech acquisition effect.
[0156] Example 4 refines the design of the sound source location prediction and positioning parameter conversion process, significantly improving the timeliness and accuracy of sound source localization in dynamic scenarios. Its core advantage lies in proposing a linear prediction method based on historical trajectory data. By calculating the movement rate along the x, y, and z axes, it proactively predicts the sound source location in the near future and converts the predicted location into parameters such as azimuth and elevation angles required for microphone array beamforming. This "look-ahead" approach to providing positioning parameters breaks through the latency bottleneck of traditional real-time speech signal-based localization, enabling the microphone array to adjust its beam direction in advance to adapt to user movement. Simultaneously, by clearly defining the calculation rules and quadrant divisions for azimuth and elevation angles, it ensures precise beam pointing in three-dimensional space, effectively improving the extraction efficiency of target speech in mobile scenarios and reducing speech signal attenuation or distortion caused by sound source movement.
[0157] As a fifth embodiment of the present invention:
[0158] In its specific implementation, compared to Embodiments 1, 2, 3, and 4, the difference between this embodiment and Embodiments 1, 2, 3, and 4 lies only in that this embodiment also proposes specific steps for establishing the "radar trajectory-voice feature" association model and activating the dynamic compensation mode:
[0159] Step 4.1: Construct a "radar trajectory-speech feature" association model
[0160] The "radar trajectory-voice feature" association model is used to associate user trajectory information (with movement speed as the key feature) and voice features acquired by millimeter-wave radar to determine whether to activate the dynamic compensation mode.
[0161] Step 4.1.1: Targeting the amplitude feature A and the trajectory velocity feature v n By fitting multiple sets of data, the correlation formula is obtained: ;
[0162] Wherein, a1 is the influence coefficient of trajectory velocity feature on amplitude feature. In this embodiment, it represents the change in amplitude for every 1 m / s increase in velocity, which is specifically determined by data fitting; b1 is the base amplitude when trajectory velocity feature is 0. In this embodiment, it is obtained by fitting static scene data.
[0163] Step 4.1.2, regarding the frequency proportion feature R voice Trajectory displacement characteristics Δs n By fitting multiple sets of data, the correlation formula is obtained: ;
[0164] Where a2 is the influence coefficient of trajectory displacement feature on frequency proportion feature, and b2 is the basic human voice proportion when trajectory displacement feature remains unchanged. In this embodiment, it is obtained by fitting user stable movement scene data.
[0165] Step 4.1.3, Comprehensive trajectory displacement characteristics Δs n and frequency proportion characteristics R voice Construct a composite risk index R risk The correlation formula is: ;
[0166] Among them, c1 and c2 are the weighting coefficients of the pre-screening. In this implementation, Rrisk can effectively reflect the possibility of speech enhancement tortuosity. The larger the value, the higher the tortuosity risk.
[0167] Step 4.1.4: Integrate these correlation formulas to form a "radar trajectory-speech feature" correlation model. This model takes radar trajectory features as input and outputs a predicted speech feature value R. voice and A and risk indicator R risk This is used to determine whether speech enhancement may have a gap problem;
[0168] Step 4.2: Activate dynamic compensation mode and shorten sound source localization response time.
[0169] Dynamic compensation mode is activated when the following conditions are met:
[0170] Velocity condition: v current >0.5m / s, where the threshold of 0.5m / s can be adjusted according to actual testing and scenario requirements.
[0171] Risk conditions: R risk,current >Rthreshold, The risk of speech enhancement tomography is considered high.
[0172] Among them, R threshold The preset risk threshold was determined through extensive testing.
[0173] Step 4.3: In dynamic compensation mode, re-optimize the sound source localization process to shorten the localization response time;
[0174] The specific steps are as follows:
[0175] Traditional sound source localization algorithms rely on static scenes, resulting in significant localization delays. This invention, in dynamic compensation mode, utilizes trajectory prediction information acquired by millimeter-wave radar to directly integrate look-ahead localization parameters into the microphone array localization process. Instead of waiting for the microphone array to perform slow localization calculations, it rapidly adjusts the microphone array beamforming based on the localization parameters predicted by the radar.
[0176] Using the look-ahead positioning parameters predicted by millimeter-wave radar, along with the predicted location coordinates of the sound source, the horizontal azimuth angle α, and the vertical elevation angle β, as the initial input parameters for the microphone array beamforming algorithm, we substitute them into the beamforming direction control formula: ;
[0177] Where w(θ) is the beamforming weighting coefficient, used to adjust the signal weights of different microphone channels; M0 is the number of microphones in the microphone array; d m ϕ is the distance between the m-th microphone and the center of the array; θ is the desired beam direction; ϕ is the actual angle at which the sound source reaches the microphone; ϕ m λ is the installation angle of the m-th microphone; λ is the wavelength of the speech signal. c is the speed of light, f voice The center frequency of the speech signal;
[0178] The beamforming weighting coefficient w(θ) is calculated, and then the signal weights of each microphone channel in the microphone array are adjusted so that the array beam quickly moves toward the predicted sound source position, thereby achieving directional enhancement of the target speech acquisition.
[0179] The response time of the microphone array for sound source localization is denoted as t. original In dynamic compensation mode, by optimizing the sound source localization process, the response time t is reduced. new Keep it within 50 milliseconds;
[0180] Example 5 focuses on optimizing the construction logic and dynamic compensation mechanism of the "radar trajectory-voice feature" correlation model, significantly improving the system's ability to predict and adaptively adjust for voice enhancement gap risks. By establishing quantitative correlation formulas between trajectory velocity and voice amplitude, and between trajectory displacement and frequency proportion, a composite risk index, Rrisk, is constructed, enabling accurate assessment of voice enhancement gap risks. Simultaneously, the activation conditions for the dynamic compensation mode are clarified, and in the compensation mode, beamforming weight calculation is directly optimized by introducing radar look-ahead positioning parameters, reducing the positioning response time to less than 50 milliseconds. This design allows the system to adaptively adjust its operating mode based on user movement status and voice quality risk, effectively avoiding voice enhancement gap problems caused by positioning delays at high movement speeds, and significantly improving the stability and reliability of far-field voice enhancement in complex dynamic scenarios.
[0181] As an embodiment of the present invention:
[0182] In specific implementation, compared with Embodiment 1, Embodiment 2, Embodiment 3, Embodiment 4 and Embodiment 5, the technical solution of this embodiment is to combine the solutions of Embodiment 1, Embodiment 2, Embodiment 3, Embodiment 4 and Embodiment 5.
[0183] Example 6 integrates all the technical solutions from Examples 1 to 5, forming a complete and closed-loop far-field voice enhancement technology system. Its core advantage lies in achieving synergistic effects among various technical modules. This example not only covers the entire process details of trajectory acquisition, feature extraction, sound source prediction, correlation modeling, and dynamic compensation, but also constructs a complete technical chain of "data acquisition - feature analysis - forward-looking prediction - risk assessment - dynamic optimization" through the organic combination of various modules. This integrated design fully leverages the trajectory perception advantages of millimeter-wave radar and the voice acquisition advantages of microphone arrays, ensuring both the accuracy of trajectory data and the quantification of feature analysis, while also achieving forward-looking sound source localization and adaptive compensation mechanisms. This comprehensively improves the voice enhancement effect of far-field video call systems in user mobile scenarios, possessing extremely high practical application value.
[0184] It should be stated that all user data collected in this application was collected with the user's consent and authorization, and the use of user data is legal and compliant, and the use and processing of user data comply with the relevant laws, regulations and standards of the relevant regions.
[0185] Furthermore, any content not described in detail in this specification is existing technology known to those skilled in the art.
[0186] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0187] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
[0188] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A far-field voice-enhanced video call system, characterized in that, include: The millimeter-wave radar module is used to continuously collect user movement-related data and obtain user movement trajectory information; Microphone array module, used for synchronous acquisition of voice signals; The trajectory processing and feature extraction module is used to extract trajectory features from the user's movement trajectory information and to extract speech features from the speech signal; The sound source location prediction and positioning parameter conversion module is used to predict the sound source location change trend based on user movement trajectory information and convert the sound source location change trend into look-ahead positioning parameters used by microphone array beamforming algorithm. The association model and dynamic compensation module are used to construct a "radar trajectory-speech feature" association model. Then, based on the association model, it is determined whether to activate the dynamic compensation mode. In the dynamic compensation mode, the sound source localization process is optimized to achieve far-field speech enhancement in mobile scenarios.
2. The far-field voice enhancement video call system according to claim 1, characterized in that: The process by which the millimeter-wave radar module collects the user's movement trajectory includes: With a fixed time interval T radar Transmit and receive millimeter-wave signals, and receive reflected signals reflected by the user; Based on the Doppler effect, the user's position coordinates (x, y) at different times are calculated by detecting the frequency change of the reflected signal and combining it with radar spatial layout parameters. n ,y n ,z n ), where n represents the nth acquisition time; The sequence of position coordinates arranged in chronological order is (x1, y1, z1), (x2, y2, z2), ..., (x... k ,y k ,z k Let L0 be the user's movement trajectory, i.e., L0 = {(x1, y1, z1), (x2, y2, z2), ..., (x...} k ,y k ,z k )}.
3. The far-field voice enhancement video call system according to claim 1, characterized in that: The trajectory features extracted by the trajectory processing and feature extraction module include trajectory displacement features and trajectory velocity features, wherein: Trajectory displacement characteristics Δs n Through formula The calculation yields, where (x) n ,y n ,z n ) and (x n+1 ,y n+1 ,z n+1 () represents the user location coordinates from two consecutive data collections; The trajectory velocity feature v n Through formula The calculation yields T, where T radar The acquisition time interval of the millimeter-wave radar is defined as the time-varying velocity feature, which constitutes a velocity sequence V0={v1,v2,……,v...}. k } 4. A far-field voice enhancement video call system according to claim 1, characterized in that: The speech features extracted by the trajectory processing and feature extraction module include amplitude features and frequency proportion features, wherein: Amplitude characteristic A is expressed by the formula The calculation yields, where s i Let f be the amplitude of the speech signal at the i-th sampling point, and N be the total number of sampling points in the speech segment, where N = f s ×T voice f s T is the speech sampling frequency. voice The time interval for extracting audio segments; Frequency proportion feature R voice The proportion of energy within the target frequency range to the total energy is determined by the formula. The calculation yields, where E voice This is the sum of the squares of the amplitudes of all sampling points within the human voice frequency band, i.e. E total It is the sum of squares of the amplitudes of all sampling points within the speech sample, i.e. The target frequency range is 300Hz–3400Hz.
5. A far-field voice enhancement video call system according to claim 1, characterized in that: The process by which the sound source location prediction and localization parameter conversion module predicts the trend of sound source location changes includes: Based on the current time t and historical trajectory data, a linear prediction method is used to predict the sound source location (x) at the next time t+T (radar). t+T(radar) p ,y t+T(radar) p ,z t+T(radar) p ); Along the x-axis, through ; Calculate the rate of change of position r along the x-axis. x : Where, x t It is the x-axis coordinate at time t, x t−T(radar) It is the x-axis coordinate at time t−T (radar), r x Indicates the x-axis direction at T radar The change in position within, i.e., the rate of movement along the x-axis; Similarly, through: and The y-axis moving speed r is obtained respectively. y and the speed of movement r in the z-axis direction z ; Then through ; Predict the x-axis coordinate at time t+T (radar). t+T(radar) p ; Similarly, through and Predict the x-axis coordinate y at time t+T (radar) respectively. t+T(radar) p and z-axis coordinates t+T(radar) p ; That is, the predicted value (x) of the sound source location at the next moment is obtained. t+T(radar) p ,y t+T(radar) p ,z t+T(radar) p ); By continuously calculating and rolling out the predicted location of the sound source at a series of future moments, a trend of sound source location change is formed.
6. A far-field voice enhancement video call system according to claim 1, characterized in that: The look-ahead positioning parameters include the distance d between the sound source and the center of the microphone array, the horizontal azimuth angle α, and the vertical elevation angle β; in: Distance d is obtained through the formula The calculation yields, where (x) s p ,y s p ,z s p ) is the predicted sound source location (x t+T(radar) p ,y t+T(radar) p ,z t+T(radar) p The relabeling of (x) m ,y m ,z m () represents the coordinates of the center position of the microphone array; The horizontal azimuth angle α is obtained through the formula The calculation shows that α is used to guide the adjustment of the microphone array beam in the horizontal direction; Vertical elevation angle β is obtained through the formula The calculated value of β is used to guide the adjustment of the microphone array beam in the vertical direction.
7. A far-field voice enhancement video call system according to claim 1, characterized in that: The "radar trajectory-voice feature" association model includes: Amplitude correlation formula: , where a1 is the influence coefficient of trajectory velocity feature on amplitude feature, and b1 is the basic amplitude when trajectory velocity feature is 0; Frequency percentage correlation formula: , where a2 is the influence coefficient of trajectory displacement feature on frequency proportion feature, and b2 is the basic human voice proportion when trajectory displacement feature remains unchanged; Composite Risk Index Formula: Where c1 and c2 are preset weighting coefficients, R risk Used to reflect the possibility of speech enhancement tomography.
8. A far-field voice enhancement video call system according to claim 1, characterized in that: The conditions for the correlation model and dynamic compensation module to activate the dynamic compensation mode include: Velocity condition: Current trajectory velocity characteristic v current >0.5m / s; Risk conditions: Current composite risk index R risk,current >R threshold , where R threshold This is a preset risk threshold.
9. A far-field voice enhancement video call system according to claim 1, characterized in that: The process by which the correlation model and dynamic compensation module optimize the sound source localization process in dynamic compensation mode includes: The look-ahead positioning parameters are directly integrated into the microphone array positioning process, and the microphone array beamforming is adjusted based on the positioning parameters predicted by the radar. The weighting coefficient w(θ) for beamforming is calculated based on the look-ahead positioning parameters. The formula is: Where M0 is the number of microphones, d m Let θ be the distance between the m-th microphone and the center of the array, θ be the desired beam direction, and ϕ be the actual angle of arrival of the sound source. m Let be the installation angle of the m-th microphone, and λ be the wavelength of the speech signal. c is the speed of light, f voice The center frequency of the speech signal; By adjusting the signal weights of each microphone channel, the array beam is directed toward the predicted sound source location, thus controlling the sound source localization response time to within 50 milliseconds.
10. A far-field voice enhancement video call system according to claim 5, characterized in that: The location of the sound source is associated with the location of the user, and the location of the user is identified as the location of the sound source. The sound source changes as the user moves.