Voice recognition control method and system for electric two-wheeled vehicle
By synchronously collecting and processing multi-source data of electric two-wheelers, establishing a wind field intensity distribution matrix, and performing asymmetric differential compensation and dynamic wind noise suppression, the problem of low voice recognition accuracy in open environments with electric two-wheelers is solved, and high-precision and fast voice recognition is achieved.
Patent Information
- Application Number
- CN202510872156.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional in-vehicle voice recognition systems have low recognition accuracy in the open environment of electric two-wheeled vehicles due to dynamic wind field interference and cannot adapt to complex riding conditions.
By integrating a three-axis wind direction sensor, a vehicle speed sensor, a six-axis attitude sensor and a stereo microphone array for synchronous data acquisition, a multi-source data set and the original voice signal are obtained; the multi-source data set is subjected to coordinate system transformation and vehicle aerodynamic calculation to obtain a wind field intensity distribution matrix; based on the wind field intensity distribution matrix, asymmetric differential compensation is performed on the original voice signal to perform dynamic wind noise suppression processing and voice recognition.
The voice recognition accuracy and response speed of electric two-wheeled vehicles under different wind direction interference are improved, ensuring the reliability and stability of voice recognition.
Smart Images

Figure CN120673757A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voice recognition control, and in particular to a voice recognition control method and system for an electric two-wheeled vehicle. Background Art
[0002] Electric two-wheelers have a strong user demand for navigation, communication, and safety control via voice commands. However, the open riding environment of electric two-wheelers differs fundamentally from the enclosed cabin of traditional cars. Riders are directly exposed to complex airflow, posing challenges to traditional in-vehicle voice recognition technology.
[0003] Existing in-vehicle speech recognition systems are primarily designed for closed vehicle environments, relying on a stable acoustic environment and a fixed microphone layout to ensure recognition accuracy. However, these technical solutions are unable to adapt to the dynamic wind field interference in the open environment of electric two-wheeled vehicles. During the riding of an electric two-wheeled vehicle, the natural wind direction and the vehicle's speed combine to form a complex relative wind field. Especially during dynamic riding conditions such as turning, tilting, accelerating, and decelerating, the impact of wind direction on microphones in different positions is significantly asymmetric and time-varying, resulting in low recognition accuracy for traditional speech recognition algorithms. Summary of the Invention
[0004] The present invention provides a voice recognition control method and system for an electric two-wheeled vehicle, which ensures the reliability of voice recognition under different wind direction interference intensities and effectively improves recognition accuracy and response speed.
[0005] In a first aspect, the present invention provides a voice recognition control method for an electric two-wheeled vehicle, the voice recognition control method for an electric two-wheeled vehicle comprising: Synchronous data acquisition is performed on the electric two-wheeled vehicle's three-axis wind direction sensor, vehicle speed sensor, six-axis attitude sensor, and stereo microphone array to obtain a multi-source data set and original speech signal; Performing coordinate system conversion and vehicle aerodynamic calculation on the multi-source data set to obtain a wind field intensity distribution matrix; Performing asymmetric differential compensation on the original voice signal based on the wind field intensity distribution matrix to obtain a compensated voice signal; Dynamic wind noise suppression processing and speech recognition are performed on the compensated speech signal to generate a target speech recognition result.
[0006] In combination with the first aspect, in a first implementation of the first aspect of the present invention, synchronously collecting data from a three-axis wind direction sensor, a vehicle speed sensor, a six-axis attitude sensor, and a stereo microphone array of the electric two-wheeled vehicle to obtain a multi-source data set and an original voice signal includes: The three-axis wind direction sensor arranged on the electric two-wheeled vehicle is sampled to obtain real-time wind direction vector data including the X-axis longitudinal wind speed component, the Y-axis transverse wind speed component and the Z-axis vertical wind speed component; The vehicle speed sensor and six-axis attitude sensor installed on the electric two-wheeled vehicle are synchronously sampled to obtain vehicle motion state data including vehicle speed vector and vehicle attitude angle; By performing parallel audio acquisition on D microphones of a stereo microphone array in an electric two-wheeled vehicle, D original speech signals are obtained; The real-time wind direction vector data and the vehicle body motion state data are fused based on a unified timestamp to obtain a multi-source data set.
[0007] In combination with the first aspect, in a second implementation of the first aspect of the present invention, performing coordinate system conversion and vehicle aerodynamic calculation on the multi-source data set to obtain a wind field intensity distribution matrix includes: Performing a rotation transformation from a ground coordinate system to a vehicle coordinate system based on the real-time wind direction vector data in the multi-source data set to obtain transformed wind direction vector data in the vehicle coordinate system; Performing Kalman filtering and state estimation on the vehicle body motion state data in the multi-source data set to obtain vehicle body motion parameters in a vehicle body coordinate system; Performing vector subtraction and vehicle body induced wind field compensation based on the transformed wind direction vector data and the vehicle speed vector in the vehicle body motion parameter to obtain relative wind direction vector data; performing a turning tilt posture correction on the relative wind direction vector data according to the vehicle body posture angle in the vehicle body motion parameter to obtain a corrected wind direction vector in a vehicle body coordinate system; Vehicle body aerodynamics calculation is performed based on the corrected wind direction vector and the vehicle body motion parameters to obtain a wind field intensity distribution matrix.
[0008] In combination with the first aspect, in a third implementation of the first aspect of the present invention, performing the electric two-wheeled vehicle turning tilt posture correction on the relative wind direction vector data based on the vehicle body posture angle in the vehicle body motion parameter to obtain the corrected wind direction vector in the vehicle body coordinate system includes: Performing a threshold judgment on the vehicle body posture angle in the vehicle body motion parameter to obtain a turning tilt state identifier and a corresponding turning tilt angle; Calculating a turning radius and a centrifugal force vector based on the turning tilt angle and the vehicle speed vector in the vehicle body motion parameters to obtain turning dynamics parameters; Perform centrifugal wind field compensation and tilt angle correction on the relative wind direction vector data according to the turning dynamics parameters to obtain turning correction wind direction vector data; The turning correction wind direction vector data and the relative wind direction vector data are selected based on the turning tilt state identifier to obtain a correction wind direction vector in a vehicle body coordinate system.
[0009] In combination with the first aspect, in a fourth implementation of the first aspect of the present invention, performing vehicle aerodynamic calculation based on the corrected wind direction vector and the vehicle body motion parameter to obtain a wind field intensity distribution matrix includes: Calculating the frontal area and drag coefficient of the vehicle body based on the corrected wind direction vector and the vehicle speed vector in the vehicle body motion parameters to obtain vehicle body aerodynamic parameters including frontal drag parameters, side drag parameters, and rear drag parameters; Calculating the directional wind field intensity of the corrected wind direction vector according to the vehicle body aerodynamic parameters to obtain directional wind field intensity data including a frontal wind field intensity component, a lateral wind field intensity component, and a rearward wind field intensity component; Performing wind field interference intensity analysis at microphone positions based on the directional wind field intensity data and the spatial position distribution of the stereo microphone array to obtain three-dimensional wind field interference intensity values corresponding to D microphone positions; The three-dimensional wind field interference intensity values corresponding to the D microphone positions are subjected to matrix conversion to obtain a wind field intensity distribution matrix.
[0010] In combination with the first aspect, in a fifth implementation of the first aspect of the present invention, performing asymmetric differential compensation on the original voice signal based on the wind field intensity distribution matrix to obtain a compensated voice signal includes: Calculating the wind direction compensation coefficient for each microphone based on the three-dimensional wind field interference intensity values corresponding to the D microphone positions in the wind field intensity distribution matrix to obtain a D-way microphone compensation coefficient group including a handlebar lateral compensation coefficient, a front vehicle front compensation coefficient, and a vehicle body rear compensation coefficient; Asymmetrically modulating the D-channel microphone compensation coefficient group based on the turning tilt state identifier to obtain a turning asymmetric compensation coefficient group; The original voice signals of the D channels are subjected to amplitude compensation calculation and phase correction at corresponding positions with the turn asymmetric compensation coefficient group to obtain D channel corrected voice signals, and the D channel corrected voice signals are subjected to differential signal synthesis to obtain a single channel compensated voice signal.
[0011] In combination with the first aspect, in a sixth implementation of the first aspect of the present invention, asymmetrically modulating the D-channel microphone compensation coefficient group based on the turning tilt state identifier to obtain the turning asymmetric compensation coefficient group includes: According to the turning tilt state identifier and the turning tilt angle, determining the inner and outer positions of the stereo microphone array to obtain microphone spatial distribution data; performing asymmetric modulation intensity calculation based on the turning tilt angle to obtain a turning modulation parameter including an inner attenuation modulation coefficient and an outer enhancement modulation coefficient; Position-matching the D-channel microphone compensation coefficient group with the microphone spatial distribution data to obtain an inner microphone compensation coefficient and an outer microphone compensation coefficient; The inner microphone compensation coefficient and the outer microphone compensation coefficient are differentially modulated according to the turning modulation parameter to obtain a turning asymmetric compensation coefficient group.
[0012] In combination with the first aspect, in a seventh implementation of the first aspect of the present invention, performing dynamic wind noise suppression processing and speech recognition on the compensated speech signal to generate a target speech recognition result includes: Performing frequency domain analysis on the compensated voice signal to obtain wind direction interference spectrum feature data including a low-frequency vehicle body vibration interference spectrum, a medium-frequency airflow turbulence interference spectrum, and a medium-high frequency high-speed riding wind noise spectrum; generating dynamic filter configuration data based on the wind direction interference spectrum characteristic data and the vehicle speed vector in the vehicle body motion parameter; constructing an adaptive multi-frequency domain filter according to the dynamic filter configuration data, and inputting the compensated speech signal into the adaptive multi-frequency domain filter to suppress wind noise, thereby obtaining a multi-channel filtered speech signal; Performing beamforming and wind direction compensation confidence calculation on the multi-path filtered speech signal to obtain a single-path enhanced speech signal and a corresponding wind direction compensation confidence; Speech recognition is performed on the single-channel enhanced speech signal according to the wind direction compensation confidence level to generate a target speech recognition result.
[0013] In combination with the first aspect, in an eighth implementation of the first aspect of the present invention, performing speech recognition on the single-channel enhanced speech signal according to the wind direction compensation confidence to generate a target speech recognition result includes: Selecting a corresponding speech recognition mode according to the wind direction compensation confidence level; Based on the speech recognition mode, the time window of the single-channel enhanced speech signal is adjusted and the recognition parameter configuration is configured to obtain speech recognition configuration data; A target speech recognition engine is constructed according to the speech recognition configuration data, and the single-channel enhanced speech signal is input into the target speech recognition engine for speech recognition to generate a target speech recognition result including navigation instructions, communication instructions and safety instructions.
[0014] In a second aspect, the present invention provides a speech recognition control system for an electric two-wheeled vehicle, the speech recognition control system for the electric two-wheeled vehicle comprising: A synchronous data acquisition module is used to synchronously collect data from the electric two-wheeled vehicle's three-axis wind direction sensor, vehicle speed sensor, six-axis attitude sensor, and stereo microphone array to obtain a multi-source data set and original voice signal; A calculation module, configured to perform coordinate system conversion and vehicle aerodynamic calculation on the multi-source data set to obtain a wind field intensity distribution matrix; a compensation module, configured to perform asymmetric differential compensation on the original voice signal based on the wind field intensity distribution matrix to obtain a compensated voice signal; The speech recognition module is used to perform dynamic wind noise suppression processing and speech recognition on the compensated speech signal to generate a target speech recognition result.
[0015] The technical solution provided by this invention integrates a three-axis wind direction sensor, a vehicle speed sensor, a six-axis attitude sensor, and a stereo microphone array to achieve the coordinated operation of multi-dimensional environmental perception and voice acquisition. This provides a comprehensive data foundation for accurate wind direction compensation in open environments for electric two-wheeled vehicles, resolving the technical challenge of a single sensor's inability to accurately describe complex riding environments. A rotational transformation mechanism from the ground coordinate system to the vehicle coordinate system is established, combined with a Kalman filter algorithm to accurately estimate the vehicle's motion state, effectively eliminating sensor noise and vibration interference, ensuring the accuracy and stability of wind direction vector calculations. Based on the unique structural characteristics of electric two-wheeled vehicles, a three-dimensional wind resistance parameter model encompassing the front, side, and rear of the vehicle is established, enabling precise calculation of the wind field distribution at each position in the stereo microphone array, addressing the existing technology's inability to accurately predict wind field distribution in open environments. In response to the fact that the inner and outer microphones are affected differently by wind when turning and leaning, an asymmetric modulation algorithm based on spatial position judgment is developed to achieve differentiated compensation processing for microphones in different positions, effectively addressing the technical deficiency of traditional uniform compensation methods in adapting to the dynamic riding conditions of electric two-wheeled vehicles. A classification and filtering mechanism has been established for low-frequency body vibration, medium-frequency airflow turbulence, and medium-high frequency high-speed riding wind noise of electric two-wheeled vehicles in open environments. By combining Butterworth, elliptic, and Chebyshev filters, precise suppression of wind interference in different frequency bands is achieved, significantly improving the signal-to-noise ratio and clarity of speech signals. Based on the confidence level of wind direction compensation, a three-level adaptive mode of standard recognition, robust enhancement, and conservative recognition is established. By dynamically adjusting the recognition time window and algorithm parameters, the reliability of speech recognition is ensured under different wind direction interference intensities, effectively improving recognition accuracy and response speed.
[0016] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or understood by practicing the present invention. The purposes and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0017] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 Schematic diagram of an embodiment of a voice recognition control method for an electric two-wheeled vehicle according to an embodiment of the present invention; Figure 2 Schematic diagram of a voice recognition control system for an electric two-wheeled vehicle according to an embodiment of the present invention. DETAILED DESCRIPTION
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0020] The terms "including," "having," and any variations thereof, as used in the embodiments of the present invention are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device comprising a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or device.
[0021] To facilitate understanding of this embodiment, a speech recognition control method for an electric two-wheeled vehicle disclosed in an embodiment of the present invention is first described in detail. Figure 1 As shown, this method includes the following steps: 101. Synchronously collect data from the three-axis wind direction sensor, vehicle speed sensor, six-axis attitude sensor, and stereo microphone array of the electric two-wheeled vehicle to obtain a multi-source data set and an original voice signal; It is understood that the execution subject of the present invention can be a voice recognition control system of an electric two-wheeled vehicle, or a terminal or a server, and the specific implementation is not limited here. The embodiment of the present invention is described by taking a server as the execution subject as an example.
[0022] Specifically, a variety of sensor nodes are arranged in the key structural parts of the electric two-wheeled vehicle, among which the three-axis wind direction sensor is installed in front of or on the top of the vehicle head to ensure that it can fully perceive the flow direction and intensity of the natural airflow in front during riding. Its sampling frequency is set at 200Hz to ensure real-time performance, and outputs three component data, namely the longitudinal wind speed along the X-axis direction of the vehicle body, the lateral wind speed along the Y-axis direction and the vertical wind speed along the Z-axis direction. These components are combined to form a three-dimensional wind direction vector to describe the current wind field state; at the same time, the speed sensor installed at the center of the electric two-wheeled vehicle frame obtains the current speed vector of the vehicle through the inertial unit or the wheel speed encoder, and the six-axis attitude sensor arranged near the center of gravity of the vehicle body includes a three-axis accelerometer and a three-axis gyroscope component, which are used to capture the electric two-wheeled vehicle respectively. The acceleration and angular velocity changes of the vehicle body are calculated in real time to calculate the key posture information such as the pitch angle, roll angle and yaw angle, and form a description vector of the vehicle body's motion state; in terms of audio collection, a stereo microphone array is formed by arranging D MEMS microphones on both sides of the handlebars, both ends of the front and the rear of the vehicle body. All microphones perform parallel audio sampling with a unified synchronous clock. These sampling channels cover the typical sound source propagation paths around electric two-wheelers and can retain as much spatial voice information as possible in a complex open environment; the above-mentioned collected data streams are subjected to unified timestamp binding processing, and the wind direction vector, vehicle speed vector, posture angle and microphone voice signal at each moment are aligned and fused according to a unified time axis to construct a multi-source perception data set at that moment.
[0023] 102. Perform coordinate system conversion and vehicle aerodynamic calculation on the multi-source data set to obtain a wind field intensity distribution matrix; Specifically, the collected wind direction vector data and vehicle motion state data are processed using a unified time series benchmark. The real-time wind direction vector data is initially obtained as a three-dimensional velocity vector based on the ground coordinate system. To align it with the vehicle's structural and motion models, a rotational transformation from the ground coordinate system to the vehicle coordinate system is performed based on the vehicle's current attitude angle. This transformation uses Euler angles or quaternion methods to map the wind speed data, including X, Y, and Z components, to an isomorphic vector in vehicle coordinates, resulting in the transformed wind direction vector data. This redefines the wind directionality relative to the vehicle's orientation. Kalman filtering is used to smooth and estimate the vehicle's velocity and attitude state data. The filter uses a recursive method to eliminate vehicle vibration, sensor jitter, and sampling noise, ultimately generating stable vehicle velocity and attitude angle curves. A vector subtraction operation is performed between the wind direction vector in the vehicle coordinate system and the estimated vehicle speed vector to calculate the wind flow direction and intensity relative to the moving vehicle, yielding the relative wind direction vector. To better reflect real-world conditions, the system considers the effects of induced wind fields caused by vehicle acceleration or turning, such as oncoming air compression caused by the vehicle's forward motion and wake disturbances induced by the vehicle's shape. These additional wind fields are then compensated for in the relative wind direction through aerodynamic modeling, resulting in more realistic flow field data. Attitude changes can cause the relative relationship between the microphone position and the airflow direction to shift, affecting the accuracy of wind noise perception. By combining the vehicle's pitch and roll angles, the relative wind direction vector data is corrected for turning and tilt, resulting in a corrected wind direction vector in the vehicle coordinate system. Vehicle aerodynamic calculations are performed based on the corrected wind direction vector and vehicle motion parameters. Combined with vehicle structural parameters (such as frontal area, drag coefficient, and relative microphone positions), the system models the airflow propagation characteristics at each microphone position to derive the three-dimensional wind intensity distribution at the stereo microphone array. This results in a wind intensity distribution matrix that quantitatively characterizes the wind field interference experienced by each microphone in the X, Y, and Z directions.
[0024] 103. Perform asymmetric differential compensation on the original voice signal based on the wind field intensity distribution matrix to obtain a compensated voice signal; Specifically, the wind direction interference intensity value corresponding to each microphone in its three-dimensional spatial position is extracted from the wind field intensity distribution matrix. The distribution matrix is a matrix with a dimension of D×3, where D represents the total number of microphones (such as the 8 MEMS microphones configured in the system), and the three components contained in each row represent the wind pressure disturbance components felt by the microphone in the X-axis (longitudinal), Y-axis (transverse) and Z-axis (vertical) directions in the vehicle coordinate system. The corresponding wind direction compensation coefficient is calculated based on the intensity of these components. The microphones on both sides of the handlebars are mainly affected by the Y-axis transverse airflow disturbance. The compensation coefficient for the microphone at the front of the vehicle faces the main direction of the airflow, so it is determined by the weighted product of the square of the X-axis wind speed and the forward wind resistance area. The microphone at the rear of the vehicle is located in the area affected by the superposition of wake disturbance and turbulence. In addition to the reference X-axis wind speed, the compensation factor also includes a weighted term for the eddy current interference factor. The compensation factors at these different locations are assembled into a D-dimensional microphone compensation coefficient group. This array specifies the wind direction amplitude correction ratio that should be applied to each microphone at its physical location. Based on this, the vehicle's current attitude state, or turning tilt indicator, is introduced. If the detected roll angle exceeds a certain threshold (e.g., 5 or 8 degrees), it indicates that the vehicle is in a lateral tilt or high-speed turning state. At this point, the wind pressure disturbances sensed by the left and right microphones will be significantly asymmetric. Therefore, the original compensation coefficient set is asymmetrically modulated. Specifically, the compensation coefficients for the microphones on the inside of the turn are proportionally decreased (e.g., by 15% to 20%), while the compensation coefficients for the microphones on the outside are proportionally increased (e.g., by 10% to 15%). This creates a set of turning asymmetric compensation coefficients that reflects the actual airflow disturbances. Amplitude compensation is performed on the original speech signal collected by each microphone and the turning asymmetric compensation coefficient at the corresponding microphone position. This involves weighting the signal amplitude by the coefficient. A phase correction term is introduced in the frequency or time domain to compensate for differences in speech signal propagation delay caused by wind direction interference or path length differences, ensuring temporal consistency across all channel signals. After completing these operations, a D-channel corrected speech signal sequence is obtained. The D-channel corrected speech signals are differentially synthesized. By constructing a beamforming model or weighted summation method, the speech information of multiple spatial channels is fused on the basis of filtering out environmental noise and wind interference, and a single-channel compensated speech signal is output.
[0025] 104. Perform dynamic wind noise suppression processing and speech recognition on the compensated speech signal to generate a target speech recognition result.
[0026] Specifically, the compensated speech signal is input into the frequency domain analysis module, which extracts the power spectral density of the speech signal in various frequency bands through fast Fourier transform (FFT) or wavelet transform (WT), thereby identifying the spectral characteristics of multiple types of wind interference contained therein. The low-frequency band, primarily concentrated between 20Hz and 200Hz, represents structural vibration interference caused by road impact or mechanical resonance during the operation of the electric two-wheeled vehicle. The mid-frequency band, approximately between 200Hz and 800Hz, corresponds to unsteady turbulence interference caused by the coupling of disturbances between the front wind and the vehicle geometry. Finally, the mid- and high-frequency bands, concentrated between 800Hz and 2000Hz, represent wind noise components generated by shear flow, boundary layer shedding, and strong disturbances in the wake region caused by high-speed riding. Based on the wind interference spectral characteristics data and the vehicle speed vector from the vehicle motion parameters, dynamic filter configuration data is generated. This configuration data automatically sets core filter parameters such as the filter order, bandwidth, stopband attenuation, and passband pass rate based on factors such as vehicle speed, wind direction angle variation, and attitude stability. This configuration data is used to construct an adaptive multi-frequency domain filter bank. This filter bank includes a 4th-order Butterworth filter for the low-frequency band, an 8th-order elliptic filter for the mid-frequency band, and a 6th-order Chebyshev filter for the mid- and high-frequency bands. These filters are designed to effectively suppress wind noise interference in specific frequency bands while preserving the effective information components of the speech spectrum as much as possible. Each compensated speech signal is then input into the filter bank for real-time filtering, resulting in multiple wind-noise-suppressed filtered speech signals. Beamforming is then performed on the multi-channel filtered speech signals. Using a weighted aggregation algorithm based on the spatial signal source direction (such as the minimum variance distortionless beamformer (MVDR)), the information from multiple spatial channels is fused into a high-quality enhanced speech signal. During this process, the wind direction compensation confidence level is calculated based on the amplitude consistency and delay matching between the channels. This confidence level is used to assess the overall reliability and recognition stability of the current speech signal in a wind noise interference environment. According to the wind direction compensation confidence, speech recognition operations are performed on the enhanced speech signal. If the confidence is higher than the set threshold (such as 0.8), the standard recognition algorithm is used to directly parse the speech content. If the confidence is in the middle range (such as 0.6 to 0.8), the robust recognition mode is started, the speech frame window length is expanded, and the recognition model's adaptability to weak and ambiguous signals is enhanced. If the confidence is lower than 0.6, the conservative recognition mode is used, and only high-confidence command words are parsed to avoid the risk of misrecognition, thereby ultimately outputting the target speech recognition result.
[0027] In a specific embodiment, the process of executing step 101 may specifically include the following steps: The three-axis wind direction sensor arranged on the electric two-wheeled vehicle is sampled to obtain real-time wind direction vector data including the X-axis longitudinal wind speed component, the Y-axis transverse wind speed component and the Z-axis vertical wind speed component; The vehicle speed sensor and six-axis attitude sensor installed on the electric two-wheeled vehicle are synchronously sampled to obtain vehicle motion state data including vehicle speed vector and vehicle attitude angle; By performing parallel audio acquisition on D microphones of a stereo microphone array in an electric two-wheeled vehicle, D original speech signals are obtained; The real-time wind direction vector data and the vehicle body motion state data are fused based on a unified timestamp to obtain a multi-source data set.
[0028] Specifically, the sensor system is arranged and the synchronization mechanism is designed on the structure of the electric two-wheeled vehicle, wherein the three-axis wind direction sensor is installed in an open position at the front end of the vehicle, such as the top of the dashboard or above the headlights, to ensure that it can sense the direction and speed of the natural airflow without obstruction when driving. The wind direction sensor samples at a high frequency of 200Hz and outputs three wind speed components, namely the longitudinal wind speed component along the X-axis direction represents the airflow intensity facing the driving direction, the transverse wind speed component in the Y-axis direction represents the degree to which the side of the vehicle is affected by the wind, and the vertical wind speed component in the Z-axis direction is used to characterize the airflow fluctuations caused by terrain undulations or turbulent disturbances. The wind direction vector composed of these three components expresses the three-dimensional wind environment of the vehicle at a certain moment; and the vehicle speed sensor is arranged on the central axis of the vehicle, and the current speed vector is derived through the wheel speed signal or inertial speed integral. The vector provides real-time speed information of the vehicle in the forward direction for subsequent use. Wind speed compensation and aerodynamic model calculation; the six-axis attitude sensor is installed near the center of gravity of the vehicle body to ensure the stability and overall representativeness of attitude perception. It contains a three-axis accelerometer and a three-axis gyroscope, which measure changes in acceleration and angular velocity along three spatial directions respectively. Through integration and coordinate transformation, the vehicle's attitude angle parameters, including pitch, roll, and yaw, are calculated in real time. These parameters are the basis for determining whether the vehicle is in non-uniform speed states such as turning, tilting, and climbing. At the same time, a total of D MEMS microphones are installed on both sides of the vehicle structure, in the front and rear, in a three-dimensional spatially balanced layout, with a typical number of 8. Each microphone collects audio signals in parallel through an independent channel at a high sampling rate, acquiring voice signals from different spatial locations during riding, ensuring that the system has sufficient spatial redundancy for subsequent signal enhancement and differential filtering processing. To ensure the temporal alignment of the collected wind direction data, vehicle speed and attitude data, and D-channel microphone voice data, a unified sampling timestamp management mechanism was established. This mechanism uses a high-precision clock source as the primary reference. All sensor modules initiate sampling via a unified trigger signal and append the current timestamp information to the data, avoiding data dislocation or mismatch caused by asynchronous sampling. Time synchronization accuracy is controlled within 1 millisecond. After data acquisition is completed, the system uses the timestamp to fuse the wind direction vector data, vehicle speed and attitude parameters, and voice signals. The specific process includes spatial transformation of the wind direction vector to align it with the vehicle coordinate system, state estimation and filtering of the attitude angle and velocity vector to obtain a continuous motion state stream, and preliminary amplitude and phase equalization of the microphone signals to match the current wind flow disturbance conditions. The final result is a fused multi-source data set.
[0029] In a specific embodiment, the process of executing step 102 may specifically include the following steps: Performing a rotation transformation from a ground coordinate system to a vehicle coordinate system based on the real-time wind direction vector data in the multi-source data set to obtain transformed wind direction vector data in the vehicle coordinate system; Performing Kalman filtering and state estimation on the vehicle body motion state data in the multi-source data set to obtain vehicle body motion parameters in a vehicle body coordinate system; Performing vector subtraction and vehicle body induced wind field compensation based on the transformed wind direction vector data and the vehicle speed vector in the vehicle body motion parameter to obtain relative wind direction vector data; performing a turning tilt posture correction on the relative wind direction vector data according to the vehicle body posture angle in the vehicle body motion parameter to obtain a corrected wind direction vector in a vehicle body coordinate system; Vehicle body aerodynamics calculation is performed based on the corrected wind direction vector and the vehicle body motion parameters to obtain a wind field intensity distribution matrix.
[0030] Specifically, the natural wind speed data collected by the three-axis wind direction sensor is converted from the ground stationary reference system to a dynamic reference system centered on the vehicle body. This coordinate transformation process depends on the current attitude angle information of the vehicle body, and is specifically implemented through a rotation matrix described by Euler angles or quaternions. That is, the pitch angle, roll angle, and yaw angle at the current moment are used as input parameters to construct a direction cosine matrix from the ground coordinate system to the vehicle coordinate system. The original wind speed vector is then multiplied by the matrix to obtain the transformed wind direction vector represented in the vehicle coordinate system. This coordinate transformation eliminates the difference between the ground wind speed and the geometric orientation of the vehicle body, providing a unified directional basis for wind field distribution calculations. To enhance the stability and timing consistency of motion parameters, a Kalman filter operation is performed on the data from the vehicle speed sensor and six-axis attitude sensor. The filter adopts a recursive update structure to jointly optimize the historical state and the current observation value, thereby achieving dynamic suppression of noise interference, short-term jitter and mechanical vibration. The output motion parameters include a smooth vehicle speed vector and a high-precision estimate of the vehicle attitude angle. The system then uniformly projects these filtered state quantities into the vehicle coordinate system to form a complete set of motion parameters including the speed vector and attitude angle. A vector subtraction operation is performed based on the transformed wind direction vector data and the vehicle speed vector in the vehicle body motion parameters to solve the actual wind direction vector perceived by the vehicle body relative to the airflow during its own motion, revealing the relative speed relationship between the airflow and the vehicle. In order to enhance the model's accurate modeling of the wind field effect induced by vehicle body motion, a vehicle-induced wind field compensation is introduced. That is, the additional disturbance component caused by air compression and wake rewind during high-speed vehicle movement, sudden acceleration or rapid deceleration is estimated based on the vehicle speed change rate, vehicle body geometry and windward area. This induced wind field compensation term is added to the aforementioned vector subtraction operation to obtain the relative wind direction vector. Linear motion parameters alone cannot fully describe the vehicle's perception of airflow in space. Therefore, the system corrects the relative wind direction vector based on the vehicle's posture angle. This is especially true in complex posture scenarios such as turning, leaning, or climbing, where the vehicle's posture significantly affects the angular relationship between each microphone and the airflow direction. Therefore, the system uses the current roll and pitch angles as the primary modulation factors to perform posture mapping and projection transformation on the relative wind direction, eliminating vector offsets caused by the vehicle's spatial rotation and ultimately obtaining the corrected wind direction vector in the vehicle's coordinate system. Vehicle aerodynamics calculations are performed based on the corrected wind direction vector and vehicle motion parameters, and the corrected wind direction is mapped to the physical position of each microphone to obtain a quantitative description of the degree of wind field interference at each point in space.A geometric model is established based on the vehicle structure, including parameters such as the front windward area, the side area of the vehicle body, the volume of the rear recirculation area, and the relative installation positions of each microphone. The wind pressure coefficient, drag coefficient, and shear force distribution are calculated based on the basic principles of aerodynamics. In combination with the corrected wind direction vector and vehicle speed state, the streamline distribution and velocity disturbance characteristics of the wind around the vehicle body are simulated, and further decomposed into three-dimensional component intensities at each microphone to construct a wind field intensity distribution matrix with a dimension of D×3, where D represents the number of microphones and each row contains the wind speed disturbance intensity felt by the microphone in the X, Y, and Z directions.
[0031] In a specific embodiment, the step of performing, based on the body posture angle in the body motion parameter, a correction of the turning tilt posture of the electric two-wheeled vehicle on the relative wind direction vector data to obtain the corrected wind direction vector in the vehicle body coordinate system may specifically include the following steps: Performing a threshold judgment on the vehicle body posture angle in the vehicle body motion parameter to obtain a turning tilt state identifier and a corresponding turning tilt angle; Calculating a turning radius and a centrifugal force vector based on the turning tilt angle and the vehicle speed vector in the vehicle body motion parameters to obtain turning dynamics parameters; Perform centrifugal wind field compensation and tilt angle correction on the relative wind direction vector data according to the turning dynamics parameters to obtain turning correction wind direction vector data; The turning correction wind direction vector data and the relative wind direction vector data are selected based on the turning tilt state identifier to obtain a correction wind direction vector in a vehicle body coordinate system.
[0032] Specifically, the posture angle information is extracted from the vehicle body motion parameters. This information is expressed in the form of Euler angles, including three dimensions: pitch, roll, and yaw. The roll angle is the key indicator for determining whether the vehicle is in a turning state. Because in actual riding, when an electric two-wheeled vehicle enters a turning action, due to the dynamic balance mechanism between centrifugal force and gravity, the vehicle body will inevitably tilt inward, and the tilt angle is directly reflected by the roll angle. Therefore, the system sets an empirical threshold, such as 5 degrees or 8 degrees as the critical value, and determines whether the current vehicle is in a turning state by continuously judging whether the roll angle exceeds the threshold. If the value is greater than the threshold, it is marked as "turning", otherwise it is "non-turning". At the same time, the roll angle value itself is also retained as the "turning tilt angle". After identifying the turning state, the turning radius is estimated in combination with the vehicle speed vector. According to the principles of kinematics, a nonlinear relationship based on gravity acceleration and vehicle speed is established between the turning radius and the roll angle in steady-state turning. After obtaining the turning radius, the centripetal acceleration formula is used to calculate the direction of the centrifugal acceleration vector acting on the vehicle body, and the three-dimensional centrifugal force vector is constructed in combination with the vehicle orientation. This vector is a dynamic parameter that describes how the vehicle generates additional disturbances relative to the air flow field during the turning process. The centrifugal force vector is superimposed on the aforementioned relative wind direction vector to construct a wind field disturbance model corrected by the centrifugal force field. The centrifugal force is directed outward along the vehicle's transverse direction, resulting in an equivalent "pushing" effect on the airflow direction, causing the perceived wind direction to shift angularly from the vehicle's reference frame. When constructing the turning correction wind direction vector, the system performs vector addition on the centrifugal wind field vector and the original relative wind direction. This resultant wind direction is then rotated and projected based on the turning bank angle, completing the full compensation correction process for the wind direction vector in the turning state. Based on the result of the turning bank state identification, the system determines which set of wind direction vectors should be used as the final input. If the current state is identified as turning, the turning correction wind direction vector, which has been corrected for centrifugal force and bank angle, is used. Otherwise, the original, uncompensated relative wind direction vector is used, ensuring that the system uses the most reasonable airflow modeling data under different dynamic conditions.
[0033] In a specific embodiment, the step of performing vehicle aerodynamic calculation based on the corrected wind direction vector and the vehicle body motion parameters to obtain a wind field intensity distribution matrix may specifically include the following steps: Calculating the frontal area and drag coefficient of the vehicle body based on the corrected wind direction vector and the vehicle speed vector in the vehicle body motion parameters to obtain vehicle body aerodynamic parameters including frontal drag parameters, side drag parameters, and rear drag parameters; Calculating the directional wind field intensity of the corrected wind direction vector according to the vehicle body aerodynamic parameters to obtain directional wind field intensity data including a frontal wind field intensity component, a lateral wind field intensity component, and a rearward wind field intensity component; Performing wind field interference intensity analysis at microphone positions based on the directional wind field intensity data and the spatial position distribution of the stereo microphone array to obtain three-dimensional wind field interference intensity values corresponding to D microphone positions; The three-dimensional wind field interference intensity values corresponding to the D microphone positions are subjected to matrix conversion to obtain a wind field intensity distribution matrix.
[0034] Specifically, the directional relationship between the corrected wind direction vector and the vehicle speed vector in the vehicle coordinate system is clarified, and this is used as a basis for estimating the windward area of the vehicle in different directions and modeling air resistance. The corrected wind direction vector has already been spatially aligned in the previous step through centrifugal compensation and attitude rotation during cornering. Therefore, in the current stage, it is used together with the vehicle speed vector as a projection reference to dynamically estimate the projected area of the three main windward surfaces of the vehicle: the front, side, and rear. The system normalizes the corrected wind direction vector and then calculates the inner product with the three direction vectors in the vehicle coordinate system to determine the projection intensity of the wind direction along each axis. These projection quantities are applied as dynamic weights to the static geometric area of each structural area. For example, the frontal area of a vehicle's front is between 0.6 and 0.8 square meters. However, as the angle between the front and the wind changes during riding, the effective area changes. The system multiplies this area by the projection of the wind direction vector on the X-axis to obtain the real-time frontal area of the vehicle. Similarly, the projection on the Y-axis is multiplied by the side area (1.2 to 1.5 square meters) to estimate the intensity of the wind on the side of the vehicle, while the change on the Z-axis is used to correct for the rearward disturbance area caused by the wind swirling around the rear. Given a known air density constant, the system uses the classic air resistance formula to calculate the corresponding frontal, side, and rear drag components. The drag coefficient is dynamically adjusted based on the vehicle's shape and airflow direction. For example, the frontal drag coefficient ranges from 0.3 to 0.45, while the lateral drag coefficient can reach as high as 0.6 due to the edge flow effect. At the rear, the drag coefficient ranges from 0.4 to 0.5 due to backflow and low-pressure areas. The frontal, side, and rearward drag parameters constructed in this way together constitute the vehicle's aerodynamic parameter set. Based on the above aerodynamic parameters, the corrected wind direction vector is spatially decomposed into directional wind intensity. The components of the corrected wind direction vector along the X (forward), Y (lateral), and Z (up and down) directions are extracted and multiplied by the corresponding wind resistance parameters respectively to obtain the wind field intensity components in the three directions. These components represent the air kinetic energy transfer rate caused by airflow disturbances on the vehicle structure in different directions, which can be equivalently understood as the aerodynamic impact intensity acting on different structural areas. Among them, the frontal wind field intensity component mainly affects the front of the vehicle and the front microphone. The lateral wind field intensity has a significant impact on the handlebars and the microphones on both sides. The rear wind field mainly reflects the interference pressure of the tail swirling airflow on the rear microphone. The system records these three types of directional wind field intensity data in a unified manner as a directional wind field intensity dataset.After obtaining the directional wind field intensity, interference analysis is performed in combination with the spatial layout of the stereo microphone array. Since the D microphones are distributed in different positions on the vehicle body, such as the left and right ends of the front, both sides of the handlebars, the rider's ears, and both sides of the rear, each microphone perceives a different disturbance intensity in the wind field. Based on the physical coordinates and spatial orientation information of each microphone, the system performs weighted projection processing on the degree of its impact on the front, side, and rear wind fields. For example, the front microphone mainly receives the frontal airflow in the X-axis direction, and its interference intensity will be dominated by the frontal wind field component; the wind direction interference of the microphones on both sides of the handlebars is most related to the Y-axis component. In addition, an additional compensation term is introduced due to the asymmetric disturbance caused by centrifugal force. For the rear microphone, the system considers the enhanced Z-axis disturbance caused by factors such as wake diffraction and turbulent interference, and therefore introduces an additional disturbance factor for wind pressure expansion modeling. The wind speed disturbance intensity values sensed by each microphone position in three directions are organized and unified to construct a D×3 interference intensity matrix, where D is the number of microphones and each row contains the wind field disturbance vector received by that microphone in the X, Y, and Z directions, forming a wind field intensity dataset. A matrix transformation operation is performed on this interference intensity matrix, normalizing the D×3 raw interference data column-wise and amplitude-normalizing them according to the spatial distribution of wind directions. While maintaining the true physical magnitude of the interference amplitude in each dimension, this ensures that each element in the matrix corresponds to the quantized intensity of the aerodynamic disturbance experienced in the actual acoustic channel. The final output is a wind field intensity distribution matrix.
[0035] In a specific embodiment, the process of executing step 103 may specifically include the following steps: Calculating the wind direction compensation coefficient for each microphone based on the three-dimensional wind field interference intensity values corresponding to the D microphone positions in the wind field intensity distribution matrix to obtain a D-way microphone compensation coefficient group including a handlebar lateral compensation coefficient, a front vehicle front compensation coefficient, and a vehicle body rear compensation coefficient; Asymmetrically modulating the D-channel microphone compensation coefficient group based on the turning tilt state identifier to obtain a turning asymmetric compensation coefficient group; The original voice signals of the D channels are subjected to amplitude compensation calculation and phase correction at corresponding positions with the turn asymmetric compensation coefficient group to obtain D channel corrected voice signals, and the D channel corrected voice signals are subjected to differential signal synthesis to obtain a single channel compensated voice signal.
[0036] Specifically, the system uses the wind field intensity distribution matrix as its core input. This matrix, derived from previous aerodynamic modeling and wind direction correction calculations, is structured as a D×3 matrix, where D is the number of microphones configured in the system. Each row represents the wind disturbance intensity perceived by a microphone in the current environment. The three dimensions correspond to the disturbance components along the X-axis (frontal wind direction), Y-axis (lateral wind direction), and Z-axis (perpendicular wind flow) in the vehicle coordinate system. The system performs a modulus calculation on the three-dimensional wind velocity vector of each microphone to obtain the total wind energy disturbance. It then independently weights the components in each direction to establish a corresponding wind direction sensitivity model. Based on this, the system classifies the microphones into three types based on their physical mounting location and directivity: "handlebar-mounted side microphones," which are primarily affected by side wind interference; "front-mounted front microphones," which are mounted on either side of the vehicle and primarily face forward and face the frontal airflow; and "rear-mounted rear microphones," which are located in the low-pressure wake region and are susceptible to vortex and backflow disturbances. The system calculates the handlebar lateral compensation coefficient by extracting and squaring the Y-axis components corresponding to each microphone in the wind field distribution matrix, combined with the drag coefficient corresponding to the handlebar microphone. The frontal compensation coefficient is calculated by extracting the X-axis components and combining them with the forward airflow drag coefficient. The rearward compensation coefficient is constructed by synthesizing the X- and Z-axis perturbations and multiplying them by the tail airflow turbulence factor. These coefficients describe the intensity of the wind pressure disturbance experienced by each microphone at its current physical location and are constructed as a complete set of D-channel microphone compensation coefficients in the form of linear amplification factors or gain control parameters. The system asymmetrically modulates this compensation coefficient set based on the turning tilt state indicator output by the turning state determination module. This process accounts for the spatial attitude deviation of the electric two-wheeled vehicle caused by the increased roll angle during high-speed cornering, particularly the disruption of wind flow symmetry around the microphones on both sides of the handlebars. Once the system detects that the roll angle exceeds a preset threshold (such as 5 or 8 degrees), it identifies the current state as a turn and uses this roll angle as a basis for adjusting the asymmetric modulation intensity. The system proportionally decreases the original compensation coefficient of the inner handlebar microphone by 15% to 20%, while increasing the compensation coefficient of the outer handlebar microphone by 10% to 15%. This differential modulation ensures that the inner and outer microphones' wind sensitivity is dynamically balanced in turning scenarios dominated by centrifugal force, maintaining phase consistency and amplitude symmetry during signal superposition. Simultaneously, small adjustments are made to the compensation coefficients of the front and rear microphones based on the vehicle's yaw angle change rate and the mean Z-axis disturbance to mitigate wind noise intensity deviations caused by streamline shifts caused by turning. This step generates a set of turning asymmetric compensation coefficients. The original D-channel voice signal is then amplitude-compensated and phase-corrected at the corresponding locations using the turning asymmetric compensation coefficient set to produce a corrected D-channel voice signal.Amplitude compensation involves adjusting the gain of each microphone's sampled signal according to its corresponding compensation coefficient, performing a linear multiplication operation to reduce energy distortion caused by wind pressure disturbances. Phase correction involves extracting the signal phase spectrum through a short-time Fourier transform based on the degree of wind field disturbance on the signal propagation path and the multipath propagation effects caused by the vehicle structure. The phase delay adjustment required for each signal is then calculated based on a phase offset model to achieve time alignment of the multi-channel signals and address the phase misalignment caused by asymmetric wind disturbances. Differential signal synthesis is performed on the D-channel corrected speech signals using weighted beamforming or minimum mean square difference aggregation algorithms. This utilizes the spatial differential information between the microphones to extract wind noise distribution characteristics and enhance the coherent components of the true speech signal across all channels, while suppressing random noise and directional interference. The result is a single compensated speech signal with higher clarity and lower background interference.
[0037] In a specific embodiment, the step of performing asymmetrical modulation on the D-channel microphone compensation coefficient group based on the turning tilt state identifier to obtain the turning asymmetric compensation coefficient group may specifically include the following steps: According to the turning tilt state identifier and the turning tilt angle, determining the inner and outer positions of the stereo microphone array to obtain microphone spatial distribution data; performing asymmetric modulation intensity calculation based on the turning tilt angle to obtain a turning modulation parameter including an inner attenuation modulation coefficient and an outer enhancement modulation coefficient; Position-matching the D-channel microphone compensation coefficient group with the microphone spatial distribution data to obtain an inner microphone compensation coefficient and an outer microphone compensation coefficient; The inner microphone compensation coefficient and the outer microphone compensation coefficient are differentially modulated according to the turning modulation parameter to obtain a turning asymmetric compensation coefficient group.
[0038] Specifically, the roll angle in the body motion parameters is used as the core basis to determine whether the vehicle is currently in a turning state and obtain its tilt direction and tilt angle. In the actual operation of the electric two-wheeled vehicle, when the vehicle enters a turn, due to the balance between inertia and gravity, the body will tilt inward. This change is directly reflected in the offset of the roll angle value. When the system determines whether the angle exceeds the set threshold (for example, 5°), it generates a "turning tilt state identifier" and uses the absolute value of the roll angle as the "turning tilt angle". Further judging the positive and negative sign of the roll angle can determine whether the current turn is left or right, that is, a positive value represents a tilt to the right, and a negative value represents a tilt to the left. This direction information helps to determine which microphones in the stereo microphone array belong to the inside and outside of the turn. After obtaining the state identifier and tilt direction, the system determines the sign of each microphone's lateral coordinate (Y-axis) based on the known physical microphone mounting structure and coordinate information, combined with the vehicle coordinate system. When the vehicle is tilted to the right (positive roll angle), all microphones located on the negative Y-axis (i.e., the left side of the vehicle) are classified as inboard, while microphones on the positive Y-axis are classified as outboard. Otherwise, the positions are reversed. This generates microphone spatial distribution data, which identifies the spatial location of each of the D microphones as a Boolean vector or index. After completing the spatial division, asymmetric modulation intensity is calculated based on the roll angle. This process dynamically adjusts the modulation level as the tilt angle increases. Therefore, an adjustable function or interpolation lookup table mechanism is established to map the tilt angle to two independent modulation coefficients: an "inboard attenuation modulation coefficient" and an "outboard boost modulation coefficient." The inboard attenuation modulation coefficient ranges from 0.8 to 0.85, indicating the degree of suppression of the original compensation value, while the outboard boost modulation coefficient ranges from 1.1 to 1.2, indicating the degree of amplification of the original value. For example, when the turning angle is 8°, the inner attenuation modulation coefficient is set to 0.82, while the outer boost coefficient is 1.18. If the angle is reduced to 5°, the inner and outer attenuation modulation coefficients are set to 0.9 and 1.1, respectively. This proportional control mechanism ensures that the dynamic symmetry between wind pressure and signal response is maintained during turns and is adjusted step-by-step, thus avoiding sudden miscompensation. The D-channel microphone compensation coefficient set is matched one-to-one with the microphone spatial distribution data. For each microphone, whether it is located inside or outside is indicated by the spatial distribution data, while its original compensation value is provided by the compensation coefficient set. The system groups all compensation values by position label, dividing them into an "inner microphone compensation coefficient set" and an "outer microphone compensation coefficient set." After completing this matching operation, different modulation parameters are applied to these two sets. The compensation values of the inner microphones are all multiplied by the inner attenuation modulation coefficient to achieve amplitude compression, while the compensation values of the outer microphones are all multiplied by the outer boost modulation coefficient to achieve relative enhancement. Finally, all modulated values are combined into a new D-dimensional vector, namely the turning asymmetric compensation coefficient set.
[0039] In a specific embodiment, the process of executing step 104 may specifically include the following steps: Performing frequency domain analysis on the compensated voice signal to obtain wind direction interference spectrum feature data including a low-frequency vehicle body vibration interference spectrum, a medium-frequency airflow turbulence interference spectrum, and a medium-high frequency high-speed riding wind noise spectrum; generating dynamic filter configuration data based on the wind direction interference spectrum characteristic data and the vehicle speed vector in the vehicle body motion parameter; constructing an adaptive multi-frequency domain filter according to the dynamic filter configuration data, and inputting the compensated speech signal into the adaptive multi-frequency domain filter to suppress wind noise, thereby obtaining a multi-channel filtered speech signal; Performing beamforming and wind direction compensation confidence calculation on the multi-path filtered speech signal to obtain a single-path enhanced speech signal and a corresponding wind direction compensation confidence; Speech recognition is performed on the single-channel enhanced speech signal according to the wind direction compensation confidence level to generate a target speech recognition result.
[0040] Specifically, high-resolution frequency domain analysis is performed on the compensated speech signal. Fast Fourier transform (FFT) is used to perform short-time window sliding analysis on each speech signal, converting the original time domain signal into a complex frequency domain feature matrix containing both amplitude and phase spectra. Within this spectrum, the system divides the spectrum into three main interference source bands based on empirical frequency bandwidth. The energy peaks in the 20Hz to 200Hz range are closely related to low-frequency structural vibrations of the vehicle body. These interference sources arise from physical processes such as tire rolling, chassis impact, and periodic motor vibration. The 200Hz to 800Hz band is dominated by air turbulence, exhibiting non-stationary, rapidly varying amplitudes, and discontinuous mid-frequency interference characteristics. These interference arise from flow around the vehicle's edges, mirror reflections, and front wind crosstalk. Finally, the 800Hz to 2000Hz and even higher frequency bands primarily reflect the sharp wind noise caused by high-speed relative wind pressure, geometric cavity resonance of the vehicle body, and helmet reflections during riding. Although this noise has a dispersed energy distribution, it can easily suppress speech clarity. After extracting frequency domain features, the system combines the spectral content with the vehicle's motion state, combining it with the vehicle speed vector from the vehicle's motion parameters. This involves determining the changing trend of wind noise structure based on the current speed and rate of acceleration. For example, when vehicle speeds are below 20 km / h, low-frequency structure-borne noise predominates. However, above 30 km / h, mid- and high-frequency wind noise becomes dominant. Therefore, the system dynamically generates filter configuration data consisting of multiple parameters, including filter type (e.g., Butterworth, elliptic, Chebyshev), order, passband range, stopband attenuation, cutoff frequency, and window function type. These parameters vary nonlinearly across different speed ranges. The filter parameter set generated through a table lookup mechanism or dynamic interpolation ensures a good balance between filter structure, amplitude modulation range, stability, and speech fidelity. The system then uses this configuration data to construct an adaptive multi-domain filter bank consisting of three independent filter channels, one for processing wind direction interference signals in the low, mid, and mid-high frequency bands, respectively. A 4th-order Butterworth band-stop filter is used in the low-frequency band to ensure phase linearity and attenuate structural vibration components. An 8th-order elliptic filter is used in the mid-frequency band to balance a steep stopband with minimal passband distortion, making it suitable for turbulent airflow scenarios. A 6th-order Chebyshev filter is used for mid- and high-frequency wind noise, suppressing interference from sharp high-frequency components on the high-frequency portion of speech through high-stopband attenuation. The system feeds each compensated speech signal into these filter channels and dynamically adjusts their rate based on vehicle speed. The filters respond in real time to changes in spectral disturbance characteristics caused by increasing or decreasing vehicle speed, outputting multiple filtered speech signals with suppressed wind noise.Beamforming is performed on multi-channel filtered speech signals. By weighting the spatial time and amplitude differences of the filtered signals from multiple microphone channels, an enhanced speech signal with directional enhancement is constructed. This process uses a minimum variance distortion-free response algorithm or a coherent weighted superposition method based on expected response optimization. Simultaneously, the system evaluates wind direction compensation confidence in real time during the synthesis process. This is achieved by calculating consistency metrics (such as cross-correlation coefficients and covariance consistency) between the multi-channel signals and combining factors such as wind direction modeling error, filter suppression residual, and vehicle speed fluctuation amplitude to construct a comprehensive confidence score. This confidence score ranges from 0 to 1. A value above 0.8 indicates sufficient wind direction modeling and ideal filter suppression, allowing the speech to directly enter the standard recognition process. If the confidence score is between 0.6 and 0.8, the input time window of the speech recognition model is expanded, and a robust endpoint detection and noise adaptation mechanism is adopted. If the confidence score is below 0.6, the system activates a conservative mode, allowing only high-priority, high-definition command words to be recognized, and reliability is improved through a secondary confirmation interaction. Under the confidence guidance mechanism, the single-channel enhanced voice signal is input into the voice recognition engine to execute the extraction and recognition of the target voice command. During the entire process, the system continuously cross-validates the confidence and recognition results. If the recognition output is seriously inconsistent with the current vehicle status, the recognition result will be rejected or the user will be prompted for confirmation, thereby ensuring that the final voice recognition result is still highly accurate, stable and controllable in open wind fields and high wind noise environments.
[0041] In a specific embodiment, the step of performing speech recognition on the single-channel enhanced speech signal according to the wind direction compensation confidence level to generate a target speech recognition result may specifically include the following steps: Selecting a corresponding speech recognition mode according to the wind direction compensation confidence level; Based on the speech recognition mode, the time window of the single-channel enhanced speech signal is adjusted and the recognition parameter configuration is configured to obtain speech recognition configuration data; A target speech recognition engine is constructed according to the speech recognition configuration data, and the single-channel enhanced speech signal is input into the target speech recognition engine for speech recognition to generate a target speech recognition result including navigation instructions, communication instructions and safety instructions.
[0042] Specifically, the wind direction compensation confidence is a trust score calculated in the previous module by combining multi-channel microphone consistency, filter residual evaluation, vehicle speed state stability and wind direction modeling error. The score value is between 0 and 1. The closer it is to 1, the more sufficient the compensation is, and the closer the enhanced speech signal is to a clean speech scene. Otherwise, it means that the wind noise suppression is incomplete or there is residual wind flow disturbance that has not been filtered out. The system divides the speech recognition mode into three levels based on the confidence. The high confidence state (such as confidence greater than 0.8) corresponds to the standard recognition mode. In this mode, the system uses the default time window length (such as 25ms window length, 10ms frame shift), and uses normal audio energy threshold and zero-crossing rate parameters for speech endpoint detection. In the process of speech feature extraction, standard MFCC (Mel-frequency cepstral coefficient) or FBANK (filter bank energy) features are used as input, and further combined with a general deep acoustic model for full-word recognition. This mode is suitable for speech with high signal-to-noise ratio, clear speech, and uniform speaking speed. uniform situation; when the confidence level is in the middle range (for example, between 0.6 and 0.8), the system enters the robust enhancement recognition mode. In this mode, the system expands the time window length, for example, increasing the window length to 30ms and the frame shift to 15ms, to enhance the energy aggregation capability of weak speech and ambiguous syllables. At the same time, the spectral subtraction residual correction mechanism is enabled in the acoustic model part to perform dimensionality reduction fitting on the irregular disturbances in the characteristic spectrum, and adapt the robust acoustic structure based on RNN or Transformer to improve the system's fault-tolerant recognition capability of incomplete syllables in wind noise environment. In this mode, the system will also relax the keyword detection threshold, and perform similarity matching on approximate expressions and partially distorted forms in voice commands, thereby improving the recall rate of actual commands; when the confidence level is less than 0.6, the system enters conservative recognition mode. In this mode, the system will extend the time window to more than 35ms, and at the same time increase the energy judgment threshold to avoid false triggering of non-voice segments. At the same time, the endpoint double confirmation mechanism is enabled. Only when two or more consecutive high-confidence feature segments meet the conditions can the recognition process be entered. In terms of acoustic modeling, a lightweight recognition model built based on a small vocabulary is used, and the target vocabulary is limited to a set of high-priority command words. It only supports key operations such as navigation (such as left turn, right turn, stop), communication (such as answering, hanging up) and safety control (such as braking, warning, avoidance), ensuring that the system will not misidentify low-confidence content and trigger wrong commands under extreme wind noise conditions. After completing the recognition mode selection, the corresponding recognition parameter configuration is executed according to the mode to generate speech recognition configuration data. The configuration data includes multiple dimensional parameters such as time window structure, speech endpoint detection strategy, feature extraction type, model structure type, vocabulary size, recognition threshold value, number of candidate recognition items, etc.The system uses this configuration data to build or call the corresponding version of the speech recognition engine. For example, it loads the full-size acoustic model and the complete language model in standard mode, loads the additional residual detection mechanism and the dual-channel front-end fusion network in robust mode, and loads the lightweight, high-priority vocabulary-restricted recognition path in conservative mode. All of these recognition engine instances are based on a single-channel enhanced speech signal as input. The input signal undergoes feature preprocessing, frame-level modeling, decoding network solution, and language model weighting to output structured speech recognition results. The system performs task semantic decoding on the recognition result, and the output includes navigation instructions (such as "turn left", "turn right", "go straight", "U-turn"), communication instructions (such as "answer", "hang up", "dial", "voice message") and safety control instructions (such as "slow down", "stop", "obstacle ahead"). These recognition results are checked for consistency with the current vehicle status before output to prevent the execution of invalid or conflicting instructions in unsafe scenarios, such as refusing to execute the "accelerate" command when parked, refusing to execute the "answer" command when turning at high speed, etc. The system realizes dynamic closed-loop control between voice command recognition, parsing and execution through a joint confidence and status verification mechanism, thereby ensuring that electric two-wheeled vehicles always maintain high reliability, high fault tolerance and high response speed voice interaction capabilities in various complex wind fields and changeable riding environments.
[0043] The above describes the voice recognition control method of the electric two-wheeled vehicle in the embodiment of the present invention. The following describes the voice recognition control system of the electric two-wheeled vehicle in the embodiment of the present invention. Figure 2 In one embodiment of the present invention, a speech recognition control system for an electric two-wheeled vehicle includes: Synchronous data acquisition module 201, for synchronously acquiring data from the three-axis wind direction sensor, vehicle speed sensor, six-axis attitude sensor, and stereo microphone array of the electric two-wheeled vehicle to obtain a multi-source data set and an original voice signal; A calculation module 202 is configured to perform coordinate system conversion and vehicle aerodynamic calculation on the multi-source data set to obtain a wind field intensity distribution matrix; A compensation module 203 is configured to perform asymmetric differential compensation on the original voice signal based on the wind field intensity distribution matrix to obtain a compensated voice signal; The speech recognition module 204 is configured to perform dynamic wind noise suppression processing and speech recognition on the compensated speech signal to generate a target speech recognition result.
[0044] Through the collaborative efforts of these components, integrating a three-axis wind direction sensor, a vehicle speed sensor, a six-axis attitude sensor, and a stereo microphone array, the system achieves multi-dimensional environmental perception and voice acquisition, providing a comprehensive data foundation for accurate wind direction compensation in open environments for electric two-wheelers, resolving the technical challenge of a single sensor failing to accurately describe complex riding environments. A rotational transformation mechanism from the ground coordinate system to the vehicle coordinate system is established, combined with a Kalman filter algorithm to accurately estimate the vehicle's motion state, effectively eliminating sensor noise and vibration interference and ensuring the accuracy and stability of wind direction vector calculations. Based on the unique structural characteristics of electric two-wheelers, a three-dimensional wind resistance parameter model encompassing the front, side, and rear of the vehicle is established, enabling precise calculation of the wind field distribution at each position in the stereo microphone array, addressing the inability of existing technologies to accurately predict wind field distribution in open environments. To address the differential wind influence on the inner and outer microphones during turns and banked turns, an asymmetric modulation algorithm based on spatial positional analysis was developed to achieve differentiated compensation for microphones in different locations, effectively addressing the technical limitation of traditional uniform compensation methods that is incapable of adapting to the dynamic riding conditions of electric two-wheelers. A classification and filtering mechanism has been established for low-frequency body vibration, medium-frequency airflow turbulence, and medium-high frequency high-speed riding wind noise of electric two-wheeled vehicles in open environments. By combining Butterworth, elliptic, and Chebyshev filters, precise suppression of wind interference in different frequency bands is achieved, significantly improving the signal-to-noise ratio and clarity of speech signals. Based on the confidence level of wind direction compensation, a three-level adaptive mode of standard recognition, robust enhancement, and conservative recognition is established. By dynamically adjusting the recognition time window and algorithm parameters, the reliability of speech recognition is ensured under different wind direction interference intensities, effectively improving recognition accuracy and response speed.
[0045] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0046] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0047] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A voice recognition control method for an electric two-wheeled vehicle, characterized in that: include: Synchronous data acquisition is performed on the electric two-wheeled vehicle's three-axis wind direction sensor, vehicle speed sensor, six-axis attitude sensor, and stereo microphone array to obtain a multi-source data set and original speech signal; Performing coordinate system conversion and vehicle aerodynamic calculation on the multi-source data set to obtain a wind field intensity distribution matrix; Performing asymmetric differential compensation on the original voice signal based on the wind field intensity distribution matrix to obtain a compensated voice signal; Dynamic wind noise suppression processing and speech recognition are performed on the compensated speech signal to generate a target speech recognition result.
2. The voice recognition control method for an electric two-wheeled vehicle according to claim 1, characterized in that: The method includes synchronously collecting data from a three-axis wind direction sensor, a vehicle speed sensor, a six-axis attitude sensor, and a stereo microphone array of the electric two-wheeled vehicle to obtain a multi-source data set and an original voice signal, including: The three-axis wind direction sensor arranged on the electric two-wheeled vehicle is sampled to obtain real-time wind direction vector data including the X-axis longitudinal wind speed component, the Y-axis transverse wind speed component and the Z-axis vertical wind speed component; The vehicle speed sensor and six-axis attitude sensor installed on the electric two-wheeled vehicle are synchronously sampled to obtain vehicle body motion state data including vehicle speed vector and vehicle body attitude angle; By performing parallel audio acquisition on D microphones of a stereo microphone array in an electric two-wheeled vehicle, D original speech signals are obtained; The real-time wind direction vector data and the vehicle body motion state data are fused based on a unified timestamp to obtain a multi-source data set.
3. The voice recognition control method for an electric two-wheeled vehicle according to claim 2, characterized in that: The coordinate system conversion and vehicle aerodynamic calculation are performed on the multi-source data set to obtain a wind field intensity distribution matrix, including: Performing a rotation transformation from a ground coordinate system to a vehicle coordinate system based on the real-time wind direction vector data in the multi-source data set to obtain transformed wind direction vector data in the vehicle coordinate system; Performing Kalman filtering and state estimation on the vehicle body motion state data in the multi-source data set to obtain vehicle body motion parameters in a vehicle body coordinate system; Performing vector subtraction and vehicle body induced wind field compensation based on the transformed wind direction vector data and the vehicle speed vector in the vehicle body motion parameter to obtain relative wind direction vector data; performing a turning tilt posture correction on the relative wind direction vector data according to the vehicle body posture angle in the vehicle body motion parameter to obtain a corrected wind direction vector in a vehicle body coordinate system; Vehicle body aerodynamic calculation is performed based on the corrected wind direction vector and the vehicle body motion parameters to obtain a wind field intensity distribution matrix.
4. The voice recognition control method for an electric two-wheeled vehicle according to claim 3, characterized in that: The method of performing a turning tilt posture correction on the relative wind direction vector data of the electric two-wheeled vehicle according to the vehicle body posture angle in the vehicle body motion parameter to obtain a corrected wind direction vector in a vehicle body coordinate system includes: Performing a threshold judgment on the vehicle body posture angle in the vehicle body motion parameter to obtain a turning tilt state identifier and a corresponding turning tilt angle; Calculating a turning radius and a centrifugal force vector based on the turning tilt angle and the vehicle speed vector in the vehicle body motion parameters to obtain turning dynamics parameters; Perform centrifugal wind field compensation and tilt angle correction on the relative wind direction vector data according to the turning dynamics parameters to obtain turning correction wind direction vector data; The turning correction wind direction vector data and the relative wind direction vector data are selected based on the turning tilt state identifier to obtain a correction wind direction vector in a vehicle body coordinate system.
5. The voice recognition control method for an electric two-wheeled vehicle according to claim 4, characterized in that: The vehicle body aerodynamic calculation is performed based on the corrected wind direction vector and the vehicle body motion parameters to obtain a wind field intensity distribution matrix, including: Calculating the frontal area and drag coefficient of the vehicle body based on the corrected wind direction vector and the vehicle speed vector in the vehicle body motion parameters to obtain vehicle body aerodynamic parameters including frontal drag parameters, side drag parameters, and rear drag parameters; Calculating the directional wind field intensity of the corrected wind direction vector according to the vehicle body aerodynamic parameters to obtain directional wind field intensity data including a frontal wind field intensity component, a lateral wind field intensity component, and a rearward wind field intensity component; Performing wind field interference intensity analysis at microphone positions based on the directional wind field intensity data and the spatial position distribution of the stereo microphone array to obtain three-dimensional wind field interference intensity values corresponding to D microphone positions; The three-dimensional wind field interference intensity values corresponding to the D microphone positions are subjected to matrix conversion to obtain a wind field intensity distribution matrix.
6. The voice recognition control method for an electric two-wheeled vehicle according to claim 5, characterized in that: The performing asymmetric differential compensation on the original voice signal based on the wind field intensity distribution matrix to obtain a compensated voice signal includes: Calculating the wind direction compensation coefficient for each microphone based on the three-dimensional wind field interference intensity values corresponding to the D microphone positions in the wind field intensity distribution matrix to obtain a D-way microphone compensation coefficient group including a handlebar lateral compensation coefficient, a front vehicle front compensation coefficient, and a vehicle body rear compensation coefficient; Asymmetrically modulating the D-channel microphone compensation coefficient group based on the turning tilt state identifier to obtain a turning asymmetric compensation coefficient group; The original voice signals of the D channels are subjected to amplitude compensation calculation and phase correction at corresponding positions with the turn asymmetric compensation coefficient group to obtain D channel corrected voice signals, and the D channel corrected voice signals are subjected to differential signal synthesis to obtain a single channel compensated voice signal.
7. The voice recognition control method for an electric two-wheeled vehicle according to claim 6, characterized in that: The asymmetrically modulating the D-channel microphone compensation coefficient group based on the turning tilt state identifier to obtain the turning asymmetric compensation coefficient group includes: According to the turning tilt state identifier and the turning tilt angle, determining the inner and outer positions of the stereo microphone array to obtain microphone spatial distribution data; performing asymmetric modulation intensity calculation based on the turning tilt angle to obtain a turning modulation parameter including an inner attenuation modulation coefficient and an outer enhancement modulation coefficient; Position-matching the D-channel microphone compensation coefficient group with the microphone spatial distribution data to obtain an inner microphone compensation coefficient and an outer microphone compensation coefficient; The inner microphone compensation coefficient and the outer microphone compensation coefficient are differentially modulated according to the turning modulation parameter to obtain a turning asymmetric compensation coefficient group.
8. The voice recognition control method for an electric two-wheeled vehicle according to claim 7, characterized in that: The performing dynamic wind noise suppression processing and speech recognition on the compensated speech signal to generate a target speech recognition result includes: Performing frequency domain analysis on the compensated voice signal to obtain wind direction interference spectrum feature data including a low-frequency vehicle body vibration interference spectrum, a medium-frequency airflow turbulence interference spectrum, and a medium-high frequency high-speed riding wind noise spectrum; generating dynamic filter configuration data based on the wind direction interference spectrum characteristic data and the vehicle speed vector in the vehicle body motion parameter; constructing an adaptive multi-frequency domain filter according to the dynamic filter configuration data, and inputting the compensated speech signal into the adaptive multi-frequency domain filter to suppress wind noise, thereby obtaining a multi-channel filtered speech signal; Performing beamforming and wind direction compensation confidence calculation on the multi-path filtered speech signal to obtain a single-path enhanced speech signal and a corresponding wind direction compensation confidence; Speech recognition is performed on the single-channel enhanced speech signal according to the wind direction compensation confidence level to generate a target speech recognition result.
9. The voice recognition control method for an electric two-wheeled vehicle according to claim 8, characterized in that: The performing speech recognition on the single-channel enhanced speech signal according to the wind direction compensation confidence to generate a target speech recognition result includes: Selecting a corresponding speech recognition mode according to the wind direction compensation confidence level; Based on the speech recognition mode, the time window of the single-channel enhanced speech signal is adjusted and the recognition parameter configuration is configured to obtain speech recognition configuration data; A target speech recognition engine is constructed according to the speech recognition configuration data, and the single-channel enhanced speech signal is input into the target speech recognition engine for speech recognition to generate a target speech recognition result including navigation instructions, communication instructions and safety instructions.
10. A voice recognition control system for an electric two-wheeled vehicle, characterized in that: A method for executing a voice recognition control method for an electric two-wheeled vehicle according to any one of claims 1 to 9, comprising: A synchronous data acquisition module is used to synchronously collect data from the electric two-wheeled vehicle's three-axis wind direction sensor, vehicle speed sensor, six-axis attitude sensor, and stereo microphone array to obtain a multi-source data set and original voice signal; A calculation module, configured to perform coordinate system conversion and vehicle aerodynamic calculation on the multi-source data set to obtain a wind field intensity distribution matrix; a compensation module, configured to perform asymmetric differential compensation on the original voice signal based on the wind field intensity distribution matrix to obtain a compensated voice signal; The speech recognition module is used to perform dynamic wind noise suppression processing and speech recognition on the compensated speech signal to generate a target speech recognition result.
Citation Information
Cited By
Crosswind early warning method for vehicle auxiliary driving
CN121201097A
Two-wheeled electric vehicle real-time interaction system based on vector voice agent
CN121331123A