Two-wheeled vehicle leasing voice service system based on user portrait and position awareness
By integrating multi-sensor data fusion and location-aware adaptive noise processing, combined with acoustic gene recognition based on user profiles and multi-dimensional contextual decision-making, the noise interference and individual differences issues in the two-wheeled vehicle rental voice service system during riding have been resolved, achieving efficient and accurate voice interaction and personalized services.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN HOT WHEELS TECHNOLOGY CO LTD
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-01
AI Technical Summary
Existing voice service systems for two-wheeled vehicle rentals are susceptible to interference from wind noise and bumps during riding, resulting in decreased voice intelligibility, lack of personalized adaptation, insufficient recognition robustness, low accuracy in intent recognition, and inability to combine vehicle movement status and geographical location.
The system employs a speech enhancement module that combines multi-sensor data fusion with motion state coupling, a location-aware adaptive noise adversarial processing module, an acoustic gene fault-tolerant recognition module that integrates user profiles, and an intent decision-making and personalized service module based on multi-dimensional context. Speech recognition is achieved through motion micro-motion event detection, geographic noise fingerprint screening channel, personalized acoustic gene library, and multi-dimensional context decision-making.
To improve voice intelligibility, enhance recognition robustness and accuracy in complex cycling scenarios, provide personalized services, and optimize user interaction experience.
Smart Images

Figure CN121963732A_ABST
Abstract
Description
Two-wheeled vehicle rental voice service system based on user profiles and location awareness Technical Field
[0001] This invention relates to the field of intelligent voice interaction technology, and more specifically, to a two-wheeled vehicle rental voice service system based on user profiles and location awareness. Background Technology
[0002] In the two-wheeled vehicle rental industry, voice interaction has become a key technology for improving riding safety and convenience due to its elimination of manual operation. With the increasing popularity of shared two-wheeled vehicles, users' demand for voice-controlled vehicle functions and navigation services is growing, making voice service systems a crucial area for technological upgrades in the industry.
[0003] Current voice service systems for two-wheeled vehicle rentals have significant technical shortcomings: Firstly, environmental interference such as wind noise and bumps during riding can cause traditional voice endpoint detection to fail, resulting in a significant decrease in voice intelligibility. Existing noise reduction technologies mostly employ full-band suppression, which is computationally expensive and has limited effectiveness. Secondly, the lack of personalized adaptation to users' acoustic characteristics makes it difficult to cope with the differences in voice among users of different ages, genders, and accents, resulting in insufficient robustness in recognition. At the same time, insufficient integration of scene information such as vehicle movement status and geographical location makes it impossible to establish a connection between actions and voice commands, leading to low accuracy in intent recognition and impacting user experience.
[0004] To address the aforementioned issues, this invention proposes a two-wheeled vehicle rental voice service system based on user profiles and location awareness. Through the collaborative work of multiple modules, it overcomes the limitations of existing technologies and achieves efficient and accurate voice interaction in complex riding scenarios. Summary of the Invention
[0005] In view of the shortcomings of existing technologies, the purpose of this invention is to provide a two-wheeled vehicle rental voice service system based on user profiles and location awareness.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a two-wheeled vehicle rental voice service system based on user profiles and location awareness, comprising a voice enhancement module that fuses multi-sensor data and couples motion states, a location-aware adaptive noise adversarial processing module, an acoustic gene fault-tolerant recognition module that fuses user profiles, and an intent decision-making and personalized service module based on multi-dimensional context. The voice enhancement module that fuses multi-sensor data and couples motion states establishes action-intent associations based on multi-dimensional sensor data collected from the front end through motion micro-motion event detection, providing a physical synchronization benchmark for voice stream segmentation and recognition. The location-aware adaptive noise adversarial processing module constructs a geographic noise fingerprint based on location data and audio signals, filters the optimal voice transmission channel, and improves voice intelligibility. The acoustic gene fault-tolerant recognition module that fuses user profiles constructs a personalized acoustic gene library and uses a fragmented voting matching algorithm to improve recognition robustness. The intent decision-making and personalized service module based on multi-dimensional context fuses multi-dimensional information to complete intent decisions and updates user profiles through an interactive closed loop to achieve personalized services.
[0007] Furthermore, the multi-sensor data fusion and motion state coupling voice enhancement module includes a motion micro-motion event detection unit and a motion event-based voice stream segmentation and context weighting unit. The motion micro-motion event detection unit uses multiple types of onboard sensors to collect multi-dimensional vehicle motion data and detects motion micro-motion events coupled with voice commands by analyzing the joint temporal features of the data. The voice stream segmentation and context weighting unit uses the detected motion micro-motion events as the basis for voice stream segmentation and simultaneously corrects the recognition probability through motion context weighting to reduce recognition deviation under noise interference.
[0008] Furthermore, the motion micro-motion events include preset event types related to steering, braking, and stable riding. The activation of an event is determined by analyzing the trend and duration of characteristic parameter changes in vehicle motion data.
[0009] Furthermore, the location-aware adaptive noise countermeasures module includes a dynamic construction and matching unit for geographic noise fingerprints and an optimal voice transmission channel selection and guidance unit. The dynamic construction and matching unit periodically collects ambient background noise, processes the signal to form a noise fingerprint for the corresponding location, and compares and updates it with a cloud-based reference noise fingerprint. The channel selection and guidance unit searches for the optimal voice transmission channel based on the current noise fingerprint, adjusts the relevant parameters for voice acquisition and recognition, and can guide the user to adjust their vocalization method to adapt to the channel.
[0010] Furthermore, the geographic noise fingerprint is obtained by performing spectral analysis on the noise signal and calculating the long-term average power spectral density, which covers the main frequency bands of human speech and is updated when the vehicle is stationary and there is no user interaction.
[0011] Furthermore, the acoustic gene fault-tolerant recognition module integrating user profiles includes a command acoustic gene library construction and personalized adaptation unit and a gene fragment voting matching algorithm unit. The construction and personalized adaptation unit extracts key acoustic gene fragments from the core command set to construct a general gene library, dynamically adjusts the general gene fragments in combination with acoustic feature parameters in the user profile, and generates personalized gene templates. The voting matching algorithm unit extracts acoustic feature sequences from the speech to be recognized, calculates similarity with personalized gene fragments, and realizes command recognition through a voting mechanism.
[0012] Furthermore, the acoustic gene fragments serve as templates for characterizing the local stable acoustic features of instructions. They are extracted from pure instruction speech samples from different speakers, with each instruction corresponding to multiple gene fragments, covering the key acoustic features of the instructions.
[0013] Furthermore, the intent decision and personalized service module includes a multi-dimensional confidence fusion and intent decision unit and a personalized interaction closed loop and profile update unit. The confidence fusion and intent decision unit fuses acoustic model probabilities, gene matching scores, and environmental context prior probabilities to calculate a joint confidence score to determine the recognition result. When the confidence is low, an active clarification interaction is triggered. The interaction closed loop and profile update unit records the complete data and associated context of each interaction, and continuously optimizes the instruction preferences, voice features, and scene adaptation strategies in the user profile based on this data.
[0014] Compared with existing technologies, this invention has the following advantages: 1. To address the problem of noise interference in cycling environments, this system adopts a location-aware adaptive noise adversarial processing method. By constructing a geographical noise fingerprint to select the optimal voice transmission channel, and combining motion context weighting to correct the recognition probability, it replaces the traditional full-band noise reduction scheme. This technology does not require complex noise suppression calculations, improves voice intelligibility with a channel selection strategy, and establishes a physical synchronization benchmark by leveraging motion micro-motion events, effectively solving the recognition deviation problem caused by bumps and wind noise; 2. To address the problems of individual differences in user voice and insufficient recognition robustness, this system constructs an acoustic gene fault-tolerant recognition mechanism that integrates user profiles. By extracting key acoustic gene fragments from commands to establish a general gene library, and combining this with user acoustic feature parameters to generate personalized templates, a fragmented voting matching algorithm is employed. Even with partial speech occlusion or accents, recognition can still be achieved through key gene fragment matching, significantly improving the adaptability and reliability of speech recognition for different user groups. 3. To address the issue of low intent recognition accuracy, this system adopts a multi-dimensional contextual fusion intent decision-making method. This method integrates acoustic model probability, gene matching scores, and prior probabilities from environmental and motion contexts. The recognition result is determined through joint confidence calculation, triggering proactive clarification interaction when confidence is low. Simultaneously, user profiles are continuously updated through an interactive closed loop, optimizing command preferences and scenario adaptation strategies to achieve continuously optimized personalized services and improve the user experience. Attached Figure Description
[0015] Figure 1 is a block diagram of a two-wheeled vehicle rental voice service system based on user profiles and location awareness; Figure 2 is a flowchart of the implementation of the voice enhancement module of the present invention, which integrates multi-sensor data fusion and motion state coupling; Figure 3 is a flowchart of the implementation of the acoustic gene fault-tolerant recognition module of the present invention, which integrates user profiles. Detailed Implementation
[0016] Referring to Figure 1, the two-wheeled vehicle rental voice service system based on user profiles and location awareness in this embodiment includes a voice enhancement module that fuses multi-sensor data and couples motion states, a location-aware adaptive noise adversarial processing module, an acoustic gene fault-tolerant recognition module that fuses user profiles, and an intent decision-making and personalized service module based on multi-dimensional context. The voice enhancement module that fuses multi-sensor data and couples motion states: Based on multi-dimensional sensor data collected from the front end, it establishes action-intent association through motion micro-motion event detection, providing a physical synchronization benchmark for voice stream segmentation and recognition. It is the core module for solving the problem of voice endpoint detection failure under riding bumps and wind noise. As shown in Figure 2, the specific implementation is as follows: S11, Motion micro-motion event detection unit: The system uses the vehicle's onboard inertial measurement unit (recommended model: MPU6050, sampling rate 100Hz), pedal frequency sensor (Hall type, accuracy ±1rpm), and brake pressure sensor (piezoelectric type, measurement range 0-1MPa) to collect multi-dimensional motion data of the vehicle in real time. By analyzing the joint temporal characteristics of these data, a series of micro-motion events highly coupled with voice commands were defined and detected. These events are physical precursors to the user's cycling intentions. The predefined event types and judgment logic are as follows, ensuring that those skilled in the art can directly reproduce them: Preparing to turn left: Yaw angle velocity Based on statistical analysis of cycling data from 100 testers in intersection turning scenarios, 95% of the yaw angular velocities during turn preparation fell within this range. This approach accurately captures turning intentions while avoiding confusion with instantaneous angular velocities caused by bumps. The duration is over 0.3 seconds; in actual testing, the average duration of the turn preparation action was 0.45 seconds. 0.3 seconds covers over 90% of effective actions while filtering out instantaneous interference. (Cadrenaline...) A drop of more than 10% means the user needs to operate the handlebars with one hand when turning, interrupting power and causing a decrease in cadence. A 10% drop is the critical value that distinguishes turning from normal riding; event activation probability threshold. The value is determined based on ROC curve analysis of 100,000 training samples. This threshold corresponds to an accuracy of 92% and a false detection rate of 5%, achieving an optimal balance between recognition accuracy and false negative rate. A typical application scenario is before the turn command is triggered at an intersection.
[0017] Start braking: Brake pressure In actual testing, the minimum pressure required for conscious braking by the user was 0.18 MPa. Using 0.2 MPa avoids false triggering due to road bumps; triaxial acceleration. The average acceleration during emergency deceleration is , It can effectively distinguish between normal deceleration and emergency braking scenarios; event activation probability threshold. The value is determined based on the following criteria: braking events require a rapid response. Sample tests show that the response time under this threshold is ≤0.1s, while the false detection rate is controlled within 8%. Typical application scenarios include emergency deceleration and stopping command scenarios.
[0018] Stable pedaling: cadence Fluctuation range During constant-speed cycling, the average fluctuation of the user's cadence is ±3 rpm, and ±5 rpm can cover 95% of stable cycling conditions; the duration of stable cycling is longer than 1 second, with a minimum duration of 0.8 seconds, and 1 second can filter out momentary stable interference scenarios; acceleration variance The statistical mean of the variance of acceleration during constant-speed cycling is , This is a threshold value; exceeding it indicates a bumpy or deceleration state. (Event activation probability threshold) The value is determined based on the following criteria: the recognition accuracy of stable riding scenarios is 94% under this threshold, which can accurately filter out routine command triggering scenarios; the typical application scenario is routine commands during constant speed riding.
[0019] Set time The sensor data vector is ,in These are triaxial accelerations. Yaw angular velocity, For cadence, Braking pressure. This is achieved using a pre-trained lightweight temporal classification model. (Using an LSTM network, with an input dimension of 6, a hidden layer dimension of 32, and an output dimension consistent with the predefined number of events), the event probability is calculated in real time. ;in, Representing the Class-predefined micro-action events; The event time window used for judgment is typically 0.5 to 1 second. The determination method is as follows: adaptively adjusted based on the action duration characteristics of the event type. For braking events, which are rapid actions, 0.5 seconds is used to ensure response speed; for stable pedaling events, which require continuous observation, 1 second is used to ensure recognition stability; other custom events can be set based on 1.2 times the average duration of the action; when a certain event... probability Exceeding the set threshold If the event occurs at time 1, then the event is determined to occur at time 2. Activated.
[0020] The model training data consisted of cycling data from 100 testers under different road conditions (asphalt, concrete, and gravel), totaling 100,000 valid samples. The training iterations were 50 rounds. During training, the model's accuracy on the validation set stabilized after 50 iterations; further iterations offered no significant performance improvement and avoided overfitting, achieving a validation set accuracy of ≥92%. This elevates the motion sensor from a traditional vehicle status monitor to an active signal source for interpreting user interaction intentions, providing a physical synchronization benchmark for subsequent voice processing.
[0021] S12. Motion Event-Based Speech Stream Segmentation and Context Weighting Unit: Based on previously detected motion micro-motion events, this unit replaces traditional acoustic silence detection to achieve accurate speech stream segmentation. Simultaneously, it corrects the recognition probability through motion context weighting, addressing recognition bias issues under noise interference. Traditional speech endpoint detection is highly susceptible to failure under cycling wind noise and bumpy conditions. This invention abandons silence detection that solely relies on acoustic features, proposing to use detected motion micro-motion events as the primary basis for speech stream segmentation.
[0022] When the system is at time Event detected Once activated, it will be within the time window. Inside, among which and The pre-set look-ahead and delay times based on the event type are as follows, to ensure segmentation accuracy; the acquired audio stream... (Sampling rate 16kHz, 16-bit mono) marked as related to the event Related audio clips. Meanwhile, the event... Provides strong semantic context for the current speech recognition process Preparing to turn left: Preview time In actual testing, the average lead time between user turning actions and voice commands was 0.15 seconds, and 0.2 seconds can cover more than 85% of command-preparation scenarios; the delay time... The average duration of turn commands is 0.6 seconds, and the voice command can be completely captured in 0.8 seconds; the weight vector adjustment logic is to enhance the weight of left turn, navigation, and deceleration commands. The values are determined by the proportion of high-frequency commands in the scenario: left turn accounts for 45%, navigation accounts for 30%, and deceleration accounts for 15%. The weights are positively correlated with the proportions to ensure that key commands are identified first.
[0023] Initiating Braking: Detective Time In case of emergency braking, users may provide an early warning via audio, with an average advance warning time of 0.25 seconds; 0.3 seconds can cover this scenario; the delay time... The average duration of braking-related commands is 0.5 seconds, and they can be fully captured in 0.7 seconds; the weight vector adjustment logic is to enhance the weight of stop, caution, and deceleration commands. The values are determined based on the following criteria: safety commands have the highest priority in this scenario, with 50% for stopping, 30% for caution, and 18% for deceleration. High weights can improve the recognition priority of emergency commands.
[0024] Stable pedaling: look-ahead time Regular commands have no obvious pre-command actions, and irrelevant speech can be filtered out in 0.2 seconds; post-delay time The average duration of regular commands is 0.8 seconds, and 1.0 seconds can cover all valid commands; the weight vector adjustment logic enhances the weight of speed control, navigation, and music playback commands. The values are determined based on the fact that entertainment and functional commands are balanced in this scenario, and the weight settings take into account both recognition accuracy and diversity.
[0025] In the acoustic feature recognition stage, the system introduces a motion context weight vector. The acoustic model output probability of different types of commands is dynamically modulated. The weight vector is obtained from a predefined policy mapping table according to the event type, and the weight coefficient ranges from 0.5 to 2.2. A value less than 1 indicates suppression of irrelevant commands, and a value greater than 1 indicates enhancement of target commands. The value is determined by testing in 1000 noisy scenarios. This range can improve the accuracy of target command recognition by 20%-30%, while avoiding missed detections caused by excessive suppression.
[0026] Assume the acoustic model uses a CNN+CTC architecture, with input being Mel-spectral features of dimension 40×100. The 40-dimensional Mel-spectral features can cover key speech features, and the 100-frame length is suitable for short command speech. For the... Candidate instructions The original output probability is The probability after motion context modulation for: ;in, It is a weight vector Corresponding instructions Weighting coefficients; This is a smoothing factor that adjusts the modulation intensity, typically 1.2, with a range of 1.0 to 1.5. The parameter is determined dynamically based on real-time noise intensity: 1.5 for noise levels ≥60dB, 1.2 for 40-60dB, and 1.0 for ≤40dB, automatically adapted to the ambient noise decibel value collected by a noise sensor. This method creatively utilizes the inherent action-intention association during cycling to correct recognition biases under noise interference at a probabilistic level, significantly improving the accuracy of semantic understanding.
[0027] Location-aware adaptive noise adversarial processing module: Based on the location data and audio signal output by the preceding module, it constructs a geographic noise fingerprint and selects the optimal voice transmission channel. It adopts a channel selection strategy to replace traditional noise reduction, reducing computational overhead while improving speech intelligibility in complex environments.
[0028] S21. Dynamic Construction and Matching Unit for Geographic Noise Fingerprints: Clarify the construction method, update logic, and matching criteria for noise fingerprints, provide data support for subsequent optimal channel selection, and ensure that noise characteristics under different geographical environments can be accurately characterized.
[0029] The system periodically collects ambient background noise via an onboard microphone, with a collection cycle of 5 seconds and a collection duration of 2 seconds. The 2-second duration covers the main frequency components of the noise, ensuring the integrity of fingerprint features; a duration shorter than 1 second results in insufficient feature extraction. For locations at coordinates... For vehicles, short-time Fourier transform (SFT) is performed on the collected noise signals: frame length 25ms, the short-time stationarity of the speech signal is approximately 20-30ms, balancing time and frequency resolution; frame shift 10ms ensures 75% inter-frame overlap, avoiding feature loss; 512 FFT points achieve a frequency resolution of 31.25Hz, meeting the requirements for speech segment recognition. The long-term average power spectral density is calculated, with a statistical window of 10 acquisition cycles, totaling 50s, to smooth instantaneous noise fluctuations, ensuring fingerprint stability and forming a noise fingerprint for that location. ,in The frequency range is 300Hz-3400Hz, covering the main frequency bands of human speech.
[0030] When the vehicle is stationary and the user is not interacting, the system will generate the current noise fingerprint. Location-based Reference noise fingerprint obtained from the cloud The system performs comparison and fusion updates. The comparison uses a cosine similarity algorithm with a similarity threshold of 0.85. This threshold is determined by analyzing the similarity distribution of noise fingerprints at the same location over different time periods; 0.85 is considered a critical value for noise feature consistency, and values higher than this value indicate similar noise environments. When the similarity is ≥0.85, a weighted average method is used to update the reference fingerprint. The current fingerprint has a weight of 0.3, and the reference fingerprint has a weight of 0.7. This is based on the fact that the reference fingerprint is more representative after multiple statistical analyses, and a higher weight ensures fingerprint stability. The current fingerprint weight is used to adapt to subtle noise changes. When the similarity is <0.85, the current fingerprint is stored as a new entry in the cloud fingerprint database. The core of this unit is that the system does not seek to suppress noise across the entire frequency band, but rather identifies quiet channels with relatively weak noise energy in the current environment.
[0031] S22, Optimal Voice Transmission Channel Selection and Guidance Unit: Based on the noise fingerprint constructed in the previous step, the optimal frequency band is selected by signal-to-noise ratio prediction, and the front-end processing and user guidance logic are clarified to ensure that the technical solution can be implemented.
[0032] When a user prepares to initiate a voice interaction (either by pressing a button or using specific voice content), the system analyzes the current noise fingerprint in real time. Define the frequency range. The signal-to-noise ratio prediction function is: ;in, The system estimates the average power spectrum of user speech based on typical acoustic features stored in user profiles. It pre-models the speech by collecting 3-5 standard user commands, covering different tones and speech rates. Within the main frequency band of human speech (300Hz–3400Hz), the system employs a sliding window method (window width of 200Hz, which captures key frequency components; too narrow a window results in fragmented features, while too wide a window fails to accurately locate quiet channels; step size of 50Hz ensures continuous window movement and avoids missing optimal channels) for searching. Maximizing continuous sub-band This channel was selected as the optimal voice transmission channel for this interaction, with a bandwidth of no less than 500Hz to ensure voice integrity; otherwise, voice information would be severely lost.
[0033] Subsequently, the system dynamically adjusts the frequency band filters of the voice acquisition front-end and recognition engine to the optimal channel, and performs corresponding frequency band compression and enhancement processing on the input voice. This embodiment uses logarithmic domain spectral subtraction, with a gain coefficient range of 0.8-1.2. The parameter determination method is based on the dynamic adjustment of the signal-to-noise ratio (SNR) of the optimal channel, specifically by automatically adapting the calculated value through the SNR prediction function, as detailed below: Among them, SNR is calculated using the formula Calculation, where Represents the speech power spectral density (STFT calculation); This represents the noise power spectral density (silent frame estimation or fingerprint database matching); when SNR=5dB (strong noise): G≈0.8+0.2×e−2≈0.87 (relatively large gain); when SNR=15dB (weak noise): G≈0.8+0.2×e0=1.0 (no additional gain); SNR<5dB: fixed gain 0.6 (avoid over-amplifying noise).
[0034] Simultaneously, by incorporating user tone preferences from their profiles, voice prompts can guide users to adjust their vocalization to better match the channel. The preset prompts include: "The current environment is noisy; please use a slightly higher / lower voice to give instructions." The prompt content is adaptively selected based on the optimal channel frequency band: higher frequencies (≥2000Hz) prompt with a slightly higher voice, and lower frequencies (≤1000Hz) prompt with a slightly lower voice. The frequency band division is based on the fact that high-frequency human speech is dominated by voiceless consonants, while low frequencies are dominated by voiced consonants and vowels; adapting the vocalization style improves channel utilization. This technology transcends the limitations of traditional noise reduction thinking, achieving improved speech intelligibility in characteristic noise environments with extremely low computational overhead (single-frame processing time ≤1ms) through a channel-optimized rather than all-encompassing strategy.
[0035] The acoustic gene fault-tolerant recognition module, which integrates user profiles, constructs a personalized acoustic gene library based on the optimized speech features output by the preceding module. It employs a fragmented voting matching algorithm to address the robustness of recognition when speech is obscured, distorted, or has an accent. Simultaneously, it leverages user profiles to achieve personalized adaptation. As shown in Figure 3, the specific process is as follows: S31, Construction and Personalized Adaptation Unit of the Instruction Acoustic Gene Library: This unit clarifies the extraction criteria for acoustic gene fragments, the gene library construction process, and the personalized adaptation method, ensuring that the gene template accurately represents instruction features and adapts to the speech habits of different users.
[0036] Core instruction set for two-wheeled vehicle rental It includes 20 high-frequency commands, such as left turn, right turn, stop, and navigate to XX, covering three major categories: safety control, navigation, and entertainment; the system assigns each command... Construct a set of key acoustic gene fragments Each instruction corresponds to 3-5 gene fragments. For short commands (such as parking), 3 segments are taken; for long commands (such as navigation to XX intersection), 5 segments are taken. The number of segments is determined according to the number of syllables in the command, following the principle of "number of syllables + 1" to ensure coverage of the key features of the command. Each segment is 0.1~0.3s long, which can ensure the capture of phoneme or syllable-level features of the speech. If it is too short, the features will be incomplete; if it is too long, the anti-masking ability will decrease.
[0037] Each gene segment It is a template representing the local stable acoustic features of an instruction. The specific extraction process is as follows: (1) Collect 1000 pure instruction speech samples from different speakers, covering different ages, genders, and accents to ensure the universality of the gene template; (2) Perform pre-emphasis processing on the samples with a coefficient of 0.97, which can effectively enhance the high-frequency speech components and compensate for the microphone's attenuation of high frequencies; perform frame processing with a frame length of 20ms to adapt to the short-term stationary characteristics of speech; shift the frame by 10ms to ensure inter-frame overlap and avoid feature loss; (3) Extract the Mel-frequency cepstral coefficients and fundamental frequency. The curve, Mel-frequency cepstral coefficient (MFCC), 12th order + 1st order energy, a total of 13 dimensions; the 12th order MFCC can cover the main spectral features of speech, and the 1st order energy improves the feature recognition; (4) Common feature segments are mined by K-means clustering, the number of clusters = the number of gene segments, and segments with intra-class similarity ≥0.9 are selected as general gene templates. Data verification shows that intra-class similarity ≥0.9 can ensure the consistency of the template and reduce the impact of individual differences on recognition. For example, the gene segments of the parking instruction include the formant trajectory of the word "stop" (the frequency ranges of the 1st to 3rd order formants are 300-700Hz, 1500-2500Hz, and 2500-3500Hz respectively; the value is based on the distribution of the formants of 1000 pronunciations of the word "stop" to determine the common interval) and the fundamental frequency drop segment feature of the word "car".
[0038] A user's personalized profile includes their unique acoustic characteristic parameters. Parameters such as average fundamental frequency, speech rate, and typical formant distribution are generated from three calibrated speech samples collected during the user's first use. During recognition, the system utilizes... Universal gene fragment template Dynamic adjustments are made to generate personalized gene fragment templates tailored to the user. The adjustment rules are as follows: The baseband profile is linearly scaled according to the ratio of the user's average baseband to the average baseband of the general template, using the scaling formula: ,in The average base frequency for users, The fundamental frequency is the average of the general template; the resonant frequencies are shifted according to the user's resonant offset, using the following formula: ,in For the user's resonant peak frequency, The formant frequency is a general template; the speech rate is adjusted according to the user's speech rate characteristics, and the segment duration is adjusted using the following formula: ,in For general template speaking speed, The user's speaking speed.
[0039] S32, Gene Fragment Voting Matching Algorithm Unit: Clarifies the similarity calculation method, voting rules, and scoring formula parameters for gene fragment matching, ensuring that the fault-tolerant recognition logic can be accurately reproduced and improving the recognition robustness in complex scenarios.
[0040] For a speech segment to be recognized Instead of performing end-to-end word-level matching, the recognition engine executes a gene fragment-level search and voting mechanism.
[0041] First, from Extracting continuous acoustic feature sequences: Dimension 14; 12th order MFCC + 1st order energy + 1st order This forms a 14-dimensional feature vector that takes into account both spectral and fundamental frequency information, and calculates its relationship with all personalized gene fragments. (For all instructions) Local similarity The similarity is calculated using the Dynamic Time Warping (DTW) algorithm, with the path constraint window set at ±5 frames to balance time warp tolerance and computational complexity, avoiding excessive matching; similarity threshold. Set to 0.75, adapting to the user's accent. Use 0.75 for standard accents and 0.7 for heavy accents. The value will be automatically adjusted based on the accent recognition results of the user's initial voice calibration.
[0042] For each candidate instruction Statistical analysis of all its gene fragments In voice The number of high similarity matches obtained, i.e. Exceeding the threshold The number of times, recorded as Simultaneously, considering the consistency of the spatial distribution of gene fragments (their relative positions in the instructions), the voting results are weighted. Ultimately, the instructions... Gene matching score for: ;in, This refers to the importance weight of the gene fragment, the core fragment. Secondary fragments Weight determination method: The segment discrimination is calculated by information gain. Segments with a discrimination of ≥0.8 are core segments, and those without are secondary segments. It is an indicator function that takes the value 1 if the condition is met, and 0 otherwise; It is a gene fragment in The detected positions are normalized to the [0,1] interval. The normalization method is as follows: ,in For the detection time point, For audio clips Total duration; It is the approximate position where it is expected to appear in the complete instruction. For example, 0.2 for the first 1 / 3 of the instruction, 0.5 for the middle section, and 0.8 for the last section. The value is determined by dividing the instruction duration into three segments and taking the midpoint of each segment as the expected position. It is a coefficient that controls the intensity of the positional deviation penalty, with a typical value of 10.0 and a range of 8.0 to 12.0. The determination method is to dynamically adjust it according to the gene fragment duration: take a larger value (such as 12.0) when the fragment duration is <0.2 seconds and take a smaller value (such as 8.0) when the fragment duration is ≥0.2 seconds, so as to balance the requirements for the positional accuracy of short fragments and the tolerance for the natural fluctuations of long fragments.
[0043] This segmented recognition strategy allows the system to reliably recognize commands even when speech is partially obscured, distorted, or has an accent, as long as it captures enough key acoustic genes, greatly improving the system's robustness and fault tolerance.
[0044] The multi-dimensional context-based intent decision-making and personalized service module integrates acoustic probability, gene matching scores, motion and environmental context data output from preceding modules to complete the final intent decision. At the same time, it updates the user profile through an interactive closed loop to achieve continuously optimized personalized services, forming a complete closed loop of the technical solution.
[0045] S41. Multi-dimensional confidence fusion and intent decision-making unit: Clarify the fusion rules of multi-dimensional data, the confidence calculation method, and the low confidence processing logic to ensure the accuracy and reliability of intent decision-making and avoid false triggering.
[0046] The system integrates multi-dimensional information from acoustic models, gene matching, motion context, and environmental context to make the final intent decision. The environmental context is obtained by classifying environmental acoustic features, such as wind and rain, noise, and quiet, using a Support Vector Machine (SVM) classifier. The classifier was chosen because SVM exhibits excellent classification performance in scenarios with small samples and high-dimensional features, making it suitable for environmental acoustic feature classification requirements.
[0047] Let the probability distribution of the acoustic model after motion context modulation be... Gene matching scores are normalized using the Softmax function and have the following distribution: This facilitates multi-dimensional data fusion and environmental context. The provided prior intent probability distribution is Based on the preset environment type, for example, the prior probability of safety commands is increased to 0.6 in windy and rainy environments; the value is determined by statistically analyzing the frequency of command usage in different environments, with safety commands accounting for 60% in windy and rainy environments, and the prior probability is set accordingly. The final joint confidence score... It will be integrated in the following ways: Among them, the fusion coefficient The weights are non-negative and satisfy the following conditions: In this embodiment, the default value is The values are determined based on the optimization using a grid search method, which yields the highest overall system recognition accuracy. The parameters are determined dynamically based on the real-time confidence level of each module. The real-time confidence level is calculated using the matching accuracy of the most recent 100 interactions, using the following formula: ,in For the module's real-time accuracy, The average accuracy of the three modules was adjusted and renormalized to maintain a sum of 3. System selection. highest command As the identification result, if the difference between the highest and second-highest scores is less than the safety threshold... (In this embodiment, a typical value of 0.3 is used, with a range of 0.2 to 0.4; the parameter determination method is as follows: adjusted according to the system's false trigger rate requirements, 0.3 is used when the false trigger rate requirement is ≤3%, and 0.4 is used when the requirement is ≤1%, adapted through historical false trigger data statistics), then it is determined to be a low confidence situation, triggering an active clarification interaction based on multimodal context, with the clarification statement being: "Do you want to instruct..." "Is that right?", wait for user confirmation before proceeding.
[0048] S42. Personalized Interaction Closed Loop and Profile Update Unit: Clearly define the storage content of interaction records, profile update rules, and learning cycle to ensure continuous optimization of personalized adaptation capabilities, while also explaining the compliance and security of data processing.
[0049] A complete record of each interaction will be recorded, along with the multi-dimensional context (motion event) at the time of triggering. ,environment ,Location The system associates user IDs with data and stores it in an encrypted cloud database (using AES-256 encryption; the storage period is 6 months, after which it is automatically anonymized). Recorded content includes: original voice segments, sensor data, recognition results, confidence scores, and user feedback (confirmation / negation).
[0050] By analyzing this data, the system continuously optimizes user profiles. The update rules are as follows: (1) Instruction preference: Statistically analyze the high-frequency instructions of users in specific scenarios, adjust the weight of the corresponding gene fragments, and adjust the formula: ,in For instructions (2) Voice features: For every 10 valid interactions, the average fundamental frequency, formant and other parameters of the user are recalibrated. The calibration method is to use the moving average method. The new parameter = 0.7 × old parameter + 0.3 × new sample parameter, balancing stability and adaptability. 0.7 and 0.3 are weights, which are determined by the technicians based on the importance of the old parameters in the actual scenario. (3) Scenario adaptation: Record the recognition accuracy in different environments and optimize the personalized adjustment strategy. The optimization method is to increase the personalized adjustment range in environments with recognition accuracy <80%, such as increasing the fluctuation range of the fundamental frequency scaling coefficient by 20%. The update cycle is a combination of real-time update and daily batch update at dawn.
[0051] This continuous learning loop enables the system to constantly adapt to users' personalized habits, achieving a continuously optimized interactive experience. By deeply binding user profile updates with specific scenarios (location, environment, actions), personalization is no longer a static label, but a dynamic, contextualized behavioral model.
[0052] Through the detailed description of the above embodiments, the two-wheeled vehicle rental voice service system of the present invention, based on user profiles and location awareness, constructs a complete voice interaction technology system through the coordinated operation of four core modules: multi-sensor data fusion, location-aware noise reduction, personalized acoustic gene recognition, and multi-dimensional intent decision-making. The system innovatively integrates vehicle motion status, geographical location, and user profiles, breaking through the technical bottlenecks of traditional voice services in riding scenarios and effectively solving problems such as environmental noise interference, individual voice differences, and insufficient scenario adaptation. Its personalized adaptation capabilities and high-precision recognition performance provide two-wheeled vehicle rental users with a safe and convenient voice interaction experience, promoting the upgrading and iteration of voice service technology in the two-wheeled vehicle rental industry, and has broad application value.
[0053] The above formulas are all dimensionless calculations, and the preset parameters in the formulas should be set by those skilled in the art according to the actual situation.
[0054] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0055] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0056] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0057] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0058] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0059] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0060] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A two-wheeled vehicle rental voice service system based on user profiles and location awareness, characterized in that: It includes a voice enhancement module that fuses multi-sensor data and couples motion states, a location-aware adaptive noise adversarial processing module, an acoustic gene fault-tolerant recognition module that fuses user profiles, and an intent decision-making and personalized service module based on multi-dimensional context. The multi-sensor data fusion and motion state coupling speech enhancement module establishes action-intent associations based on multi-dimensional sensor data collected from the front end through motion micro-motion event detection, providing a physical synchronization benchmark for speech stream segmentation and recognition. The location-aware adaptive noise adversarial processing module constructs a geographic noise fingerprint based on location data and audio signals, selects the optimal speech transmission channel, and improves speech intelligibility. The acoustic gene fault-tolerant recognition module that integrates user profiles constructs a personalized acoustic gene library and uses a fragmented voting matching algorithm to improve recognition robustness. The intent decision and personalized service module based on multi-dimensional context integrates multi-dimensional information to complete intent decisions and updates user profiles through interactive closed loops to achieve personalized services.
2. The two-wheeled vehicle rental voice service system based on user profiles and location awareness according to claim 1, characterized in that, The multi-sensor data fusion and motion state coupling voice enhancement module includes a motion micro-motion event detection unit and a motion event-based voice stream segmentation and context weighting unit. The motion micro-motion event detection unit uses multiple types of onboard sensors to collect multi-dimensional vehicle motion data and detects motion micro-motion events coupled with voice commands by analyzing the joint temporal features of the data. The voice stream segmentation and context weighting unit uses the detected motion micro-motion events as the basis for voice stream segmentation and corrects the recognition probability through motion context weighting to reduce recognition deviation under noise interference.
3. The two-wheeled vehicle rental voice service system based on user profiles and location awareness according to claim 2, characterized in that, The motion micro-motion events include preset event types related to steering, braking, and stable riding. The activation of an event is determined by analyzing the trend and duration of changes in characteristic parameters in the vehicle motion data.
4. The two-wheeled vehicle rental voice service system based on user profiles and location awareness according to claim 1, characterized in that, The location-aware adaptive noise countermeasures module includes a dynamic construction and matching unit for geographic noise fingerprints and an optimal voice transmission channel selection and guidance unit. The dynamic construction and matching unit periodically collects ambient background noise, processes the signal to form a noise fingerprint for the corresponding location, and compares and updates it with a cloud-based reference noise fingerprint. The channel selection and guidance unit searches for the optimal voice transmission channel based on the current noise fingerprint, adjusts the relevant parameters for voice acquisition and recognition, and can guide the user to adjust their vocalization method to adapt to the channel.
5. The two-wheeled vehicle rental voice service system based on user profiles and location awareness according to claim 4, characterized in that, The geographic noise fingerprint is obtained by performing spectral analysis on the noise signal and calculating the long-term average power spectral density. It covers the main frequency bands of human speech and is updated when the vehicle is stationary and there is no user interaction.
6. The two-wheeled vehicle rental voice service system based on user profiles and location awareness according to claim 1, characterized in that, The acoustic gene fault-tolerant recognition module that integrates user profiles includes a command acoustic gene library construction and personalized adaptation unit and a gene fragment voting matching algorithm unit. The construction and personalized adaptation unit extracts key acoustic gene fragments from the core command set to construct a general gene library, dynamically adjusts the general gene fragments in combination with acoustic feature parameters in the user profile, and generates personalized gene templates. The voting matching algorithm unit extracts acoustic feature sequences from the speech to be recognized, calculates similarity with personalized gene fragments, and realizes command recognition through a voting mechanism.
7. The two-wheeled vehicle rental voice service system based on user profiles and location awareness according to claim 6, characterized in that, The acoustic gene fragments serve as templates for characterizing the local stable acoustic features of instructions. They are extracted from pure instruction speech samples from different speakers, with each instruction corresponding to multiple gene fragments, covering the key acoustic features of the instruction.
8. The two-wheeled vehicle rental voice service system based on user profiles and location awareness according to claim 1, characterized in that, The intent decision and personalized service module includes a multi-dimensional confidence fusion and intent decision unit and a personalized interaction closed loop and profile update unit. The confidence fusion and intent decision unit fuses acoustic model probability, gene matching score, and environmental context prior probability to calculate a joint confidence score to determine the recognition result. When the confidence score is low, an active clarification interaction is triggered. The interaction closed loop and profile update unit records the complete data and associated context of each interaction, and continuously optimizes the instruction preferences, voice features, and scene adaptation strategies in the user profile based on this data.