Electric bicycle control method based on multi-mode voice recognition and related equipment

By integrating real-time voice, vehicle status, and environmental information using multimodal speech recognition technology, decision-making information is generated to control the functional components of electric bicycles, solving the problem of low intelligence in electric bicycles and realizing safe and convenient voice control during riding.

CN121506128APending Publication Date: 2026-02-10GUANGDONG YITONG LIANYUN INTELLIGENT INFORMATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511515045.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Electric bicycles have a low level of intelligence in terms of user interaction, and cannot provide a good interactive experience during riding, which requires users to operate them manually, posing safety hazards and causing inconvenience.

Method used

An electric bicycle control method based on multimodal speech recognition is adopted. By detecting real-time speech, vehicle status and environmental information, multimodal fusion is performed to generate decision information and control the functional components of the electric bicycle. This method includes feature extraction, weighted fusion, reinforcement learning and safety verification.

Benefits of technology

This technology enables voice control of electric bicycles during riding, enhancing their intelligence, improving the riding experience, and increasing their adaptability and safety in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506128A_ABST
    Figure CN121506128A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an electric bicycle control method and related equipment based on multi-modal speech recognition, and the method controls the functional parts of an electric bicycle through the control information generated through multi-modal fusion, and can improve the adaptability of the control of the electric bicycle to a complex environment. A complete technical scheme is provided for intellectualization of the electric bicycle; by executing the electric bicycle control method based on multi-mode voice recognition, a user can express the control intention of parts of the electric bicycle in a voice mode in the riding process, and control information is generated by automatically combining the real-time bicycle state and the real-time environment of the electric bicycle; therefore, on the basis of realizing voice control, facilitating man-bicycle interaction and improving the riding experience, the control of the parts of the electric bicycle is matched and coordinated with the real-time bicycle state and the real-time environment of the electric bicycle, and the intelligent degree of the electric bicycle is improved. The method is widely applied to the technical field of battery management.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of battery management, and particularly to an electric bicycle control method based on multi-modal speech recognition and related equipment. BACKGROUND

[0002] Electric bicycles have become an important tool for short-distance travel, and have made a great contribution to solving the problem of people's travel.

[0003] At present, the intelligence level of electric bicycles in user interaction is relatively low. For example, some electric bicycles are provided with control chips, but only have basic functions such as motor driving and battery management, and users still need to rely on traditional interaction methods such as manual operation to control speed and light control, and the interaction method is fragmented, which is difficult to provide users with a good use experience. Moreover, the current electric bicycle cannot provide good interaction during the user's riding of the electric bicycle, and the user needs to manually operate when using navigation, vehicle condition query, speed adjustment and other functions, which has safety hazards if operated during riding, and it is inconvenient to delay the journey when parking. SUMMARY

[0004] In view of the technical problem that the intelligence level of the current electric bicycle in user interaction is low, the purpose of the embodiments of the present application is to provide an electric bicycle control method based on multi-modal speech recognition and related equipment.

[0005] In one aspect, the embodiments of the present application include an electric bicycle control method based on multi-modal speech recognition, which comprises the following steps: detecting real-time voice information, real-time vehicle state information and real-time environment information of an electric bicycle; performing multi-modal fusion according to the real-time voice information, the real-time vehicle state information and the real-time environment information to obtain fusion feature information; generating decision information according to the fusion feature information; generating control information according to the decision information; controlling the functional components of the electric bicycle according to the control information.

[0006] Further, the multi-modal fusion according to the real-time voice information, the real-time vehicle state information and the real-time environment information to obtain fusion feature information comprises: extracting features from the real-time voice information to obtain voice feature information; extracting features from the real-time vehicle state information to obtain vehicle state feature information; characteristic extraction is performed on the real-time environment information to obtain environment characteristic information; The voice characteristic information, the vehicle state characteristic information, and the environment characteristic information are weighted and fused to obtain the fusion characteristic information.

[0007] Further, the voice characteristic information, the vehicle state characteristic information, and the environment characteristic information are weighted and fused to obtain the fusion characteristic information, including: The first weight, the second weight, and the third weight are dynamically set; the first weight is a weight corresponding to the voice characteristic information, the second weight is a weight corresponding to the vehicle state characteristic information, and the third weight is a weight corresponding to the environment characteristic information; According to the first weight, the second weight, and the third weight, the voice characteristic information, the vehicle state characteristic information, and the environment characteristic information are weighted and fused to obtain the fusion characteristic information.

[0008] Further, the first weight, the second weight, and the third weight are dynamically set, including: A first change rate is calculated by tracking changes of the vehicle state characteristic information over time; A second change rate is calculated by tracking changes of the environment characteristic information over time; When the first change rate and the second change rate are both less than a threshold value, the first weight is set to a first value; When the first change rate and / or the second change rate is greater than or equal to a threshold value, the first weight is set to a second value; the second value is greater than the first value; The sum of the first weight, the second weight, and the third weight is set to a constant value.

[0009] Further, the decision information is generated based on the fusion characteristic information, including: Expert knowledge of electric bicycle driving is obtained; The expert knowledge is coded into deterministic rules; A reinforcement learning agent is established; The fusion characteristic information is converted into a state vector; The state vector is input into the reinforcement learning agent for processing to obtain action information output by the reinforcement learning agent; The action information is constrained and processed according to the deterministic rules to obtain the decision information.

[0010] Further, the electric bicycle control method based on multi-modal speech recognition further includes: Collecting execution result data of the function components of the electric bicycle controlled according to the control information; According to the execution result data, updating the model parameters of the reinforcement learning agent.

[0011] Further, the generating control information according to the decision information comprises: Obtaining performance limit information of the electric bicycle; According to the performance limit information, performing safety verification on the decision information; When the safety verification passes, converting the decision information into the control information.

[0012] Further, the generating control information according to the decision information further comprises: When the safety verification fails, performing the following steps: Obtaining preset alternative information, and converting the alternative information into the control information; Generating abnormal prompt information, and pushing the abnormal prompt information; According to the decision information, updating the deterministic rule.

[0013] In another aspect, the embodiments of the present application also include a computer device comprising a memory and a processor, the memory being used to store at least one program, and the processor being used to load the at least one program to execute the electric bicycle control method based on multi-modal speech recognition in the embodiments.

[0014] In another aspect, the embodiments of the present application also include a computer program product comprising a computer program, which, when executed by a processor, implements the electric bicycle control method based on multi-modal speech recognition in the embodiments.

[0015] The embodiment of the present application has the following advantages: the electric bicycle control method based on multi-modal speech recognition in the embodiment realizes a multi-modal fusion three-dimensional decision model, fuses three dimensions of real-time voice information, real-time vehicle state information and real-time environmental information, and closely combines the hardware characteristics in algorithm design. The control information generated by multi-modal fusion is used to control the functional components of the electric bicycle, which can improve the adaptability of electric bicycle control to complex environments and provide a complete technical solution for electric bicycle intelligence. By executing the electric bicycle control method based on multi-modal speech recognition, the user can express the control intention of the components of the electric bicycle through voice during riding, and automatically generate control information combined with the real-time vehicle state and real-time environment of the electric bicycle, so as to realize voice control, facilitate human-vehicle interaction, improve the riding experience, and make the control of the components of the electric bicycle match and coordinate with the real-time vehicle state and real-time environment of the electric bicycle, which is conducive to improving the intelligence of the electric bicycle. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 A schematic diagram of a hardware platform to which the electric bicycle control method based on multi-modal speech recognition in the embodiment can be applied; Figure 2 A schematic diagram of the steps of the electric bicycle control method based on multi-modal speech recognition in the embodiment; Figure 3 A schematic diagram of the overall flow of the electric bicycle control method based on multi-modal speech recognition in the embodiment; Figure 4 And Figure 5 A schematic diagram of the flow of safety verification in the embodiment. DETAILED DESCRIPTION

[0017] The electric bicycle in the embodiment mainly refers to a two-wheeled or three-wheeled vehicle driven by electricity, also known as an electric bicycle. Specifically, it can be an electric bicycle purchased and used by a user personally, or an electric bicycle purchased and operated by an enterprise, which is rented or shared by users in the mode of "shared bicycle".

[0018] In the embodiment, the electric bicycle control method based on multi-modal speech recognition can be applied to the hardware platform shown in Figure 1 Referring to Figure 1 , the hardware platform includes the following components: • CPU module: uses Allwinner V3S chip (1GHz dual-core ARM Cortex-A7), integrates 128MB DDR3 memory, supports edge AI inference, and has a computing power of 0.5TOPS; • Speech processing module: Main chip: WT2003HX voice dedicated chip, supporting 8kHz~16kHz sampling rate, built-in hardware noise reduction unit; Microphone array: dual MEMS microphone, 8cm apart, achieving ±30° sound source positioning, suppressing environmental noise (SNR improved by 15dB); • RTK positioning module: Module: Dream Chip MXT906B, supporting Beidou + GPS dual mode, achieving centimeter-level positioning through differential technology (horizontal accuracy ±1cm, vertical accuracy ±2cm); Reference station communication: supports 433MHz wireless reception of reference station differential data, with a coverage radius of 10km; • Communication module: 4G module: China Mobile ML307, supporting LTE Cat.1, with an uplink speed of 5Mbps and a downlink speed of 10Mbps; Bluetooth module: BK3633, supporting BLE 5.0, used for mobile phone network configuration and firmware upgrade; • Sensor interface: RJ45 interface: connects external camera (1080P, 30fps), used for abnormal situation video collection; RS485 interface: access vehicle speed sensor (accuracy ±0.5km / h), power sensor (sampling rate 10Hz).

[0019] Figure 1 Among the hardware platforms shown, the part other than the cloud AI service platform can be installed on the electric bicycle, and the cloud AI service platform can be run by the manufacturer of the electric bicycle or a dedicated service provider. The electric bicycle establishes a connection with the cloud AI service platform through the 4G network with its 4G module.

[0020] Before each use of the electric bicycle, the hardware platform can be system started and initialized before executing the electric bicycle control method based on multi-modal speech recognition. Specifically, the system startup and initialization includes the following processes: 1. After the electric bicycle is powered on, the power module (5V / 3V voltage stabilizer + backup battery) starts, and the CPU module loads the Linux operating system (kernel version 4.19); 2. The voice processing module initializes the microphone array and starts the noise reduction preprocessing thread; 3. The RTK module searches for satellite signals (cold start time <30s), and the 4G module connects to the operator network at the same time; 4. The controller broadcasts device information through Bluetooth, and waits for the mobile phone APP to configure the network (first use needs to scan the code to bind).

[0021] After the system startup and initialization of the hardware platform are completed, the electric bicycle control method based on multi-modal speech recognition is executed.

[0022] In this embodiment, with reference to Figure 2 The electric bicycle control method based on multi-modal speech recognition comprises the following steps: S1. detecting real-time speech information, real-time vehicle state information of the electric bicycle, and real-time environment information; S2. performing multi-modal fusion according to the real-time speech information, the real-time vehicle state information, and the real-time environment information to obtain fusion feature information; S3. generating decision information according to the fusion feature information; S4. generating control information according to the decision information; S5. controlling the functional components of the electric bicycle according to the control information.

[0023] In this embodiment, each step in the electric bicycle control method based on multi-modal speech recognition can be executed by the CPU module in the hardware platform, including steps S1-S5. The CPU module can call or control other components when executing a specific step.

[0024] In this embodiment, the CPU module can execute steps S1-S5 in a loop. For example, the CPU module can first execute a round of steps S1-S5, in which the CPU module executes steps S2-S4 to process the real-time speech information, the real-time vehicle state information of the electric bicycle, and the real-time environment information obtained by executing step S1, generates control information, and finally executes step S5 to control the functional components of the electric bicycle according to the control information, thereby completing the execution of this round of steps S1-S5; then, the next round of steps S1-S5 is triggered to detect new real-time speech information, real-time vehicle state information of the electric bicycle, and real-time environment information, generate new control information to control the functional components of the electric bicycle. Since the execution speed of the CPU module is fast enough, the CPU module can continuously detect real-time speech information, real-time vehicle state information of the electric bicycle, and real-time environment information, and control the functional components of the electric bicycle in real time, thereby realizing real-time response to the speech spoken by the user, the vehicle state of the electric bicycle, and the environment in which the electric bicycle is located.

[0025] In this embodiment, one round of steps S1-S5 is taken as an example for illustration. The overall flow of steps S1-S5 is shown in Figure 3 .

[0026] In step S1, the CPU module calls the microphone array to record audio. The sound collected by the microphone array may include the voice of the user driving the electric bicycle, the noise of the environment where the electric bicycle is located (including the voice of others talking, the sound of cars driving on the road, and wind noise, etc.). The microphone array converts the collected sound into raw sound data that can be recognized by the computer and sends the raw sound data to the voice processing module.

[0027] In this embodiment, the speech processing module runs a speech interaction engine. The speech interaction engine includes an acoustic model and a language model, and its specific principles include: Acoustic model: An end-to-end ASR (Automatic Speech Recognition) model based on the Transformer architecture, pre-trained using the LibriSpeech dataset, fine-tuned with 1000 hours of road test data for electric bicycle scenarios (engine noise, wind noise), with a word error rate (WER) of <8%; the acoustic model can convert raw sound data into text form representing sound content; Language Model: Lightweight LLM (500M parameters), supports multi-turn dialogue and contextual understanding. The language model can perform semantic understanding on raw audio data in text form; for example, a parsing process performed by the language model is as follows: # Instruction parsing example def parse_command(context, speech): If "navigation" and "home" are both in speech: return { "action": "navigate", "target": "home", "context": context.last_route } elif "speed up" in speech and context.speed<20: return {"action": "speed_up", "value": 5} In step S1, the CPU module calls sensors such as vehicle speed and battery level sensors to detect and obtain status data such as vehicle speed and battery level as real-time vehicle status information.

[0028] In step S1, the CPU module can call the RTK positioning module to locate the electric bicycle and obtain its current location information. The CPU module communicates with the cloud AI service platform through the 4G module, sending the current location information to the cloud AI service platform to query information such as weather and road conditions at the electric bicycle's location as real-time environmental information.

[0029] In this embodiment, when the CPU module calls the RTK positioning module for positioning, it can execute an RTK / GPS dynamic switching strategy, intelligently selecting the positioning mode based on vehicle speed and signal quality. Specifically, as follows... Figure 3 As shown, the dynamic switching strategy between RTK and GPS is as follows: When the vehicle speed is greater than 10km / h or the RTK signal is lost, it will automatically switch to GPS (accuracy of 10 meters). When the vehicle speed is ≤10km / h and the RTK signal is good, RTK centimeter-level positioning is enabled for precise parking of charging stations.

[0030] By implementing a dynamic RTK / GPS switching strategy, positioning accuracy can be improved, and the need for precise positioning can be met even in urban canyons, in severe weather, and other conditions.

[0031] By executing step S1, the CPU module obtains the real-time voice information, real-time vehicle status information, and real-time environmental information to be processed in this round of steps S1-S5.

[0032] Steps S2-S4 belong to the steps of the multimodal fusion algorithm. That is, the CPU module implements the three-dimensional decision model based on the multimodal fusion algorithm by executing steps S2-S4.

[0033] In this embodiment, when performing step S2, which involves multimodal fusion based on real-time voice information, real-time vehicle status information, and real-time environmental information to obtain fused feature information, the following steps can be specifically performed: S201. Extract features from real-time speech information to obtain speech feature information; S202. Extract features from real-time vehicle status information to obtain vehicle status feature information; S203. Extract features from real-time environmental information to obtain environmental feature information; S204. Weighted fusion of voice feature information, vehicle status feature information and environmental feature information is performed to obtain fused feature information.

[0034] Before executing steps S201-S204, the CPU module can preprocess the real-time voice information, real-time vehicle status information, and real-time environmental information obtained in step S1. Specifically, the CPU module can perform text cleaning, word segmentation, and part-of-speech tagging on the real-time voice information; perform outlier filtering (3σ principle) and normalization on the real-time vehicle status information; and perform format unification and timestamp alignment on the real-time environmental information. In this embodiment, the real-time voice information, real-time vehicle status information, and real-time environmental information processed in steps S201-S204 are preprocessed real-time voice information, real-time vehicle status information, and real-time environmental information.

[0035] In step S201, the CPU module can extract intent vectors and emotion features from real-time speech information through LLM encoding to obtain speech feature information. .

[0036] In step S202, the CPU module can encode real-time vehicle status information into the form of State of Charge (SOC), Power Status, and Fault Codes to obtain vehicle status characteristic information. .

[0037] In step S203, the CPU module can encode real-time environmental information into the form of location features, road condition features, and weather features to obtain environmental feature information. .

[0038] In this embodiment, when performing step S204, which involves weighted fusion of voice feature information, vehicle state feature information, and environmental feature information to obtain fused feature information, the following steps can be specifically performed: S20401. Dynamically set the first weight, second weight, and third weight; S20402. Based on the first weight, the second weight, and the third weight, the speech feature information, vehicle state feature information, and environmental feature information are weighted respectively, and then concatenated to obtain the fused feature information.

[0039] In step S20401, voice feature information can be set. The corresponding weight is the first weight. Vehicle status characteristic information The corresponding weight is the second weight. Environmental characteristic information The corresponding weight is the third weight. .

[0040] Specifically, when executing step S20401, the first method can be used: setting each weight according to a fixed size, that is, setting the first weight... Second weight and third weight Each is set to a fixed size.

[0041] Specifically, when performing step S20401, a second method can also be used: first, process the speech feature information... Vehicle status characteristic information and environmental characteristic information Feature concatenation is performed to obtain the concatenated vector. In other words:

[0042] Transformer is used to concatenate vectors Perform attention weight calculation to obtain the weight vector. That is

[0043] From the weight vector Extract the first weight Second weight and third weight .

[0044] Specifically, during step S20401, a third method can also be performed: calculating vehicle state feature information. The corresponding first rate of change, for example, can be calculated by first calculating the vehicle state feature information obtained in step S2 of this round. The vehicle state feature information obtained in step S2 of the previous round Differences between Obtain the time difference between the current execution step S2 and the previous execution step S2. ,calculate As the first rate of change; similarly, the environmental feature information obtained in step S2 of this round can be calculated first. Compared with the environmental feature information obtained in step S2 of the previous round. Differences between Obtain the time difference between the current execution step S2 and the previous execution step S2. ,calculate As the second rate of change; a first threshold can be set to evaluate the magnitude of the first rate of change, and a second threshold can be set to evaluate the magnitude of the second rate of change. When the first rate of change is less than the first threshold and the second rate of change is less than the second threshold, that is, when both the first and second rates of change are small, the first weight is... Set a smaller first value (e.g., a value between 0 and 0.2); conversely, when the first rate of change is greater than or equal to the first threshold, and / or the second rate of change is greater than or equal to the second threshold, i.e., at least one of the first and second rates of change is larger, the first weight is adjusted. Set a larger second value (e.g., a value between 0.8 and 1); and set a first weight. Second weight and third weight The sum is a constant value, for example, set to + + =1 A second weight can also be set. With the third weight They are equal, thus determining the first weight. After obtaining the value, the second weight can be further calculated. and third weight Their respective values.

[0045] In this embodiment, the principle of executing step S20401 in the third way is based on: vehicle state feature information Environmental characteristic information These respectively represent the state of the electric bicycle itself and the state of its surrounding environment during the user's operation of the electric bicycle; that is, both represent external states that affect the interactive operation of driving the electric bicycle. When vehicle state characteristic information... The smaller first rate of change indicates that the electric bicycle's speed and battery level are stable, meaning the electric bicycle is in a stable operating state. Similarly, environmental characteristic information... The smaller second rate of change also indicates that the electric bicycle is in a stable driving state, and the user's need for voice interaction is low. Therefore, the voice feature information... The corresponding first weight Setting a smaller first value can reduce the impact of sound collected by the microphone array on the final control information, thereby reducing interference from environmental noise collected by the microphone array on the control of this electric bicycle. This aligns with the characteristics of electric bicycles, which are characterized by an open driving environment and susceptibility to wind noise, the voices of bystanders, and the sounds of other vehicles. Conversely, when the vehicle status characteristic information... A larger first rate of change indicates a significant change in the electric bicycle's speed or battery level, meaning a substantial shift in the bicycle's operating status. Similarly, environmental characteristic information... The larger corresponding second rate of change also indicates a significant change in the electric bicycle's driving status, suggesting a high demand for voice interaction from the user. Therefore, the voice feature information... The corresponding first weight Setting it to a larger second value can extract speech feature information. More of the data is transferred to the final control information to control the functional components of the electric bicycle, thereby promptly meeting the user's high demands for voice interaction.

[0046] In step S20401, the first weight is obtained through the first, second, or third method. Second weight and third weight Next, step S20402 can be executed to map different modal features to the same vector space. That is, according to the first weight, the second weight, and the third weight, the speech feature information, vehicle state feature information, and environmental feature information are weighted respectively, and then concatenated to obtain the fused feature information.

[0047] Specifically, during step S20402, the following steps can be performed for weighted fusion:

[0048]

[0049]

[0050]

[0051] Thus, fused feature information is obtained. Fusion of feature information It integrates features from multiple modalities, including speech features, vehicle status features, and environmental features.

[0052] In this embodiment, when performing step S3, which is to generate decision information based on the fused feature information, the following steps can be specifically performed: S301. Obtain expert knowledge on electric bicycle driving; S302. Encode expert knowledge into deterministic rules; S303. Establish a reinforcement learning agent; S304. Convert the fused feature information into a state vector; S305. Input the state vector into the reinforcement learning agent for processing to obtain the action information output by the reinforcement learning agent; S306. Based on deterministic rules, constrain the action information to obtain decision information.

[0053] In step S301, the expert knowledge on electric bicycle driving to be obtained can be expert knowledge edited by the electric bicycle manufacturer or expert knowledge edited by the traffic management department.

[0054] In step S302, the expert knowledge obtained in step S301 is encoded into the form of deterministic rules applicable to reinforcement learning.

[0055] In step S303, a reinforcement learning agent (RL agent) is established. The RL agent can specifically take the form of a neural network.

[0056] In step S304, the fused feature information obtained in step S2 is executed. It is converted into a state vector form that can be processed by a reinforcement learning agent (RL agent).

[0057] In step S305, the fused feature information is... The transformed state vector is input to the reinforcement learning agent (RL Agent) for processing. The RL Agent outputs action information, which is represented by the fused feature information. In the context of including voice feature information, vehicle status feature information, and environmental feature information, the actions that functional components on an electric bicycle need to perform include: actions that the voice feature information commands functional components on the electric bicycle to perform (e.g., if the voice feature information is "drive faster," then the action information could be "motor accelerate"); actions that satisfy the vehicle status indicated by the vehicle status feature information (e.g., if the vehicle status feature information is "low battery," then the action information could be "motor speed limit"); and actions that satisfy the environmental features indicated by the environmental feature information (e.g., if the environmental feature information is "foggy weather," then the action information could be "motor speed limit").

[0058] In step S306, the parameters such as the motion amplitude of the motion information are constrained according to deterministic rules to obtain decision information. That is, in this embodiment, the decision information is the motion information with parameters such as the motion amplitude constrained.

[0059] In this embodiment, when performing step S4, which is the step of generating control information based on decision information, the following steps can be specifically performed: S401. Obtain performance limitation information for electric bicycles; S402. Perform security verification on decision information based on performance limitation information; S403. Once the security check passes, the decision information is converted into control information; S404. When the security check fails, perform the following steps: Obtain preset alternative solution information and convert the alternative solution information into control information; Generate and push error message alerts; Update the deterministic rules based on the decision information.

[0060] In this embodiment, the process of steps S401-S404 is as follows: Figure 4 As shown.

[0061] In step S401, the performance limitation information of the electric bicycle can be either constant information such as "maximum motor speed" or variable information such as "battery charge".

[0062] In step S402, the decision information is subjected to a safety verification based on the performance limit information to determine whether the decision information is within the safe range corresponding to the performance limit information. For example, if the content of the decision information is "motor accelerates to 80% of maximum speed", and the content of the performance limit information is "medium battery level (between 30% and 80%)", then it can be determined that the decision information is within the safe range corresponding to the performance limit information, and thus the safety verification is passed; if the content of the decision information is "motor accelerates to 80% of maximum speed", and the content of the performance limit information is "low battery level (below 20%)", then it can be determined that the decision information is outside the safe range corresponding to the performance limit information, and thus the safety verification is failed.

[0063] In this embodiment, the process of steps S403-S404 is as follows: Figure 5 As shown.

[0064] Reference Figure 4 and Figure 5 If the security verification of the decision information passes, step S403 is executed to convert the decision information into control information. Specifically, the decision information can be directly used as control information, or the decision information can be converted into control information through format conversion or other operations.

[0065] Reference Figure 4 and Figure 5If the security verification of the decision information fails, step S404 is executed to obtain preset alternative solution information and convert it into control information. Specifically, the preset alternative solution information may include the message "find the nearest charging station location and make a recommendation." Simultaneously, the CPU module can generate an anomaly alert and push it to the cloud-based AI service platform. The CPU module can also update deterministic rules based on the decision information that failed the security verification. For example, the updated deterministic rules can impose greater constraints on the action information during subsequent executions of step S306, thereby reducing the likelihood of newly obtained decision information failing the security verification, increasing the probability of obtaining decision information that can be converted into control information, and improving the real-time performance of electric bicycle control.

[0066] In this embodiment, a simplified code for the multimodal fusion algorithm executed in steps S2-S4 is shown below: import torchfrom transformers import AutoModel class MultiModalFusion(torch.nn.Module): def __init__(self): super(MultiModalFusion, self).__init__() # Voice intent encoder to obtain real-time voice information self.text_encoder = AutoModel.from_pretrained("bert-base-chinese") # Vehicle status encoder, to obtain real-time vehicle status information self.vehicle_encoder = torch.nn.Sequential( torch.nn.Linear(5, 32), # Input: 5 features including vehicle speed and battery level torch.nn.ReLU(), torch.nn.Linear(32, 64) ) # Environment encoder, to obtain real-time environmental information self.environment_encoder = torch.nn.Sequential( torch.nn.Linear(8, 32), # Input: 8 features including location and weather torch.nn.ReLU(), torch.nn.Linear(32, 64) ) # Attention Fusion Layer self.attention = torch.nn.Sequential( torch.nn.Linear(192, 3), # Real-time voice information, real-time vehicle status information, and real-time environmental information (three modalities) torch.nn.Softmax(dim=1) ) # Decision-making level self.decision_maker = torch.nn.Sequential( torch.nn.Linear(192, 128), torch.nn.ReLU(), torch.nn.Linear(128, 32), torch.nn.ReLU(), torch.nn.Linear(32, 5) # Outputs 5 types of control information ) def forward(self, text_input, vehicle_input, env_input): # Feature Extraction text_features = self.text_encoder(**text_input).pooler_output vehicle_features = self.vehicle_encoder(vehicle_input) env_features = self.environment_encoder(env_input) # Feature splicing concat_features = torch.cat([text_features, vehicle_features, env_features], dim=1) # Attention Weight Calculation attn_weights = self.attention(concat_features) # Weighted fusion weighted_text = text_features * attn_weights[:, 0:1] weighted_vehicle = vehicle_features * attn_weights[:, 1:2] weighted_env = env_features * attn_weights[:, 2:3] # Fusion Features fused_features = torch.cat([weighted_text, weighted_vehicle,weighted_env], dim=1) # Decision Generation control_signal = self.decision_maker(fused_features) return control_signal When steps S401-S403 are executed, the control information obtained by converting the decision information that has passed the safety verification is obtained. When step S5 is executed, the CPU module can send the control information to the functional components pointed to by the control information, including the motor, trunk, headlights, etc., so as to control these functional components to perform corresponding actions, such as motor acceleration and deceleration, trunk opening or closing, headlight switching or brightness adjustment, etc.

[0067] In this embodiment, a typical scenario for executing steps S1-S5 is shown below: • Complex command processing: The user says, "Navigate to the company and adjust the speed to 20km / h," and the system responds: a. Analyzing the dual intent of "navigation" and "speed adjustment"; b. Use RTK to locate the current position and plan a route; c. Check the battery level (if ≥30%), send a PWM signal to the motor controller, and stabilize the vehicle speed at 20km / h; d. The voice feedback says, "You have been navigated to the company. Current speed is 20km / h."

[0068] • Abnormal situation handling: When the battery level is detected to be <15% and the user issues an "accelerate" command, the system: e. Reject the acceleration command and announce "Low battery, we suggest switching to power saving mode"; f. Automatically limits the motor power to 50% of its rated power; g. Mark the nearest charging station on the navigation screen.

[0069] In this embodiment, a multimodal fusion three-dimensional decision model is realized by executing steps S1-S5. The multimodal fusion three-dimensional decision model uses a "Transformer + cross-modal attention mechanism" as its core architecture, integrating real-time voice information, real-time vehicle status information, and real-time environmental information. The algorithm design is closely integrated with hardware characteristics, using control information generated through multimodal fusion to control the functional components of the electric bicycle. This improves the adaptability of electric bicycle control to complex environments and provides a complete technical solution for the intelligentization of electric bicycles. Based on this, the specific effects of the electric bicycle control method based on multimodal voice recognition are reflected in: 1. Multimodal interaction integration: By integrating speech recognition with AI big data models, it enables real-time parsing and execution of composite commands such as "voice + vehicle status + environment". Combined with a three-dimensional decision model of speech semantics, vehicle status and environmental data, it solves the limitations of single threshold control. 2. High-precision positioning and noise reduction: The positioning mode is dynamically switched according to vehicle speed and signal quality to realize a collaborative mechanism of "GPS as a backup in high-speed scenarios + RTK precision in low-speed scenarios". Combined with RTK differential positioning technology and adaptive noise reduction algorithm, the positioning drift and voice misrecognition problems in complex urban environments are solved. 3. Cross-vehicle compatibility design: Through built-in sensors (non-vehicle computer communication) and standardized interfaces, an independent sensor acquisition architecture that eliminates the need for a vehicle computer is achieved, enabling non-destructive installation across all vehicle models; 4. Technological Integration Innovation: For the first time, RTK centimeter-level positioning, multi-command voice recognition, and a lightweight AI large model are integrated into the e-bike controller, improving functional integration by 300% compared to traditional solutions; 5. Breakthrough in positioning accuracy: RTK differential technology improves positioning accuracy from the 10-meter level of ordinary GPS to the centimeter level, meeting the needs of scenarios such as automatic docking of charging piles; 6. Improved noise resistance: The dual-microphone array combined with the adaptive spectral reduction algorithm still achieves a speech recognition accuracy of 92% in an 80dB noise environment (compared to only 65% ​​for traditional solutions). 7. Cross-vehicle compatibility: With built-in sensors (not onboard computer communication), it is compatible with more than 95% of electric bicycle models, reducing modification costs by 60% (existing solutions are only compatible with specific models). 8. Real-time advantage: Edge AI inference latency <100ms, which is more suitable for cycling safety control than cloud-based solutions (latency >500ms).

[0070] In this embodiment, refer to Figure 4 In addition to executing steps S1-S5, the CPU module can also execute the following steps: S6. Collect execution result data of functional components controlling the electric bicycle based on control information; S7. Update the model parameters of the reinforcement learning agent based on the execution result data.

[0071] By executing steps S6-S7, incremental learning can be achieved for the reinforcement learning agent, thereby optimizing the agent's decision-making strategy.

[0072] In this embodiment, a computer device can be used, including a memory and a processor. The memory is used to store at least one program, and the processor is used to load at least one program to execute the electric bicycle control method based on multimodal speech recognition, thereby obtaining the effect of the electric bicycle control method based on multimodal speech recognition.

[0073] In this embodiment, a computer program product, including a computer program, can be used. When the computer program is executed by a processor, it implements the electric bicycle control method based on multimodal speech recognition in this embodiment.

[0074] It should be noted that, unless otherwise specified, when a feature is referred to as "fixed" or "connected" to another feature, it can be directly fixed or connected to the other feature, or indirectly fixed or connected to the other feature. Furthermore, the descriptions of "upper," "lower," "left," and "right" used in this disclosure are only relative to the relative positional relationships of the components of this disclosure in the accompanying drawings. The singular forms "a" and "the" used in this disclosure are also intended to include the plural forms, unless the context clearly indicates otherwise. Moreover, unless otherwise defined, all technical and scientific terms used in this embodiment have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this embodiment specification is only for describing specific embodiments and is not intended to limit the embodiments of the invention. The term "and / or" as used in this embodiment includes any combination of one or more of the associated listed items.

[0075] It should be understood that although the terms first, second, third, etc., may be used to describe various elements in this disclosure, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, a first element may also be referred to as a second element without departing from the scope of this disclosure, and similarly, a second element may also be referred to as a first element. The use of any and all instances or exemplary language (“e.g.,” “such as,” etc.) provided in this embodiment is intended only to better illustrate embodiments of the invention and, unless otherwise required, does not impose a limitation on the scope of embodiments of the invention.

[0076] It should be recognized that embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable storage medium. The method can be implemented using standard programming techniques—including a non-transitory computer-readable storage medium configured with a computer program, wherein such a storage medium causes the computer to operate in a specific and predefined manner—according to the methods and drawings described in the specific embodiments. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. Furthermore, for this purpose, the program can run on a programmed application-specific integrated circuit (ASIC).

[0077] Furthermore, the procedures described in this embodiment can be performed in any suitable order unless otherwise indicated by this embodiment or otherwise obviously contradict the context. The procedures (or variations and / or combinations thereof) described in this embodiment can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented by hardware or a combination thereof as code (e.g., executable instructions, one or more computer programs, or one or more applications) that commonly executes on one or more processors. A computer program includes a plurality of instructions executable by one or more processors.

[0078] Furthermore, the method can be implemented in any suitable type of computing platform, including but not limited to personal computers, minicomputers, mainframes, workstations, networked or distributed computing environments, standalone or integrated computer platforms, or in communication with charged particle tools or other imaging devices, etc. Aspects of embodiments of the invention can be implemented as machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it is readable by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the processes described herein. Furthermore, the machine-readable code, or portions thereof, can be transmitted via wired or wireless networks. The invention of this embodiment includes these and other different types of non-transitory computer-readable storage media when such media comprises instructions or programs that implement the steps above in conjunction with a microprocessor or other data processor. Embodiments of the invention also include the computer itself when programmed according to the methods and techniques of embodiments of the invention.

[0079] A computer program can be applied to input data to perform the functions of this embodiment, thereby transforming the input data to generate output data stored in non-volatile memory. The output information can also be applied to one or more output devices, such as a display. In a preferred embodiment of the invention, the transformed data represents physical and tangible objects, including a specific visual depiction of physical and tangible objects generated on the display.

[0080] The above are merely preferred embodiments of the present invention. The embodiments of the present invention are not limited to the above-described implementations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the embodiments of the present invention, as long as they achieve the same technical effects, should be included within the scope of protection of the embodiments of the present invention. Within the scope of protection of the embodiments of the present invention, the technical solutions and / or implementation methods can have various modifications and variations.

Claims

1. A method for controlling an electric bicycle based on multimodal speech recognition, characterized in that, The electric bicycle control method based on multimodal speech recognition includes: It detects real-time voice information, real-time vehicle status information of electric bicycles, and real-time environmental information. Multimodal fusion is performed based on the real-time voice information, the real-time vehicle status information, and the real-time environmental information to obtain fused feature information; Decision information is generated based on the fused feature information; Based on the decision information, control information is generated; The functional components of the electric bicycle are controlled according to the control information.

2. The electric bicycle control method based on multimodal speech recognition according to claim 1, characterized in that, The step of performing multimodal fusion based on the real-time voice information, the real-time vehicle status information, and the real-time environmental information to obtain fused feature information includes: The real-time voice information is subjected to feature extraction to obtain voice feature information; Feature extraction is performed on the real-time vehicle status information to obtain vehicle status feature information; Feature extraction is performed on the real-time environmental information to obtain environmental feature information; The voice feature information, the vehicle state feature information, and the environmental feature information are weighted and fused to obtain the fused feature information.

3. The electric bicycle control method based on multimodal speech recognition according to claim 2, characterized in that, The weighted fusion of the voice feature information, the vehicle state feature information, and the environmental feature information to obtain the fused feature information includes: The first weight, the second weight, and the third weight are dynamically set; the first weight is the weight corresponding to the voice feature information, the second weight is the weight corresponding to the vehicle state feature information, and the third weight is the weight corresponding to the environmental feature information. The fused feature information is obtained by weighting the speech feature information, the vehicle state feature information, and the environmental feature information according to the first weight, the second weight, and the third weight, respectively, and then concatenating them.

4. The electric bicycle control method based on multimodal speech recognition according to claim 3, characterized in that, The dynamic setting of the first weight, the second weight, and the third weight includes: The first rate of change is calculated by tracking the changes in the vehicle state feature information over time; The second rate of change is calculated by tracking the changes in the environmental feature information over time; When both the first rate of change and the second rate of change are less than the threshold, the first weight is set to the first value; When the first rate of change and / or the second rate of change is greater than or equal to a threshold, the first weight is set to a second value; the second value is greater than the first value. The sum of the first weight, the second weight, and the third weight is set to a constant value.

5. The electric bicycle control method based on multimodal speech recognition according to claim 1, characterized in that, The step of generating decision information based on the fused feature information includes: Gain expert knowledge on riding electric bicycles; The expert knowledge is encoded into deterministic rules; Establish reinforcement learning agents; The fused feature information is converted into a state vector; The state vector is input into the reinforcement learning agent for processing to obtain the action information output by the reinforcement learning agent. The action information is constrained according to the deterministic rules to obtain the decision information.

6. The electric bicycle control method based on multimodal speech recognition according to claim 5, characterized in that, The electric bicycle control method based on multimodal speech recognition also includes: Collect execution result data of the functional components of the electric bicycle controlled according to the control information; The model parameters of the reinforcement learning agent are updated based on the execution result data.

7. The electric bicycle control method based on multimodal speech recognition according to claim 1, characterized in that, The step of generating control information based on the decision information includes: Obtain performance limitation information for electric bicycles; The decision information is security-verified based on the performance limitation information; Once the security verification passes, the decision information is converted into the control information.

8. The electric bicycle control method based on multimodal speech recognition according to claim 7, characterized in that, The step of generating control information based on the decision information further includes: If the security check fails, perform the following steps: Obtain preset alternative solution information and convert the alternative solution information into the control information; Generate and push out the error message; The deterministic rule is updated based on the decision information.

9. A computer device, characterized in that, The device includes a memory and a processor, wherein the memory is used to store at least one program, and the processor is used to load at least one program to execute the electric bicycle control method based on multimodal speech recognition as described in any one of claims 1-8.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the electric bicycle control method based on multimodal speech recognition as described in any one of claims 1-8.