Voice interaction control method and system for electric scooter

By using a microphone array and vehicle sensors to locate the user's mouth space, valid voice input is filtered out. Combined with the electric scooter's operating status, an interactive control relationship is established, solving the problem of environmental interference affecting the voice recognition of electric scooters and improving the accuracy of voice commands and operational stability.

CN120998196AActive Publication Date: 2025-11-21WUYI RUIJIANG TECH CO LTD

Patent Information

Application Number
CN202511214938.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-11-21
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing voice interaction solutions for electric scooters are susceptible to environmental noise interference and have low voice recognition accuracy, resulting in insufficient operational stability and safety.

Method used

By combining a microphone array with vehicle sensors, the system locates the user's mouth position, filters valid voice input information, and establishes an interactive control relationship based on the electric scooter's operating status information to perform safety assessments and control operations.

Benefits of technology

It improves the accuracy of voice command recognition and the stability of the interaction process, ensuring the safe operation of electric scooters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998196A_ABST
    Figure CN120998196A_ABST
Patent Text Reader

Abstract

The invention discloses a voice interaction control method and system for an electric scooter, and relates to the technical field of voice control. The method comprises the following steps: when a preset voice activation instruction is detected, starting a voice recognition channel to collect a current voice signal; carrying out source positioning according to the collected current voice signal, and screening effective voice input information according to a source result and a direction result; acquiring current running state information of the electric scooter; establishing an interaction control relationship between the current operation state information and the control intention identified by the effective voice input information, performing security judgment on the interaction control relationship, and converting the interaction control relationship into a control execution instruction based on a security judgment result; and performing control operation on the electric scooter according to the control execution instruction. The technical problem that in the prior art, voice recognition is interfered by the environment, and consequently the voice instruction recognition accuracy is low is solved, and the technical effects of improving the voice instruction recognition accuracy and guaranteeing the stability of the interaction process are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of voice control, in particular to a voice interaction control method and system for an electric scooter. BACKGROUND

[0002] With the popularity of electric scooters in urban transportation and short-distance travel, users have higher requirements for the convenience and intelligence of their operation. Voice control, as a natural human-computer interaction method, is gradually applied to the control of intelligent transportation tools. However, the existing voice interaction solutions for electric scooters generally have problems such as voice recognition being easily disturbed by environmental noise, low accuracy of voice input instruction recognition, and lack of running state perception, which leads to insufficient stability and safety of voice control during the user's riding process, causing operation misjudgment or delay and affecting the travel experience. SUMMARY

[0003] The present application provides a voice interaction control method and system for an electric scooter, which solves the technical problem of low accuracy of voice instruction recognition due to environmental interference in the prior art.

[0004] In a first aspect, the present application provides a voice interaction control method for an electric scooter, which comprises: When a preset voice activation instruction is detected, a voice recognition channel is started to collect a current voice signal; the source of the collected current voice signal is located, and valid voice input information is screened according to the source result and the direction result; the current running state information of the electric scooter, including the speed, the running mode and the power, is obtained; an interaction control relationship between the current running state information and the control intention recognized by the valid voice input information is established, the interaction control relationship is judged for safety, the interaction control relationship is converted into a control execution instruction based on the safety judgment result, and the electric scooter is controlled according to the control execution instruction.

[0005] In a second aspect, the present application provides a voice interaction control system for an electric scooter, which comprises: The voice signal collection module: when a preset voice activation instruction is detected, a voice recognition channel is started to collect a current voice signal; the voice information screening module: according to the collected current voice signal, source positioning is performed, and according to the source result and the direction result, valid voice input information is screened; the running state acquisition module: current running state information of the electric scooter is acquired, including vehicle speed, running mode and power; the control relationship conversion module: an interactive control relationship between the current running state information and a control intention recognized by the valid voice input information is established, the interactive control relationship is subjected to safety determination, and the interactive control relationship is converted into a control execution instruction based on the safety determination result; and the scooter control module: the electric scooter is controlled according to the control execution instruction.

[0006] One or more technical solutions provided in the present application have at least the following technical effects or advantages: When a preset voice activation instruction is detected, a voice recognition channel is started to collect a current voice signal; the voice information screening module: according to the collected current voice signal, source positioning is performed, and according to the source result and the direction result, valid voice input information is screened; the running state acquisition module: current running state information of the electric scooter is acquired, including vehicle speed, running mode and power; the control relationship conversion module: an interactive control relationship between the current running state information and a control intention recognized by the valid voice input information is established, the interactive control relationship is subjected to safety determination, and the interactive control relationship is converted into a control execution instruction based on the safety determination result; and the scooter control module: the electric scooter is controlled according to the control execution instruction. The technical problem of low voice command recognition accuracy caused by environmental interference in the prior art is solved, and the technical effects of improving voice command recognition accuracy and ensuring the stability of the interaction process are achieved. BRIEF DESCRIPTION OF DRAWINGS

[0007] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0008] Figure 1 A flow chart of a voice interaction control method for an electric scooter provided by the embodiment of the present application is shown in the figure. Figure 2 A structure diagram of a voice interaction control system for an electric scooter provided by the embodiment of the present application is shown in the figure.

[0009] The reference signs are explained as follows: voice signal collection module 11, voice information screening module 12, running state acquisition module 13, control relationship conversion module 14, and scooter control module 15. DETAILED DESCRIPTION

[0010] The application provides a voice interaction control method and system for an electric scooter, and solves the technical problem of low voice command recognition accuracy caused by environmental interference in voice recognition in the prior art.

[0011] The technical solutions in the embodiments of the application will be clearly and completely described in combination with the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the application.

[0012] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server comprising a series of steps or units need not be limited to only those steps or units clearly listed, but can include other steps or modules not clearly listed or inherent to the process, method, product or device.

[0013] Embodiment one, as shown in the application provides a voice interaction control method for an electric scooter, wherein the method comprises: Figure 1 When a preset voice activation instruction is detected, starting a voice recognition channel to collect a current voice signal. When a preset voice activation instruction is detected, starting a voice recognition channel to collect a current voice signal.

[0014] In the embodiments of the application, when the system detects a preset voice activation instruction (for example, "hello scooter" or any user-set wake-up word) issued by a user, the system enters a voice recognition working mode, automatically starts a voice recognition channel to collect a voice signal issued by the current user. Specifically, a front-end collection circuit of a microphone array or a headset microphone is called, a voice input hardware port is activated, and a real-time voice collection path is established according to a preset voice collection sensitivity parameter.

[0015] Further, starting the voice recognition channel to collect the current voice signal comprises: Acquiring a spatial position range of a user's mouth, establishing a corresponding voice collection conical region, and performing voice recognition channel labeling filtering based on the voice collection conical region to collect a voice signal.

[0016] Preferably, the spatial position range of the user's mouth is obtained by activating the previously established body information and combining the posture detection result of the vehicle body sensor; on this basis, a voice collection conical region is constructed with the user's mouth as the center, the voice collection conical region covers the main direction of sound wave propagation when the user normally speaks, which is used as the effective space for voice collection. Subsequently, the input channels of the microphone array are labeled and filtered according to the voice collection conical region, only allowing the microphone signals located in the voice collection conical region to enter the voice recognition channel, and the channel signals outside the region are subjected to weight reduction or shielding processing, thereby effectively reducing the influence of environmental noise, distant human voice and other interference factors. Finally, under the action of the voice collection conical region labeling and filtering, the system stores the collected voice signals in real time.

[0017] Further, the spatial position range of the user's mouth is obtained, including: The user's pre-input height information is obtained by the pre-collection module before the riding activation, and the user's posture information detected by the vehicle body sensor is combined to analyze the user's body state and establish a user body state model; the spatial position range of the user's mouth is determined based on the user body state model, and the position is matched with the microphone setting position.

[0018] Before the riding activation, the system calls the user's pre-input or stored height information through the pre-collection module, and combines the real-time detection of the user's posture information in the standing riding state by the vehicle body sensor (including acceleration sensor, gyroscope, pressure sensor, etc.) to obtain the data of the user's body inclination angle, head height and shoulder and neck position. Then, the system analyzes the user's body state based on the height information and the posture information, establishes a user body state model corresponding to the current riding state, which is used to reflect the approximate spatial distribution of the user's head and mouth in the three-dimensional coordinate system. Next, the system further locates the spatial position range of the user's mouth based on the user body state model, and establishes a local space envelope region with the mouth space position as the center. Finally, the mouth space position range is matched with the microphone position set on the vehicle body to obtain the effective space position of the mouth relative to the microphone array.

[0019] The source of the collected current voice signal is located, and the effective voice input information is selected according to the source result and the direction result.

[0020] After the current voice signal is collected, the system performs source positioning on the voice signal to determine whether the signal is from the user's vocalization area. Specifically, the system performs sound source positioning processing on the multi-channel voice signal collected by the microphone array installed on the electric scooter, calculates the spatial position coordinates of the sound source by comparing the time difference, energy difference or phase difference between different microphone signals, and obtains the source result of the voice signal. The system combines the source result with the voice collection cone area corresponding to the user's mouth space position range, compares whether the propagation direction of the voice signal and the mouth area are consistent, and obtains the orientation result of the voice signal. When the voice signal meets both the source positioning near the mouth space and the orientation result matching the preset mouth direction, the voice signal is marked as valid voice input information; for voice signals not in the mouth range or not matching the orientation, the system performs weight reduction processing or directly shields to avoid misjudging environmental noise or irrelevant voice as control input.

[0021] Further, according to the collected current voice signal, the source positioning is performed, and according to the source result and the orientation result, the valid voice input information is screened, including: According to the voice signal source positioning, the position of the microphone is determined, and the channel label corresponding to the voice collection cone area is identified. The voice signal falling within the channel label range is marked as valid input, and the voice signal beyond the range is weighted or shielded. The energy feature, frequency spectrum feature and zero-crossing rate feature of the collected voice signal are extracted to determine whether the voice signal is a near-end voice or a far-end voice. When it is determined to be a near-end voice, the orientation is determined based on the relative feature comparison of the microphone signal, and the orientation result is matched with the preset mouth space position range. The voice signal that meets both the near-end voice feature and the orientation matching condition is regarded as the valid voice input information.

[0022] The system determines the position of the microphone according to the voice signal source positioning result, and identifies whether the microphone is located within the voice collection cone area. If the collection channel belongs to the channel label within the cone area, the voice signal collected by the channel is marked as valid input. If the collection channel is located outside the voice collection cone area, the signal is weighted or directly shielded to avoid irrelevant environmental sound entering the subsequent processing link.

[0023] The collected voice signal is subjected to feature extraction to obtain multiple acoustic feature parameters including energy features, spectral features and zero-crossing rate features; the system determines the voice signal to be near-end voice or far-end voice based on these features, ensuring that the subsequent analysis is only directed to the voice signal near the user's mouth. When the determination result is near-end voice, the system calls the relative feature comparison method of the microphone signal, such as calculating the time difference of arrival, relative phase difference or energy difference between the signals collected by different microphones, calculating the sound source direction and generating the orientation result; finally, the orientation result is matched with the preset mouth space position range, and only when the voice signal meets the near-end voice feature and mouth orientation matching conditions at the same time, it is determined as valid voice input information.

[0024] Further, the orientation determination based on the relative feature comparison of the microphone signal includes: The number and installation position of the collection microphones are identified, including the vehicle-mounted microphone and / or the peripheral microphone; when the number of collection microphones is two or more, the time difference of arrival, energy difference or phase difference between the signals collected by different microphones is calculated; the sound source direction is analyzed based on the difference result, and the orientation result of the voice signal is determined in combination with the relative feature comparison of the microphone signal.

[0025] Specifically, the number of microphones participating in voice collection and their installation positions are identified, wherein the microphones can be vehicle-mounted microphones arranged on the body of the electric scooter, or peripheral microphones connected to the body through wireless or wired means.

[0026] When the number of collection microphones is two or more, the system pre-processes the signals of the same voice event in different microphones, including band-pass filtering and amplitude normalization, to eliminate environmental noise interference. Then, the time difference of arrival, energy difference or phase difference between different microphones is calculated, which specifically includes: performing cross-correlation operation on the signals collected by two microphones to obtain the delay corresponding to the maximum correlation coefficient as the propagation time difference of the voice signal between the two microphones; at the same time, the energy difference and phase difference of the signals collected by different microphones are calculated, wherein the energy difference is obtained by integrating the short-time energy of each signal, and the phase difference is obtained by extracting the phase of the main frequency component through fast Fourier transform and calculating the difference.

[0027] The system calls the sound source direction analysis algorithm based on the three types of feature parameters: time difference of arrival, energy difference and phase difference; when the spatial coordinates of the microphones are known, the orientation angle of the sound source on the horizontal plane and the vertical plane is calculated by the trilateration method or the least squares solution method. The obtained orientation angle result is compared with the spatial position range of the user's mouth: if the included angle between the calculated orientation angle and the preset mouth direction is less than the preset threshold, the orientation result of the voice signal is determined as valid source, otherwise it is marked as invalid source.

[0028] Further, the screening of the valid voice input information further comprises: When the voice recognition channel is an earphone microphone or a head-mounted near-mouth microphone, the microphone voice channel is set as the highest priority, and the voice signal collected by the priority channel is directly taken as the valid voice input information.

[0029] When the voice recognition channel is an earphone microphone or a head-mounted near-mouth microphone, the system sets this type of channel as the highest priority channel in the voice channel priority management module. Specifically, the earphone microphone or the head-mounted near-mouth microphone is close to the user's mouth due to the physical installation position, and the voice signal collected by it has a high signal-to-noise ratio and is less disturbed by environmental noise and far-end sound sources. Therefore, when it is detected that the current voice recognition channel belongs to the above type, the system directly marks the voice signal collected by the priority channel as the valid voice input information, skipping the secondary screening step based on the spatial position range, energy feature and direction matching, so as to reduce the delay and improve the interactive response speed.

[0030] The current running state information of the electric scooter is obtained, including the speed, running mode and power.

[0031] The system collects the running data of the electric scooter in real time through the vehicle body controller and the built-in sensor module. The speed information is obtained by the wheel speed sensor or the motor speed detection module, and can be converted into the actual driving speed through the wheel radius and speed parameters. The running mode information is obtained by the mode selection module of the vehicle body control unit, which can identify different riding modes preset by the user, such as energy-saving mode, standard mode or sports mode. The power information is detected in real time by the battery management system (BMS), which calculates the remaining power percentage through the battery voltage, current and capacity parameters, and generates a battery status report.

[0032] An interactive control relationship between the current running state information and the control intention identified by the valid voice input information is established, the safety of the interactive control relationship is determined, and the interactive control relationship is converted into a control execution instruction based on the safety determination result.

[0033] The system performs semantic analysis on the valid voice input information, extracts the control intent, such as "accelerate", "decelerate", "switch mode", "view power", and other instruction semantics, and then associates the control intent with the current running state information of the electric scooter, such as the current speed, power, and running mode constraints, retrieves the corresponding candidate control actions in the preset strategy library, and modifies the action parameters and action timing, thereby generating an interactive control relationship that meets the current running state. The system performs safety judgment on the generated interactive control relationship, including: risk classification of the control intent, scene threshold judgment based on the current speed and running mode, calculation of the comprehensive judgment score combined with the voice recognition confidence, user identity matching degree, and position matching degree, and comparison with the adaptive threshold to obtain the safety judgment result. Finally, the system converts the interactive control relationship into a control execution instruction according to the safety judgment result: when the judgment is to allow execution, the control execution instruction is directly output; when the judgment is to delay execution, a buffer parameter is set, and the instruction is issued after the confirmation condition is met; when the judgment is to reject execution, an alternative prompt information is generated, which is fed back to the user through voice broadcast or interface.

[0034] Further, the interactive control relationship between the current running state information and the control intent recognized by the valid voice input information is established, including: Based on the control intent, the corresponding candidate control action and its action timing template are retrieved in the preset strategy library; the candidate control action and the action timing template are constrained and analyzed according to the current running state information, the feasibility verification, parameter clipping, and timing rearrangement are completed, and the interactive control relationship is generated.

[0035] The system retrieves in the preset strategy library based on the control intent, calls the candidate control action matched with the control intent, and at the same time calls the action timing template associated with the candidate control action, which is used to describe the execution order and duration of the control action at different time stages. Subsequently, the system performs constraint analysis on the candidate control action and its action timing template according to the current running state information of the electric scooter, such as prohibiting or parameter limiting the "accelerate" type candidate control action when the vehicle speed is higher than the safety threshold, and constraining the action such as "switching to sports mode" when the power is lower than the set lower limit. The process of constraint analysis includes: performing feasibility verification on the candidate control action to determine whether it has execution conditions under the current running state; clipping the control parameters of the candidate control action to ensure that the action amplitude and timing meet the safety requirements; rearranging the action timing template to adjust the order and interval time of each execution step to ensure that the action is adapted to the running state in timing. Finally, after the above analysis and modification, a complete interactive control relationship is generated, which serves as the basis for subsequent safety judgment and control execution.

[0036] Further, the security determination of the interaction control relationship includes: The control intention is risk graded, the scene threshold is determined based on the current running state information, the interaction control relationship with a risk label not meeting a safety condition is determined, the recognition confidence of the effective voice input information, the speaker identity matching degree, and the position matching degree are calculated, the comprehensive determination score is calculated by fusing the confidence, the identity matching degree, and the position matching degree, and compared with an adaptive threshold to determine the security determination result.

[0037] The system grades the control intention into low risk, medium risk, and high risk levels, for example, "querying power" belongs to low risk, "switching running mode" belongs to medium risk, and "emergency acceleration / urgent braking" belongs to high risk. The scene threshold is determined based on the current running state information, for example, in the case that the vehicle speed is higher than a safety threshold, the road slope exceeds a preset range, or the power is lower than the minimum safe power, the corresponding interaction control relationship is risk labeled, and it is determined that it does not meet the safety condition. The system analyzes the quality and source reliability of the collected effective voice input information, and calculates the recognition confidence, the speaker identity matching degree, and the position matching degree of the voice source, respectively. The recognition confidence is used to measure the voice transcription accuracy, the identity matching degree is used to determine whether the input voice is from an authorized user, and the position matching degree is used to determine whether the voice source direction falls within the user's mouth space range.

[0038] The system fuses the recognition confidence, the identity matching degree, and the position matching degree by weighting to obtain a comprehensive determination score, and compares it with an adaptive threshold. When the comprehensive determination score is higher than the adaptive threshold, the security determination result of allowing execution is outputted; when the comprehensive determination score is close to or lower than the threshold, the security determination results of delaying execution or refusing execution are respectively outputted.

[0039] Further, the interaction control relationship is converted into a control execution instruction based on the security determination result, including: The security determination result is divided into multiple scene levels of allowing execution, delaying execution, and refusing execution, and the execution mode of the interaction control relationship is selected according to the determination scene level category; when the security determination result is allowing execution, the corresponding control execution instruction is directly generated according to the interaction control relationship; when the security determination result is delaying execution, the buffer parameter is set based on the current running state information, and the control execution instruction is issued when the confirmation condition is met; when the security determination result is refusing execution, the alternative prompt instruction is generated, and is fed back to the user through voice broadcast or display interface.

[0040] The system divides the security judgment result into three scene levels of allowing execution, delaying execution and rejecting execution, and selects the corresponding execution mode according to the judgment scene level category. When the security judgment result is allowing execution, the system does not need additional delay or correction, directly generates the corresponding control execution instruction according to the interaction control relationship, and issues it to the vehicle body control unit of the electric scooter for action execution; when the security judgment result is delaying execution, the system sets the buffer parameters according to the current running state information of the electric scooter, such as buffer time, vehicle speed threshold or power supplement condition, and then converts the interaction control relationship into control execution instruction and issues it for execution after confirming that the condition is met, to ensure safety and feasibility of execution; when the security judgment result is rejecting execution, the system no longer converts the interaction control relationship into actual control instruction, but generates an alternative prompt instruction, such as "the current state does not allow this operation" or "please reduce the vehicle speed and try again", and feeds back to the user through the vehicle-mounted voice broadcast module or display interface.

[0041] Control the electric scooter according to the control execution instruction.

[0042] The system issues the control execution instruction generated after security judgment to the vehicle body control unit of the electric scooter, and the vehicle body control unit calls the corresponding execution module to complete the specific operation. For example, when the control execution instruction is "accelerate", the motor drive module adjusts the output power according to the instruction to increase the vehicle speed; when the control execution instruction is "decelerate" or "brake", the brake module or energy recovery module receives the instruction and executes the deceleration action; when the control execution instruction is "switch running mode", the running mode management module switches to the energy saving mode, standard mode or sports mode according to the instruction; when the control execution instruction is "query power", the battery management system (BMS) feeds back the current power information to the user interface or outputs it through voice broadcast.

[0043] In summary, the embodiments of the present application have at least the following technical effects: When the preset voice activation instruction is detected, the voice recognition channel is started to collect the current voice signal; the source is located according to the collected current voice signal, and the effective voice input information is filtered according to the source result and the direction result. Then, the current running state information of the electric scooter is obtained, including vehicle speed, running mode and power. Then, the interaction control relationship between the current running state information and the control intention recognized by the effective voice input information is established, the security of the interaction control relationship is judged, and the interaction control relationship is converted into control execution instruction based on the security judgment result. Finally, the electric scooter is controlled according to the control execution instruction. The technical problem of low voice command recognition accuracy caused by environmental interference in the prior art is solved, and the technical effects of improving voice command recognition accuracy and ensuring the stability of the interaction process are achieved.

[0044] Embodiment two, based on the same inventive concept as the voice interaction control method for electric scooters in the preceding embodiments, as Figure 2 shown, the present application provides a voice interaction control system for electric scooters, wherein the system comprises: A voice signal acquisition module 11: when a preset voice activation instruction is detected, a voice recognition channel is started to acquire the current voice signal; a voice information screening module 12: according to the source positioning of the acquired current voice signal, according to the source result and the direction result, screening the effective voice input information; a running state acquisition module 13: acquiring the current running state information of the electric scooter, including speed, running mode, power; a control relationship conversion module 14: establishing the interaction control relationship between the current running state information and the control intention recognized by the effective voice input information, judging the safety of the interaction control relationship, and converting the interaction control relationship into a control execution instruction based on the safety judgment result; a scooter control module 15: controlling the electric scooter according to the control execution instruction.

[0045] Further, the voice signal acquisition module 11 is used to execute the following method: Acquire the spatial position range of the user's mouth, establish the corresponding voice acquisition conical area; based on the voice acquisition conical area, perform voice recognition channel labeling filtering, and collect voice signals.

[0046] Further, the voice signal acquisition module 11 is used to execute the following method: Acquire the user's pre-imported height information through the pre-acquisition module before riding activation, combine the user's posture information detected by the vehicle body sensor, and perform user body analysis to establish a user body model; determine the spatial position range of the user's mouth based on the user body model, and match the position with the microphone setting position.

[0047] Further, the voice information screening module 12 is used to execute the following method: According to the voice signal source positioning acquisition microphone position, and recognizing the channel labeling corresponding to the voice acquisition conical area, marking the voice signal falling within the channel labeling range as effective input, and reducing the weight or shielding the voice signal beyond the range; extract the energy features, spectral features and zero-crossing rate features of the collected voice signals to determine whether the voice signals are near-end voice or far-end voice; when it is determined to be near-end voice, the direction is determined based on the relative feature comparison of the microphone signal, and the direction result is matched with the preset mouth spatial position range, and the voice signal that meets the near-end voice feature and direction matching conditions at the same time is taken as the effective voice input information.

[0048] Further, the voice information screening module 12 is configured to perform the following method: The number and installation position of the collection microphones are identified, the microphones including vehicle-mounted microphones and / or peripheral microphones; when the number of collection microphones is two or more, the time difference, energy difference or phase difference between signals collected by different microphones is calculated; based on the difference result, the sound source direction is analyzed, and the relative characteristic comparison of the microphone signals is combined to determine the orientation result of the voice signal.

[0049] Further, the voice information screening module 12 is configured to perform the following method: When the voice recognition channel is a headset microphone or a head-mounted near-mouth microphone, the microphone voice channel is set as the highest priority, and the voice signal collected by the priority channel is directly used as the effective voice input information.

[0050] Further, the control relationship conversion module 14 is configured to perform the following method: Based on the control intention, a corresponding candidate control action and its action timing template are retrieved from a preset strategy library; the candidate control action and the action timing template are analyzed based on the current running state information to complete feasibility verification, parameter pruning and timing rearrangement, and the interactive control relationship is generated.

[0051] Further, the control relationship conversion module 14 is configured to perform the following method: The control intention is risk graded; based on the current running state information, a scene threshold is judged, and the interactive control relationship with a risk label that does not meet the safety condition is determined; the recognition confidence of the effective voice input information, the speaker identity matching degree and the orientation matching degree are calculated; the comprehensive judgment score is calculated by fusing the confidence, the identity matching degree and the orientation matching degree, and compared with an adaptive threshold to determine the safety judgment result.

[0052] Further, the control relationship conversion module 14 is configured to perform the following method: The safety judgment result is divided into multiple scene levels of allowed execution, delayed execution and refused execution, and the execution mode of the interactive control relationship is selected according to the judgment scene level category; when the safety judgment result is allowed execution, the corresponding control execution instruction is directly generated according to the interactive control relationship; when the safety judgment result is delayed execution, the buffer parameter is set based on the current running state information, and the control execution instruction is issued when the confirmation condition is met; when the safety judgment result is refused execution, a replacement prompt instruction is generated, and the user is fed back through voice broadcast or a display interface.

[0053] It should be noted that the above-mentioned embodiment sequences of the present application are merely for description only, but not for representing the advantages and disadvantages of the embodiments. And the above-mentioned embodiment sequences of the present application are described in the specification. The processes depicted in the drawings do not necessarily require the particular sequence or continuous sequence shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or can be advantageous.

[0054] The above only describes the preferred embodiments of the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0055] The specification and drawings are merely exemplary of the present application, and any and all modifications, variations, combinations or equivalents that are within the scope of the present application should be included. Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the present application and its equivalents, the present application is intended to include these modifications and variations.

Claims

1. A voice interaction control method for electric scooters, characterized in that, The method includes: When a preset voice activation command is detected, the voice recognition channel is activated to collect the current voice signal; The source of the collected current speech signal is located, and valid speech input information is filtered based on the source and location results. Obtain the current operating status information of the electric scooter, including speed, operating mode, and battery level; Establish an interactive control relationship between the current operating status information and the control intent identified by the valid voice input information, perform a security determination on the interactive control relationship, and convert the interactive control relationship into a control execution command based on the security determination result; The electric scooter is controlled according to the control execution command.

2. The voice interaction control method for electric scooters according to claim 1, characterized in that, Start the speech recognition channel to collect the current speech signal, including: Obtain the spatial range of the user's mouth and establish a corresponding cone-shaped area for voice acquisition; Based on the cone-shaped area for voice acquisition, voice recognition channel labeling and filtering are performed to acquire voice signals.

3. The voice interaction control method for electric scooters according to claim 2, characterized in that, Obtain the spatial range of the user's mouth, including: The system obtains the user's pre-imported height information through the pre-collection module before cycling activation, and combines it with the user's posture information detected by the vehicle's sensors to perform user posture analysis and establish a user posture model. Based on the user's body shape model, the spatial range of the user's mouth is determined, and position matching is performed in conjunction with the microphone's set position.

4. The voice interaction control method for an electric scooter according to claim 2, characterized in that, Based on the collected current speech signal, the source is located. Based on the source and location results, valid speech input information is filtered, including: The microphone position is located based on the voice signal source, and the channel label corresponding to the voice acquisition cone area is identified. Voice signals falling within the channel label range are marked as valid inputs, and voice signals outside the range are downweighted or masked. Extract the energy characteristics, spectral characteristics, and zero-crossing rate characteristics of the acquired speech signal to determine whether the speech signal is near-end speech or far-end speech; When the speech is determined to be near-end speech, the orientation is determined based on the relative feature comparison of the microphone signal, and the orientation result is matched with the preset mouth space position range. The speech signal that simultaneously meets the near-end speech feature and orientation matching conditions is taken as the effective speech input information.

5. The voice interaction control method for an electric scooter according to claim 4, characterized in that, Location determination based on relative feature comparison of microphone signals includes: The number and installation location of the acquisition microphones are identified, including vehicle-mounted microphones and / or external microphones; When there are two or more microphones, calculate the arrival time difference, energy difference or phase difference between the signals collected by different microphones; Based on the difference results, the direction of the sound source is analyzed, and combined with the comparison of the relative features of the microphone signal, the directional result of the speech signal is determined.

6. The voice interaction control method for an electric scooter according to claim 5, characterized in that, Filtering valid voice input information also includes: When the voice recognition channel is an earphone microphone or a headset near-mouth microphone, the microphone voice channel is set to the highest priority, and the voice signal collected by the priority channel is directly used as the valid voice input information.

7. The voice interaction control method for an electric scooter according to claim 1, characterized in that, Establishing an interactive control relationship between the current operating status information and the control intent identified by the valid voice input information includes: Based on the control intent, retrieve the corresponding candidate control actions and their timing templates from the preset strategy library; Based on the current operating status information, the candidate control actions and the action timing templates are constrained and parsed to complete feasibility verification, parameter trimming and timing rearrangement, and generate the interactive control relationship.

8. The voice interaction control method for an electric scooter according to claim 7, characterized in that, The security determination of the interactive control relationship includes: The control intent is risk-classified; Based on the current operating status information, a scenario threshold judgment is made, and the risk is marked as an interaction control relationship that does not meet the safety conditions. Calculate the recognition confidence, speaker identity matching degree, and location matching degree of the effective voice input information; The confidence level, identity matching degree, and location matching degree are combined to calculate a comprehensive judgment score, which is then compared with an adaptive threshold to determine the security judgment result.

9. The voice interaction control method for an electric scooter according to claim 8, characterized in that, Based on the security determination results, the interactive control relationship is transformed into control execution instructions, including: The security determination results are divided into multiple scenario levels: allowed execution, delayed execution, and denied execution. The execution mode of the interaction control relationship is selected according to the scenario level category. When the security determination result indicates that execution is allowed, the corresponding control execution instruction is directly generated according to the interaction control relationship; When the security determination result is delayed execution, buffer parameters are set based on the current running status information, and a control execution command is issued when the confirmation condition is met. When the security determination result is "reject execution", an alternative prompt instruction is generated and fed back to the user through voice broadcast or display interface.

10. A voice interaction control system for electric scooters, characterized in that, The system is used to implement the voice interaction control method for an electric scooter according to any one of claims 1-9, the system comprising: Voice signal acquisition module: When a preset voice activation command is detected, the voice recognition channel is activated to acquire the current voice signal; Voice information filtering module: Based on the current voice signal collected, the source is located, and based on the source and location results, valid voice input information is filtered. Operating status acquisition module: Acquires the current operating status information of the electric scooter, including speed, operating mode, and battery level; Control Relationship Conversion Module: Establishes an interactive control relationship between the current running status information and the control intent identified by the valid voice input information, performs a security determination on the interactive control relationship, and converts the interactive control relationship into a control execution command based on the security determination result; Scooter control module: Controls the electric scooter according to the control execution instructions.

Citation Information

Patent Citations

  • Control system of electric scooter and control method of electric scooter applying control system

    CN111661227A

  • Vehicle control method based on voice instruction and related device

    CN116994577A

  • Voice control method and device

    CN120260569A

  • Aircraft voice instruction security analysis method and system based on natural language processing

    CN120412564A

  • On-board voice command identification method and apparatus, and storage medium

    US20180190283A1

Cited By

  • Control method, device and equipment of four-foot machine horse in entertainment scene and medium

    CN121477902A