A voice command recognition method and system for an electric wheelchair
By recognizing ambiguous commands, generating specific clarifying questions, and optimizing attribution strategies, the problem of information loss and inaccurate execution caused by the single clarification method in the voice command recognition system of electric wheelchairs has been solved, achieving higher accuracy and intelligence.
Patent Information
- Application Number
- CN202511311217.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-09-15
AI Technical Summary
Existing voice command recognition systems for electric wheelchairs suffer from problems such as information loss, inaccurate command execution, and incorrect system attribution when faced with ambiguous commands due to their simplistic clarification methods.
By receiving user voice commands, recognizing ambiguous commands, acquiring environmental awareness data, generating embodying clarification questions, receiving voice responses and generating command parameters, recording interaction logs, optimizing attribution strategies, processing negative semantics and generating new clarification questions, and combining user identity and contextual information to optimize strategies.
It significantly improves the accuracy of voice command recognition and user interaction experience, ensuring precise execution of commands and the system's intelligence level.
Smart Images

Figure CN120808767B_ABST
Abstract
Description
Technical Field
[0001] This application relates to voice command recognition technology for electric wheelchairs, and more specifically, to a voice command recognition method and system for electric wheelchairs. Background Technology
[0002] In the field of mobility assistive devices, particularly electric wheelchairs, voice command recognition technology aims to provide users with a convenient operating experience. When a voice command recognition system detects uncertainty or ambiguity in a command, it typically initiates a clarifying inquiry process to obtain more information. However, if this clarifying inquiry is too simplistic or rigid, such as being limited to closed-ended questions (e.g., simple "yes" or "no" confirmations), it may prevent the system from effectively recognizing, parsing, and utilizing further precise information provided by the user in their response (e.g., supplementary explanations regarding the object, method, or spatial location of the operation).
[0003] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this application provides a voice command recognition method and system for electric wheelchairs, which can effectively solve the problems of information loss, inaccurate command execution, and system error attribution caused by improper clarification of ambiguous commands, and significantly improve the accuracy of voice command recognition for electric wheelchairs.
[0005] This application provides a voice command recognition method for an electric wheelchair, the method comprising:
[0006] Receive user voice commands and identify ambiguous commands within them;
[0007] Based on fuzzy commands, acquire perception data related to the current environment of the electric wheelchair;
[0008] Identify physical targets in the perceived data, as well as the relative distance between the physical targets and the electric wheelchair;
[0009] Based on the preset attribution strategy, a embodied clarification query containing the target name and location information is generated from the physical target and its relative distance.
[0010] It receives the user's voice response to the personalized clarification inquiry, generates command parameters based on the voice response, and controls the electric wheelchair to perform corresponding actions based on the command parameters;
[0011] Voice commands, embodied clarification questions, voice responses, and command parameters are recorded as interaction logs, and attribution strategies are optimized based on these interaction logs.
[0012] The above solution solves the problem in existing technologies where the lack of a single clarification method leads to the inability to effectively identify and utilize supplementary user information, resulting in inaccurate command execution and incorrect attribution. This improves the accuracy of voice command recognition and the user interaction experience.
[0013] Furthermore, command parameters are generated based on the voice response, and the electric wheelchair is controlled to perform corresponding actions based on the command parameters, including:
[0014] When a user's voice response contains negative semantics, the voice response is parsed to obtain new target information and its relative distance;
[0015] Based on the new target information and its relative distance, a new embodied clarification question is generated, and after receiving the user's voice response, the command parameters are extracted. Based on the command parameters, the electric wheelchair is controlled to perform the corresponding actions.
[0016] Furthermore, when a user's voice response contains negative semantics, the voice response is parsed to obtain new target information and its relative distance, including:
[0017] Based on the new target information and its relative distance, multiple candidate physical targets are matched from the detected and identified physical targets;
[0018] Acquire the spatial relationships between multiple candidate physical targets and the electric wheelchair, including relative distance, spatial position, and orientation angle;
[0019] Based on spatial relationships, multiple candidate physical targets are prioritized.
[0020] If there is only one candidate physical target with the highest priority, select the candidate physical target with the highest priority as the new target information;
[0021] If there are multiple candidate physical targets with the highest priority, descriptive information containing the differentiating characteristics among the candidate physical targets is generated, which is used to generate a specific clarification query in the future.
[0022] Furthermore, generating command parameters based on voice responses and controlling the electric wheelchair to perform corresponding actions based on these parameters also includes:
[0023] When the voice response is affirmative, instruction parameters are generated based on the target name and location information contained in the embodied clarification question.
[0024] The parking distance of the physical target is calculated based on the location information, and this parking distance is used as the execution boundary of the command parameter to control the electric wheelchair to move along the direction of the physical target corresponding to the target name until it stops at a preset safe distance from its parking position.
[0025] Furthermore, voice commands, embodied clarification inquiries, voice responses, and command parameters are recorded as interaction logs. Attribution strategies are optimized based on these interaction logs, including:
[0026] Obtain the current user's identity and context information;
[0027] Based on identity and contextual information, it records interaction logs including voice commands, embodying clarification inquiries, voice responses, and command parameters;
[0028] Filter out historical interaction records from the interaction log that match the current user's identity and context information;
[0029] Adjust the attribution strategy based on historical interaction records;
[0030] When a change in user identity information or contextual information is detected, the attribution strategy corresponding to the new identity information or contextual information is loaded.
[0031] Furthermore, attribution strategies include:
[0032] Based on the perception data, identify all identified physical targets located in front of the electric wheelchair and their corresponding relative distance, spatial position, and orientation angle;
[0033] According to the preset exclusion rules, the identified physical targets are screened: identified physical targets that are obstructed by obstacles from the electric wheelchair, identified physical targets whose passage path width to the electric wheelchair is less than a preset threshold, and / or identified physical targets whose relative distance to the electric wheelchair is less than a safe distance threshold are excluded, and candidate physical targets are obtained.
[0034] The candidate physical target with the smallest angle to the current direction of travel of the electric wheelchair is selected as the source of the target name and location information to be referenced when generating the embodiment clarification query.
[0035] Furthermore, attribution strategies include:
[0036] Based on the physical target and its relative distance, determine whether the physical target belongs to the preset sensitive target type, and whether the relative distance of the physical target is less than the preset safe distance;
[0037] If the physical target is a sensitive target type, or the relative distance is less than the safe distance, then generate a definite clarification query containing the target name and location information;
[0038] If the physical target is not a sensitive target and the relative distance is not less than the safe distance, then determine whether the physical target is within the preset movement path of the electric wheelchair;
[0039] If the physical target is within the movement path, a specific clarification query is generated that specifies the target's name and location information.
[0040] Furthermore, it identifies ambiguous instructions within voice commands, including:
[0041] The voice commands are parsed to extract parameter information including the target of the operation, the operation method, and the spatial location.
[0042] Determine whether the voice command has ambiguous features based on the parameter information;
[0043] If fuzzy features are identified, the corresponding fuzzy instructions in the voice commands are identified based on the fuzzy features.
[0044] The above scheme defines the identification process of fuzzy instructions in detail. By parsing parameter information and judging fuzzy features, the system can more accurately identify instructions that need clarification, laying the foundation for subsequent embodiment clarification.
[0045] Furthermore, based on the parameter information, it is determined whether the voice command has ambiguous features, including:
[0046] Obtain the current physical state information and environmental context information of the electric wheelchair;
[0047] By combining parameter information, current physical state information, and environmental context information, the executability and semantic integrity of voice commands are evaluated.
[0048] When the evaluation results indicate that the voice command has an unclear execution target, an unclear path goal, and / or an uncertain execution scope, the voice command is judged to be an ambiguous command.
[0049] Furthermore, this application also discloses an electric wheelchair voice command recognition system, the system comprising:
[0050] The receiving and recognition module is used to receive the user's voice commands and recognize ambiguous commands in the voice commands;
[0051] The perception data acquisition module is used to acquire perception data related to the current environment of the electric wheelchair based on fuzzy commands;
[0052] The target distance recognition module is used to identify physical targets in the perception data, as well as the relative distance between the physical targets and the electric wheelchair;
[0053] The query generation module is used to generate a definite clarifying query containing the target name and location information based on the physical target and its relative distance, according to a preset attribution strategy.
[0054] The execution module is used to receive the user's voice response to the personalized clarification inquiry, generate instruction parameters based on the voice response, and control the electric wheelchair to perform corresponding actions based on the instruction parameters;
[0055] The recording optimization module is used to record voice commands, embodied clarification inquiries, voice responses, and command parameters as interaction logs, and optimize the attribution strategy based on the interaction logs.
[0056] The above scheme provides a system for implementing the above method. Through modular design, the method can be effectively implemented, providing hardware and software support for the intelligent control of electric wheelchairs.
[0057] In summary, the electric wheelchair voice command recognition method and system provided in this application effectively solves the problems of information loss, inaccurate command execution, and system error attribution caused by improper clarification of ambiguous commands in the prior art by receiving ambiguous commands, acquiring environmental perception data, generating personalized clarification questions, receiving voice responses and controlling actions, and recording interaction logs and optimizing attribution strategies. It significantly improves the accuracy of electric wheelchair voice command recognition, the level of intelligence of interaction, and user experience. Attached Figure Description
[0058] Figure 1 A flowchart illustrating a voice command recognition method for an electric wheelchair provided in this application.
[0059] Figure 2 A flowchart of a voice command recognition system for an electric wheelchair provided in this application.
[0060] In the diagram: 1. Receiving and recognition module; 2. Sensing data acquisition module; 3. Target distance recognition module; 4. Inquiry generation module; 5. Execution module; 6. Recording and optimization module. Detailed Implementation
[0061] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0062] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0063] Reference Figure 1 This application proposes a voice command recognition method for electric wheelchairs, including:
[0064] S1000: Receives user voice commands and recognizes ambiguous commands within the voice commands;
[0065] S2000: Based on fuzzy commands, acquire perception data related to the current environment of the electric wheelchair;
[0066] S3000: Identifies physical targets in the sensing data, as well as the relative distance between the physical targets and the electric wheelchair;
[0067] S4000: Based on a preset attribution strategy, generate a definite clarifying query containing the target name and location information from the physical target and its relative distance;
[0068] S5000: Receives the user's voice response to the personalized clarification inquiry, generates command parameters based on the voice response, and controls the electric wheelchair to perform corresponding actions based on the command parameters;
[0069] S6000: Records voice commands, embody clarification inquiries, voice responses, and command parameters as interaction logs, and optimizes attribution strategies based on these interaction logs.
[0070] Fuzzy commands refer to voice commands issued by users that contain parts that are difficult to directly parse into precise control parameters, such as unclear operation objects, ambiguous path targets, or uncertain execution ranges. These commands can be implemented using natural language processing techniques, semantic analysis models, or pre-set fuzzy vocabularies. For example, lexical and syntactic analysis combined with contextual information can be used for judgment. The main purpose is to identify the uncertainty in the user's intent in order to initiate subsequent clarification processes. Perceptual data refers to real-time information related to the current environment acquired by the sensor system mounted on the electric wheelchair. This can be achieved using multiple sensor fusion technologies such as LiDAR, cameras, ultrasonic sensors, or depth sensors. For example, LiDAR can be used to acquire point cloud data, and cameras can capture image information. The main purpose is to provide objective and comprehensive spatial information about the electric wheelchair's surrounding environment as a contextual basis for understanding fuzzy commands.
[0071] Physical targets refer to entities with specific shapes, locations, and semantics identified in perceived data. This can be achieved using object detection algorithms, image recognition technology, or point cloud clustering analysis, such as identifying tables, chairs, doors, or walls. The primary purpose is to extract specific objects potentially related to user commands from complex environmental information, serving as reference points for subsequent clarification inquiries. Attribution strategies refer to a pre-defined set of rules used to infer a user's potential intentions and generate clarification inquiries based on environmental perception information and ambiguous user commands. This can be implemented using rule-based inference engines, machine learning models, or expert systems. For example, selecting the most likely target based on the physical target's location, type, and relative relationship to a wheelchair. The main purpose is to intelligently select the most relevant physical targets and their location information under ambiguous commands to construct embodied clarification inquiries. Embodying clarification queries refer to interactive queries that include specific target names and location information, guiding users to provide precise command input. These can be implemented using speech synthesis technology combined with preset templates, such as "Do you want to go to the coffee table?" or "Do you want to move closer to the sofa?" Their main purpose is to transform abstract, vague commands into specific questions related to physical targets in the real environment, thereby guiding users to provide more precise command information and avoiding information loss. Interaction logs refer to the data collection of a series of interactive events recorded by the system during the processing of user voice commands, including voice commands, embodying clarification queries, voice responses, and command parameters. These can be implemented using structured databases or unstructured file storage, recording, for example, the timestamp of each interaction, user ID, original command text, clarification content issued by the system, user response text, and the final executed command parameters. Their main purpose is to accumulate user interaction data, providing a data foundation for subsequent attribution strategy optimization and enabling the system's adaptive learning.
[0072] This application's solution achieves intelligent recognition and execution of voice commands for electric wheelchairs through a series of collaborative steps, demonstrating unique advantages, particularly in handling ambiguous commands. When the system receives a user's voice command, it first parses the command to identify any ambiguous elements. Upon detecting ambiguity, the system immediately initiates an environmental perception process to acquire sensory data related to the electric wheelchair's current environment. This sensory data is processed to identify specific physical targets in the environment and calculate the relative distances between these targets and the electric wheelchair. Based on this, the system, according to a preset attribution strategy, uses the identified physical targets and their relative distance information to generate a definite clarifying question containing the target's name and location information. This questioning method differs from simple affirmative or negative questions; it associates the user's ambiguous intent with specific objects in the actual environment, guiding the user to provide more precise feedback. Subsequently, the system receives the user's voice response to this definite clarifying question and performs in-depth analysis to extract precise command parameters. Based on these command parameters, the electric wheelchair is controlled to perform corresponding actions, thereby ensuring accurate command execution. To continuously improve the system's intelligence and adaptability, all voice commands, embodied clarification inquiries, voice responses, and the final generated command parameters throughout the interaction process are recorded as interaction logs. Based on these logs, the system continuously optimizes its pre-defined attribution strategies. This optimization mechanism allows the system to learn from historical interactions, constantly adjusting its understanding and clarification of ambiguous user commands. This avoids incorrect attributions caused by repeated interaction failures and gradually adapts to different users' language habits and usage scenarios, forming a positive self-learning cycle.
[0073] In some preferred embodiments, this application is implemented as follows:
[0074] When a user issues a vague voice command like "Wheelchair, move forward a little," the electric wheelchair's voice recognition module receives and identifies "move forward a little" as a vague command. Subsequently, the wheelchair's lidar and camera work together to acquire perception data of the current environment. For example, the lidar scans and generates a point cloud map of the surrounding environment, while the camera captures images. By processing this perception data, the system identifies a "coffee table" and a "sofa" in front, calculating the relative distance between the coffee table and the wheelchair to be 2 meters, and the relative distance between the sofa and the wheelchair to be 3 meters. Based on a preset attribution strategy, which may prioritize the closest physical target in the wheelchair's direction of travel, the system determines that the user is more likely pointing towards the coffee table. Therefore, the system generates a specific clarifying question, such as a voice synthesis announcement: "Do you want to move closer to the coffee table?" When the user responds, "Yes, closer to the coffee table, slowly," the system receives and parses this voice response, recognizing the affirmative semantics and the speed limit for "slowly." Based on this information, the system generates instruction parameters, such as "Move towards the coffee table at a low speed until you stop at a preset safe distance from the coffee table." The electric wheelchair then performed the actions according to these parameters. At the same time, the original command, clarifying questions, user responses, and final command parameters of this interaction were all recorded in the interaction log for subsequent attribution strategy optimization, such as adjusting the priority attribution weight for "coffee table" under similar ambiguous commands, or learning the speed parameters corresponding to "slow down".
[0075] In another embodiment of this application, the sub-step of S5000—generating instruction parameters based on the voice response and controlling the electric wheelchair to perform corresponding actions based on the instruction parameters—comprising:
[0076] S5210: When a user's voice response contains negative semantics, parse the voice response to obtain new target information and its relative distance;
[0077] S5220: Based on the new target information and its relative distance, generate a new embodied clarification question, and after receiving the user's voice response, extract the instruction parameters and control the electric wheelchair to perform the corresponding actions according to the instruction parameters.
[0078] The process of analyzing voice responses to obtain new target information and their relative distance involves semantic analysis and information extraction from user voice responses containing negation. Specifically, this can be achieved through natural language processing (NLP) technology to identify new, desired target entities mentioned by the user while negating the original target, and their spatial relationship to the electric wheelchair's current location. The aim is to accurately capture the user's true intention from their negation feedback, providing foundational data for subsequent command correction. New target information refers to the physical entity or location that the electric wheelchair needs to re-identify or focus on, explicitly or implicitly indicated by the user's voice response after negating the original command target. This could be a new object name, a more precise location description, or a directional indication, providing the electric wheelchair with a new target or destination. Relative distance refers to the spatial measurement between the physical entity or location indicated by the new target information and the electric wheelchair's current location. This can be inferred from distance descriptive words (e.g., "nearer," "farr," "next to") or spatial orientation words (e.g., "left," "right," "front") contained in the voice response, providing precise location references for the electric wheelchair's movement commands. Among them, generating new embodied clarification questions refers to constructing an interactive question that includes the specific target name and location description, which can be intuitively understood and confirmed by the user, based on the parsed new target information and its relative distance. Specifically, this can be done by combining visual or auditory cues, such as highlighting the new target on the screen or describing its specific location by voice. The purpose is to reconfirm the system's understanding of the new intent with the user, ensure the accuracy of the instructions, and reduce misunderstandings.
[0079] This solution proposes a specialized processing mechanism for negative voice responses from users during the voice command recognition process of electric wheelchairs, significantly improving the system's ability to understand complex user intentions. When the system receives a negative response from a user to a clarification question, it no longer simply interrupts the command flow but further analyzes the new target information and its relative distance contained in the negative statement. For example, if a user expresses "not the sofa, but the chair next to it," the system can extract "chair" as the new target and generate a new clarification question based on spatial relationships, such as "Do you want to go to the chair next to the sofa?" This new clarification statement is dynamically generated based on user feedback, making it more precise and specific, and helping to guide the user to give a clear response. The system then extracts the final command parameters based on the new response to complete the control action of the electric wheelchair. Through this iterative and context-sensitive processing method, the system can continuously learn and correct its original understanding from the user's negative responses, avoiding command execution errors caused by ignoring negative information. Compared to the traditional single-round clarification process, this solution constructs a more flexible and robust interactive loop, enabling electric wheelchairs to accurately grasp the user's true intentions through multiple rounds of interaction, even when faced with ambiguous initial commands or misunderstandings, thereby enhancing the reliability of voice control and user experience.
[0080] In another embodiment of this application, S5210 further includes:
[0081] S5211: Based on the new target information and its relative distance, match multiple candidate physical targets from the detected and identified physical targets;
[0082] S5212: Obtain the spatial relationship between multiple candidate physical targets and the electric wheelchair, including relative distance, spatial position, and orientation angle;
[0083] S5213: Prioritize multiple candidate physical targets based on spatial relationships;
[0084] S5214: If there is only one candidate physical target with the highest priority, select the candidate physical target with the highest priority as the new target information;
[0085] S5215: If there are multiple candidate physical targets with the highest priority, generate descriptive information containing the differentiated features among the candidate physical targets for subsequent generation of embodied clarification queries.
[0086] The process involves several key elements: First, matching multiple candidate physical targets. This involves filtering potential targets from all physical targets detected and identified by the perception module of the electric wheelchair based on the newly analyzed target information and their relative distances. Second, spatial relationships refer to the relative positions of the candidate physical targets and the electric wheelchair in three-dimensional space, including their straight-line distance, specific coordinates in a particular coordinate system, and their respective orientation angles. This provides multi-dimensional data for a comprehensive evaluation of the candidate targets. Third, priority ranking involves arranging the matched candidate physical targets according to preset rules or algorithms based on their importance or probability. This aims to identify the potential target that best matches the user's intent. Fourth, descriptive information of differentiated features refers to the system automatically extracting and generating distinguishable visual or physical attribute descriptions, such as color, shape, material, or specific markings, when multiple candidate physical targets with the same priority exist. This provides clear distinguishing criteria for subsequent clarification inquiries, guiding the user to make an accurate selection.
[0087] This application addresses the problem of accurately extracting new target information from user negative responses through a multi-stage, data-driven inference mechanism. When a user's voice response to a clarification inquiry contains negative semantics, the system first preliminarily parses new target information and its relative distance from the voice response. However, relying solely on this preliminary parsing may be insufficient to determine a unique and accurate target. Therefore, this application further utilizes the physical target data already detected and identified by the electric wheelchair, matching multiple candidate physical targets from these identified physical targets based on the new target information and their relative distance obtained from the preliminary parsing. This process associates the user's ambiguous intent with specific objects in the actual environment, avoiding blind guessing and utilizing existing environmental perception data. To more finely evaluate these candidate physical targets, the spatial relationship between each candidate physical target and the electric wheelchair, including relative distance, spatial position, and orientation angle, is obtained. This multi-dimensional spatial data provides a comprehensive understanding of the relative state between the target and the electric wheelchair; for example, one target may be close but not in the user's line of sight, while another may be far away but directly facing the user. Based on these spatial relationships, the system prioritizes multiple candidate physical targets. The ranking criteria can comprehensively consider factors such as distance, angle, and user historical preferences to filter out the targets most likely to match the user's true intent. After prioritization, the system makes a decision based on the results. If there is only one candidate physical target with the highest priority, that target is explicitly selected as the new target information for subsequent generation of embodying clarification queries. This ensures that the system can respond directly and accurately when the user's intent is clear. However, if there are multiple candidate physical targets with the highest priority, it indicates that the user's negative response still has a certain degree of ambiguity and cannot be completely determined by a single dimension. In this case, the system generates descriptive information containing the differentiating characteristics among these candidate physical targets. These differentiating characteristics, such as color, size, or specific identifiers, will serve as key components of subsequent embodying clarification queries, guiding the user to make more precise references.
[0088] Through the steps described above, the proposed solution effectively extracts more precise new target information from the user's negative response and combines this information with environmental perception data from the electric wheelchair for inference. This allows the system to proactively and intelligently narrow down the target range and provide more guided clarification questions when faced with negative user feedback, rather than simply repeating questions or erroneously discarding information. This mechanism, combined with the electric wheelchair's ability to acquire perception data and identify physical targets, creates an effective synergy, making the entire interaction process smoother and more accurate. This avoids negative interaction loops caused by information loss or misattribution, significantly improving user experience and the accuracy of command execution.
[0089] In another embodiment of this application, the sub-step of S5000—generating instruction parameters based on the voice response and controlling the electric wheelchair to perform corresponding actions based on the instruction parameters—further includes:
[0090] S5230: When the voice response is affirmative, generate instruction parameters based on the target name and location information contained in the embodied clarification inquiry;
[0091] S5240: Calculate the parking distance of the physical target based on the location information, and use the parking distance as the execution boundary of the command parameter to control the electric wheelchair to move along the direction of the physical target corresponding to the target name until it stops at a preset safe distance from its parking position.
[0092] In this context, a voice response with affirmative semantics refers to a user's confirmation of the specific clarification inquiry issued by the system. Specifically, this can be achieved by using voice recognition technology to identify words or phrases such as "yes," "right," "okay," or "confirm" that indicate agreement or affirmation. The purpose is to clarify the user's acceptance and recognition of the current clarification content.
[0093] The target name and location information included in the embodied clarification inquiry refers to the physical entity that the user intends to point to and its specific spatial coordinates or relative position in the environment, which the system determines after recognizing a vague instruction, perceiving the environment, and interacting with the user for confirmation. Specifically, it can be generated by acquiring environmental perception data through LiDAR, visual sensors, etc., and combining it with semantic understanding. Its purpose is to provide clear target guidance for the subsequent movement of the electric wheelchair.
[0094] Command parameters refer to a set of quantified commands used to control the electric wheelchair to perform specific actions. These can include movement direction, speed, acceleration, and, in this solution, a stopping distance. The purpose is to translate the user's intentions into precise control commands that the electric wheelchair can execute. The stopping distance of the physical target refers to the distance at which the electric wheelchair needs to begin decelerating and eventually stop while moving towards the target. This distance can be calculated based on factors such as the size, shape, and material of the physical target, as well as the electric wheelchair's current speed and braking performance. The purpose is to ensure that the electric wheelchair can stop safely and smoothly when approaching the target. The execution boundary of the command parameters refers to the critical value that limits the range of movement or stopping conditions when the electric wheelchair executes movement commands. Specifically, the calculated stopping distance of the physical target can be used as the trigger condition for stopping the electric wheelchair. The purpose is to provide a safe and reliable stopping mechanism for the autonomous movement of the electric wheelchair. The preset safety distance refers to the minimum safe interval that should be maintained between the electric wheelchair and the physical target when the electric wheelchair finally stops. This distance can be set based on factors such as the size of the electric wheelchair, the user's safety needs, and environmental obstacles. The purpose is to avoid collisions between the electric wheelchair and the target, ensuring the safety of the user and the surrounding environment.
[0095] This application's solution achieves safe and intelligent movement of an electric wheelchair towards a target location through precise semantic understanding of the user's voice response, combined with environmental perception data. When the electric wheelchair system receives a user's voice response to a clarification inquiry, and the response is recognized as affirmative, the system immediately extracts the user-confirmed target name and location information from the clarification inquiry. This information is the result of the system's initial understanding of ambiguous commands and environmental perception, confirmed through interaction with the user in previous steps, thus ensuring the accuracy of the commands. Based on this, the system uses the extracted location information to accurately calculate the stopping distance required for the electric wheelchair to approach the physical target. This stopping distance calculation considers the electric wheelchair's motion characteristics and environmental factors, aiming to provide a key safety parameter for subsequent movement control. Subsequently, this calculated stopping distance is set as the execution boundary of the command parameters, meaning that the electric wheelchair's movement will be strictly constrained by this boundary. The electric wheelchair will then begin moving along the direction of the physical target corresponding to the target name. During the movement, the system continuously monitors the real-time distance between the electric wheelchair and the target. Once the distance between the electric wheelchair and the target reaches or approaches the preset safe distance, the system will trigger a stopping mechanism, bringing the electric wheelchair to a smooth stop. This mechanism ensures that the electric wheelchair does not approach the target excessively, thus avoiding potential collision risks and significantly improving the safety of movement. In this way, this solution is closely integrated with the basic voice command recognition and clarification mechanism, forming a complete, closed-loop intelligent interaction and control process. The basic mechanism provides clear user intent and target information, while this solution, based on this, introduces refined distance calculation and safe parking control, transforming abstract "affirmative" intents into precise and safe physical movement. This combination enables the electric wheelchair not only to understand the user's explicit commands but also to execute these commands in a safe and controllable manner, solving the problem of how to achieve precise and safe movement to the target location after the user confirms the target, and avoiding safety hazards caused by a lack of precise parking distance control.
[0096] In some preferred embodiments, this application is implemented as follows:
[0097] Suppose a user issues a vague command, "Go there." The electric wheelchair system senses the environment and generates a clarifying question: "Do you want to go to the sofa?" When the user's voice response is affirmative, such as "Yes, the sofa," the system immediately recognizes this affirmative meaning. The system then extracts the target name "sofa" and its corresponding location information from the previously generated clarifying question, such as the sofa's three-dimensional coordinates or relative orientation within the electric wheelchair's current coordinate system. Based on this location information, the system activates a distance calculation module. This module uses real-time environmental data acquired by LiDAR or a depth camera, combined with the electric wheelchair's current speed and braking performance model, to calculate the stopping distance at the physical target where the electric wheelchair needs to begin decelerating and eventually stop. For example, if there is sufficient space in front of the sofa, the system might calculate a distance that allows the electric wheelchair to decelerate smoothly. The calculated stopping distance is then set as the execution boundary of the command parameters. This means that when the electric wheelchair moves towards the sofa, its internal control system uses this distance as the trigger condition for stopping. The electric wheelchair's navigation and control module plans a path and controls the electric wheelchair to move in the direction of the sofa. During movement, the system continuously monitors the real-time distance between the electric wheelchair and the sofa. Once the distance reaches a preset safe distance, such as 0.3 meters, the wheelchair's drive system receives a stop command and immediately applies the brakes, ensuring the wheelchair comes to a smooth stop while maintaining a safe distance from the sofa. This allows users to safely approach the sofa without worrying about a collision.
[0098] In another embodiment of this application, S6000 is further proposed to include:
[0099] S6100: Obtain the current user's identity information and context information;
[0100] S6200: Based on identity information and contextual information, it records interaction logs including voice commands, embodying clarification inquiries, voice responses, and command parameters;
[0101] S6300: Filter out historical interaction records from the interaction log that match the current user's identity and context information;
[0102] S6400: Adjust the attribution strategy content based on historical interaction records;
[0103] S6500: When a change in user identity information or context information is detected, load the attribution strategy corresponding to the new identity information or context information.
[0104] Among them, identity information refers to relevant data used to uniquely identify users, specifically user ID, age, gender, health status, or personalized preference settings. Its purpose is to distinguish different users so that the system can personalize processing based on individual user characteristics. Contextual information refers to the environmental context data in which the user issues commands, specifically the current time, geographical location (e.g., indoor or outdoor, specific room type), weather conditions, the distribution of obstacles in the surrounding environment, or the current operating status of the electric wheelchair (e.g., speed, battery level). Its purpose is to provide richer background information for the system to understand user intentions, because the same command may have different meanings in different contexts. Interaction logs refer to the structured data collection that records all interactions between the user and the electric wheelchair, specifically text, voice, system responses, etc., stored in time series format. Its purpose is to provide comprehensive historical data support for subsequent attribution strategy optimization. Attribution strategy content refers to the set of rules used by the system to map fuzzy commands to specific physical targets or operating parameters, specifically a set of preset logical rules, a decision tree model, or the parameters and weights of a machine learning model.
[0105] This application's solution refines the recording of interaction logs and the optimization of attribution strategies by introducing user identity information and contextual information, thereby improving the accuracy of the system's understanding of user intent. Specifically, this solution achieves refined optimization of voice interaction log recording and attribution strategies by introducing user identity information and contextual information, significantly improving the system's accuracy in understanding user intent. The system first acquires contextual information such as the current user's identity characteristics and environment as the semantic basis for subsequent processing, moving beyond the literal meaning of the instructions themselves. Based on this, the system records interaction logs including voice commands, embodied clarification inquiries, voice responses, and command parameters, along with complete identity and contextual metadata, forming structured contextual behavioral data. Subsequently, the system filters interaction content highly relevant to the current identity and context from historical records, identifies the expression habits and response patterns of specific users in similar environments, and then personalizes the attribution strategy. For example, if a user consistently gives a negative response to clarification inquiries in a certain context, the trigger frequency of such clarifications should be reduced. The system also supports real-time detection of identity or contextual changes; once a change occurs, it dynamically switches the matching attribution strategy to maintain continuous adaptation between the interaction strategy and the user's state. This solution reconstructs the attribution logic from a complete perspective of "user-context-command-response," avoiding one-sided judgments based on static logs, effectively reducing misjudgments and redundant clarifications, and improving the intelligence and user experience of the electric wheelchair voice command recognition system.
[0106] In some preferred embodiments, this application is implemented as follows:
[0107] First, when the electric wheelchair is started or the user logs in, the system can obtain the current user's identity and contextual information. For example, identity information can be read from the user's logged-in account information, including the user's unique ID, preset language preferences, and historical usage habit tags. Contextual information can be obtained through the electric wheelchair's built-in sensor array, such as obtaining geographical location information (indoor / outdoor) through the GPS module, obtaining light intensity (day / night) through the ambient light sensor, analyzing background noise (quiet / noisy) through the microphone array, and obtaining current speed and tilt angle through the wheelchair's own odometer and posture sensors. Next, based on this obtained identity and contextual information, the system records an interaction log containing voice commands, embodied clarification queries, voice responses, and command parameters. These logs can be stored in structured JSON or XML format on the electric wheelchair's local storage or a cloud server. Each log record, in addition to containing the original voice command text, the system-generated embodied clarification query text, the user's voice response text, and the finally parsed command parameters (such as target name, location coordinates, movement speed, etc.), also includes contextual metadata such as timestamps, user ID, geographical location tags, and ambient noise levels. Subsequently, when the attribution strategy needs optimization, the system filters historical interaction records from the interaction log database that match the current user's identity and contextual information. For example, if the current user is "User A" and the current context is "indoor, quiet environment," the system will query the logs to find all historical commands issued by "User A" in the "indoor, quiet environment" and their corresponding clarification and response records. The filtering can be based on a preset similarity threshold; for example, a contextual information matching degree of 80% or higher is considered a match. Then, the system adjusts the attribution strategy content based on these filtered historical interaction records. For example, the system can run a machine learning model (e.g., a decision tree-based or neural network-based classifier) that uses historical interaction records as training data to learn which ambiguous commands tend to be attributed to which specific targets, or which clarification inquiries are more likely to lead to effective responses, in specific user and context situations. By analyzing patterns of attribution errors in historical data, the model can automatically update its internal weights or rules to reduce future attribution errors. For example, if it's found that "User A" is in an "indoor" environment, and when saying "go over there," they usually mean "go to the coffee table," then the attribution strategy can be adjusted to prioritize attributing "go over there" to "coffee table" in this user and context. Finally, to ensure the dynamic adaptability of the attribution strategy, the system continuously monitors changes in user identity information or contextual information. For example, when the GPS module detects that the electric wheelchair has moved from an "indoor" area to an "outdoor" area, or when the user switches "driving mode" via voice command (e.g., from "slow mode" to "fast mode," which can be considered a contextual change), the system will immediately detect these changes.Once a change is detected, the system loads the attribution strategy corresponding to the new identity information or new contextual information from a pre-stored attribution strategy library. For example, if a user moves from indoors to outdoors, the system can load an attribution strategy optimized for the complex outdoor environment (such as more obstacles or a wider space). This strategy may interpret "going there" differently than in the indoor context, thus better adapting to the new environmental challenges.
[0108] In another embodiment of this application, the attribution strategy further includes the following steps:
[0109] A1: Based on the perception data, identify all identified physical targets located in front of the electric wheelchair and their corresponding relative distance, spatial position, and orientation angle;
[0110] A2: Based on the preset exclusion rules, the identified physical targets are filtered: identified physical targets that are obstructed by obstacles from the electric wheelchair, identified physical targets whose passage path width to the electric wheelchair is less than a preset threshold, and / or identified physical targets whose relative distance to the electric wheelchair is less than a safe distance threshold are excluded, thus obtaining the candidate physical targets;
[0111] A3: Select the candidate physical target with the smallest angle to the current direction of travel of the electric wheelchair as the source of target name and location information to be referenced when generating the embodied clarification inquiry.
[0112] Among them, the preset exclusion rules refer to a set of conditions used to filter and exclude physical objects that are not suitable as targets for clarification inquiries. These rules can be implemented using techniques such as geometric analysis, path planning algorithms, or safe distance judgment. Their purpose is to ensure that the targets selected subsequently are reachable, safe, and relevant to the user's intent.
[0113] This application's solution effectively solves the problem of generating inaccurate or irrelevant clarification queries in multi-target scenarios by performing refined perception and intelligent filtering of the environment in front of the electric wheelchair. Specifically, firstly, the system comprehensively identifies all identified physical targets in front of the electric wheelchair based on perception data, and obtains the relative distance, spatial position, and orientation angle of these targets. This step lays the data foundation for subsequent intelligent filtering, ensuring that all potential interaction objects are considered. Based on this, the system further rigorously filters these identified physical targets according to preset exclusion rules. By excluding physical targets with obstructions, insufficient path width, or excessively close relative distances, the solution effectively eliminates interference items that are impassable, pose safety hazards, or are unsuitable as interaction objects, thus obtaining a more feasible set of candidate physical targets. Finally, from these candidate physical targets, the system selects the target with the smallest angle to the electric wheelchair's current direction of travel as the source of target name and location information referenced when generating personalized clarification queries. This selection mechanism enables the generated clarification queries to be highly focused on the target that the user is most likely to be concerned about or intends to reach, significantly improving the accuracy and relevance of the clarification queries. By organically combining the above steps, this solution, building upon the generation of embodied clarification queries proposed in electric wheelchair voice command recognition methods, no longer simply selects from all identified physical targets, but introduces a multi-stage intelligent filtering mechanism. This mechanism enables the system to accurately identify targets that highly match the user's intent and are actually achievable from complex and ever-changing environments, thereby avoiding invalid or misleading clarification queries caused by inappropriate target selection. This not only significantly reduces the user's cognitive burden and improves the efficiency and fluency of human-computer interaction, but also ensures that the electric wheelchair can more accurately understand and execute the user's ambiguous commands, effectively avoiding the negative interaction loop caused by misattribution in the background technology, enabling the system to assist users in mobility more intelligently and safely.
[0114] In another embodiment of this application, the attribution strategy further includes the following steps:
[0115] A4: Based on the physical target and its relative distance, determine whether the physical target belongs to the preset sensitive target type, and whether the relative distance of the physical target is less than the preset safe distance;
[0116] A5: If the physical target is a sensitive target type, or the relative distance is less than the safe distance, generate a definite clarification query containing the target name and location information;
[0117] A6: If the physical target is not a sensitive target type and the relative distance is not less than the safe distance, then determine whether the physical target is within the preset movement path range of the electric wheelchair;
[0118] A7: If the physical target is within the movement path, generate a specific clarification query that specifies the target name and location information of the physical target.
[0119] Sensitive target types refer to categories of physical targets that are of particular importance or pose a potential risk to the electric wheelchair user or the surrounding environment. These can include, but are not limited to, pedestrians, pets, children, fragile items, hazardous materials, or specific area boundaries. The purpose is to enable the system to prioritize and handle targets that directly impact user safety or operational experience. Safe distance refers to the minimum permissible distance threshold between the electric wheelchair and physical targets. It can be dynamically or statically set based on factors such as the wheelchair's speed, braking performance, environmental complexity, and user preferences. Its purpose is to ensure the electric wheelchair has sufficient reaction time or space when approaching a specific target, avoiding collisions or unnecessary risks. The movement path range refers to the spatial area covered by the electric wheelchair's expected or planned trajectory when executing commands or autonomous navigation. This range can be dynamically calculated or preset by the electric wheelchair's navigation system based on current location, target point, environmental map, and obstacle avoidance strategies. Its purpose is to limit the system to clarifying questions only about targets related to the electric wheelchair's actual movement intentions, avoiding unnecessary interference with the user.
[0120] This solution proposes a hierarchical and prioritized attribution strategy, significantly improving the accuracy and safety of handling ambiguous voice commands in complex environments for electric wheelchairs. When the system receives an ambiguous command and identifies physical targets and their relative distances in the environment, it first assesses the safety and importance of the target. Based on the target type and its distance from the user, the system determines whether it belongs to a pre-defined sensitive target or is too close. If the target is sensitive or within a safe distance, the system immediately generates a clarification query containing the target's name and location, ensuring the user can promptly identify potential risks and avoid safety hazards. For non-sensitive targets at sufficient distance, the system proceeds to a second level of judgment, triggering a clarification query only if the target is on the electric wheelchair's path. This hierarchical approach avoids querying all targets, reducing invalid interactions and improving interaction efficiency. The overall strategy balances safety and user experience. By accurately filtering targets and dynamically adjusting the clarification strategy, the system can make judgments based on more accurate target information that aligns with the user's true intentions before generating command parameters and control actions, thereby improving the accuracy and reliability of the entire ambiguous voice command recognition and execution process.
[0121] In some preferred embodiments, this application is implemented as follows:
[0122] When an electric wheelchair receives a vague voice command, such as "go there," the system first acquires perception data of the current environment. For example, it identifies surrounding physical targets using LiDAR, cameras, or ultrasonic sensors and calculates their relative distances to the wheelchair. Suppose the system detects a pedestrian (physical target A) 1.5 meters ahead of the wheelchair; a trash can (physical target B) 2 meters to the right; and a pet (physical target C) 1 meter to the left. The system then makes a judgment based on a preset attribution strategy. First, it determines whether physical target A (the pedestrian) belongs to a preset sensitive target type. Assuming "pedestrian" is defined as a sensitive target type, and the preset safe distance is 2 meters, since the relative distance of pedestrian A (1.5 meters) is less than the safe distance of 2 meters, the system immediately generates a specific clarifying question containing "pedestrian A, 1.5 meters ahead," such as a voice announcement, "Do you mean the pedestrian 1.5 meters ahead?" If there are no sensitive targets or targets within the safe distance at this time, the system continues to assess other targets. For example, for physical target B (trash can), it is not a sensitive target type, and its relative distance of 2 meters is not less than the safe distance of 2 meters. The system will further determine whether trash can B is within the preset movement path of the electric wheelchair. If the electric wheelchair's current movement path is planned to pass near trash can B, the system will generate a specific clarification question containing "trash can B, 2 meters to the right," such as "Do you mean the trash can 2 meters to the right?". For physical target C (pet), it is a sensitive target type, and its relative distance of 1 meter is less than the safe distance of 2 meters. The system will prioritize generating a specific clarification question containing "pet C, 1 meter to the left front," such as "Do you mean the pet 1 meter to the left front?". Through this implementation, the system can prioritize targets based on their type and distance, ensuring that critical or potentially risky targets are clarified first, while avoiding unnecessary inquiries about irrelevant or off-path targets, thereby improving interaction efficiency and security.
[0123] In another embodiment of this application, the sub-step of S1000, identifying ambiguous instructions in voice commands, includes:
[0124] S1210: Parse the voice commands and extract parameter information including the operation object, operation method, and spatial location;
[0125] S1220: Determine whether the voice command has ambiguous features based on parameter information;
[0126] S1230: If fuzzy features are identified, the corresponding fuzzy instruction in the voice command is identified based on the fuzzy features.
[0127] Parsing refers to in-depth linguistic and semantic analysis of the input voice commands. This can be achieved using Natural Language Processing (NLP) technology, speech recognition post-processing algorithms, or rule-based semantic analysis engines. Its purpose is to transform unstructured voice input into structured data that machines can understand and process. Parameter information refers to key data points extracted from the voice commands that describe the core elements of the command. Specifically, this can include the target entity affected by the command, the specific action performed, and the spatial orientation or location information related to the command. Its purpose is to provide a basis for subsequent judgment of the command's clarity. Fuzzy features refer to linguistic or semantic attributes in the voice commands that cause their semantics to be unclear, incomplete, or uncertain. Specifically, this can manifest as a lack of key information, imprecise information description, or ambiguity. Its purpose is to indicate that the command needs further clarification or refinement. Fuzzy commands are voice commands that cannot be directly executed due to the presence of one or more fuzzy features. Specifically, they can be classified and identified based on the type of fuzzy features, such as fuzzy object, fuzzy operation method, or fuzzy spatial location. Its purpose is to provide a clear classification basis for subsequent targeted clarification or processing.
[0128] The solution proposed in this application systematically analyzes received voice commands. First, the voice commands are analyzed to extract key parameter information including the object of operation, the operation method, and the spatial location. This process transforms the original unstructured voice commands into structured data. Based on this extracted parameter information, the system evaluates the voice commands to determine whether there are any ambiguities. This judgment mechanism can identify situations where the commands may contain semantic ambiguity, incomplete information, or unclear direction. Once ambiguity is identified in the voice command, the system further identifies the corresponding ambiguity command type based on these specific ambiguities. For example, if the command lacks an object of operation, it is identified as an ambiguous command about the object of operation; if the description of the operation method is unclear, it is identified as an ambiguous command about the operation method. Through this progressive analysis, judgment, and recognition process, this solution can accurately locate and classify the uncertainty in voice commands. This precise ambiguity command recognition mechanism is closely integrated with the overall electric wheelchair voice command recognition method of this application. After receiving a user's voice command, it no longer simply judges whether it is ambiguous, but deeply analyzes the specific type of ambiguity. This detailed recognition result provides more accurate input for subsequent steps such as acquiring environmental perception data based on fuzzy commands, identifying physical targets, and generating specific clarifying queries. For example, when a spatial location fuzzy command is identified, the system can more specifically acquire perception data related to the spatial location and generate more targeted clarifying queries, avoiding invalid clarifications or information loss caused by inaccurate fuzzy command recognition in traditional methods. In this way, this solution can effectively avoid negative interaction loops caused by inaccurate judgment of command fuzziness, improving the efficiency and accuracy of interaction between the electric wheelchair and the user.
[0129] In some preferred embodiments, specifically, when the electric wheelchair system receives the voice command "go there" from the user, the system first activates the voice command parsing module. This module can utilize a deep learning model, such as a semantic parser based on the Transformer architecture, to analyze the text content of the voice command. Through analysis, the system can identify that "go" is an operation method, but "there," as a spatial location parameter, has an ambiguous specific target and lacks a clear operation object. Therefore, the parsing module extracts the parameter information that the operation method is "go," the spatial location is "there," and the operation object is empty. Next, the system determines whether the voice command has ambiguity features based on these parameter information. Since the operation object is missing and the spatial location "there" is ambiguous, the system determines that the voice command has ambiguity features. Subsequently, based on the identified ambiguity features, the system further identifies the corresponding ambiguity command type in the voice command. In this case, due to the missing operation object and the ambiguous spatial location, the system can identify the command as an "ambiguous operation object command" and a "ambiguous spatial location command." This identification result will serve as the basis for generating subsequent clarifying questions. For example, the system can generate specific clarifying questions such as "Which destination do you want to go to? Do you want to go to the table?" to guide users to provide more accurate information.
[0130] In another embodiment of this application, S1220 further includes:
[0131] S1221: Obtain the current physical state information and environmental context information of the electric wheelchair;
[0132] S1222: Combine parameter information, current physical state information, and environmental context information to evaluate the executability and semantic integrity of voice commands;
[0133] S1223: When the evaluation results indicate that the voice command has an unclear execution object, unclear path target and / or uncertain execution scope, the voice command is determined to be an ambiguous command.
[0134] The current physical state information of the electric wheelchair refers to the internal operating parameters and state data of the electric wheelchair at a specific moment. This can be achieved by real-time collection and reporting of data from sensors built into the wheelchair (such as speed sensors, posture sensors, battery sensors, odometers, etc.), aiming to provide the wheelchair's own executability constraints. Environmental context information refers to the real-time perception data and situational description of the external environment in which the electric wheelchair is located. This can be achieved by acquiring data such as the distribution of surrounding obstacles, spatial structure, lighting conditions, and location coordinates from external sensing devices mounted on the wheelchair (such as cameras, LiDAR, ultrasonic sensors, GPS modules, etc.), aiming to provide external environmental constraints and semantic references for command execution. Evaluating the executability and semantic integrity of voice commands refers to the process of comprehensively judging whether voice commands can be effectively executed under the current wheelchair state and environmental conditions, and whether their meaning is clear and unambiguous. This can be achieved using rule-based reasoning engines, machine learning models, or expert systems, combined with pre-set execution logic. The system uses semantic rules to make judgments, aiming to comprehensively consider the actual feasibility and clarity of the instruction. Unclear execution object refers to the lack of a clearly defined entity or target in the voice instruction, making it impossible for the system to determine the specific operation object. This can manifest as the use of vague pronouns or generic terms without specific names in the instruction. The purpose is to identify situations where the operation target is missing or unclear in the instruction. Unclear path target refers to the failure to clearly specify the endpoint or direction of movement or operation in the voice instruction, making it impossible for the system to plan a specific execution path. This can manifest as the use of vague spatial indicators or the lack of specific coordinates or landmark information in the instruction. The purpose is to identify situations where the endpoint or direction of movement in the instruction is vague. Uncertain execution range refers to the failure to clearly specify the degree, quantity, or duration of the operation in the voice instruction, making it impossible for the system to determine the specific operation quantity. This can manifest as the use of vague quantifiers or the lack of specific numerical values or range limitations in the instruction. The purpose is to identify situations where quantitative information about the operation is missing in the instruction.
[0135] This application's solution, when determining whether a voice command has ambiguity, no longer relies solely on parameter information extracted after parsing the voice command, but further acquires the current physical state information and environmental context information of the electric wheelchair. This information, such as the wheelchair's battery level, speed, tilt angle, location, distribution of surrounding obstacles, and light intensity, provides crucial background data for the actual executability of the command. Subsequently, the system comprehensively considers this acquired current physical state information and environmental context information along with the parameter information extracted from the voice command, thereby evaluating the actual executability and semantic completeness of the voice command. This evaluation process goes beyond simple grammatical or lexical analysis; it delves into the feasibility of the command in the real world. For example, even if the command appears semantically complete, if the wheelchair's battery is too low to execute, or if there are obstacles blocking the path, the system can still identify its inexecutability. Finally, when this comprehensive evaluation indicates that the voice command has an unclear execution object, an unclear path target, and / or an uncertain execution range, the system determines that the voice command is an ambiguous command. This judgment mechanism enables the system to more accurately identify instructions that are truly impossible to execute or are ambiguous in a specific context, avoiding misjudging context-limited instructions as clear instructions and avoiding unnecessary clarification of context-understandable instructions. In this way, the solution of this application can more comprehensively consider the actual context of the instruction when recognizing ambiguous instructions in voice commands, thereby improving the accuracy of ambiguous instruction judgment, reducing unnecessary interactions, and enabling the voice interaction system of the electric wheelchair to respond more intelligently to the user's true intentions.
[0136] Reference Figure 2 In another embodiment of this application, a voice command recognition system for an electric wheelchair is further proposed, the system comprising:
[0137] The receiving and recognition module 1 is used to receive the user's voice commands and recognize ambiguous commands in the voice commands.
[0138] The perception data acquisition module 2 is used to acquire perception data related to the current environment of the electric wheelchair based on fuzzy commands;
[0139] The target distance recognition module 3 is used to identify physical targets in the perception data, as well as the relative distance between the physical targets and the electric wheelchair;
[0140] The query generation module 4 is used to generate a definite clarifying query containing the target name and location information based on the physical target and its relative distance, according to a preset attribution strategy.
[0141] The execution module 5 is used to receive the user's voice response to the personalized clarification inquiry, generate instruction parameters based on the voice response, and control the electric wheelchair to perform corresponding actions based on the instruction parameters.
[0142] The recording optimization module 6 is used to record voice commands, embodied clarification inquiries, voice responses, and command parameters as interaction logs, and optimize the attribution strategy based on the interaction logs.
[0143] Specifically, the receiving and recognition module 1 can consist of a speech recognition engine and a natural language processing unit. The speech recognition engine is responsible for converting the user's voice commands into text, such as using deep learning models for acoustic and language modeling. The natural language processing unit performs semantic parsing on the text commands, identifying parameters such as the object of operation, the operation method, and the spatial location, and determining whether there are fuzzy features, such as identifying ambiguous expressions like "a little forward" or "to the left" through keyword matching, syntactic analysis, or intent recognition algorithms.
[0144] The perception data acquisition module 2 can be configured with various sensors, such as lidar for acquiring high-precision environmental point cloud data, stereo cameras for capturing visual image information, and ultrasonic sensors for near-range obstacle detection. The data from these sensors is processed through a data fusion algorithm to construct a real-time 3D map or obstacle distribution map of the environment surrounding the electric wheelchair.
[0145] The target distance recognition module 3 can use computer vision algorithms to identify specific physical targets from visual images, such as "tables," "chairs," and "doors." Simultaneously, by combining LiDAR or ultrasonic data, it accurately measures the relative distance and spatial position between these physical targets and the electric wheelchair through point cloud clustering, target tracking, or geometric calculations.
[0146] The query generation module 4 can have a built-in knowledge graph or rule engine to store preset attribution strategies. Upon recognizing fuzzy instructions and physical targets in the environment, this module selects one or more candidate targets based on the attribution strategy and generates a clarifying query containing the target's name and location information. For example, if there is a table and a chair in front of you, the system might generate questions like, "Do you mean to move towards the table two meters ahead?" or "Do you want to move closer to the chair on your left?".
[0147] The execution module 5 may include a speech synthesizer for broadcasting clarification inquiries, as well as an instruction parser and a motion control unit. The instruction parser receives the user's voice response to the clarification inquiries and extracts the user-confirmed target information or new instruction intent through speech recognition and semantic understanding. The motion control unit then generates specific motor control signals based on the parsed instruction parameters to drive the hub motors of the electric wheelchair to perform actions such as forward movement, backward movement, turning, or stopping.
[0148] The recording optimization module 6 can be a database system used to store interaction logs. Each complete voice interaction is recorded as a log entry, including a timestamp, user ID, original voice command, clarification question issued by the system, user's voice response, and final command parameters. A background machine learning model can periodically analyze these interaction logs to identify user behavior patterns and the effectiveness of attribution strategies. For example, if a certain attribution strategy frequently leads to negative responses from users, the system can adjust the weight or rules of that strategy to improve the accuracy of subsequent clarification questions and user satisfaction.
[0149] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A voice command recognition method for an electric wheelchair, characterized in that, The method includes: Receive user voice commands and identify ambiguous commands within the voice commands; Based on the fuzzy instructions, obtain perception data related to the current environment of the electric wheelchair; Identify physical targets in the perceived data, and the relative distance between the physical targets and the electric wheelchair; Based on a preset attribution strategy, a definite clarifying query containing the target name and location information is generated from the physical target and its relative distance. Receive the user's voice response to the personalized clarification inquiry, generate instruction parameters based on the voice response, and control the electric wheelchair to perform corresponding actions based on the instruction parameters; The voice commands, embodied clarification questions, voice responses, and command parameters are recorded as interaction logs, and the attribution strategy is optimized based on the interaction logs. The attribution strategy includes: Based on the physical target and its relative distance, determine whether the physical target belongs to a preset sensitive target type, and whether the relative distance of the physical target is less than a preset safe distance; If the physical target belongs to the sensitive target type, or the relative distance is less than the safe distance, then the embodied clarification query containing the target name and location information is generated; If the physical target is not a sensitive target type and the relative distance is not less than the safe distance, then it is determined whether the physical target is within the preset movement path range of the electric wheelchair; If the physical target is located within the movement path, then a specific clarification query is generated that specifies the target name and location information of the physical target.
2. The voice command recognition method for electric wheelchairs according to claim 1, characterized in that, The step of generating instruction parameters based on the voice response and controlling the electric wheelchair to perform corresponding actions based on the instruction parameters includes: When the user's voice response contains negative semantics, the voice response is parsed to obtain new target information and its relative distance; Based on the new target information and its relative distance, a new embodied clarification question is generated, and after receiving the user's voice response, the instruction parameters are extracted. Based on the instruction parameters, the electric wheelchair is controlled to perform corresponding actions.
3. The voice command recognition method for electric wheelchairs according to claim 2, characterized in that, The step of parsing the voice response to obtain new target information and its relative distance when the user's response contains negative semantics includes: Based on the new target information and its relative distance, multiple candidate physical targets are matched from the detected and identified physical targets; Obtain the spatial relationship between multiple candidate physical targets and the electric wheelchair, including relative distance, spatial position, and orientation angle; Based on the spatial relationships, the multiple candidate physical targets are prioritized. If there is only one candidate physical target with the highest priority, select the candidate physical target with the highest priority as the new target information; If there are multiple candidate physical targets with the highest priority, descriptive information containing the differentiated features among the candidate physical targets is generated for subsequent generation of embodied clarification queries.
4. The voice command recognition method for electric wheelchairs according to claim 1 or 2, characterized in that, The method further includes generating instruction parameters based on the voice response and controlling the electric wheelchair to perform corresponding actions based on the instruction parameters, and also includes: When the voice response is affirmative, instruction parameters are generated based on the target name and location information contained in the embodied clarification inquiry; The parking distance of the physical target is calculated based on the location information, and this parking distance is used as the execution boundary of the command parameter to control the electric wheelchair to move along the direction of the physical target corresponding to the target name until it stops at a preset safe distance from its parking position.
5. The voice command recognition method for electric wheelchairs according to claim 1, characterized in that, The step of recording the voice commands, embodied clarification inquiries, voice responses, and command parameters as an interaction log, and optimizing the attribution strategy based on the interaction log, includes: Obtain the current user's identity and context information; Based on the identity information and context information, an interaction log is recorded, including the voice command, the embodying clarification inquiry, the voice response, and the command parameters. Filter out historical interaction records from the interaction logs that match the current user's identity and context information; Adjust the attribution strategy content based on the historical interaction records; When a change in user identity information or contextual information is detected, the attribution strategy corresponding to the new identity information or contextual information is loaded.
6. The voice command recognition method for electric wheelchairs according to claim 1, characterized in that, The attribution strategy includes: Based on the perception data, identify all identified physical targets located in front of the electric wheelchair and their corresponding relative distance, spatial position, and orientation angle; According to the preset exclusion rules, the identified physical targets are screened: identified physical targets that are obstructed by obstacles from the electric wheelchair, identified physical targets whose passage path width to the electric wheelchair is less than a preset threshold, and / or identified physical targets whose relative distance to the electric wheelchair is less than a safe distance threshold are excluded, and candidate physical targets are obtained. The candidate physical target with the smallest angle to the current direction of travel of the electric wheelchair is selected as the source of the target name and location information referenced when generating the embodiment clarification inquiry.
7. The voice command recognition method for electric wheelchairs according to claim 1, characterized in that, The process of recognizing ambiguous commands in the voice commands includes: The voice command is parsed to extract parameter information including the operation object, operation method, and spatial location; Based on the parameter information, determine whether the voice command has any ambiguity features; If a fuzzy feature is identified, the corresponding fuzzy instruction in the voice command is identified based on the fuzzy feature.
8. The voice command recognition method for electric wheelchairs according to claim 7, characterized in that, The step of determining whether the voice command has ambiguous features based on the parameter information includes: Obtain the current physical state information and environmental context information of the electric wheelchair; By combining the parameter information, the current physical state information, and the environmental context information, the executability and semantic integrity of the voice command are evaluated. When the evaluation results indicate that the voice command has an unclear execution target, an unclear path goal, and / or an uncertain execution scope, the voice command is judged to be an ambiguous command.
9. A voice command recognition system for an electric wheelchair, characterized in that, The system includes: The receiving and recognition module is used to receive the user's voice commands and recognize ambiguous commands in the voice commands; The perception data acquisition module is used to acquire perception data related to the current environment of the electric wheelchair based on the fuzzy instructions. The target distance recognition module is used to identify physical targets in the perceived data, as well as the relative distance between the physical targets and the electric wheelchair. The query generation module is used to generate a definite clarifying query containing the target name and location information based on the physical target and its relative distance, according to a preset attribution strategy. The execution module is used to receive the user's voice response to the personalized clarification inquiry, generate instruction parameters based on the voice response, and control the electric wheelchair to perform corresponding actions based on the instruction parameters. The recording optimization module is used to record the voice commands, embodied clarification inquiries, voice responses and command parameters as interaction logs, and optimize the attribution strategy based on the interaction logs. It is also used to determine, based on the physical target and its relative distance, whether the physical target belongs to a preset sensitive target type, and whether the relative distance of the physical target is less than a preset safe distance; If the physical target belongs to the sensitive target type, or the relative distance is less than the safe distance, then the embodied clarification query containing the target name and location information is generated; If the physical target is not a sensitive target type and the relative distance is not less than the safe distance, then it is determined whether the physical target is within the preset movement path range of the electric wheelchair; If the physical target is located within the movement path, then a specific clarification query is generated that specifies the target name and location information of the physical target.
Citation Information
Patent Citations
Voice interaction method, device, equipment of intelligent voice equipment, medium and product
CN112767916A
Automatic driving wheelchair navigation method and system based on intelligent voice interaction
CN119268707A
Humanoid robot multi-mode instruction analysis system
CN120516701A
Robot Natural Language Term Disambiguation and Entity Labeling
US20190102377A1