Electric wheelchair voice instruction recognition method and system

By recognizing ambiguous commands from electric wheelchairs, generating specific clarifying questions, and optimizing attribution strategies, the system solves the problems of information loss and inaccurate command execution caused by the single clarification method in electric wheelchair voice command recognition systems, thereby improving the system's intelligence level and user experience.

CN120808767AActive Publication Date: 2025-10-17深圳复成医疗科技有限公司

Patent Information

Application Number
CN202511311217.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-10-17
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

Existing voice command recognition systems for electric wheelchairs suffer from limited clarification methods when faced with ambiguous commands. This results in the inability to effectively recognize and utilize supplementary information provided by the user, leading to inaccurate command execution and incorrect system attribution.

Method used

By receiving user voice commands, recognizing ambiguous commands, acquiring environmental awareness data, generating embodying clarifying questions, receiving voice responses and generating command parameters, recording interaction logs to optimize attribution strategies, processing negative semantics and generating new clarifying questions, and combining user identity and contextual information to optimize strategies.

Benefits of technology

It significantly improves the accuracy of voice command recognition and user interaction experience, ensures the precision and security of command execution, and avoids information loss and misattribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808767A_ABST
    Figure CN120808767A_ABST
Patent Text Reader

Abstract

The invention relates to a voice instruction recognition technology of an electric wheelchair, in particular to a voice instruction recognition method and system of the electric wheelchair. The method comprises the following steps: receiving a voice instruction of a user, and identifying a fuzzy instruction in the voice instruction; on the basis of the fuzzy instruction, sensing data related to the current environment of the electric wheelchair are obtained; identifying a physical target in the sensing data and a relative distance between the physical target and the electric wheelchair; according to a preset attribution strategy, generating a customized clarification inquiry according to the physical target and the relative distance thereof; receiving a voice response, generating an instruction parameter according to the voice response, and controlling the electric wheelchair to execute a corresponding action according to the instruction parameter; and recording an interaction log, and optimizing the attribution strategy based on the interaction log. The problems of information loss, inaccurate instruction execution and system error attribution caused by improper clarification of fuzzy instructions can be effectively solved, and the accuracy of voice instruction recognition of the electric wheelchair is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the voice instruction recognition technology of electric wheelchairs, in particular to an electric wheelchair voice instruction recognition method and system. BACKGROUND

[0002] In the field of action-assisting devices, especially electric wheelchairs, the voice instruction recognition technology aims to provide users with a convenient operation experience. When the voice instruction recognition system detects uncertainty or ambiguity in the instruction, it usually starts a clarifying inquiry process to obtain more information. However, if this clarifying inquiry method is too single or rigid, such as being limited to closed-ended questions (e.g., simple "yes" or "no" confirmations), it may lead to the situation that the further refined information provided by the user in the reply (e.g., supplementary descriptions about the operation object, method, and spatial location) cannot be effectively recognized, analyzed, and utilized by the system.

[0003] In view of the above problems, the existing technology needs to be improved. SUMMARY

[0004] In order to solve the problems of the prior art, the present application provides an electric wheelchair voice instruction recognition method and system, which can effectively solve the problems of information loss, inaccurate instruction execution, and system error attribution caused by improper clarification of ambiguous instructions, and significantly improve the accuracy of electric wheelchair voice instruction recognition.

[0005] The present application provides an electric wheelchair voice instruction recognition method, which comprises: receiving a user's voice instruction, and identifying ambiguous instructions in the voice instruction; based on the ambiguous instructions, obtaining perception data related to the current environment of the electric wheelchair; identifying a physical target in the perception data and the relative distance between the physical target and the electric wheelchair; according to a preset attribution strategy, generating a embodied clarifying inquiry containing the target name and location information from the physical target and its relative distance; receiving the user's voice response to the embodied clarifying inquiry, generating instruction parameters from the voice response, and controlling the electric wheelchair to perform corresponding actions according to the instruction parameters; recording the voice instruction, the embodied clarifying inquiry, the voice response, and the instruction parameters as an interaction log, and optimizing the attribution strategy based on the interaction log.

[0006] Through the above-mentioned scheme, the problem of the prior art that the user's supplementary information cannot be effectively recognized and utilized due to the single clarifying method, and the problems of inaccurate instruction execution and error attribution caused thereby, are solved, and the accuracy of voice instruction recognition and the user interaction experience are improved.

[0007] Further, the instruction parameter is generated according to the voice response, and the electric wheelchair is controlled to perform a corresponding action according to the instruction parameter, including: When the voice response contains negative semantics, new target information and its relative distance are obtained by analyzing the voice response; According to the new target information and its relative distance, a new embodied clarification inquiry is generated, and after receiving the user's voice response, the instruction parameter is extracted, and the electric wheelchair is controlled to perform a corresponding action according to the instruction parameter.

[0008] Further, when the voice response contains negative semantics, new target information and its relative distance are obtained by analyzing the voice response, including: According to the new target information and its relative distance, a plurality of candidate physical targets are matched from the detected and recognized physical targets; The spatial relationship between the plurality of candidate physical targets and the electric wheelchair includes relative distance, spatial position and orientation angle; According to the spatial relationship, the plurality of candidate physical targets are prioritized; If the number of candidate physical targets with the highest priority is only one, the candidate physical target with the highest priority is selected as the new target information; If there are multiple candidate physical targets with the highest priority, description information containing the differentiated characteristics of the candidate physical targets is generated for subsequent generation of embodied clarification inquiry.

[0009] Further, the instruction parameter is generated according to the voice response, and the electric wheelchair is controlled to perform a corresponding action, further including: When the voice response is positive, the instruction parameter is generated according to the target name and position information contained in the embodied clarification inquiry; The parking distance of the physical target is calculated according to the position information, and the parking distance is taken as the execution boundary of the instruction parameter, and the electric wheelchair is controlled to move in the direction of the physical target corresponding to the target name until it stops at a predetermined safe distance from the parking position.

[0010] Further, the voice instruction, embodied clarification inquiry, voice response and instruction parameter are recorded as an interaction log, and the attribution strategy is optimized based on the interaction log, including: The identity information and context information of the current user are obtained; Based on the identity information and context information, the interaction log containing the voice instruction, embodied clarification inquiry, voice response and instruction parameter is recorded; The historical interaction records matching the identity information and context information of the current user are filtered out from the interaction log; According to the historical interaction records, the attribution strategy content is adjusted; When it is detected that the user identity information or the context information changes, a attribution strategy corresponding to the new identity information or the new context information is loaded.

[0011] Further, the attribution strategy comprises: According to the perception data, all identified physical targets located in front of the electric wheelchair and their corresponding relative distances, spatial positions, and orientation angles are identified. According to the preset exclusion rules, the identified physical targets are screened: the identified physical targets that are blocked by obstacles between the electric wheelchair and the identified physical targets, the identified physical targets whose passage width between the electric wheelchair is less than a preset threshold, and / or the identified physical targets whose relative distance between the electric wheelchair is less than a safety distance threshold are excluded, to obtain selected physical targets; The selected physical target with the smallest included angle with the current travel direction of the electric wheelchair is selected as the source of the target name and position information referred to when generating the embodied clarification inquiry.

[0012] Further, the attribution strategy comprises: According to the physical target and its relative distance, it is determined whether the physical target belongs to a preset sensitive target type and whether the relative distance of the physical target is less than a preset safety distance. If the physical target belongs to the sensitive target type or the relative distance is less than the safety distance, an embodied clarification inquiry containing the target name and position information is generated. If the physical target does not belong to the sensitive target type and the relative distance is not less than the safety distance, it is determined whether the physical target is located within a preset movement path range of the electric wheelchair. If the physical target is located within the movement path range, an embodied clarification inquiry of the target name and position information pointing to the physical target is generated.

[0013] Further, the fuzzy instruction in the voice instruction is identified, comprising: The voice instruction is parsed to extract parameter information containing an operation object, an operation method, and a spatial position. According to the parameter information, it is determined whether the voice instruction has a fuzzy feature. If the fuzzy feature is identified, the corresponding fuzzy instruction in the voice instruction is identified according to the fuzzy feature.

[0014] Through the above scheme, the identification process of the fuzzy instruction is defined in detail. By parsing the parameter information and judging the fuzzy feature, the system can more accurately identify the instruction that needs to be clarified, laying a foundation for subsequent embodiment clarification.

[0015] Further, according to the parameter information, it is determined whether the voice instruction has a fuzzy feature, comprising: The current physical state information and environmental context information of the electric wheelchair are obtained. evaluate the executability and semantic integrity of the voice instruction in combination with the parameter information, the current physical state information and the environmental context information; determine that the voice instruction is an ambiguous instruction when the evaluation result indicates that the voice instruction has the condition of unclear execution object, unclear path target and / or uncertain execution range.

[0016] Further, the application also discloses an electric wheelchair voice instruction recognition system, which comprises: a receiving recognition module, configured to receive a voice instruction of a user and recognize an ambiguous instruction in the voice instruction; a perception data acquisition module, configured to acquire perception data related to a current environment of the electric wheelchair based on the ambiguous instruction; a target distance recognition module, configured to recognize a physical target in the perception data and a relative distance between the physical target and the electric wheelchair; an inquiry generation module, configured to generate an embodied clarification inquiry containing target name and location information from the physical target and the relative distance according to a preset attribution strategy; an execution module, configured to receive a voice response of the user to the embodied clarification inquiry, generate an instruction parameter according to the voice response, and control the electric wheelchair to perform a corresponding action according to the instruction parameter; a record optimization module, configured to record the voice instruction, the embodied clarification inquiry, the voice response and the instruction parameter as an interaction log, and optimize the attribution strategy based on the interaction log.

[0017] Through the above scheme, a system for implementing the above method is provided, and the modular design makes the method be effectively implemented, thereby providing hardware and software support for intelligent control of the electric wheelchair.

[0018] In summary, the electric wheelchair voice instruction recognition method and system provided by the application effectively solve the problems of information loss, inaccurate instruction execution and system error attribution caused by improper clarification of ambiguous instructions in the prior art by receiving an ambiguous instruction, acquiring environmental perception data, generating an embodied clarification inquiry, receiving a voice response and controlling an action, and recording an interaction log and optimizing an attribution strategy, thereby significantly improving the accuracy of electric wheelchair voice instruction recognition, the intelligent level of interaction and user experience. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 A flowchart of an electric wheelchair voice instruction recognition method provided by the application is shown.

[0020] Figure 2 A program block diagram of an electric wheelchair voice instruction recognition system provided by the application is shown.

[0021] In the figure: 1, receiving identification module; 2, perception data acquisition module; 3, target distance identification module; 4, query generation module; 5, execution module; 6, record optimization module. DETAILED DESCRIPTION

[0022] The technical solutions in the present application will be described in detail below with reference to the accompanying drawings in the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. The components of the present application described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0023] It should be noted that: similar numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, the terms "first", "second" and the like are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.

[0024] Reference Figure 1 The present application proposes an electric wheelchair voice instruction recognition method, comprising: S1000: receiving the user's voice instruction, identifying the ambiguous instruction in the voice instruction; S2000: based on the ambiguous instruction, obtaining the perception data related to the current environment of the electric wheelchair; S3000: identifying the physical target in the perception data and the relative distance between the physical target and the electric wheelchair; S4000: according to the preset attribution strategy, generating the embodied clarification query containing the target name and position information from the physical target and its relative distance; S5000: receiving the user's voice response to the embodied clarification query, generating the instruction parameters according to the voice response, and controlling the electric wheelchair to execute the corresponding action according to the instruction parameters; S6000: record the voice instruction, embodied clarification query, voice response and instruction parameters as interaction log, and optimize the attribution strategy based on the interaction log.

[0025] Among them, the fuzzy instruction refers to the part in the voice instruction issued by the user that is difficult to directly parse into accurate control parameters due to unclear operation object, unclear path target or uncertain execution range. This instruction can be implemented by using natural language processing technology, semantic analysis model or pre-set fuzzy vocabulary library, for example, by lexical analysis, syntactic analysis combined with context information for judgment, which is mainly to identify the uncertainty in the user's intention in order to start the subsequent clarification process. The perception data refers to the real-time information related to the current environment obtained by the sensor system carried by the electric wheelchair, which can be implemented by using various sensor fusion technologies such as laser radar, camera, ultrasonic sensor or depth sensor, for example, by laser radar scanning to obtain point cloud data, by camera to capture image information, which is mainly to provide objective and comprehensive spatial information of the environment around the electric wheelchair as the context basis for understanding the fuzzy instruction.

[0026] The physical target refers to the entity identified in the perception data that has a specific shape, position and semantics, which can be implemented by using target detection algorithm, image recognition technology or point cloud clustering analysis, for example, to identify tables, chairs, doors or walls, etc., which is mainly to extract specific objects that may be related to the user's instruction from complex environmental information as a reference for subsequent clarification inquiry. The attribution strategy refers to a set of pre-set rules for inferring the user's potential intention and generating clarification inquiries according to environmental perception information and user's fuzzy instruction, which can be implemented by using rule-based reasoning engine, machine learning model or expert system, for example, to select the most likely target according to the position, type and relative relationship with the wheelchair of the physical target, which is mainly to intelligently select the most relevant physical target and its position information under the fuzzy instruction to construct embodied clarification inquiry. The embodied clarification inquiry refers to an interactive inquiry that contains specific target name and position information to guide the user to input accurate instruction, which can be implemented by using speech synthesis technology combined with pre-set templates, for example, "Do you want to go to the side of the coffee table?" or "Do you want to approach the sofa?", which is mainly to convert the abstract fuzzy instruction into a specific question related to the physical target in the actual environment, so as to guide the user to provide more accurate instruction information and avoid information loss. The interaction log refers to the data set of a series of interactive events such as voice instruction, embodied clarification inquiry, voice response and instruction parameters recorded by the system in the process of processing user voice instruction, which can be implemented by using structured database or unstructured file storage, for example, to record the timestamp, user ID, original instruction text, clarification content issued by the system, user's response text and finally executed instruction parameters of each interaction, which is mainly to accumulate user interaction data to provide data basis for subsequent attribution strategy optimization and realize adaptive learning of the system.

[0027] The scheme of the present application realizes the intelligent recognition and execution of voice instructions for electric wheelchairs through a series of steps working in synergy, especially showing its unique advantages in handling ambiguous instructions. When the system receives a user's voice instruction, it first analyzes the instruction to identify whether there is an ambiguous instruction. Once the ambiguity of the instruction is detected, the system will immediately start the environmental perception process to obtain perception data related to the current environment of the electric wheelchair. These perception data, after processing, can identify specific physical targets in the environment and calculate the relative distances between these physical targets and the electric wheelchair. On this basis, the system will generate a embodied clarification query containing the target name and location information according to the preset attribution strategy, using the identified physical targets and their relative distance information. This query method is different from a simple yes or no question, as it associates the user's ambiguous intention with specific objects in the actual environment, guiding the user to provide more accurate feedback. Subsequently, the system receives the user's voice response to the embodied clarification query and deeply analyzes the response to extract accurate instruction parameters. According to these instruction parameters, the electric wheelchair is controlled to perform corresponding actions, ensuring the accurate execution of the instruction. In order to continuously improve the intelligence and adaptability of the system, the voice instructions, embodied clarification queries, voice responses, and finally generated instruction parameters in the entire interaction process are recorded as interaction logs. The system will continuously optimize the preset attribution strategy based on these interaction logs. This optimization mechanism enables the system to learn from historical interactions, continuously adjust its understanding and clarification of ambiguous instructions, thereby avoiding incorrect attribution due to repeated interaction failures, and gradually adapting to different users' language habits and usage scenarios, forming a positive self-learning cycle.

[0028] In some preferred embodiments, the present application is implemented as follows: When a user issues an ambiguous voice command like "Wheelchair, move forward a little," the electric wheelchair's speech recognition module receives and recognizes "move forward a little" as an ambiguous command. The wheelchair's onboard lidar and camera then work together to acquire sensory data about the current environment. For example, the lidar scans and generates a point cloud map of the surroundings, while the camera captures images. By processing this sensory data, the system identifies a coffee table and a sofa in front of it, calculating the relative distance between the coffee table and the wheelchair as 2 meters and 3 meters between the sofa and the wheelchair. Based on a pre-set attribution strategy, which may prioritize the closest physical object in the wheelchair's direction of travel, the system determines that the user is more likely pointing at the coffee table. Therefore, the system generates an embodied clarification query, such as "Do you want to move closer to the coffee table?" via speech synthesis. When the user responds, "Yes, move closer to the coffee table, slower," the system receives and interprets the voice response, identifying the affirmative meaning and the speed constraint of "slower." Based on this information, the system generates instruction parameters, such as "Move toward the coffee table at a low speed until you stop at a preset safe distance from the coffee table." The electric wheelchair then executes the action according to these parameters. The original instruction, clarification query, user response, and final instruction parameters of this interaction are all recorded in the interaction log. This is used to optimize subsequent attribution strategies, such as adjusting the priority attribution weight for "coffee table" in similar ambiguous instructions or learning the speed parameters corresponding to "slow down."

[0029] In another embodiment of the present application, it is further proposed that the sub-step of S5000: generating instruction parameters according to the voice response, and controlling the electric wheelchair to perform corresponding actions according to the instruction parameters includes: S5210: When the user's voice response contains negative semantics, the voice response is parsed to obtain new target information and its relative distance; S5220: Generate a new embodied clarification inquiry based on the new target information and its relative distance, extract the instruction parameters after receiving the user's voice response, and control the electric wheelchair to perform corresponding actions based on the instruction parameters.

[0030] Wherein, the parsing voice response obtains new target information and its relative distance refers to the semantic analysis and information extraction of the voice response containing negative semantics of the user, which can be specifically through natural language processing technology, identifying the new and expected target entity mentioned by the user while denying the original target and its spatial relationship with the current position of the electric wheelchair, and the purpose is to accurately capture the real intention of the user from the negative feedback of the user, and to provide basic data for subsequent instruction modification. Wherein, the new target information refers to the physical entity or position that the electric wheelchair needs to re-identify or pay attention to, which is explicitly or implicitly expressed by the user through voice response after denying the original instruction target, which can be specifically a new object name, a more accurate location description or a direction indication, and the purpose is to provide a new operation object or destination for the electric wheelchair. Wherein, the relative distance refers to the spatial measure between the physical entity or position indicated by the new target information and the current position of the electric wheelchair, which can be specifically inferred through distance description words (such as "a little closer", "a little farther", "next to") or spatial orientation words (such as "left", "right", "front") contained in the voice response, and the purpose is to provide accurate position reference for the moving instruction of the electric wheelchair. Wherein, the generation of new embodied clarification inquiry refers to constructing an interactive question containing specific target name and location description that can be intuitively understood and confirmed by the user according to the new target information and its relative distance obtained by parsing, which can be specifically through visual or auditory prompts, such as highlighting the new target on the screen or describing its specific location through voice, and the purpose is to confirm the system's understanding of the new intention to the user again, to ensure the accuracy of the instruction and to reduce misunderstanding.

[0031] The scheme proposes a special processing mechanism for the user's negative voice response in the voice instruction recognition process of the electric wheelchair, which significantly improves the system's understanding ability of complex user intentions. When the system receives the user's negative response to the embodied clarification inquiry, it no longer simply interrupts the instruction process, but further analyzes the new target information and its relative distance implied in the negative statement. For example, the user expresses "not the sofa, but the chair next to it", and the system can extract "chair" as the new target, and generate a new embodied clarification inquiry such as "do you want to go to the chair next to the sofa?". This new clarification sentence is dynamically generated based on user feedback, which is more accurate and specific, and helps to guide the user to give a clear response. The system then extracts the final instruction parameters according to the new response and completes the control action of the electric wheelchair. Through this iterative and context-sensitive processing method, the system can continuously learn and correct the original understanding from the user's negative answer, avoiding the execution error caused by the neglect of negative information. Compared with the traditional single-round clarification process, the scheme constructs a more flexible and robust interaction cycle, so that the electric wheelchair can accurately grasp the real intention of the user through multiple interactions even if the initial instruction is ambiguous or the understanding is biased, thereby enhancing the reliability and user experience of voice control.

[0032] In another embodiment of the present application, it is further proposed that S5210 includes: S5211: Match multiple candidate physical targets from the detected and identified physical targets based on the new target information and its relative distance; S5212: Acquire spatial relationships between multiple candidate physical targets and the electric wheelchair, including relative distances, spatial positions, and orientation angles; S5213: Prioritizing multiple candidate physical targets based on spatial relationships; S5214: If there is only one candidate physical target with the highest priority, select the candidate physical target with the highest priority as the new target information; S5215: If there are multiple candidate physical targets with the highest priority, descriptive information containing differentiated features between the candidate physical targets is generated for subsequent generation of embodied clarification inquiries.

[0033] Matching multiple candidate physical targets refers to selecting a set of potential targets that match these descriptions, based on the new target information and relative distances obtained from the initial analysis, from all physical targets detected and identified by the electric wheelchair's perception module. Spatial relationships refer to the relative positional information between the candidate physical targets and the electric wheelchair in three-dimensional space, including the straight-line distance between them, their specific coordinate positions in a specific coordinate system, and their respective orientation angles. This provides multi-dimensional data for a comprehensive evaluation of the candidate targets. Prioritization refers to ranking the multiple matched candidate physical targets by importance or likelihood based on pre-set rules or algorithms, with the goal of identifying the potential target that best matches the user's intent. Descriptive information for differentiating features refers to the system automatically extracting and generating distinguishable visual or physical attribute descriptions, such as color, shape, material, or specific markings, when there are multiple candidate physical targets with the same priority. This provides a clear basis for differentiation in subsequent clarification inquiries, guiding the user to make a precise selection.

[0034] The solution of the present application solves the problem of accurately extracting new target information from user's negative response through a multi-stage, data-driven inference mechanism. When the user's voice response to the embodied clarification query contains negative semantics, the system first preliminarily analyzes the new target information and its relative distance from the voice response. However, relying solely on this preliminary analysis may not be sufficient to determine a unique and accurate target. Therefore, the present application further utilizes the physical target data detected and identified by the electric wheelchair, and according to the new target information and its relative distance obtained by preliminary analysis, matches multiple candidate physical targets from these identified physical targets. This process associates the user's ambiguous intention with specific objects in the actual environment, avoiding blind guessing and utilizing existing environmental perception data. In order to more finely evaluate these candidate physical targets, the spatial relationship between each candidate physical target and the electric wheelchair is obtained, including relative distance, spatial position and orientation angle. These multi-dimensional spatial data provide a comprehensive understanding of the relative state between the target and the electric wheelchair, for example, one target may be close but not in the user's line of sight, and another may be far but directly opposite the user. Based on these spatial relationships, the system prioritizes multiple candidate physical targets. The basis for sorting can consider distance, angle, user historical preferences and other factors, thereby filtering out the target that is most likely to meet the user's real intention. After the priority ranking is completed, the system will make a decision based on the results. If there is only one candidate physical target with the highest priority, this target is explicitly selected as the new target information for subsequent generation of embodied clarification queries. This ensures that when the user's intention is clear, the system can directly and accurately respond. However, if there are multiple candidate physical targets with the highest priority, it indicates that the user's negative response still has a certain degree of ambiguity and cannot be completely determined by a single dimension. At this time, the system will generate description information containing the differentiated features of these candidate physical targets. These differentiated features, such as color, size or specific identification, will be key components of subsequent embodied clarification queries, guiding the user to make more accurate references.

[0035] Through the above steps, the solution of the present application can effectively extract more accurate new target information from the user's negative response and combine the environmental perception data of the electric wheelchair for inference. This makes the system, when facing user's negative feedback, no longer simply repeat the query or mistakenly discard the information, but can actively and intelligently narrow down the target range and provide more guiding clarification queries. This mechanism effectively cooperates with the electric wheelchair's ability to obtain perception data and identify physical targets, making the entire interaction process more smooth and accurate, thereby avoiding negative interaction cycles caused by information loss or incorrect attribution, and significantly improving user experience and the accuracy of instruction execution.

[0036] In another embodiment of the present application, it is further proposed that the sub-step of S5000: generating instruction parameters according to the voice response, and controlling the electric wheelchair to perform corresponding actions according to the instruction parameters also includes: S5230: When the voice response is affirmative, generate instruction parameters according to the target name and location information contained in the embodied clarification inquiry; S5240: Calculate the parking distance of the physical target based on the position information, and use the parking distance as the execution boundary of the instruction parameter to control the electric wheelchair to move in the direction of the physical target corresponding to the target name until it stops at a preset safe distance from its parking position.

[0037] Among them, a voice response with affirmative semantics means that the user gives a confirmatory reply to the embodied clarification inquiry issued by the system. Specifically, it can be recognized by voice recognition technology, such as "yes", "right", "okay", "confirm", etc., which express agreement or affirmation. Its purpose is to clarify the user's acceptance and recognition of the current clarification content.

[0038] The target name and location information contained in the embodied clarification inquiry refers to the physical entity that the user intends to point to and its specific spatial coordinates or relative position in the environment, which is determined by the system after recognizing the ambiguous instruction, perceiving the environment and confirming it through interaction with the user. Specifically, it can be obtained through lidar, visual sensors, etc. Environmental perception data is generated in combination with semantic understanding. Its purpose is to provide clear target guidance for the subsequent movement of the electric wheelchair.

[0039] The command parameters refer to a set of quantized instructions used to control the electric wheelchair to perform specific actions. Specifically, they may include movement direction, speed, acceleration, and the stopping distance specifically introduced in this solution. Their purpose is to convert the user's intention into precise control commands that the electric wheelchair can execute. The stopping distance of a physical target refers to the distance at which the electric wheelchair needs to start decelerating and eventually stop when moving towards the target. It can be calculated based on factors such as the size, shape, and material of the physical target, as well as the current speed and braking performance of the electric wheelchair. Its purpose is to ensure that the electric wheelchair can stop safely and smoothly when approaching the target. The execution boundary of the command parameters refers to the critical value that limits the movement range or stopping condition of the electric wheelchair when executing the movement command. Specifically, it can be used as the calculated stopping distance of the physical target as the trigger condition for the electric wheelchair to stop. Its purpose is to provide a safe and reliable stopping mechanism for the autonomous movement of the electric wheelchair. The preset safety distance refers to the minimum safe distance that should be maintained between the electric wheelchair and the physical target when the electric wheelchair finally stops. It can be set based on factors such as the size of the electric wheelchair, the user's safety requirements, and environmental obstacles. Its purpose is to avoid collision between the electric wheelchair and the target, thereby ensuring the safety of the user and the surrounding environment.

[0040] The solution of the present application realizes the safe and intelligent movement of the electric wheelchair to the target location through the accurate semantic understanding of the user's voice response combined with environmental perception data. When the electric wheelchair system receives the user's voice response to the embodiment clarification inquiry, and the response is identified as a positive semantic, the system will immediately extract the user-confirmed target name and location information from the embodiment clarification inquiry. These information is the result of the system's interaction with the user in the previous step based on the preliminary understanding of the ambiguous instruction and environmental perception, thus ensuring the accuracy of the instruction. On this basis, the system will use the extracted location information to accurately calculate the parking distance required by the electric wheelchair when approaching the physical target. This calculation of parking distance takes into account the motion characteristics of the electric wheelchair and environmental factors, aiming to provide a key safety parameter for subsequent movement control. Subsequently, this calculated parking distance is set as the execution boundary of the instruction parameter, which means that the movement of the electric wheelchair will be strictly constrained by this boundary. The electric wheelchair will start moving in the direction of the physical target corresponding to the target name. During the movement, the system will continuously monitor the real-time distance between the electric wheelchair and the target. Once the distance between the electric wheelchair and the target reaches or approaches the preset safety distance, the system will trigger the stopping mechanism to make the electric wheelchair stop smoothly. This mechanism ensures that the electric wheelchair will not approach the target excessively, thus avoiding potential collision risks and significantly improving the safety of movement. In this way, the present solution is closely combined with the basic voice instruction recognition and clarification mechanism to form a complete and closed-loop intelligent interaction and control process. The basic mechanism provides clear user intent and target information, while the present solution introduces fine distance calculation and safe parking control on this basis, transforming the abstract "positive" intent into precise and safe physical movement. This combination enables the electric wheelchair not only to understand the user's explicit instructions, but also to execute these instructions in a safe and controllable manner, solving the problem of how to accurately and safely move to the target location after the user confirms the target, and avoiding the safety hazards caused by the lack of precise parking distance control.

[0041] In some preferred embodiments, the present application is implemented as follows: Assuming the user issues the ambiguous instruction "go that way", the electric wheelchair system perceives the environment and generates an embodied clarifying query: "Do you want to go to the sofa?". When the user's voice response is a positive semantic, such as replying "yes, the sofa", the system immediately recognizes the positive semantic. Subsequently, the system extracts the target name "sofa" and its corresponding location information, such as the sofa's three-dimensional coordinates or relative orientation in the electric wheelchair's current coordinate system, from the previously generated embodied clarifying query. Based on this location information, the system initiates a distance calculation module that utilizes real-time environmental data obtained from a laser radar or depth camera, combined with the electric wheelchair's current speed and brake performance model, to calculate the stopping distance of the physical target that the electric wheelchair needs to start decelerating and eventually stop. For example, if there is enough space in front of the sofa, the system may calculate a distance that allows the electric wheelchair to decelerate smoothly. The calculated stopping distance is then set as the execution boundary of the instruction parameter. This means that the electric wheelchair's internal control system will use this distance as the trigger condition for stopping when moving towards the sofa. The electric wheelchair's navigation and control module plans a path and controls the electric wheelchair to move in the direction of the sofa. During the movement, the system continuously monitors the real-time distance between the electric wheelchair and the sofa. Once the distance between the electric wheelchair and the sofa reaches a pre-set safe distance, such as 0.3 meters, the electric wheelchair's drive system will receive a stop command and immediately execute the brake, ensuring that the electric wheelchair stops smoothly while maintaining a safe distance from the sofa. In this way, the user can safely approach the sofa without worrying about collisions.

[0042] In another embodiment of the present application, it is further proposed that S6000 comprises: S6100: Obtain the identity information and situational information of the current user; S6200: Based on the identity information and situational information, record the interaction log containing voice instructions, embodied clarifying queries, voice responses, and instruction parameters; S6300: From the interaction log, filter out the historical interaction records matching the identity information and situational information of the current user; S6400: According to the historical interaction records, adjust the attribution strategy content; S6500: When detecting changes in user identity information or situational information, load the attribution strategy corresponding to the new identity information or new situational information.

[0043] The identity information refers to relevant data for uniquely identifying a user, which can be a user ID, age, gender, health status, or personalized preference setting of the user, and aims to distinguish different users so that the system can perform personalized processing according to the characteristics of individual users; the context information refers to environmental context data when the user issues an instruction, which can be the current time, geographical location (for example, indoor or outdoor, specific room type), weather condition, obstacle distribution in the surrounding environment, or the current running state (for example, speed, power) of the electric wheelchair, and aims to provide richer background information for the system to understand the user's intention, because the same instruction can have different meanings in different contexts; the interaction log refers to a structured data set recorded by the system for all interaction processes between the user and the electric wheelchair, which can be text, voice, system response and other data stored in a time sequence, and aims to provide comprehensive historical data support for subsequent attribution strategy optimization; the attribution strategy content refers to a rule set used by the system to map ambiguous instructions to specific physical targets or operation parameters, which can be a set of preset logical rules, a decision tree model, or parameters and weights of a machine learning model.

[0044] The scheme of the present application improves the accuracy of the system in understanding the user's intention by introducing user identity information and context information, and performing fine processing on the recording of the interaction log and the optimization of the attribution strategy. By introducing user identity information and context information, the present application realizes fine optimization of the recording of the voice interaction log and the attribution strategy, and significantly improves the accuracy of the system in understanding the user's intention. The system first obtains the identity characteristics of the current user and the context information such as the environment, which serves as the semantic basis for subsequent processing, so that it is no longer limited to the surface meaning of the instruction. On this basis, the system records the interaction log containing voice instructions, embodied clarification inquiries, voice responses and instruction parameters, and attaches complete identity and context metadata to form structured context behavior data. Subsequently, the system selects the interaction content highly related to the current identity and context from the historical records, identifies the expression habits and response rules of the specific user in similar environments, and then adjusts the attribution strategy individually. For example, if a user always gives a negative answer to a clarification inquiry in a certain context, the frequency of triggering such clarifications should be reduced. The system also supports real-time detection of changes in identity or context, and will dynamically switch the matched attribution strategy once a change occurs, so as to maintain the continuous adaptation of the interaction strategy to the user's state. The scheme reconstructs the attribution logic from the complete perspective of "user-context-instruction-response", avoids one-sided judgment based on static logs, effectively reduces misjudgment and redundant clarification, and improves the intelligence and user experience of the electric wheelchair voice instruction recognition system.

[0045] In some preferred embodiments, the present application is implemented as follows: First, upon the electric wheelchair's startup or user login, the system can acquire the current user's identity information and situational information. For example, identity information can be read from the user's login account information, including the user's unique ID, preset language preferences, and historical usage habit tags; situational information can be acquired through the wheelchair's built-in sensor array, such as obtaining geographic location information (indoor / outdoor) through the GPS module, obtaining light intensity (day / night) through the ambient light sensor, analyzing background noise (quiet / noisy) through the microphone array, and obtaining current speed and tilt angle through the wheelchair's own odometer and attitude sensor. Then, based on these acquired identity information and situational information, the system records an interaction log containing voice commands, embodied clarification questions, voice responses, and command parameters. These logs can be stored in a structured JSON or XML format in the electric wheelchair's local storage or on a cloud server. Each log record contains not only the original voice command text, system-generated embodied clarification question text, user voice response text, and finally parsed command parameters (such as target name, location coordinates, movement speed, etc.), but also timestamps, user IDs, geographic location tags, and environmental noise levels. Subsequently, when optimizing the attribution strategy is needed, the system filters historical interaction records from the interaction log database that match the current user's identity information and situational information. For example, if the current user is "User A" and the current situation is "indoor, quiet environment," the system will query the logs to find all historical commands issued by "User A" in "indoor, quiet environment" and their corresponding clarification and response records. Filtering can be based on a preset similarity threshold, for example, a situational information match of 80% or higher is considered a match. Then, the system adjusts the attribution strategy content based on these filtered historical interaction records. For example, the system can run a machine learning model (such as a decision tree or neural network-based classifier) that uses historical interaction records as training data to learn which ambiguous commands tend to be attributed to which specific targets or which clarification questions are more likely to result in effective responses in a particular user and situation. By analyzing patterns of attribution errors in historical data, the model can automatically update its internal weights or rules to reduce future attribution errors. For example, if it is found that "User A" in an "indoor" environment often refers to "the coffee table" when saying "go there," the attribution strategy can be adjusted to prioritize attributing "go there" to "the coffee table" in that user and situation. Finally, to ensure the dynamic adaptability of the attribution strategy, the system continuously monitors changes in user identity information or situational information. For example, when the GPS module detects that the electric wheelchair has moved from an "indoor" area to an "outdoor" area, or when the user switches "driving mode" (such as from "slow mode" to "fast mode," which can be considered a situational change) through voice commands, the system will immediately detect these changes.Once the change is detected, the system loads the attribution strategy corresponding to the new identity information or new context information from the pre-stored attribution strategy library. For example, if the user moves from indoors to outdoors, the system can load a set of attribution strategies optimized for complex outdoor environments (e.g., more obstacles, more spacious space), which may have different explanations for "go there" in the indoor context, thus better adapting to the new environmental challenges.

[0046] It is further proposed in another embodiment of the present application that the steps of the attribution strategy include: A1: According to the perception data, identify all identified physical targets in front of the electric wheelchair and their corresponding relative distance, spatial position and orientation angle; A2: According to the pre-set exclusion rule, screen the identified physical targets: exclude the identified physical targets that are blocked by obstacles between the electric wheelchair, the identified physical targets whose path width between the electric wheelchair is less than the pre-set threshold and / or the identified physical targets whose relative distance between the electric wheelchair is less than the safety distance threshold, to obtain the selected physical targets; A3: Select the selected physical target with the smallest angle with the current travel direction of the electric wheelchair as the source of the target name and position information referred to in generating the embodied clarification inquiry.

[0047] Among them, the pre-set exclusion rule refers to a set of conditions for filtering and excluding physical objects that are not suitable as clarification inquiry targets, which can be implemented based on geometric analysis, path planning algorithm or safety distance judgment, etc. The purpose is to ensure that the selected target is reachable, safe and related to the user's intention.

[0048] The scheme of the present application effectively solves the problem of generating inaccurate or irrelevant clarification inquiries in a multi-target scenario by fine perception and intelligent screening of the environment in front of the electric wheelchair. Specifically, first, the system comprehensively identifies all identified physical targets in front of the electric wheelchair according to the perception data, and obtains the relative distance, spatial position and orientation angle of these targets. This step lays a data foundation for subsequent intelligent screening, ensuring that all potential interactive objects are included in the consideration range. On this basis, the system further screens these identified physical targets according to the pre-set exclusion rules. By excluding physical targets that are obstructed by obstacles, have insufficient path width or are too close in relative distance, the scheme effectively eliminates those interference items that cannot pass, have safety hazards or are not suitable as interactive objects, thereby obtaining a set of more feasible candidate physical targets. Finally, from these candidate physical targets, the system selects the target with the smallest angle with the current direction of travel of the electric wheelchair as the source of the target name and position information referred to in generating the embodied clarification inquiry. This selection mechanism enables the generated clarification inquiry to be highly focused on the target that the user is most likely to pay attention to or intend to go to, significantly improving the accuracy and relevance of the clarification inquiry. Through the organic combination of the above steps, the present scheme, based on the generation of embodied clarification inquiry in the electric wheelchair voice command recognition method, no longer simply selects from all identified physical targets, but introduces a multi-stage intelligent screening mechanism. This mechanism enables the system to accurately identify targets that are highly matched with the user's intention and actually reachable from a complex and variable environment, thereby avoiding ineffective or misleading clarification inquiries due to improper target selection. This not only greatly reduces the user's cognitive burden, improves the efficiency and fluency of human-computer interaction, but also ensures that the electric wheelchair can more accurately understand and execute the user's ambiguous instructions, effectively avoiding the negative interaction cycle caused by false attribution in the background technology, enabling the system to more intelligently and safely assist the user in moving.

[0049] In another embodiment of the present application, it is further proposed that the step of the attribution strategy comprises: A4: judging whether the physical target belongs to a pre-set sensitive target type and whether the relative distance of the physical target is less than a pre-set safety distance according to the physical target and its relative distance; A5: if the physical target belongs to the sensitive target type or the relative distance is less than the safety distance, generating an embodied clarification inquiry containing the target name and position information; A6: if the physical target does not belong to the sensitive target type and the relative distance is not less than the safety distance, judging whether the physical target is located within a pre-set movement path range of the electric wheelchair; A7: if the physical target is located within the movement path range, generating an embodied clarification inquiry of the target name and position information pointing to the physical target.

[0050] wherein the sensitive target type refers to a category of physical targets that have special importance or potential risks to the electric wheelchair user or the surrounding environment, which can include but not limited to pedestrians, pets, children, fragile objects, hazardous materials, or specific area boundaries, aiming to enable the system to prioritize identifying and handling targets that have direct impact on user safety or operation experience. The safety distance refers to the minimum allowable distance threshold that the electric wheelchair maintains from the physical target, which can be dynamically or statically set according to the electric wheelchair's travel speed, braking performance, environmental complexity, and user preferences, aiming to ensure that the electric wheelchair has sufficient reaction time or space when approaching certain targets to avoid collisions or unnecessary risks. The movement path range refers to the spatial area covered by the electric wheelchair's intended or planned trajectory when executing instructions or autonomous navigation, which can be dynamically calculated or preset by the electric wheelchair's navigation system according to the current location, target point, environmental map, and obstacle avoidance strategy, aiming to limit the system to only clarify inquiries for targets related to the electric wheelchair's actual movement intention, avoiding unnecessary interference to the user.

[0051] The present scheme proposes a hierarchical and prioritized attribution strategy, significantly improving the accuracy and safety of electric wheelchair in handling ambiguous voice instructions in complex environments. When the system receives ambiguous instructions and identifies the physical targets in the environment and their relative distances, it first evaluates the safety and importance of the targets. The system determines whether the target belongs to the pre-set sensitive target or is too close to the user based on the target type and distance from the user. Once the target is of a sensitive type or within the safety distance, the system generates a embodied clarification inquiry containing the target name and location, ensuring that the user can timely confirm potential risks and avoid safety hazards. For non-sensitive targets with sufficient distance, the system enters the second level of judgment, and only when the target is on the electric wheelchair's travel path, does it trigger an embodied clarification inquiry. This hierarchical processing method avoids initiating inquiries for all targets, reducing ineffective interactions and improving interaction efficiency. The overall strategy balances safety assurance and user experience, by precisely filtering targets and dynamically adjusting clarification strategies, enabling the system to make judgments based on more accurate and user-intended target information before generating instruction parameters and control actions, thereby improving the accuracy and reliability of the entire ambiguous voice instruction recognition and execution process.

[0052] In some preferred embodiments, the present application is implemented as follows: When the electric wheelchair receives an ambiguous voice instruction, such as "go there", the system will first obtain the perception data of the current environment, such as identifying the surrounding physical targets through lidar, camera or ultrasonic sensor, and calculating their relative distance from the electric wheelchair. Assuming that the system identifies a pedestrian (physical target A) in front of the electric wheelchair, 1.5 meters away from the electric wheelchair; a garbage can (physical target B) on the right side, 2 meters away from the electric wheelchair; and a pet (physical target C) in front of the left side, 1 meter away from the electric wheelchair. At this time, the system will make a judgment according to the preset attribution strategy. First, determine whether physical target A (pedestrian) belongs to the preset sensitive target type. Assuming that "pedestrian" is defined as a sensitive target type, and the preset safety distance is 2 meters. Since the relative distance of pedestrian A is 1.5 meters, which is less than the safety distance of 2 meters, the system will immediately generate a embodied clarification inquiry containing "pedestrian A, 1.5 meters in front", such as "Are you referring to the pedestrian 1.5 meters in front?". If there is no sensitive target or target with a distance less than the safety distance at this time, the system will continue to judge other targets. For example, for physical target B (garbage can), it does not belong to the sensitive target type, and the relative distance of 2 meters is not less than the safety distance of 2 meters. The system will further determine whether the garbage can B is located within the preset moving path range of the electric wheelchair. If the current moving path of the electric wheelchair is planned to pass near the garbage can B, the system will generate an embodied clarification inquiry containing "garbage can B, 2 meters to the right", such as "Are you referring to the garbage can 2 meters to the right?". For physical target C (pet), it belongs to the sensitive target type, and the relative distance of 1 meter is less than the safety distance of 2 meters. The system will preferentially generate an embodied clarification inquiry containing "pet C, 1 meter in front left", such as "Are you referring to the pet 1 meter in front left?". Through such an implementation, the system can prioritize according to the type and distance of the target, ensuring that key targets or potential risk targets are clarified first, while avoiding unnecessary inquiries on irrelevant or non-path targets, thereby improving interaction efficiency and safety.

[0053] In another embodiment of the present application, it is further proposed that the sub-step S1000 of identifying ambiguous instructions in the voice instruction comprises: S1210: analyzing the voice instruction to extract parameter information containing operation object, operation method and spatial position; S1220: determining whether the voice instruction has ambiguous features according to the parameter information; S1230: if the ambiguous features are identified, identifying the corresponding ambiguous instructions in the voice instruction according to the ambiguous features.

[0054] Among them, analysis refers to in-depth linguistic and semantic analysis of the input voice instruction, which can be specifically implemented by natural language processing (NLP) technology, speech recognition post-processing algorithm or rule-based semantic analysis engine, and the purpose is to convert unstructured voice input into structured data that can be understood and processed by machines; parameter information refers to the key data points extracted from the voice instruction, which are used to describe the core elements of the instruction, and can specifically include the target entity of the instruction action, the specific action of the instruction execution and the spatial orientation or position information related to the instruction, and the purpose is to provide a basis for subsequent judgment of the definiteness of the instruction; fuzzy characteristics refer to the linguistic or semantic attributes in the voice instruction that cause its semantics to be unclear, incomplete or uncertain, which can specifically manifest as lack of key information, inaccurate information description or ambiguity, and the purpose is to indicate that the instruction needs to be further clarified or refined; fuzzy instruction refers to a voice instruction that cannot be directly executed due to one or more fuzzy characteristics, which can be specifically classified according to the type of fuzzy characteristics, such as operation object fuzzy, operation mode fuzzy or spatial position fuzzy, and the purpose is to provide clear classification basis for subsequent targeted clarification or processing.

[0055] The scheme of the present application can be summarized as follows. First, the voice instruction is analyzed to extract key parameter information including the operation object, operation method and spatial position. This process converts the original unstructured voice instruction into structured data. Based on these extracted parameter information, the system evaluates the voice instruction to determine whether there are ambiguous features. This judgment mechanism can identify the possible ambiguous, incomplete or unclear situations in the instruction. Once the ambiguous features in the voice instruction are identified, the system further identifies the corresponding ambiguous instruction type according to these specific ambiguous features. For example, if the operation object is missing, it is identified as an operation object ambiguous instruction; if the operation method description is not clear, it is identified as an operation method ambiguous instruction. Through this layer-by-layer analysis, judgment and identification process, the present scheme can accurately locate the uncertainty in the voice instruction and classify it. This accurate ambiguous instruction identification mechanism is closely integrated with the overall voice instruction recognition method of the electric wheelchair. After receiving the user's voice instruction, it is no longer simply determined whether it is ambiguous, but the specific type of ambiguity is analyzed in depth. This detailed identification result can provide more accurate input for subsequent steps such as obtaining environment perception data based on ambiguous instructions, identifying physical targets, and generating embodied clarification inquiries. For example, when a spatial position ambiguous instruction is identified, the system can more specifically obtain perception data related to the spatial position and generate a more specific clarification inquiry, avoiding the invalid clarification or information loss caused by inaccurate ambiguous instruction recognition in traditional methods. In this way, the present scheme can effectively avoid the negative interaction cycle caused by inaccurate instruction ambiguity judgment, improving the efficiency and accuracy of the interaction between the electric wheelchair and the user.

[0056] In some preferred embodiments, specifically when the electric wheelchair system receives a voice instruction "go there" issued by the user, the system first starts a voice instruction parsing module. This module can use a deep learning model, such as a semantic parser based on a Transformer architecture, to analyze the text content of the voice instruction. Through analysis, the system can identify that "go" is an operation mode, but "there" as a spatial location parameter, its specific direction is not clear, and it lacks a clear operation object. Therefore, the parsing module extracts the operation mode as "go", the spatial location as "there", and the operation object as an empty parameter information. Then, the system determines whether the voice instruction has ambiguous features according to these parameter information. Since the operation object is missing and the spatial location "there" is not clear, the system determines that the voice instruction has ambiguous features. Subsequently, the system further identifies the corresponding ambiguous instruction type in the voice instruction according to the identified ambiguous features. In this case, since the operation object is missing and the spatial location is not clear, the system can identify that the instruction is an "operation object ambiguous instruction" and a "spatial location ambiguous instruction". This identification result will be the basis for generating a subsequent clarification inquiry, for example, the system can generate a somatized clarification inquiry such as "Which target do you want to go to? Is it the side of the table?" to guide the user to provide more accurate information.

[0057] In another embodiment of the present application, it is further proposed that S1220 comprises: S1221: obtaining the current physical state information and environmental context information of the electric wheelchair; S1222: combining the parameter information, the current physical state information and the environmental context information, evaluating the executability and semantic integrity of the voice instruction; S1223: when the evaluation result shows that the voice instruction has the situation of unclear execution object, unclear path target and / or uncertain execution range, determining that the voice instruction is an ambiguous instruction.

[0058] The current physical state information of the electric wheelchair refers to the internal operating parameters and the state data of the electric wheelchair at a specific time, which can be achieved by real-time collection and reporting of data by the built-in sensors of the wheelchair (such as speed sensor, attitude sensor, power sensor, odometer, etc.), and its purpose is to provide the executable constraints of the wheelchair itself; the environmental context information refers to the real-time perception data and situation description of the external environment where the electric wheelchair is located, which can be achieved by obtaining the surrounding obstacle distribution, spatial structure, lighting conditions, position coordinates, etc. Data such as cameras, laser radars, ultrasonic sensors, global positioning system modules carried by the wheelchair, and its purpose is to provide external environmental restrictions and semantic references for instruction execution; evaluating the executability and semantic integrity of the voice instruction refers to the process of comprehensively judging whether the voice instruction can be effectively executed under the current wheelchair state and environmental conditions and whether its meaning is clear and explicit, which can be achieved by using a rule-based reasoning engine, a machine learning model or an expert system, combined with the pre-set execution logic and semantic rules for judgment, and its purpose is to comprehensively consider the actual feasibility and clarity of the instruction; the execution object is not clear refers to the lack of a clear entity or target in the voice instruction, which causes the system to be unable to determine the specific operation object, which can be achieved by using vague pronouns or general words without specifying specific names in the instruction, and its purpose is to identify the lack of operation target or unclear reference in the instruction; the path target is not clear refers to the lack of clear specification of the end position or direction of movement or operation in the voice instruction, which causes the system to be unable to plan a specific execution path, which can be achieved by using vague spatial instructions or lacking specific coordinates, landmark information in the instruction, and its purpose is to identify the ambiguity of the movement end point or direction in the instruction; the execution range is not determined refers to the lack of clear specification of the degree, quantity or duration of operation in the voice instruction, which causes the system to be unable to determine the specific operation amount, which can be achieved by using vague quantifiers or lacking specific numerical values, range limits in the instruction, and its purpose is to identify the lack of operation quantification information in the instruction.

[0059] The scheme of the present application further acquires the current physical state information and environmental context information of the electric wheelchair when judging whether the voice instruction has ambiguous features. These information, such as the power of the wheelchair, speed, tilt angle, location, surrounding obstacle distribution, and light intensity, provides important background data for the actual executability of the instruction. Subsequently, the system comprehensively considers these acquired current physical state information and environmental context information with the parameter information extracted from the voice instruction, thereby evaluating the actual executability and semantic integrity of the voice instruction. This evaluation process goes beyond simple syntax or vocabulary analysis, and it goes deep into the feasibility level of the instruction in the real world. For example, even if the instruction seems complete in semantics, if the wheelchair power is too low to execute, or there are obstacles in front of the wheelchair, the system can identify its unexecutability. Ultimately, when the results of this comprehensive evaluation show that the voice instruction has unclear execution objects, unclear path targets, and / or uncertain execution ranges, the system determines that the voice instruction is an ambiguous instruction. This judgment mechanism enables the system to more accurately identify instructions that are indeed unexecutable or ambiguous in a specific context, avoiding the misjudgment of context-limited instructions as clear instructions, and avoiding unnecessary clarification of understandable instructions in the context. In this way, the scheme of the present application can more comprehensively consider the actual context of the instruction when identifying ambiguous instructions in the voice instruction, thereby improving the accuracy of ambiguous instruction judgment, reducing unnecessary interaction, and enabling the voice interaction system of the electric wheelchair to more intelligently respond to the user's true intention.

[0060] Reference Figure 2 In another embodiment of the present application, an electric wheelchair voice instruction recognition system is further proposed, which comprises: A receiving and identifying module 1 for receiving a voice instruction of a user and identifying ambiguous instructions in the voice instruction; A perception data acquisition module 2 for acquiring perception data related to the current environment of the electric wheelchair based on the ambiguous instruction; A target distance identification module 3 for identifying physical targets in the perception data and the relative distance between the physical targets and the electric wheelchair; An inquiry generation module 4 for generating a somatized clarification inquiry containing target name and location information from the physical targets and their relative distances according to a preset attribution strategy; An execution module 5 for receiving a voice response of the user to the somatized clarification inquiry, generating instruction parameters from the voice response, and controlling the electric wheelchair to perform corresponding actions according to the instruction parameters; A record optimization module 6 for recording the voice instruction, the somatized clarification inquiry, the voice response, and the instruction parameters as an interaction log, and optimizing the attribution strategy based on the interaction log.

[0061] Specifically, the receiving recognition module 1 can be composed of a speech recognition engine and a natural language processing unit. The speech recognition engine is responsible for converting the user's voice instructions into text, such as using deep learning models for acoustic modeling and language modeling. The natural language processing unit then performs semantic analysis on the text instructions, identifies the operation object, operation method and spatial position parameters, and judges whether there are ambiguous features, such as identifying "a little further", "to the left" and other ambiguous expressions through keyword matching, syntax analysis or intent recognition algorithms.

[0062] The perception data acquisition module 2 can be configured with various sensors, such as laser radar for obtaining high-precision environmental point cloud data, stereo camera for capturing visual image information, and ultrasonic sensor for close-range obstacle detection. The data of these sensors is processed through data fusion algorithms to construct a real-time three-dimensional map or obstacle distribution map of the environment around the electric wheelchair.

[0063] The target distance recognition module 3 can use computer vision algorithms to identify specific physical targets from visual images, such as "table", "chair", "door", etc. At the same time, combined with laser radar or ultrasonic data, through point cloud clustering, target tracking or geometric calculation, the relative distance and spatial position between these physical targets and the electric wheelchair are accurately measured.

[0064] The query generation module 4 can be built-in with a knowledge graph or rule engine, storing pre-set attribution strategies. When ambiguous instructions and physical targets in the environment are identified, the module will select one or more candidate targets according to the attribution strategy, and generate embodied clarification inquiries containing target name and location information. For example, if there is a table and a chair in front, the system may generate "Are you referring to the table two meters in front of you?" or "Do you want to approach the chair on the left side?".

[0065] The execution module 5 can include a speech synthesizer for broadcasting clarification inquiries, as well as an instruction parser and a motion control unit. The instruction parser receives the user's voice response to the clarification inquiry, extracts the user's confirmed target information or new instruction intent through speech recognition and semantic understanding. The motion control unit then generates specific motor control signals according to the parsed instruction parameters to drive the electric wheelchair's hub motor to perform forward, backward, turning or stopping actions, etc.

[0066] The record optimization module 6 can be a database system for storing interaction logs. Each complete voice interaction is recorded as a log entry, containing a timestamp, user ID, original voice instruction, clarifying questions issued by the system, voice responses by the user, and final instruction parameters. A background machine learning model can periodically analyze these interaction logs to identify user behavior patterns and effectiveness of attribution strategies. For example, if a certain attribution strategy often leads to negative responses by the user, the system can adjust the weight or rules of that strategy to improve the accuracy of subsequent clarifying questions and user satisfaction.

[0067] The above only describes the embodiments of the present application and is not used to limit the protection scope of the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for recognizing voice commands for an electric wheelchair, characterized in that: The method comprises: receiving a user's voice command and identifying ambiguous commands in the voice command; Based on the fuzzy instruction, acquiring perception data related to the current environment of the electric wheelchair; identifying a physical target in the perception data and a relative distance between the physical target and the electric wheelchair; According to a preset attribution strategy, an embodied clarification query including the target name and location information is generated from the physical target and its relative distance; receiving a user's voice response to the embodied clarification inquiry, generating instruction parameters according to the voice response, and controlling the electric wheelchair to perform corresponding actions according to the instruction parameters; The voice commands, embodied clarification inquiries, voice responses, and command parameters are recorded as an interaction log, and an attribution strategy is optimized based on the interaction log.

2. The electric wheelchair voice command recognition method according to claim 1, characterized in that: Generating instruction parameters according to the voice response and controlling the electric wheelchair to perform corresponding actions according to the instruction parameters includes: When the user's voice response contains negative semantics, parsing the voice response to obtain new target information and its relative distance; Based on the new target information and its relative distance, a new embodied clarification inquiry is generated, and after receiving the user's voice response, the instruction parameters are extracted, and the electric wheelchair is controlled to perform corresponding actions according to the instruction parameters.

3. The electric wheelchair voice command recognition method according to claim 2, characterized in that: When the user's voice response includes a negative semantic, parsing the voice response to obtain new target information and its relative distance includes: Matching multiple candidate physical targets from the detected and identified physical targets based on the new target information and the relative distance thereof; Acquire a spatial relationship between a plurality of candidate physical objects and the electric wheelchair, including relative distances, spatial positions, and orientation angles; Prioritizing the plurality of candidate physical targets according to the spatial relationship; If there is only one candidate physical target with the highest priority, select the candidate physical target with the highest priority as the new target information; If there are multiple candidate physical targets with the highest priority, descriptive information containing differentiated features between the candidate physical targets is generated for subsequent generation of embodied clarification inquiries.

4. The electric wheelchair voice command recognition method according to claim 1 or 2, characterized in that: Generating instruction parameters according to the voice response, and controlling the electric wheelchair to perform corresponding actions according to the instruction parameters, further comprising: When the voice response is affirmative, generating instruction parameters according to the target name and location information included in the embodiment clarification inquiry; The parking distance of the physical target is calculated based on the position information, and the parking distance is used as the execution boundary of the instruction parameter to control the electric wheelchair to move in the direction of the physical target corresponding to the target name until it stops at a preset safety distance from its parking position.

5. The electric wheelchair voice command recognition method according to claim 1, characterized in that: The recording of the voice command, the embodied clarification inquiry, the voice response, and the command parameters as an interaction log, and optimizing the attribution strategy based on the interaction log, includes: Get the current user's identity information and context information; Based on the identity information and context information, recording an interaction log including the voice command, the embodied clarification query, the voice response, and the command parameters; Filtering historical interaction records that match the identity information and context information of the current user from the interaction log; Adjusting the attribution strategy content based on the historical interaction records; When a change in user identity information or context information is detected, an attribution strategy corresponding to the new identity information or new context information is loaded.

6. The electric wheelchair voice command recognition method according to claim 1, characterized in that: The attribution strategy includes: Identify all identified physical targets in front of the electric wheelchair and their corresponding relative distances, spatial positions, and orientation angles based on the perception data; According to the preset exclusion rules, the identified physical targets are screened: the identified physical targets that are blocked by obstacles and the electric wheelchair, the identified physical targets whose passage width to the electric wheelchair is less than a preset threshold, and / or the identified physical targets whose relative distance to the electric wheelchair is less than a safety distance threshold are excluded to obtain candidate physical targets; The candidate physical target with the smallest angle with the current direction of travel of the electric wheelchair is selected as the source of the target name and position information referenced when generating the embodied clarification query.

7. The electric wheelchair voice command recognition method according to claim 1, characterized in that: The attribution strategy includes: According to the physical target and its relative distance, determine whether the physical target belongs to a preset sensitive target type, and whether the relative distance of the physical target is less than a preset safety distance; If the physical target belongs to the sensitive target type, or the relative distance is less than the safe distance, generating the embodied clarification query including the target name and location information; If the physical target is not a sensitive target type and the relative distance is not less than the safety distance, determining whether the physical target is within the preset moving path of the electric wheelchair; If the physical object is within the range of the moving path, the embodiment clarification query directed to the object name and location information of the physical object is generated.

8. The electric wheelchair voice command recognition method according to claim 1, characterized in that: The identifying of ambiguous instructions in the voice instructions includes: Parsing the voice command to extract parameter information including the operation object, operation method, and spatial position; determining whether the voice command has fuzzy features according to the parameter information; If a fuzzy feature is identified, a corresponding fuzzy instruction in the voice instruction is identified according to the fuzzy feature.

9. The electric wheelchair voice command recognition method according to claim 8, characterized in that: The determining, based on the parameter information, whether the voice instruction has a fuzzy feature includes: Get the current physical state information and environmental context information of the electric wheelchair; evaluating the executability and semantic integrity of the voice command by combining the parameter information, the current physical state information, and the environmental context information; When the evaluation result shows that the voice instruction has an unclear execution object, an unclear path target and / or an uncertain execution scope, the voice instruction is determined to be an ambiguous instruction.

10. A voice command recognition system for an electric wheelchair, characterized in that: The system comprises: A receiving and identifying module, configured to receive a user's voice command and identify ambiguous commands in the voice command; a perception data acquisition module, configured to acquire perception data related to the current environment of the electric wheelchair based on the fuzzy instruction; a target distance recognition module, configured to recognize a physical target in the sensing data and a relative distance between the physical target and the electric wheelchair; a query generation module, configured to generate, based on a preset attribution strategy, an embodied clarification query including target name and location information from the physical target and its relative distance; an execution module, configured to receive a user's voice response to the embodied clarification inquiry, generate instruction parameters according to the voice response, and control the electric wheelchair to perform corresponding actions according to the instruction parameters; A record optimization module is used to record the voice instructions, embodied clarification inquiries, voice responses and instruction parameters as an interaction log, and optimize the attribution strategy based on the interaction log.

Citation Information

Patent Citations

  • Voice interaction method, device, equipment of intelligent voice equipment, medium and product

    CN112767916A

  • Automatic driving wheelchair navigation method and system based on intelligent voice interaction

    CN119268707A

  • Voice instruction recognition method, device and equipment, medium and vehicle

    CN119446130A

  • Humanoid robot multi-mode instruction analysis system

    CN120516701A

  • Robot Natural Language Term Disambiguation and Entity Labeling

    US20190102377A1

Cited By

  • Voice control method for robot

    CN121393441A