Assisted walking device and control method based on multimodal perception and interaction
Through a multimodal sensing and interactive walking assist device, combined with environmental perception, eye tracking and voice interaction, intelligent path planning and power assistance for obstacles such as stairs are achieved, solving the safety issues of existing walking assist tools on stairs and providing safe and intelligent walking support.
Patent Information
- Application Number
- CN202511028294.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Existing walking aids are difficult to provide intelligent safety support at vertical obstacles such as stairs, lack the ability to interact with the environment intelligently, and are unable to perform real-time adaptive path planning, increasing the risk of falls and injuries for people with lower limb mobility difficulties.
The assisted walking device adopts multimodal perception and interaction, integrates artificial intelligence algorithms, augmented reality technology and eye tracking mechanism, identifies stairs and obstacles through the environmental perception module, combines eye tracking and voice interaction, uses generative adversarial networks and A* algorithms for path planning, and provides power assistance through the lower limb exoskeleton.
It realizes intelligent perception and interaction of the user environment, generates the optimal walking route, provides safe and intelligent walking assistance, improves the system response speed and the accuracy of path planning, and reduces the risk of falls.
Smart Images

Figure CN120516665B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of augmented reality and intelligent assisted walking technology, and in particular to an assisted walking device based on multimodal perception and interaction and a control method thereof. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] With the acceleration of global population aging, the number of elderly people with disabilities and lower limb mobility difficulties is increasing. For these people, performing daily activities and independent mobility safely and effectively is one of the main challenges they face, especially in the current living environment. Since most residential buildings, especially old buildings, may not be equipped with elevators or lifting devices, users need to overcome these terrain obstacles on their own, and the elderly group faces difficulties in climbing stairs. Although the mobility aids currently on the market, such as crutches, walkers and wheelchairs, can provide users with basic mobility support on flat ground, they are often powerless when faced with vertical obstacles such as stairs. That is, when going up and down stairs, people with lower limb mobility difficulties need additional assistance to ensure safety. However, traditional mobility aids are difficult to provide sufficient support and lack intelligent assistance, which greatly increases the potential risk of falls and injuries for people with lower limb mobility difficulties.
[0004] While smart exoskeletons are already available on the market, their primary function is to enhance the user's strength and endurance. They lack the ability to intelligently interact with the environment and are unable to effectively handle specific scenarios like stairways. Furthermore, existing devices typically rely on mechanical assistance and lack intelligent navigation capabilities, failing to provide real-time adaptive path planning and guidance. Summary of the Invention
[0005] In order to address the deficiencies of the above-mentioned existing technologies, the present invention provides an assisted walking device based on multimodal perception and interaction and a control method thereof, which integrates artificial intelligence algorithms, augmented reality technology, eye tracking mechanism and other technologies, and applies them to lower limb exoskeleton equipment. It can perceive environmental characteristics and interact naturally with the user, thereby realizing optimal walking route planning and adaptive autonomous adjustment of the equipment, and can provide a customized user experience for disabled elderly people and people with lower limb walking difficulties, providing users with safe and intelligent walking assistance.
[0006] In a first aspect, the present invention provides an auxiliary walking device based on multimodal perception and interaction.
[0007] An assisted walking device based on multimodal perception and interaction, comprising a lower limb exoskeleton and an augmented reality device; wherein the augmented reality device is equipped with:
[0008] Eye tracking module, used to obtain the real-time focus area of the user's eyes;
[0009] The environmental perception module is used to obtain real-time surrounding environment information and use an improved recognition model to identify stairs and obstacles in the current environment;
[0010] The path planning module is used to perform local and global path searches based on surrounding environment information, identified stairs and obstacles, and the user's real-time eye focus area, using the generative adversarial network algorithm and the A* algorithm to generate the optimal planned path;
[0011] The lower limb exoskeleton is equipped with a power assistance and drive module, which is used to obtain the user's leg posture in real time, and adaptively and dynamically adjust the output force and torque according to the user's leg posture and the optimal planned path to assist the user in walking.
[0012] As a further technical solution, the augmented reality device is further equipped with:
[0013] A voice interaction module is used to obtain user voice commands, wherein user voice commands include navigation control commands, mode switching commands and emergency commands;
[0014] The display module is used to display the real environment scene, generate and annotate virtual reality navigation prompts in the scene according to the optimal planned path, and dynamically adjust the navigation prompts according to the user's leg posture;
[0015] Among them, the voice interaction module sends the acquired user voice commands to the path planning module. The path planning module determines the navigation target position based on the navigation control instructions in the user voice commands, and then uses the navigation target position as the target. According to the surrounding environment information and the identified stairs and obstacles, combined with the real-time attention area of the user's eyes, the generative adversarial network algorithm and A* algorithm are used to perform local and global path searches to generate the optimal planned path.
[0016] According to a further technical solution, the improved recognition model includes an input network, a backbone network, a neck network and a prediction network arranged in sequence:
[0017] The input network is used to input the surrounding environment image;
[0018] The backbone network is used to extract multi-scale features of the input image; wherein the backbone network includes a Focus layer, a convolution block, an SCPNet module, a CBAM module, and an SPP module; the Focus layer is used to convert the spatial information of the input image into channel information, thereby increasing the number of channels of the feature; the convolution block is used for downsampling; the CSPNet module is used to extract deep features by using separate convolutions and splicing feature maps of different paths; the CBAM module is used to extract and adjust the weights of key features in the image by using a channel and spatial attention mechanism, thereby generating channel and spatial attention weighted features; the SPP module is used to extract multi-scale features by using multi-scale pooling;
[0019] The neck network adopts FPN and PAN structures to perform top-down semantic fusion on the extracted multi-scale features, and then transfers the spatial positioning features from bottom to top to generate the final features;
[0020] The prediction network is used to output a recognition result.
[0021] A further technical solution is to use the improved GIoU loss function when training the improved recognition model. This function is used to measure the difference between the predicted box and the true box. The calculation formula is:
[0022] ;
[0023] ;
[0024] ;
[0025] in, is the prediction box With real box The intersection area, is the prediction box With real box The union area of , Is the prediction box and real frame The minimum enclosing frame.
[0026] A further technical solution is that in the path planning module:
[0027] A generative adversarial network algorithm based on a generator and a discriminator is used. The generator inputs environmental information, identified stairs and obstacles, and the user's real-time eye focus area, and outputs a region of interest. During the training of the generative adversarial network, a gradient penalty mechanism is introduced. The generator optimizes the discriminator's generated region of interest by minimizing the confidence loss of the generated region of interest. The discriminator is updated by maximizing the difference between the real region of interest and the generated region of interest and combining it with a gradient penalty.
[0028] According to the region of interest generated by the generator, the A* algorithm is used to perform local path exploration. If the target point is not explored in the region of interest, it is expanded to a global map search to generate the optimal planning path.
[0029] In a further technical solution, in the eye tracking module:
[0030] Using an infrared camera to capture the user's eye movement data in real time; the eye movement data includes the coordinates of the gaze point, pupil diameter, number of eye saccades, and gaze duration;
[0031] According to the eye movement data, after filtering by the filtering algorithm, the fixation point extraction algorithm is used to divide the fixation and saccade states, and output the eye attention area, fixation heat map, and trajectory sequence data.
[0032] According to a further technical solution, the lower limb exoskeleton includes an exoskeleton main body, and the exoskeleton main body is provided with a protective member and an adaptive adjustment member;
[0033] The protective parts include: a transverse rotation axis for expanding the freedom of thigh abduction and adjusting the hip joint angle; a spine protection plate, a waist protection holster, a thigh protection holster and a calf protection holster for wrapping the waist, thigh and calf respectively;
[0034] The adaptive adjustment components include a thigh height adjustment slot, a calf adjustment rod, and a foot adjustment buckle, which are used to adjust the height of the exoskeleton's thigh and calf and the space for the feet respectively.
[0035] As a further technical solution, the lower limb exoskeleton is also equipped with a communication module for two-way data interaction with an augmented reality device;
[0036] Among them, the lower limb exoskeleton transmits the user's leg posture obtained in real time to the augmented reality device; the augmented reality device transmits the user's voice commands and optimal planned path to the lower limb exoskeleton.
[0037] In a second aspect, the present invention provides a control method for an auxiliary walking device based on multimodal perception and interaction.
[0038] A control method for an auxiliary walking device based on multimodal perception and interaction, comprising:
[0039] Wearing the lower limb exoskeleton and augmented reality device in the walking assistance device proposed in the first aspect on the user's lower limbs and head respectively;
[0040] Use the eye tracking module to obtain the real-time focus area of the user's eyes;
[0041] Use the environmental perception module to obtain real-time surrounding environment information and identify stairs and obstacles in the current environment through an improved recognition model;
[0042] The path planning module uses the generative adversarial network algorithm and the A* algorithm to perform local and global path searches based on the surrounding environment information, identified stairs and obstacles, and the user's real-time eye focus area to generate the optimal planned path.
[0043] The power assistance and drive module is used to obtain the user's leg posture in real time, and according to the user's leg posture and the optimal planned path, the output force and torque are adaptively and dynamically adjusted to assist the user in walking.
[0044] Further technical solutions also include:
[0045] Perform voice interaction between the user and the device through the voice interaction module to obtain user voice commands;
[0046] The path planning module determines the navigation target location based on the navigation control instructions in the user's voice command. Then, based on the navigation target location, the generative adversarial network algorithm and the A* algorithm are used to perform local and global path searches based on the surrounding environment information, the identified stairs and obstacles, and the real-time focus area of the user's eyes to generate the optimal planned path.
[0047] The real environment scene is displayed through the display module, and virtual reality navigation prompts are generated and marked in the scene according to the optimal planned path, and the navigation prompts are dynamically adjusted according to the user's leg posture.
[0048] One or more of the above technical solutions have the following beneficial effects:
[0049] 1. The present invention provides an assisted walking device based on multimodal perception and interaction and its control method. It integrates artificial intelligence algorithms, augmented reality technology, eye tracking mechanism and other technologies, and applies them to lower limb exoskeleton equipment. It can perceive environmental characteristics and interact naturally with the user, thereby achieving optimal walking route planning and adaptive and autonomous adjustment of the equipment. It can provide a customized user experience for disabled elderly people and people with lower limb walking difficulties, and provide users with safe and intelligent walking assistance.
[0050] 2. The present invention introduces a generative adversarial network algorithm based on a generator and a discriminator in path planning. By introducing a gradient penalty mechanism in the training process of the generative adversarial network, the generator optimizes the confidence loss of the generated region of interest by minimizing the confidence loss of the discriminator. The discriminator is updated by maximizing the difference between the real region of interest and the generated region of interest and combining the gradient penalty. In this way, the trained generator can be used to generate the region of interest, and then the focus weighting mechanism is used to prioritize the environmental information of the user's line of sight focus area, that is, the region of interest, triggering the dynamic path planning algorithm to adjust the navigation strategy, thereby effectively improving the response speed and accuracy of the system.
[0051] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0053] Figure 1 Schematic diagram of the structure of an auxiliary walking device based on multimodal perception and interaction in an embodiment of the present invention;
[0054] Figure 2 Schematic diagrams of the structure of the lower limb exoskeleton in an embodiment of the present invention; (a) is an overall schematic diagram of the lower limb exoskeleton, and (b) is a front view of the lower limb exoskeleton;
[0055] Figure 3 Flowchart of stair and obstacle identification and path planning in an embodiment of the present invention;
[0056] Figure 4 Schematic diagram of the structure of the improved recognition model in an embodiment of the present invention.
[0057] Among them, 1. Control unit; 2. Horizontal rotation axis; 3. Torque motor; 4. Calf adjustment rod; 5. Foot adjustment leather buckle; 6. Spine protection plate; 7. Waist protection leather cover; 8. Thigh protection leather cover; 9. Aluminum alloy support frame; 10. Calf protection leather cover; 11. Display screen; 12. Emergency stop button; 13. Expansion hook; 14. Signal indicator light; 15. Adjustment button; 16. Thigh height adjustment slot. DETAILED DESCRIPTION
[0058] It should be noted that the following detailed descriptions are exemplary only and are intended to describe specific embodiments and provide further explanation of the present invention, and are not intended to limit the exemplary embodiments according to the present invention. Unless otherwise indicated, all technical and scientific terms used herein have the same meanings as those commonly understood by those of ordinary skill in the art to which the present invention belongs. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0059] Example 1
[0060] This embodiment provides an auxiliary walking device based on multimodal perception and interaction, including a lower limb exoskeleton and an augmented reality device, such as Figure 1As shown, the augmented reality device is equipped with an eye tracking module, an environmental perception module, a path planning module, a voice interaction module and a display module, and the lower limb exoskeleton is equipped with a power assistance and drive module and a communication module.
[0061] In this embodiment, the augmented reality device uses the HoloLens 2 produced by Microsoft, which is used to provide high-precision environmental perception and user interaction functions, realize the seamless integration of virtual and real information in the 3D environment, and realize humanized interaction functions. The device provides the hardware foundation required for the above-mentioned environmental perception module, eye tracking module, path planning module, display module, and voice interaction module. The specific functions of these modules are as follows:
[0062] (A) The Environmental Perception Module acquires real-time information about the surrounding environment and uses an improved recognition model to identify stairs and obstacles, ensuring safe navigation in complex environments. Specifically, the Environmental Perception Module includes a depth sensor and stereo camera, and uses the SANC algorithm to identify stairs and obstacles. These detection results are used to dynamically adjust the user's navigation path, providing real-time environmental feedback.
[0063] The above SANC algorithm is a Stair Adaptive Navigation & Climbing algorithm, such as Figure 3 As shown, the algorithm includes a stair target recognition algorithm based on an improved recognition model (ie, a stair detection and recognition model) and a path planning algorithm.
[0064] In this embodiment, Figure 4 As shown in FIG, the improved recognition model includes an input network (Input), a backbone network (Backbone), a neck network (Neck), and a prediction network (Prediction) arranged in sequence, wherein:
[0065] (1) The input network is used to input the surrounding environment image.
[0066] The Mosaic data augmentation strategy is used for image input. This strategy achieves data diversification by randomly cropping and scaling four different images and then stitching them into a new image. Specifically, four images are randomly selected from the dataset and randomly cropped and scaled. Then, they are arranged in a 2×2 layout around a central point. Finally, the four images are stitched together into a complete image, with necessary color and brightness adjustments made to ensure the naturalness of the stitched image. Finally, the original label information is adjusted based on the position of the stitched image to ensure its accuracy. This strategy significantly improves the model's detection capabilities under complex backgrounds and occlusion conditions.
[0067] (2) The backbone network is used to extract multi-scale features of the input image. The backbone network mainly includes the Focus layer, Conv convolution block, SCPNet module, CBAM module, SPP module, and CBL module, among which:
[0068] The Focus layer converts the spatial information of the input image into channel information, increasing the number of feature channels and extracting richer detail features. This feature helps accurately capture edges in staircase detection.
[0069] Convolutional blocks are used for downsampling;
[0070] The CSPNet module is used to extract more expressive deep features by using separate convolutions and splicing feature maps from different paths, making the model more suitable for processing images with complex background environments and improving recognition accuracy;
[0071] The CBAM module uses channel and spatial attention mechanisms to extract and adjust the weights of key features in the image, generating channel and spatial attention-weighted features to strengthen the feature weights of key areas such as stair edges, thereby enhancing the model's recognition performance under complex backgrounds or lighting conditions.
[0072] The SPP (Spatial Pyramid Pooling) module is used to utilize multi-scale pooling to extract multi-scale features, enabling the model to effectively cope with the different sizes and perspective changes of stairs.
[0073] The CBL module, as the basic unit in a convolutional neural network, includes a convolution layer, a batch normalization layer, and an activation layer (LeakyReLU), which are used to extract local features, accelerate training convergence, and enhance nonlinear expression capabilities. Preferably, it also includes: add (addition module), maxpool (maximum pooling module), concat (connection module), etc.
[0074] (3) The neck network adopts FPN and PAN structures to perform top-down semantic fusion on the extracted multi-scale features, and then transfers the spatial positioning features from bottom to top to generate the final features;
[0075] (4) The prediction network is used to output the recognition results.
[0076] When training the improved recognition model, the loss function used is the improved GIoU loss function. GIoU loss is an extension of IoU (Intersection over Union). This function is used to measure the difference between the predicted box and the true box. The calculation formula is:
[0077] ;
[0078] ;
[0079] ;
[0080] in, is the prediction box With real box The intersection area, is the prediction box With real box The union area of , Is the prediction box and real frame The minimum enclosing box, that is, the minimum enclosing rectangle.
[0081] The optimized YOLOv5 model can effectively improve the performance of stair detection in complex and occluded environments.
[0082] (B) The eye tracking module is used to obtain the user's real-time eye attention area. In this embodiment, the eye tracking module uses an infrared camera to capture the user's eye movement data in real time. This eye movement data includes gaze point coordinates, pupil diameter, number of saccades, and gaze duration. Based on this eye movement data, the module analyzes gaze direction and visual attention patterns to determine the user's real-time area of interest (AOI), thereby understanding their intent and optimizing the interaction logic.
[0083] The above-mentioned eye tracking module is implemented by hardware such as infrared tracking cameras and 3D spatial positioning sensors and real-time data processing algorithm software. At the hardware level, the original eye movement signals (i.e., eye movement data) are collected. At the software level, the original eye movement data collected is subjected to noise reduction by filtering algorithms (such as Kalman filtering or Savitzky-Golay filtering, etc.) to eliminate noise; the gaze point extraction algorithm is then used to divide the gaze state and the saccade state, that is, the eye movement speed is calculated according to the gaze point coordinates and gaze duration, and the judgment is made based on the eye movement speed. If the speed is lower than the set threshold, it is determined to be a gaze state, otherwise it is determined to be a saccade state. In the gaze state, the user's line of sight is stable, and the acquired eye movement data is valid, which represents the user's actual focus area. In the saccade state, rapid switching of the gaze point will generate noise data, which needs to be filtered to avoid interference. By dividing the gaze state and the saccade state, the focus area can be improved. The accuracy of area of interest (AOI) analysis is improved; clustering algorithms such as K-means are then used to merge adjacent gaze points to generate gaze heatmaps and areas of interest (AOIs), thereby outputting high-level data including gaze heatmaps, trajectory sequences, and AOI statistics. Among them, the above-mentioned gaze point aggregation only retains data points in the gaze state for aggregation, ensuring that the generated gaze heatmaps and AOIs truly reflect the user's intentions; finally, intention inference is performed. According to the position and duration of the area of interest (AOI), the user's visual attention distribution is analyzed, and the environmental information of the gaze area is prioritized. That is, in complex navigation scenarios, based on the dynamic data such as the area of interest (AOI) obtained by the above analysis, the environmental information of the user's focus area (such as obstacles, stair positions, etc.) is prioritized through the attention point weighting mechanism, and the dynamic path planning algorithm is triggered to adjust the navigation strategy, thereby improving the system's response speed and accuracy.
[0084] The process of adjusting the navigation strategy using the focus weighting mechanism and dynamic path planning algorithm is as follows:
[0085] First, weight distribution: assigning higher weight to the environmental information of the user's focus area;
[0086] Secondly, local planning is prioritized: the path planning module prioritizes generating regions of interest (RoIs) within the focus area based on weight distribution, and screens feasible paths through a generative adversarial network (GAN);
[0087] Finally, the global search trigger condition: If the target point (such as the top of the stairs) is not within the RoI, the global map is switched and the full path search is performed using the A* algorithm. For example, if the user is looking at a stair step, the path planning module first calculates the path in the step area and then expands to the entire staircase.
[0088] (C) The path planning module is used to generate the optimal planned path by performing local and global path searches based on the surrounding environment information, identified stairs and obstacles, and the user's real-time eye focus area using the Generative Adversarial Network (GAN) algorithm and the A* algorithm.
[0089] This embodiment introduces a GAN-based algorithm into the path planning module, specifically implemented as a gradient-penalized Wasserstein GAN (WGAN-GP). This algorithm significantly improves the stability and convergence speed of training by introducing a gradient penalty mechanism instead of traditional weight clipping.
[0090] Specifically, first, using the GAN algorithm based on the generator and discriminator, the surrounding environment information, the identified stairs and obstacles, and the real-time focus area of the user's eyes are input into the generator, and the area of interest is output.
[0091] Among them, the input of the generator includes a random noise vector and environmental information (such as maps, obstacle locations, starting and target points, etc.), and generates a region of interest (RoI) through a fully connected layer or a convolutional neural network, and the output is a probability distribution of areas where feasible paths may exist; the discriminator distinguishes between real RoIs and generated RoIs through a convolutional network, and introduces a gradient penalty mechanism (GP) to ensure Lipschitz continuity. This Lipschitz continuity is a stronger mathematical constraint condition than conventional continuity. It is mainly used to limit the rate of change of the function, ensuring that the slope (speed of change) of the function between any two points does not exceed a fixed constant, so that the discriminator satisfies Lipschitz continuity to stabilize adversarial training.
[0092] Specifically, during the training process of the generative adversarial network, a gradient penalty mechanism is introduced. The generator optimizes the confidence loss of the generated RoI by minimizing the discriminator, while the discriminator is updated by maximizing the difference between the true RoI and the generated RoI combined with the gradient penalty. The gradient penalty can force the gradient norm of the discriminator not to exceed the threshold, thereby avoiding gradient explosion or disappearance, which is expressed as:
[0093] ;
[0094] in, is the random interpolation of real and generated samples, λ is the penalty coefficient, represents the gradient, represents the discriminator.
[0095] Secondly, in the path planning stage, the A* algorithm is guided to perform local path exploration based on the region of interest (RoI) generated by the generator. If the target point (i.e., the final location the user needs to reach, such as the top or bottom of the stairs) cannot be explored in the RoI, it is expanded to a global map search to generate the optimal planned path.
[0096] This approach significantly improves the efficiency and accuracy of path planning by generating high-quality RoIs and combining local and global search strategies. It can especially plan safe and optimized paths in complex environments (such as stair climbing), ensuring that users can successfully and safely complete their goals.
[0097] (D) A voice interaction module, configured to obtain user voice commands, wherein user voice commands include navigation control commands, mode switching commands, and emergency commands.
[0098] By setting up a voice interaction module, the acquired user voice commands are sent to the path planning module. The path planning module determines the navigation target position based on the navigation control instructions in the user voice commands. Then, with the navigation target position as the target, based on the surrounding environment information and the identified stairs and obstacles, combined with the real-time attention area of the user's eyes, the generative adversarial network algorithm and A* algorithm are used to perform local and global path searches to generate the optimal planned path.
[0099] Specifically, the voice interaction module collects user voice signals through a multi-microphone array, suppresses ambient noise through beamforming technology, and then uses an end-to-end voice recognition model to convert the voice into text instructions. In this embodiment, three types of command semantic templates are preset: (1) Navigation control instructions, such as "Start navigating to the second floor" and "Pause moving", which trigger the path planning module to recalculate the target point and update the AR interface guidance; (2) Mode switching instructions, such as "Switch stair mode" and "Switch night mode", which dynamically adjust the exoskeleton motor joint torque coefficient and AR display brightness. For example, in stair mode, when stairs are detected, the knee joint torque is increased by 30% to assist climbing; in flat mode, the default torque is restored to optimize walking efficiency; in night mode, the torque output is reduced to match the user's gait stability requirements, and night lighting is turned on; (3) Emergency instructions, such as "Emergency stop" and "Request help", which interrupt the current task first and activate the emergency stop mechanism of the protection module.
[0100] After recognizing the above instructions, they are allocated to the environmental perception module (for target point update), the path planning module (for path re-planning) or the exoskeleton's power assistance and drive module (for power parameter adjustment) through a dynamic priority queue. Specifically, the above-mentioned dynamic priority queue actually refers to the processing mechanism of multimodal input (voice, eye movement, etc.): emergency instructions (such as "emergency stop") have the highest priority and directly interrupt the current task; the gaze area has the second highest weight and gives priority to updating the navigation path; voice instructions (such as switching modes) trigger parameter adjustment, and the priority is dynamically allocated according to the context. At the same time, the execution status is fed back in real time through the speech synthesis engine, such as "the path to the second floor has been planned, and it is expected to take 3 minutes", etc., to achieve closed-loop interaction. Preferably, the collaborative priority of voice instructions, eye tracking, and gesture input is dynamically allocated by the attention mechanism to ensure the consistency of multimodal interaction.
[0101] (E) Display module, used to display the real environment scene, generate and annotate virtual reality navigation prompts in the scene based on the optimal planned path, and dynamically adjust the navigation prompts according to the user's leg posture.
[0102] In this embodiment, the display module is responsible for overlaying virtual information and navigational prompts within the user's field of view, providing an immersive mixed reality experience. Specifically, through a high-definition perspective display and a graphics processing unit, this module utilizes 2D and 3D graphics generation and optical lens mapping technology to seamlessly overlay virtual arrows, path guidance, and precautions onto the user's view of the real environment.
[0103] Through the above-mentioned augmented reality display method, users can focus their attention on safe and effective path planning, and effectively improve navigation efficiency in complex environments.
[0104] In this embodiment, the lower limb exoskeleton is equipped with a power assistance and drive module, which is used to obtain the user's leg posture in real time, and adaptively and dynamically adjust the output force and torque according to the user's leg posture and the optimal planned path to assist the user in walking.
[0105] Specifically, the power assistance and drive module includes a control unit 1, a torque motor 3 and an aluminum alloy support frame 9, which are used to provide the necessary physical support and driving force when the user walks and makes other posture changes, especially to reduce the user's physical burden when overcoming height differences (such as going up and down stairs) or walking for a long time. The module is equipped with four high-torque motors (models: Haorui 57, BLDC115), with a single motor torque of up to 250N / M, and a posture sensor integrated inside the module to obtain the user's leg posture data in real time. The module performs dynamic calculations based on the user's leg posture, environmental perception results and the optimal planned path, and automatically adjusts the output force and torque to achieve force assistance and shock absorption effects, thereby improving the user's walking efficiency and comfort. For example, when a user needs to climb a 15cm high staircase, the environmental perception module first identifies the step height, and the path planning module generates the landing point coordinates; secondly, the power assistance and drive module uses the inverse kinematics model of robot motion control to calculate the torque required for the knee joint based on the step height and hip joint angle; then the motor output is adjusted, and the exoskeleton provides additional thrust during the user's leg lifting phase, while limiting the stride to avoid stepping on air; finally, the posture sensor provides real-time feedback on leg posture data. If the center of gravity is detected to be shifting backward, the ankle joint torque is increased to stabilize the posture.
[0106] In the above process, the calculation of the torque required for the knee joint is as follows: Based on the coordinates of the landing point (determined by the previous path planning step) and the hip joint angle, the gravity torque that the joint needs to overcome is calculated, which can be expressed as:
[0107] ;
[0108] in, is the joint torque, is the force to be applied, is the lever arm, is the angle between the force and the lever arm, is the damping term.
[0109] Preferably, during the above-mentioned torque adjustment process, a dynamic adjustment method is adopted, that is, the torque output is corrected in real time through a PID controller to ensure that the auxiliary force is smooth and matches the user's action.
[0110] like Figure 2As shown, the lower limb exoskeleton includes an exoskeleton body, which is equipped with protective components (i.e., protective modules) and adaptive adjustment components (i.e., adaptive modules). The protective components are designed to comprehensively protect the user's spine, waist, thigh, calf, and hip joints. The protective components include: a transverse rotation axis 2 for expanding the degree of freedom of thigh abduction and flexibly adjusting the angle of the hip joint to reduce damage to the hip joint during exercise; a spine protection plate 6, a waist protection holster 7, a thigh protection holster 8, and a calf protection holster 10, which respectively wrap around the waist, thigh, and calf, providing additional support for multiple key areas to prevent sprains and impacts.
[0111] The tilt angle designed to align with the body's natural curvature protects the spine, ensuring correct posture during various activities and reducing spinal stress and potential injury. Furthermore, the module features an emergency stop function. This includes an emergency stop button 12 that quickly locks the motor upon detecting an abnormal movement pattern, preventing accidents such as falls. The operator can also automatically actuate the button to terminate all operations.
[0112] The exoskeleton's adaptive adjustment components include an adjustment button 15, a thigh height adjustment slot 16, a calf adjustment lever 4, and a foot adjustment buckle 5. These adjust the exoskeleton's thigh and calf heights, as well as the foot clearance, respectively. These components can be adjusted based on the user's age, height, weight, and gender, catering to a wider range of users. This adaptive module allows the exoskeleton's height to be flexibly adjusted based on the user's height, ensuring the device adapts to the individual user's physical characteristics, providing optimal support and comfort.
[0113] Furthermore, the lower-limb exoskeleton is equipped with a display screen 11, a signal indicator light 14, and a communication module. The display screen 11 is used to display exoskeleton-related motion data, the signal indicator light 14 is used to display the current operating status, and the communication module is used for two-way data exchange with an augmented reality device. The lower-limb exoskeleton transmits the user's leg posture acquired in real time to the augmented reality device, and the augmented reality device transmits the user's voice commands and the optimal planned path to the lower-limb exoskeleton. Preferably, the lower-limb exoskeleton is also equipped with an expansion hook 13 to facilitate the expansion of other functions.
[0114] Specifically, the communication module realizes two-way data interaction between HoloLens 2 and the intelligent exoskeleton through the wireless communication interface: the AR device sends the environmental perception results (such as stair height, obstacle distribution) and dynamic navigation path to the exoskeleton, and the exoskeleton pre-adjusts the motor torque parameters accordingly (such as increasing the knee joint torque by 30% when stairs are detected to assist climbing); at the same time, the exoskeleton feeds back the user's leg posture data (such as hip joint angle, plantar pressure distribution) to the AR device in real time, which is used to dynamically calibrate the projection position of the virtual navigation arrow and trigger safety warnings.
[0115] Furthermore, the environmental perception module works in conjunction with the exoskeleton protection module. When an unexpected obstacle (such as an approaching moving object) is detected, the AR interface immediately updates the obstacle avoidance path, and the exoskeleton simultaneously adjusts the direction of support force to maintain the user's balance. Through a data synchronization protocol, the two devices achieve coordinated motion assistance and navigation guidance, ensuring safe and smooth operation in complex terrain.
[0116] Example 2
[0117] This embodiment provides a control method for an auxiliary walking device based on multimodal perception and interaction, the method comprising:
[0118] The lower limb exoskeleton and augmented reality device in the walking assistance device proposed in Example 1 are worn on the user's lower limbs and head respectively;
[0119] Use the eye tracking module to obtain the real-time focus area of the user's eyes;
[0120] Use the environmental perception module to obtain real-time surrounding environment information and identify stairs and obstacles in the current environment through an improved recognition model;
[0121] The path planning module uses the generative adversarial network algorithm and the A* algorithm to perform local and global path searches based on the surrounding environment information, identified stairs and obstacles, and the user's real-time eye focus area to generate the optimal planned path.
[0122] The power assistance and drive module is used to obtain the user's leg posture in real time, and according to the user's leg posture and the optimal planned path, the output force and torque are adaptively and dynamically adjusted to assist the user in walking.
[0123] Furthermore, it also includes:
[0124] Perform voice interaction between the user and the device through the voice interaction module to obtain user voice commands;
[0125] The path planning module determines the navigation target location based on the navigation control instructions in the user's voice command. Then, based on the navigation target location, the generative adversarial network algorithm and the A* algorithm are used to perform local and global path searches based on the surrounding environment information, the identified stairs and obstacles, and the real-time focus area of the user's eyes to generate the optimal planned path.
[0126] The real environment scene is displayed through the display module, and virtual reality navigation prompts are generated and marked in the scene according to the optimal planned path, and the navigation prompts are dynamically adjusted according to the user's leg posture.
[0127] The steps involved in the above embodiment 2 correspond to those in embodiment 1. For the specific implementation method, please refer to the relevant description part of embodiment 1.
[0128] Those skilled in the art will appreciate that the modules or steps of the present invention described above can be implemented using a general-purpose computer device. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. Alternatively, they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0129] The above description is only a preferred embodiment of the present invention. Although the specific implementation of the present invention is described in conjunction with the accompanying drawings, it does not limit the scope of protection of the present invention. Those skilled in the art should understand that on the basis of the technical solution of the present invention, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present invention.
Claims
1. An auxiliary walking device based on multimodal perception and interaction, characterized in that: The device comprises a lower limb exoskeleton and an augmented reality device; wherein the augmented reality device is equipped with: Eye tracking module, used to obtain the real-time focus area of the user's eyes; The environmental perception module is used to obtain real-time surrounding environment information and use an improved recognition model to identify stairs and obstacles in the current environment; The path planning module is used to perform local and global path searches based on surrounding environment information, identified stairs and obstacles, and the user's real-time eye focus area, using the generative adversarial network algorithm and the A* algorithm to generate the optimal planned path; The lower limb exoskeleton is equipped with a power assistance and drive module, which is used to obtain the user's leg posture in real time. Based on the user's leg posture and the optimal planned path, it dynamically and adaptively adjusts the output force and torque to assist the user in walking. The improved recognition model includes an input network, a backbone network, a neck network, and a prediction network arranged in sequence: The input network is used to input the surrounding environment image; The backbone network is used to extract multi-scale features of the input image; wherein the backbone network includes a Focus layer, a convolution block, an SCPNet module, a CBAM module, and an SPP module; the Focus layer is used to convert the spatial information of the input image into channel information, thereby increasing the number of channels of the feature; the convolution block is used for downsampling; the CSPNet module is used to extract deep features by using separate convolutions and splicing feature maps of different paths; the CBAM module is used to extract and adjust the weights of key features in the image by using a channel and spatial attention mechanism, thereby generating channel and spatial attention weighted features; the SPP module is used to extract multi-scale features by using multi-scale pooling; The neck network adopts FPN and PAN structures to perform top-down semantic fusion on the extracted multi-scale features, and then transfers the spatial positioning features from bottom to top to generate the final features; The prediction network is used to output a recognition result.
2. The walking assist device based on multimodal perception and interaction according to claim 1, characterized in that: The augmented reality device also includes: A voice interaction module is used to obtain user voice commands, wherein user voice commands include navigation control commands, mode switching commands and emergency commands; The display module is used to display the real environment scene, generate and annotate virtual reality navigation prompts in the scene according to the optimal planned path, and dynamically adjust the navigation prompts according to the user's leg posture; Among them, the voice interaction module sends the acquired user voice commands to the path planning module. The path planning module determines the navigation target position based on the navigation control instructions in the user voice commands, and then uses the navigation target position as the target. According to the surrounding environment information and the identified stairs and obstacles, combined with the real-time attention area of the user's eyes, the generative adversarial network algorithm and A* algorithm are used to perform local and global path searches to generate the optimal planned path.
3. The walking assist device based on multimodal perception and interaction according to claim 1, characterized in that: When training the improved recognition model, the loss function used is the improved GIoU loss function, which is used to measure the difference between the predicted box and the true box. The calculation formula is: ; ; ; in, is the prediction box With real box The intersection area, is the prediction box With real box The union area of , Is the prediction box and real frame The minimum enclosing frame.
4. The walking assist device based on multimodal perception and interaction according to claim 1, characterized in that: In the path planning module: A generative adversarial network algorithm based on a generator and a discriminator is used. The generator inputs environmental information, identified stairs and obstacles, and the user's real-time eye focus area, and outputs a region of interest. During the training of the generative adversarial network, a gradient penalty mechanism is introduced. The generator optimizes the discriminator's generated region of interest by minimizing the confidence loss of the generated region of interest. The discriminator is updated by maximizing the difference between the real region of interest and the generated region of interest and combining it with a gradient penalty. According to the region of interest generated by the generator, the A* algorithm is used to perform local path exploration. If the target point is not explored in the region of interest, it is expanded to a global map search to generate the optimal planning path.
5. The walking assist device based on multimodal perception and interaction according to claim 1, characterized in that: In the eye tracking module: Using an infrared camera to capture the user's eye movement data in real time; the eye movement data includes the coordinates of the gaze point, pupil diameter, number of eye saccades, and gaze duration; According to the eye movement data, after filtering by the filtering algorithm, the fixation point extraction algorithm is used to divide the fixation and saccade states, and output the eye attention area, fixation heat map, and trajectory sequence data.
6. The walking assist device based on multimodal perception and interaction according to claim 1, characterized in that: The lower limb exoskeleton comprises an exoskeleton main body, on which a protective member and an adaptive adjustment member are provided; The protective parts include: a transverse rotation axis for expanding the freedom of thigh abduction and adjusting the hip joint angle; a spine protection plate, a waist protection holster, a thigh protection holster and a calf protection holster for wrapping the waist, thigh and calf respectively; The adaptive adjustment components include a thigh height adjustment slot, a calf adjustment rod, and a foot adjustment buckle, which are used to adjust the height of the exoskeleton's thigh and calf and the space for the feet respectively.
7. The walking assist device based on multimodal perception and interaction according to claim 1, characterized in that: The lower limb exoskeleton is also equipped with a communication module for two-way data interaction with augmented reality devices; Among them, the lower limb exoskeleton transmits the user's leg posture obtained in real time to the augmented reality device; the augmented reality device transmits the user's voice commands and optimal planned path to the lower limb exoskeleton.
8. A control method for an auxiliary walking device based on multimodal perception and interaction, characterized in that: include: Wearing the lower limb exoskeleton and augmented reality device in the assistive walking device according to any one of claims 1 to 7 on the user's lower limbs and head respectively; Use the eye tracking module to obtain the real-time focus area of the user's eyes; Use the environmental perception module to obtain real-time surrounding environment information and identify stairs and obstacles in the current environment through an improved recognition model; The path planning module uses the generative adversarial network algorithm and the A* algorithm to perform local and global path searches based on the surrounding environment information, identified stairs and obstacles, and the user's real-time eye focus area to generate the optimal planned path. The power assistance and drive module is used to obtain the user's leg posture in real time, and according to the user's leg posture and the optimal planned path, the output force and torque are adaptively and dynamically adjusted to assist the user in walking.
9. The control method of the auxiliary walking device based on multimodal perception and interaction according to claim 8, characterized in that: Also includes: Perform voice interaction between the user and the device through the voice interaction module to obtain user voice commands; The path planning module determines the navigation target location based on the navigation control instructions in the user's voice command. Then, based on the navigation target location, the generative adversarial network algorithm and the A* algorithm are used to perform local and global path searches based on the surrounding environment information, the identified stairs and obstacles, and the real-time focus area of the user's eyes to generate the optimal planned path. The real environment scene is displayed through the display module, and virtual reality navigation prompts are generated and marked in the scene according to the optimal planned path, and the navigation prompts are dynamically adjusted according to the user's leg posture.
Citation Information
Patent Citations
Human motion intention recognition method and system
CN111652155A
Trajectory planning system and method for robot
CN112720462A