Wheelchair navigation method, apparatus, device, and medium
By acquiring and analyzing environmental images of the wheelchair, distinguishing between indoor and outdoor environments and generating corresponding navigation feature maps or 3D point clouds, the problem of separation between perception and decision-making in wheelchair navigation is solved, adaptive navigation control is realized, and the stability and real-time performance of navigation are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2026-06-26
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies lack a unified representation mechanism in wheelchair navigation, resulting in the separation of perception, mapping, and decision-making. They also lack continuous environmental updates and memory capabilities, have weak spatial geometric representation capabilities, and are unable to meet real-time navigation requirements.
By acquiring current and historical environmental images of the wheelchair, environmental structural features are extracted to distinguish between indoor and outdoor environments. Based on these features, a bird's-eye view feature map or a 3D point cloud is generated. Navigation is then performed in conjunction with navigation commands to achieve adaptive navigation control for both indoor and outdoor environments.
It improves the accuracy and adaptability of environmental understanding, enhances the ability to express spatial structure and the accuracy of path planning in indoor and outdoor environments, improves the stability and real-time performance of navigation, and solves the problem of poor unified representation of perception, mapping and decision-making.
Smart Images

Figure CN122448233A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of embodied intelligence and autonomous navigation technology, and in particular to a wheelchair navigation method, device, equipment and medium. Background Technology
[0002] The practical application scenarios of intelligent wheelchairs cover both indoor and outdoor environments, which places high demands on the environmental adaptability and safety of navigation systems. Therefore, it is urgent to study autonomous navigation methods suitable for complex and continuous scenarios.
[0003] Related technologies typically separate perception, mapping, and decision-making, lacking a unified representation mechanism and failing to continuously update and remember the environment during long-term operation. While visual-language-action models can be used for command understanding and action planning, their spatial geometric representation capabilities are weak, making it difficult to support high-precision navigation control. Although high-fidelity 3D reconstruction methods can provide dense geometric information, their computational complexity is high, making it difficult to meet the requirements of real-time operation. Summary of the Invention
[0004] This application provides a wheelchair navigation method, device, equipment, and medium to address the problems of poor unified representation of perception, mapping, and decision-making capabilities, insufficient spatial geometric expression and real-time performance, and lack of continuous environmental updates and memory capabilities in related technologies.
[0005] The first aspect of this application provides a wheelchair navigation method, comprising the following steps: acquiring a current environment image, navigation instructions, and historical environment images of the wheelchair; extracting environmental structural features from the current environment image, determining the target environment in which the wheelchair is located based on the environmental structural features, the target environment including an indoor environment and an outdoor environment; if the target environment is an indoor environment, generating a bird's-eye view feature map based on the current environment image, and navigating the wheelchair in the indoor environment based on the bird's-eye view feature map and navigation instructions; if the target environment is an outdoor environment, predicting the three-dimensional point cloud, relative pose, and driving trajectory of the wheelchair based on the current environment image and historical environment images, and navigating the wheelchair in the outdoor environment based on the three-dimensional point cloud, relative pose, driving trajectory, and navigation instructions.
[0006] Optionally, in one embodiment of this application, determining the target environment where the wheelchair is located based on environmental structural features includes: extracting environmental structural features from the current environment image, wherein the environmental structural features include at least one of spatial constraint features, geometric continuity features, and semantic distribution features; determining the environment type of the current environment based on the environmental structural features; and determining the target environment where the wheelchair is located based on the environmental structural features.
[0007] Optionally, in one embodiment of this application, generating a bird's-eye view feature map based on the current environmental image includes: generating a three-dimensional point cloud set based on the current environmental image; projecting the three-dimensional point cloud set onto a bird's-eye view space centered on the wheelchair; extracting image features from the current environmental image; and fusing the image features and the projection results of the bird's-eye view space to obtain the bird's-eye view feature map.
[0008] Optionally, in one embodiment of this application, navigating a wheelchair in an indoor environment based on a bird's-eye view feature map and navigation instructions includes: inputting the bird's-eye view feature map and navigation instructions into a visual language inference model, outputting at least one of the wheelchair's target movement position and layout planning path through the visual language inference model; and navigating the wheelchair in an indoor environment based on at least one of the target movement position and layout planning path.
[0009] Optionally, in one embodiment of this application, predicting the 3D point cloud, relative pose, and driving trajectory of a wheelchair based on current environmental images and historical environmental images includes: generating visual information based on the current environmental image; acquiring pose information, trajectory information, and navigation command information; fusing the visual information, pose information, trajectory information, and navigation command information to generate fused features; extracting historical environmental structural features from historical environmental images; inputting the fused features and historical environmental structural features into a visual geometric inference model; outputting the visual features, pose features, and trajectory features of the wheelchair through the visual geometric inference model; and inputting the visual features, pose features, and trajectory features into a spatiotemporal inference model, which predicts the 3D point cloud, relative pose, and driving trajectory of the wheelchair.
[0010] Optionally, in one embodiment of this application, navigating a wheelchair in an outdoor environment based on a 3D point cloud, relative pose, driving trajectory, and navigation instructions includes: navigating a wheelchair in an outdoor environment based on a 3D point cloud, relative pose, driving trajectory, and navigation instructions, which includes: obtaining at least one of the wheelchair's target movement position and layout planning path based on the 3D point cloud, relative pose, driving trajectory, and navigation instructions; and navigating the wheelchair in an outdoor environment based on at least one of the target movement position and layout planning path.
[0011] Optionally, in one embodiment of this application, the expression for the three-dimensional point cloud of the wheelchair is:
[0012] in, The current moment; For visual modules; 3D point cloud for the wheelchair; For 3D point cloud prediction head; To process features using a 3D point cloud prediction head; Visual features; The expression for relative pose is:
[0013] in, For pose module; The relative position of the wheelchair; For pose prediction head; To process features using a pose prediction head; Positional features; The expression for the driving trajectory is:
[0014] in, For trajectory module; The trajectory of the wheelchair; For trajectory prediction head; To process features using a trajectory prediction head; For trajectory features.
[0015] A second aspect of this application provides a wheelchair navigation device, comprising: an acquisition module for acquiring a current environmental image, navigation instructions, and historical environmental images of the wheelchair; an extraction module for extracting environmental structural features from the current environmental image and determining the target environment in which the wheelchair is located based on the environmental structural features, the target environment including an indoor environment and an outdoor environment; a first navigation module for generating a bird's-eye view feature map based on the current environmental image if the target environment is an indoor environment, and navigating the wheelchair in the indoor environment based on the bird's-eye view feature map and navigation instructions; and a second navigation module for predicting the three-dimensional point cloud, relative pose, and driving trajectory of the wheelchair based on the current environmental image and historical environmental images if the target environment is an outdoor environment, and navigating the wheelchair in the outdoor environment based on the three-dimensional point cloud, relative pose, driving trajectory, and navigation instructions.
[0016] Optionally, in one embodiment of this application, determining the target environment where the wheelchair is located based on environmental structural features includes: extracting environmental structural features from the current environment image, wherein the environmental structural features include at least one of spatial constraint features, geometric continuity features, and semantic distribution features; determining the environment type of the current environment based on the environmental structural features; and determining the target environment where the wheelchair is located based on the environmental structural features.
[0017] Optionally, in one embodiment of this application, generating a bird's-eye view feature map based on the current environmental image includes: generating a three-dimensional point cloud set based on the current environmental image; projecting the three-dimensional point cloud set onto a bird's-eye view space centered on the wheelchair; extracting image features from the current environmental image; and fusing the image features and the projection results of the bird's-eye view space to obtain the bird's-eye view feature map.
[0018] Optionally, in one embodiment of this application, navigating a wheelchair in an indoor environment based on a bird's-eye view feature map and navigation instructions includes: inputting the bird's-eye view feature map and navigation instructions into a visual language inference model, outputting at least one of the wheelchair's target movement position and layout planning path through the visual language inference model; and navigating the wheelchair in an indoor environment based on at least one of the target movement position and layout planning path.
[0019] Optionally, in one embodiment of this application, visual information is generated based on the current environmental image; pose information, trajectory information, and navigation command information are acquired, and the visual information, pose information, trajectory information, and navigation command information are fused to generate fused features; historical environmental structure features of historical environmental images are extracted, and the fused features and historical environmental structure features are input into a visual geometric inference model, which outputs the visual features, pose features, and trajectory features of the wheelchair; the visual features, pose features, and trajectory features are input into a spatiotemporal inference model, which predicts the three-dimensional point cloud, relative pose, and driving trajectory of the wheelchair.
[0020] Optionally, in one embodiment of this application, navigating a wheelchair in an outdoor environment based on a 3D point cloud, relative pose, driving trajectory, and navigation instructions includes: obtaining at least one of the wheelchair's target movement position and layout planning path based on the 3D point cloud, relative pose, driving trajectory, and navigation instructions; and navigating the wheelchair in an outdoor environment based on at least one of the target movement position and layout planning path.
[0021] Optionally, in one embodiment of this application, the expression for the three-dimensional point cloud of the wheelchair is:
[0022] in, The current moment; For visual modules; 3D point cloud for the wheelchair; For 3D point cloud prediction head; To process features using a 3D point cloud prediction head; Visual features; The expression for relative pose is:
[0023] in, For pose module; The relative position of the wheelchair; For pose prediction head; To process features using a pose prediction head; Positional features; The expression for the driving trajectory is:
[0024] in, For trajectory module; The trajectory of the wheelchair; For trajectory prediction head; To process features using a trajectory prediction head; For trajectory features.
[0025] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the wheelchair navigation method described above.
[0026] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the steps of the wheelchair navigation method described above.
[0027] Therefore, this application has the following beneficial effects: First, the current environment image, navigation commands, and historical environment images of the wheelchair are acquired. Then, environmental structural features are extracted from the current environment image, and based on these features, the target environment of the wheelchair is determined to be either indoor or outdoor. This effectively distinguishes complex scene types and improves the accuracy and adaptability of environmental understanding. If the target environment is indoor, a bird's-eye view feature map is generated based on the current environment image. This map, combined with navigation commands, is used to navigate the wheelchair indoors, enhancing the spatial structure representation within the indoor environment and improving navigation stability and path planning accuracy. If the target environment is outdoor... The system predicts the wheelchair's 3D point cloud, relative pose, and driving trajectory based on current and historical environmental images. It then navigates the wheelchair in outdoor environments based on the 3D point cloud, relative pose, driving trajectory, and navigation commands. This enables continuous geometric modeling and dynamic trajectory prediction for complex outdoor scenes, improving the completeness of environmental perception and temporal modeling capabilities. Through the coordinated execution of indoor and outdoor navigation modes, it achieves adaptive navigation control for different environments. This solves problems related to the poor unified representation of perception, mapping, and decision-making capabilities, insufficient spatial geometric expression and real-time performance, and lack of continuous environmental updates and memory capabilities.
[0028] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0029] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1A flowchart of a wheelchair navigation method according to an embodiment of this application; Figure 2 This is a schematic diagram of a navigation method based on a visual geometric reasoning model according to an embodiment of this application; Figure 3 This is an example diagram of a wheelchair navigation device according to an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0030] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0031] The wheelchair navigation method, apparatus, device, and medium of this application are described below with reference to the accompanying drawings. Addressing the problems mentioned in the background art, this application provides a wheelchair navigation method. In this method, firstly, the current environment image, navigation commands, and historical environment images of the wheelchair are acquired. Secondly, environmental structural features are extracted from the current environment image, and based on these features, the target environment of the wheelchair is determined to be either an indoor or outdoor environment, thereby effectively distinguishing complex scene types and improving the accuracy and adaptability of environmental understanding. If the target environment is an indoor environment, a bird's-eye view feature map is generated based on the current environment image, and the bird's-eye view feature map is combined with the navigation commands to navigate the wheelchair in the indoor environment, thereby enhancing the spatial structure representation capability in the indoor environment and improving navigation stability. Qualitative and path planning accuracy; if the target environment is an outdoor environment, the system predicts the wheelchair's 3D point cloud, relative pose, and driving trajectory based on the current and historical environmental images. Navigation is then performed on the wheelchair in the outdoor environment based on the 3D point cloud, relative pose, driving trajectory, and navigation commands, thereby achieving continuous geometric modeling and dynamic trajectory prediction for complex outdoor scenes, improving the completeness of environmental perception and temporal modeling capabilities. Through the coordinated execution of indoor and outdoor navigation modes, adaptive navigation control for different environments is achieved, thus solving problems related to poor unified representation capabilities of perception, mapping, and decision-making, insufficient spatial geometric expression and real-time performance, and lack of continuous environmental updates and memory capabilities.
[0032] Specifically, Figure 1 This is a schematic flowchart of a wheelchair navigation method provided in an embodiment of this application.
[0033] like Figure 1 As shown, the wheelchair navigation method includes the following steps: In step S101, the current environmental image, navigation instructions, and historical environmental images of the wheelchair are acquired.
[0034] Among them, the current environment image is the visual observation image acquired by the wheelchair through the onboard image acquisition device at the current time step, which in this application represents real-time perception data used to characterize the spatial structure and semantic information of the current environment; the navigation command is the path planning command issued by the user to the wheelchair through voice input, text input or preset target point, which in this application represents task constraint information used to guide the wheelchair to perform target-oriented movement; the historical environment image is the image sequence continuously acquired and stored by the wheelchair within a preset time window before the current time step, which in this application represents a set of historical visual information used to describe the temporal changes of the environment and assist in the construction of spatial memory.
[0035] It is understood that this application obtains real-time perception information of the current environment, semantic constraint information of the task objective, and temporal memory information of the historical scene by uniformly acquiring and coordinating current environmental images, navigation instructions, and historical environmental images. The current environmental images are used to reflect the immediate spatial structure and semantic state of the wheelchair's location, the navigation instructions are used to provide clear target guidance and behavioral constraints, and the historical environmental images are used to supplement the temporal changes and spatial continuity information of the environment. Through the joint input of the three, it helps to improve the perception integrity of this application in complex dynamic environments.
[0036] In step S102, environmental structure features are extracted from the current environmental image, and the target environment where the wheelchair is located is determined based on the environmental structure features. The target environment includes indoor and outdoor environments.
[0037] In this application, environmental structural features represent structured descriptive information reflecting whether the environment has obvious geometric constraints, spatial continuity, and semantic layout differences; the target environment is an environmental category obtained by judging the scene type of the wheelchair based on environmental structural features, and in this application, it represents different categories used to distinguish between indoor confined structural spaces and outdoor open unstructured spaces; the indoor environment is a scene environment with clear artificial structural constraints, clear spatial boundaries, and relatively fixed functional areas, and in this application, it represents structured spaces such as corridors, rooms, and medical functional areas; the outdoor environment is a scene environment with open space characteristics, irregular geometric structure, and large environmental changes, and in this application, it represents unstructured spaces such as roads, ramps, and unstructured passage areas.
[0038] It is understood that by extracting environmental structural features from the current environmental image and distinguishing between indoor and outdoor environments based on these structural features, rapid type identification and structured classification of complex scenes are achieved, thus providing a clear basis for the selection of subsequent navigation strategies. By dividing the environment into two typical scenarios, indoor and outdoor, this application can adopt navigation strategies based on semantic understanding or navigation strategies based on geometric modeling for different environmental characteristics, thereby enhancing the adaptability and stable operation of this application in complex continuous scenes.
[0039] In one embodiment of this application, determining the target environment where the wheelchair is located based on environmental structural features includes: extracting environmental structural features from the current environment image, wherein the environmental structural features include at least one of spatial constraint features, geometric continuity features, and semantic distribution features; determining the environment type of the current environment based on the environmental structural features; and determining the target environment where the wheelchair is located based on the environmental structural features.
[0040] Among them, spatial constraint features are feature information used to characterize the boundaries of traversable space, the distribution of obstacles, and the degree of traversal restrictions in the environment. In this application, they represent the description of the constraint relationship between the wheelchair's degrees of freedom of movement and the traversable area. Geometric continuity features are feature information used to characterize the continuity and connectivity trends of spatial structures in the environment. In this application, they represent the description of whether the passage path is continuous and whether there are broken or abrupt structural changes in the space. Semantic distribution features are feature information used to characterize the spatial distribution patterns of different semantic category targets and functional areas in the environment. In this application, they represent the description of the spatial layout relationship of semantic elements such as corridors, rooms, roads, and obstacles.
[0041] Understandably, this application introduces spatial constraint features, geometric continuity features, and semantic distribution features to perform multi-dimensional environmental structure modeling of the current environmental image. On this basis, through an environment type discrimination mechanism based on multi-feature fusion, complex continuous scenes are effectively divided into indoor structured environments and outdoor unstructured environments, achieving adaptive environmental recognition without relying on pre-built maps or manual rules. Furthermore, by directly using environmental structure features as the discrimination criterion, a mapping from low-level visual information to high-level environmental semantic categories is achieved, which helps to improve the system's generalization ability in complex dynamic scenes.
[0042] This application continuously analyzes current visual observations and spatial representations during operation, extracting key information reflecting environmental structural characteristics, such as spatial constraint degree, geometric continuity, and semantic distribution features. Based on this, the application calls upon existing large language models (such as Qianwen) to discriminate the current environment and selects the corresponding navigation model for decision-making based on the discrimination results. When environmental characteristics change, a unified spatial representation and control interface enables a smooth transition between different navigation methods, avoiding control abrupt changes or path discontinuities caused by mode switching. For example, when environmental characteristics change, this application uses a unified spatial representation and control interface to achieve a smooth transition between different navigation methods. During environment switching, indoor and outdoor navigation methods can participate in navigation decision-making simultaneously within a preset time window, ensuring the effectiveness of indoor navigation methods during the transition to outdoor environments, thereby avoiding control abrupt changes or path discontinuities caused by mode switching.
[0043] In step S103, if the target environment is an indoor environment, a bird's-eye view feature map is generated based on the current environment image, and the wheelchair is navigated in the indoor environment based on the bird's-eye view feature map and navigation instructions.
[0044] Among them, the bird's-eye view feature map is a two-dimensional or pseudo-three-dimensional spatial representation map generated by mapping the current environmental image to the overhead space centered on the wheelchair through three-dimensional reconstruction and spatial projection, and integrating visual semantic features. In this application, it represents a navigation memory representation used to uniformly express the indoor environmental spatial structure and passable areas.
[0045] Understandably, this application, after determining that the target environment is an indoor environment, maps the current environment image to generate a bird's-eye view feature map, and combines navigation instructions to perform path planning and decision-making on a unified top-down spatial representation, significantly improving the spatial understandability and path planning feasibility in indoor environments. The bird's-eye view feature map can compress complex 3D indoor scenes into a clearly structured layout representation, effectively enhancing the expressive ability of spatial topological relationships and passable areas. At the same time, fusing navigation instructions with the bird's-eye view feature map helps improve the alignment ability between semantic targets and spatial structures.
[0046] In one embodiment of this application, generating a bird's-eye view feature map based on a current environmental image includes: generating a three-dimensional point cloud set based on the current environmental image; projecting the three-dimensional point cloud set onto a bird's-eye view space centered on the wheelchair; extracting image features from the current environmental image; and fusing the image features and the projection results of the bird's-eye view space to obtain the bird's-eye view feature map.
[0047] In this application, the three-dimensional point cloud set represents a three-dimensional representation of the relationship between the geometric structure and spatial distribution of the environment; the bird's-eye view space is a top-down coordinate reference space constructed with the current position of the wheelchair as the center; and projection is the process of mapping the three-dimensional point cloud set to the bird's-eye view space according to a preset coordinate transformation relationship.
[0048] Understandably, this application first generates a 3D point cloud set from the current environmental image, then projects the 3D point cloud set onto a bird's-eye view space centered on the wheelchair, and further fuses image features with the projection results to generate a bird's-eye view feature map. This can preserve the 3D geometric structure information of the environment while introducing semantic and appearance information from the image features, thereby enhancing the ability to express the distribution of obstacles, spatial connectivity, and passable areas in complex indoor scenes. At the same time, through the unified mapping of the bird's-eye view space, information that was originally scattered in the image view is fused under a unified reference coordinate system, which helps to reduce the uncertainty caused by changes in viewpoint and improve the consistency and stability of spatial understanding.
[0049] Specifically, during the navigation process, the embodiments of this application specify the time step. The internal receiver displays current environmental images (i.e., RGB image observations) from the wheelchair, along with navigation commands. (Including text descriptions or target locations). First, the current environment image is processed using online synchronous localization and mapping methods (including but not limited to MASt3R-based synchronous localization and mapping methods) to reconstruct the 3D point cloud set of the current scene in real time:
[0050] in, This represents the reconstruction process using the simultaneous localization and mapping (SLAM) method. The current 3D point cloud set; This is the initial environment image; Image of the current environment; This is the sequence of images from the initial image to the current image.
[0051] Subsequently, the 3D point cloud set is uniformly projected onto a bird's-eye view space centered on the wheelchair. The projection process can be represented as follows:
[0052] in, The projection operator represents the projection from three-dimensional space to a bird's-eye view plane; Current wheelchair position; This is a bird's-eye view obtained through projection.
[0053] At the same time, existing visual backbone networks (including but not limited to GroundingDINO) are used to extract image features of the current environment image. Then, it is aligned and projected onto the bird's-eye view coordinate system to construct a bird's-eye view feature map that integrates geometric structure and visual semantic information:
[0054] in, This represents the fused bird's-eye view feature map; Indicates the feature fusion function; This indicates the fusion feature.
[0055] This bird's-eye view feature map serves as spatial memory during navigation, and can be continuously updated and accumulate historical observation information over time.
[0056] In one embodiment of this application, navigating a wheelchair in an indoor environment based on a bird's-eye view feature map and navigation instructions includes: inputting the bird's-eye view feature map and navigation instructions into a visual language inference model, outputting at least one of the wheelchair's target movement position and layout planning path through the visual language inference model; and navigating the wheelchair in an indoor environment based on at least one of the target movement position and layout planning path.
[0057] Among them, the visual language reasoning model is a model used to fuse visual information and language instructions and perform cross-modal semantic reasoning; the target movement position is the spatial position that the wheelchair is expected to reach next, predicted by the visual language reasoning model based on the bird's-eye view feature map and navigation instructions, and in this application, it represents the target coordinate information used to guide the local motion control of the wheelchair; the layout planning path is a continuous passage path from the current position to the target position generated by the visual language reasoning model based on the environmental spatial structure and navigation instructions, and in this application, it represents the path planning result used to describe the overall movement trajectory of the wheelchair in the indoor environment.
[0058] Understandably, this application achieves cross-modal fusion and joint reasoning between visual spatial representation and semantic instructions by uniformly inputting bird's-eye view feature maps and navigation instructions into a visual language reasoning model. It simultaneously understands environmental structural information and task objective constraints in a unified representation space, thereby improving decision consistency and semantic alignment capabilities in indoor environments. By directly outputting the target movement position or layout planning path from the model, the navigation process is transformed from the traditional phased perception-planning mode to an end-to-end reasoning generation mode. At the same time, the bird's-eye view feature map provides clearly structured spatial layout information, enabling this application to more effectively capture the distribution of obstacles and the relationship between passable areas, thereby enhancing obstacle avoidance capabilities and path feasibility judgment capabilities in complex indoor environments.
[0059] During the decision-making stage, this application encodes navigation instructions based on existing visual language models (including but not limited to the VILA open-source architecture) and performs cross-modal reasoning in the bird's-eye view feature space to directly predict the next target movement position or local path. The decision-making process can be represented as follows:
[0060] in, Represents a visual language reasoning model; Indicates the current navigation action or target location; Indicates navigation instructions.
[0061] Furthermore, this application performs geometric relationship calculations based on the predicted target movement position and the current posture of the wheelchair to generate corresponding movement direction and speed control quantities.
[0062] Because the decision-making process takes place in a unified bird's-eye view space, the visual language model can simultaneously utilize historical geometric information and current semantic information to achieve a stable understanding and planning of complex interior structures (such as corridors, room connections, etc.), thereby improving the robustness and continuity of navigation.
[0063] In step S104, if the target environment is an outdoor environment, the three-dimensional point cloud, relative pose, and driving trajectory of the wheelchair are predicted based on the current environment image and historical environment images. The wheelchair is then navigated in the outdoor environment based on the three-dimensional point cloud, relative pose, driving trajectory, and navigation instructions.
[0064] Among them, the historical environment image is a sequence of images continuously acquired and stored by the wheelchair within a preset time window before the current time step, which in this application represents a set of information used to describe the temporal changes of the environment and assist in constructing spatial memory; the relative pose is the spatial position and posture change information of the wheelchair at the current time step relative to the previous time step or the local coordinate system, which in this application represents geometric transformation parameters used to describe the changes in the wheelchair's motion state; the driving trajectory is a spatial sequence of the continuous movement positions of the wheelchair within the preset time window, which in this application represents trajectory information used to characterize the wheelchair's historical movement path and future movement trend.
[0065] Understandably, when the target environment is an outdoor environment, this application introduces a joint modeling approach combining current and historical environmental images to achieve the synergistic utilization of temporal and spatial geometric information of the environment. This enables a more stable environmental understanding in dynamically changing unstructured scenarios. By predicting 3D point clouds, the geometric perception accuracy of complex outdoor terrain (such as ramps, curbs, and obstacles) can be improved. By introducing relative pose estimation, the continuous tracking capability of the wheelchair's own motion state can be enhanced, thereby improving positioning stability. By generating driving trajectories, short-term motion trends can be further characterized, providing prior constraints for subsequent path planning. At the same time, by fusing 3D point clouds, relative poses, driving trajectories, and navigation commands, this application can make navigation decisions within a unified geometric and semantic framework, enhancing obstacle avoidance and safe driving capabilities in complex open scenarios.
[0066] In one embodiment of this application, predicting the 3D point cloud, relative pose, and driving trajectory of a wheelchair based on current environmental images and historical environmental images includes: generating visual information based on the current environmental image; acquiring pose information, trajectory information, and navigation command information; fusing the visual information, pose information, trajectory information, and navigation command information to generate fused features; extracting historical environmental structural features from historical environmental images; inputting the fused features and historical environmental structural features into a visual geometric inference model; and outputting the 3D point cloud, relative pose, and driving trajectory of the wheelchair through the visual geometric inference model.
[0067] Among them, visual information is a high-dimensional feature representation extracted from the current environmental image through a visual backbone network to characterize the appearance features and semantic information of the environment, and in this application, it represents the structured encoding of the visual perception result of the environment at the current moment; pose information is data used to describe the spatial position and posture of the wheelchair relative to the global or local coordinate system at the current moment or historical moment, and in this application, it represents the geometric state information used to characterize the changes in the wheelchair's motion state; trajectory information is the temporal sequence information composed of the continuous motion positions of the wheelchair within a preset time window, and in this application, it represents the temporal motion information used to describe the historical motion path of the wheelchair and its changing trend; navigation command information is user input The input semantic information used to describe the target location or path constraints represents, in this application, the control constraints used to guide the wheelchair to perform target-oriented movement; the fusion feature is a unified feature representation obtained by multimodal fusion of visual information, pose information, trajectory information and navigation command information; the historical environment structure feature is a set of features extracted from historical environment images to characterize the temporal changes and spatial structural stability of the environment, and in this application, it represents historical perception information used to assist in restoring the geometric consistency and spatial continuity of the environment; the visual geometric reasoning model is a model used to jointly model the multimodal fusion feature and the historical environment structure feature, and to infer and generate the three-dimensional geometry and motion state of the environment.
[0068] Understandably, this application jointly models current and historical environmental images and incorporates pose, trajectory, and navigation command information for multimodal fusion to construct a unified fusion feature representation. This enables the application to simultaneously characterize environmental appearance, motion state, and task constraints, thereby enhancing the overall understanding of complex dynamic outdoor scenes. Furthermore, by introducing historical environmental structural features, it supplements the constraints on temporal consistency and spatial structural continuity, helping to enhance stable perception under conditions of viewpoint changes and environmental occlusion. Moreover, through a visual geometric inference model, it jointly infers the fusion features and historical structural features, achieving a unified mapping from multimodal semantic information to 3D point clouds, relative pose, and driving trajectory. This allows the application to simultaneously obtain environmental geometric structure, motion state changes, and short-term trajectory trends, significantly improving positioning accuracy, path prediction capability, and motion continuity in outdoor environments. Therefore, this step effectively enhances the real-time performance, robustness, and safe navigation capabilities of wheelchairs in unstructured open environments.
[0069] The visual geometric inference model architecture used in this application includes a multi-view feature extraction module, a spatiotemporal cross-attention layer, and a multi-task prediction head, achieving nonlinear mapping through the ReLU activation function. It is trained using a large-scale real-world navigation and driving dataset, employing a joint loss function of L1 smoothing loss and geometric consistency loss, and utilizing the Adam optimizer for parameter updates. Key hyperparameters include a temporal window size of 4 to 8 frames, 6 attention layers, and a feature dimension of 256. Experimental data comparison shows that this scheme improves the 3D reconstruction accuracy and trajectory prediction accuracy in dynamic outdoor scenes by approximately 15% and 20%, respectively, compared to the baseline model, significantly enhancing the system's spatial geometric representation and real-time decision-making capabilities.
[0070] Specifically, at the current time step This application receives current environment images from the wheelchair (i.e., multi-view image input). ), navigation instructions and including the past Historical context structural features of frames The historical environmental structural features are extracted from historical environmental images and jointly predicted to form a 3D point cloud in the current wheelchair local coordinate system. The relative pose of the wheelchair compared to the previous frame And the smooth driving trajectory in the future To reduce computational overhead, this application employs strategies including but not limited to sliding window streaming, performing geometric modeling only in the local coordinate system of the current frame and reusing features through a fixed-length history cache, avoiding repeated computation of the complete history sequence, thereby reducing the overall computational complexity to constant level.
[0071] Specifically, the current environment image is first encoded into visual information using the pre-trained visual base model DINOv3:
[0072] in, For the current moment Visual features; .
[0073] The visual foundation model used in this application employs a self-supervised learning-based teacher-student network architecture, including a multi-layer Transformer encoder and a GELU activation function. Its training data covers a massive unlabeled image set, preprocessed through random cropping and color dithering. During training, the AdamW optimizer is used, and an improved self-distillation loss function is introduced to optimize high-level semantic feature representation. Key hyperparameters include approximately 7 billion model parameters, a base learning rate set to 1e-4, and a cosine annealing decay strategy. Experiments show that this model significantly outperforms traditional visual models in dense feature extraction and zero-shot generalization, effectively improving the accuracy of image representation in complex environments.
[0074] Next, this application adds a set of learnable pose information to the visual information of each viewpoint. and trajectory information and connect to the corresponding navigation instruction information. :
[0075] The navigation instruction information is obtained by encoding navigation instructions through a multi-layer fully connected network.
[0076] Then the fusion feature With historical environmental structural characteristics Input as follows Figure 2 The visual geometric reasoning model shown Joint modeling is performed to obtain the corresponding visual features. Posture characteristics Trajectory features :
[0077] in, The size of the modeling time window; The environmental structure characteristics at time tW; The environmental structural features at time t-1; The historical environmental structural characteristics from time tW to time t-1; This represents a visual geometric reasoning model.
[0078] This application achieves multi-source information fusion through in-view attention, cross-view spatial attention, and temporal causal attention, and utilizes relative temporal position encoding to improve the reuse efficiency of historical features. Finally, a multi-task prediction head outputs the 3D point cloud of the wheelchair, its relative pose, and the future trajectory of the wheelchair. Based on the predicted target movement position and the current pose of the wheelchair, geometric relationship calculations are performed to generate the corresponding movement direction and speed control variables.
[0079] In one embodiment of this application, navigating a wheelchair in an outdoor environment based on a 3D point cloud, relative pose, driving trajectory, and navigation instructions includes: inputting the 3D point cloud, relative pose, driving trajectory, and navigation instructions into a visual geometric inference model, outputting at least one of the wheelchair's target movement position and layout planning path through the visual geometric inference model; and navigating the wheelchair in an outdoor environment based on at least one of the target movement position and layout planning path.
[0080] Understandably, this application achieves deep integration of environmental geometric information, motion state information, and semantic task information by uniformly inputting 3D point clouds, relative poses, driving trajectories, and navigation commands into the visual geometric inference model. This enables the application to simultaneously perform spatial understanding and path decision-making within a unified inference framework, thereby improving decision consistency and overall coordination in complex outdoor environments. By introducing relative pose and driving trajectory information, the ability to model the wheelchair's own motion state and historical motion trends can be enhanced, thereby improving positioning stability and trajectory prediction accuracy. At the same time, 3D point clouds provide a detailed representation of obstacle distribution and terrain structure in unstructured outdoor scenes, which helps improve the feasibility assessment of path planning and obstacle avoidance capabilities.
[0081] In one embodiment of this application, the expression for the three-dimensional point cloud of the wheelchair is:
[0082] in, The current moment; For visual modules; 3D point cloud for the wheelchair; For 3D point cloud prediction head; To process features using a 3D point cloud prediction head; Visual features; The expression for relative pose is:
[0083] in, For pose module; The relative position of the wheelchair; For pose prediction head; To process features using a pose prediction head; Positional features; The expression for the driving trajectory is:
[0084] in, For trajectory module; The trajectory of the wheelchair; For trajectory prediction head; To process features using a trajectory prediction head; For trajectory features.
[0085] Understandably, this application uses a unified mathematical expression structure for the generation process of three-dimensional spatial structure, motion state and trajectory information, which helps to enhance the structural consistency and interpretability between the output branches of the model; at the same time, by using visual features, pose features and trajectory features as input representations, a clear correspondence is established between high-dimensional perception features and specific physical quantities (point cloud, pose and trajectory), so that different functional branches can be modeled independently while maintaining a unified interface.
[0086] According to the wheelchair navigation method proposed in this application, the current environment image, navigation commands, and historical environment images of the wheelchair are first acquired. Then, environmental structure features are extracted from the current environment image, and the target environment of the wheelchair is determined to be either indoor or outdoor based on these features. This effectively distinguishes between complex scene types and improves the accuracy and adaptability of environmental understanding. If the target environment is indoor, a bird's-eye view feature map is generated based on the current environment image. This bird's-eye view feature map is then combined with the navigation commands to navigate the wheelchair within the indoor environment, thereby enhancing the spatial structure representation capability within the indoor environment and improving the stability and path planning accuracy of the navigation. If the target environment is an outdoor environment, the system predicts the wheelchair's 3D point cloud, relative pose, and driving trajectory based on the current and historical environmental images. Navigation is then performed on the wheelchair in the outdoor environment based on the 3D point cloud, relative pose, driving trajectory, and navigation commands. This enables continuous geometric modeling and dynamic trajectory prediction for complex outdoor scenes, improving the completeness of environmental perception and temporal modeling capabilities. Through the coordinated execution of indoor and outdoor navigation modes, adaptive navigation control for different environments is achieved, thus solving problems related to poor unified representation capabilities of perception, mapping, and decision-making, insufficient spatial geometric expression and real-time performance, and a lack of continuous environmental updates and memory capabilities.
[0087] Next, the wheelchair navigation device according to an embodiment of this application is described with reference to the accompanying drawings.
[0088] Figure 3 This is a block diagram of a wheelchair navigation device according to an embodiment of this application.
[0089] like Figure 3As shown, the wheelchair navigation device 10 includes: an acquisition module 100, an extraction module 200, a first navigation module 300, and a second navigation module 400.
[0090] The system includes: an acquisition module 100 for acquiring the current environment image, navigation instructions, and historical environment images of the wheelchair; an extraction module 200 for extracting environmental structural features from the current environment image and determining the target environment of the wheelchair based on these features, including both indoor and outdoor environments; a first navigation module 300 for generating a bird's-eye view feature map based on the current environment image if the target environment is indoor, and navigating the wheelchair in the indoor environment based on the bird's-eye view feature map and navigation instructions; and a second navigation module 400 for predicting the wheelchair's 3D point cloud, relative pose, and driving trajectory based on the current and historical environment images if the target environment is outdoor, and navigating the wheelchair in the outdoor environment based on the 3D point cloud, relative pose, driving trajectory, and navigation instructions.
[0091] In one embodiment of this application, the extraction module 200 is further configured to extract environmental structural features from the current environmental image, wherein the environmental structural features include at least one of spatial constraint features, geometric continuity features, and semantic distribution features; to determine the environmental type of the current environment based on the environmental structural features; and to determine the target environment in which the wheelchair is located based on the environmental structural features.
[0092] In one embodiment of this application, the first navigation module 300 is further configured to generate a three-dimensional point cloud set based on the current environmental image; project the three-dimensional point cloud set onto a bird's-eye view space centered on the wheelchair; extract image features from the current environmental image; and fuse the image features and the projection results of the bird's-eye view space to obtain a bird's-eye view feature map.
[0093] In one embodiment of this application, the first navigation module 300 is further configured to input a bird's-eye view feature map and navigation instructions into a visual language reasoning model, and output at least one of the target movement position and layout planning path of the wheelchair through the visual language reasoning model; and navigate the wheelchair in an indoor environment based on at least one of the target movement position and layout planning path.
[0094] In one embodiment of this application, the second navigation module 400 is further configured to generate visual information based on the current environmental image; acquire pose information, trajectory information, and navigation command information; fuse the visual information, pose information, trajectory information, and navigation command information to generate fused features; extract historical environmental structural features from historical environmental images; input the fused features and historical environmental structural features into a visual geometric inference model; output the visual features, pose features, and trajectory features of the wheelchair through the visual geometric inference model; input the visual features, pose features, and trajectory features into a spatiotemporal inference model; and predict the three-dimensional point cloud, relative pose, and driving trajectory of the wheelchair through the spatiotemporal inference model.
[0095] In one embodiment of this application, the second navigation module 400 is further configured to obtain at least one of the target movement position and layout planning path of the wheelchair based on the three-dimensional point cloud, relative pose, driving trajectory and navigation instructions; and to navigate the wheelchair in an outdoor environment based on at least one of the target movement position and layout planning path.
[0096] In one embodiment of this application, the expression for the three-dimensional point cloud of the wheelchair is:
[0097] in, The current moment; For visual modules; 3D point cloud for the wheelchair; For 3D point cloud prediction head; To process features using a 3D point cloud prediction head; Visual features; The expression for relative pose is:
[0098] in, For pose module; The relative position of the wheelchair; For pose prediction head; To process features using a pose prediction head; Positional features; The expression for the driving trajectory is:
[0099] in, For trajectory module; The trajectory of the wheelchair; For trajectory prediction head; To process features using a trajectory prediction head; For trajectory features.
[0100] It should be noted that the foregoing explanation of the wheelchair navigation method embodiment also applies to the wheelchair navigation device of this embodiment, and will not be repeated here.
[0101] According to the wheelchair navigation device proposed in this application, the device first acquires the current environment image, navigation commands, and historical environment images of the wheelchair. Then, it extracts environmental structural features from the current environment image and determines whether the target environment of the wheelchair is indoor or outdoor based on these features. This effectively distinguishes between complex scene types and improves the accuracy and adaptability of environmental understanding. If the target environment is indoor, a bird's-eye view feature map is generated based on the current environment image. The bird's-eye view feature map is then combined with the navigation commands to navigate the wheelchair within the indoor environment, thereby enhancing the spatial structure representation capability within the indoor environment and improving the stability and path planning accuracy of the navigation. If the target environment is an outdoor environment, the system predicts the wheelchair's 3D point cloud, relative pose, and driving trajectory based on the current and historical environmental images. Navigation is then performed on the wheelchair in the outdoor environment based on the 3D point cloud, relative pose, driving trajectory, and navigation commands. This enables continuous geometric modeling and dynamic trajectory prediction for complex outdoor scenes, improving the completeness of environmental perception and temporal modeling capabilities. Through the coordinated execution of indoor and outdoor navigation modes, adaptive navigation control for different environments is achieved, thus solving problems related to poor unified representation capabilities of perception, mapping, and decision-making, insufficient spatial geometric expression and real-time performance, and a lack of continuous environmental updates and memory capabilities.
[0102] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include: The memory 401, the processor 402, and the computer program stored on the memory 401 and capable of running on the processor 402.
[0103] When the processor 402 executes the program, it implements the wheelchair navigation method provided in the above embodiments.
[0104] Furthermore, electronic devices also include: Communication interface 403 is used for communication between memory 401 and processor 402.
[0105] The memory 401 is used to store computer programs that can run on the processor 402.
[0106] The memory 401 may include high-speed RAM (Random Access Memory) memory, and may also include non-volatile memory, such as at least one disk storage.
[0107] If the memory 401, processor 402, and communication interface 403 are implemented independently, then the communication interface 403, memory 401, and processor 402 can be interconnected via a bus to complete communication between them. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0108] Optionally, in a specific implementation, if the memory 401, processor 402, and communication interface 403 are integrated on a single chip, then the memory 401, processor 402, and communication interface 403 can communicate with each other through an internal interface.
[0109] Processor 402 may be a CPU (Central Processing Unit), an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of this application.
[0110] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the wheelchair navigation method described above.
[0111] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0112] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0113] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0114] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any of the following techniques known in the art, or a combination thereof: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (FPGAs), field-programmable gate arrays (FPGAs), etc.
[0115] Those skilled in the art will understand that all or part of the steps of the methods implementing the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0116] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A wheelchair navigation method, characterized in that, Includes the following steps: Acquire the current environmental image, navigation instructions, and historical environmental image of the wheelchair; Extract environmental structural features from the current environment image, and determine the target environment where the wheelchair is located based on the environmental structural features. The target environment includes indoor and outdoor environments. If the target environment is the indoor environment, a bird's-eye view feature map is generated based on the current environment image, and the wheelchair is navigated in the indoor environment based on the bird's-eye view feature map and the navigation instructions; If the target environment is the outdoor environment, then the three-dimensional point cloud, relative pose, and driving trajectory of the wheelchair are predicted based on the current environment image and the historical environment image, and the wheelchair is navigated in the outdoor environment based on the three-dimensional point cloud, the relative pose, the driving trajectory, and the navigation instructions.
2. The wheelchair navigation method according to claim 1, characterized in that, Determining the target environment where the wheelchair is located based on the environmental structural features includes: Extract environmental structural features from the current environmental image, wherein the environmental structural features include at least one of spatial constraint features, geometric continuity features, and semantic distribution features; Based on the aforementioned environmental structural characteristics, the current environment type is determined; The target environment in which the wheelchair is located is determined based on the environmental structural characteristics.
3. The wheelchair navigation method according to claim 1, characterized in that, The step of generating a bird's-eye view feature map based on the current environment image includes: Generate a 3D point cloud set based on the current environment image; The three-dimensional point cloud set is projected onto a bird's-eye view space centered on the wheelchair; The image features of the current environment image are extracted, and the image features and the projection results of the bird's-eye view space are fused to obtain the bird's-eye view feature map.
4. The wheelchair navigation method according to claim 1 or 3, characterized in that, The navigation of the wheelchair in the indoor environment based on the bird's-eye view feature map and the navigation instructions includes: The bird's-eye view feature map and the navigation instructions are input into the visual language reasoning model, and the visual language reasoning model outputs at least one of the target movement position and layout planning path of the wheelchair. The wheelchair is navigated in the indoor environment based on at least one of the target movement location and layout planning path.
5. The wheelchair navigation method according to claim 1, characterized in that, The step of predicting the 3D point cloud, relative pose, and driving trajectory of the wheelchair based on the current environmental image and the historical environmental image includes: Generate visual information based on the current environment image; Acquire pose information, trajectory information, and navigation command information, and fuse the visual information, pose information, trajectory information, and navigation command information to generate fused features; Extract the historical environment structure features from the historical environment image, input the fused features and the historical environment structure features into the visual geometric inference model, and output the visual features, pose features and trajectory features of the wheelchair through the visual geometric inference model; The visual features, pose features, and trajectory features are input into a spatiotemporal inference model, which predicts the three-dimensional point cloud, relative pose, and driving trajectory of the wheelchair.
6. The wheelchair navigation method according to claim 1, characterized in that, The method of navigating the wheelchair in the outdoor environment based on the 3D point cloud, the relative pose, the driving trajectory, and the navigation instructions includes: At least one of the target movement position and layout planning path of the wheelchair is obtained based on the three-dimensional point cloud, the relative pose, the driving trajectory and the navigation command; The wheelchair is navigated in the outdoor environment based on at least one of the target movement location and layout planning path.
7. The wheelchair navigation method according to claim 1, characterized in that, The expression for the three-dimensional point cloud of the wheelchair is: in, The current moment; For visual modules; 3D point cloud for the wheelchair; For 3D point cloud prediction head; To process features using a 3D point cloud prediction head; Visual features; The expression for the relative pose is: in, For pose module; The relative position of the wheelchair; For pose prediction head; To process features using a pose prediction head; Positional features; The expression for the driving trajectory is: in, For trajectory module; The trajectory of the wheelchair; For trajectory prediction head; To process features using a trajectory prediction head; For trajectory features.
8. A wheelchair navigation device, characterized in that, Includes the following steps: The acquisition module is used to acquire the current environmental image, navigation instructions, and historical environmental images of the wheelchair; The extraction module is used to extract environmental structure features from the current environment image and determine the target environment where the wheelchair is located based on the environmental structure features. The target environment includes indoor and outdoor environments. The first navigation module is used to generate a bird's-eye view feature map based on the current environment image if the target environment is the indoor environment, and to navigate the wheelchair in the indoor environment based on the bird's-eye view feature map and the navigation instructions; The second navigation module is used to predict the three-dimensional point cloud, relative pose, and driving trajectory of the wheelchair based on the current environment image and the historical environment image if the target environment is the outdoor environment, and to navigate the wheelchair in the outdoor environment based on the three-dimensional point cloud, the relative pose, the driving trajectory, and the navigation instructions.
9. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the steps of the wheelchair navigation method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the wheelchair navigation method as described in any one of claims 1-7.