Human-machine collaborative navigation system and method based on visual-tactile-linguistic-action model

The human-machine collaborative navigation system based on visual-tactile-language-motion models solves the problem of insufficient human-machine motion control in the navigation process of guide robots, realizes dynamic synchronization and safe collaborative navigation between guide robots and blind people, and improves the safety and interpretability of the navigation process.

CN122108167AActive Publication Date: 2026-05-29TONGJI UNIV +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TONGJI UNIV
Filing Date
2026-04-30
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing guide robots for the blind suffer from insufficient motion control under human-machine coupling systems during navigation, making it impossible to adjust the motion state in real time, leading to motion conflicts and safety hazards. Furthermore, the lack of two-way interaction and the information gap between perception and control reduce the trust of blind people.

Method used

A human-machine collaborative navigation system based on a visual-tactile-language-action model is adopted. Through a multimodal perception alignment module, a heterogeneous volume perception inference module, and a hierarchical asynchronous control module, human-machine collaborative motion control is achieved. This includes tactile feature extraction, heterogeneous volume perception, dual-stream parallel output, and hierarchical asynchronous control, providing semantic feedback and action commands, and optimizing navigation path planning and robot motion control.

Benefits of technology

It improves the compliance and response accuracy of human-computer interaction control during guide navigation, actively identifies and avoids asymmetric spatial obstacles, enhances the safety and interpretability of guide navigation control, realizes dynamic synchronization between the robot and the blind person, and ensures the smoothness and consistency of the navigation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122108167A_ABST
    Figure CN122108167A_ABST
Patent Text Reader

Abstract

The present application relates to the field of robot body intelligence and human-computer interaction technology, and particularly relates to a human-computer collaborative navigation system and method based on a visual-tactile-linguistic-motion model, the system being a human-computer collaborative motion closed-loop control system for a blind guiding mobile robot, the human-computer collaborative navigation system comprising: a multi-modal perception alignment module configured to provide human-computer interaction intention input for robot motion control; a heterogeneous volume perception inference module configured to provide an obstacle avoidance decision basis for robot motion control; a double-flow parallel output module configured with an action instruction output end and a semantic feedback output end; a hierarchical asynchronous control module, which is a core motion control execution module of the system, configured to maintain human-robot dynamics synchronization based on a high-level planning unit and a low-level collaborative motion unit. The present application can realize human-robot collaborative closed-loop motion control of a blind guiding robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of robot embodied intelligence and human-computer interaction technology, and particularly to the fields of autonomous motion control of mobile robots and human-machine collaborative dynamic closed-loop control of guide robots. Specifically, it relates to a human-machine collaborative navigation system and method based on a visual-tactile-language-action model. Background Technology

[0002] In existing technologies, most solutions for guide robots focus on pure navigation and measurement technologies such as environmental spatial positioning, path mapping, and position parameter calculation. Their core improvements are only to enhance the accuracy and precision of position positioning, without optimizing the core of guide robot tasks—the closed-loop control of robot motion within the human-machine coupling system. This results in guide robots being unable to adjust their motion state in real time according to the user's interaction intentions, making it difficult to achieve human-machine dynamic synchronization and flexible collaborative motion, and posing serious human-machine motion conflicts and safety hazards.

[0003] With the integration of large language models (LLMs) and embodied AI, vision-language-action (VLA) models (such as RT-2, NaVILA, etc.) have given robots the ability to understand natural language and plan long-term tasks.

[0004] However, the existing VLA paradigm suffers from a serious “embodied cognitive bias” when applied to assisting visually impaired individuals (i.e., robotic guide dogs), which treats the robot as an isolated solitary actor and ignores the fact that the guide dog task is essentially a human-machine coupled system.

[0005] The existing technology has three main drawbacks: 1. Safety hazards in motion control caused by geometric asymmetry: Quadruped robots are typically short (about 0.5m), while humans are taller (about 1.7m). Traditional navigation algorithms or isolated VLA models often only consider the robot's own mobility, planning paths solely based on the robot's dimensions, without optimizing the robot's motion control strategy for the joint volume constraints of the human-robot coupling system. Overhanging obstacles that robots can easily pass through (such as low branches, construction scaffolding, and clotheslines) pose a fatal head collision risk to accompanying blind individuals. 2. Motion control conflicts caused by one-way interaction: Traditional guide robots for the blind only execute one-way "human-to-machine" commands or only output motion control signals, lacking two-way communication. When a blind person tightens the tether due to fear or hearing unusual sounds, if the robot relies solely on visual navigation, it may interpret this pulling force as terrain resistance and increase the output torque, leading to a "human-machine tug-of-war" phenomenon, which seriously affects the experience and may even cause danger. 3. The information gap between perception and control leads to a lack of trust in control: Robots are not only the legs of the blind, but also their eyes. Existing end-to-end models typically only output joint velocities, making them a "black box." Blind people cannot know why the robot stops (whether it's due to obstacles, traffic lights, or other reasons), and this information asymmetry greatly reduces their trust in robots. Summary of the Invention

[0006] In view of the shortcomings of the prior art described above, the purpose of this invention is to provide a human-machine collaborative navigation system and method based on a visual-tactile-language-action model. The core solution is to address the problem of insufficient human-machine collaborative motion control capability of existing guide robots. It can understand the tactile feedback reflected by the guide device, and has the ability to perceive heterogeneous human-machine volume and provide semantic feedback, ultimately realizing human-machine collaborative closed-loop motion control and safe collaborative navigation of guide robots.

[0007] The first aspect of this invention provides a human-machine collaborative navigation system based on a visual-tactile-language-action model. The system is a closed-loop human-machine collaborative motion control system for a guide robot for the visually impaired, comprising: A multimodal perception alignment module is configured to collect environmental visual information and real-time human-computer interaction force signals from guide devices for the visually impaired, providing human-computer interaction intention input for robot motion control; the multimodal perception alignment module includes at least a tactile feature extraction unit, which is used to extract temporal features from the interaction force signals and convert the extracted temporal features into tactile tokens aligned with the semantic space of a pre-set large language model through a feature mapping layer, so as to represent the physical interaction intention of the accompanying person; A heterogeneous volume perception and reasoning module, integrated into the pre-set large language model, is configured to construct a joint geometric envelope of the human-machine coupling system during path planning, providing obstacle avoidance decision-making basis for robot safe motion control. The heterogeneous volume perception and reasoning module extracts the geometric attributes and ground clearance of environmental obstacles based on the environmental visual information, and compares the geometric attributes and ground clearance with the accompanying personnel size parameters in the system prompts in three-dimensional space. This drives the pre-set large language model to perform chain-like reasoning with heterogeneous volume constraints when generating navigation decisions, in order to identify and / or avoid asymmetric spatial obstacles. A dual-stream parallel output module is connected to the heterogeneous volume perception inference module and is configured with an action command output terminal and a semantic feedback output terminal. The action command output terminal is at least used to output a navigation landmark sequence containing spatial coordinates. The navigation landmark sequence is used to drive the robot to perform trajectory tracking and closed-loop motion control. The semantic feedback output terminal is at least used to generate semantic voice feedback for accompanying personnel in real time based on the current environmental state and inference results. The hierarchical asynchronous control module includes a high-level planning unit operating at a first frequency and a low-level collaborative motion unit operating at a second frequency, wherein the second frequency is higher than the first frequency. The high-level planning unit runs the preset large language model to generate the navigation sign sequence, and the low-level collaborative motion unit receives the navigation sign sequence and outputs motion control commands for the robot joints based on real-time perceived interactive force signals, dynamically adjusting the robot's motion state to maintain human-machine dynamic synchronization between the robot and the accompanying person, thereby realizing human-machine collaborative closed-loop motion control of the guide robot.

[0008] As an optional implementation, the tactile feature extraction unit employs a one-dimensional convolutional neural network or a temporal attention network architecture to process the historical sequence of the interactive force signals in the form of a sliding window; The tactile feature extraction unit is pre-trained through contrastive learning to map the pulling, relaxing, or high-frequency shaking features in the interactive force signal into corresponding text semantic embedding vectors, so that the pre-set large language model can recognize the tactile token as the hesitation, fear, and / or sudden stop intention of the accompanying person.

[0009] As an optional implementation, the heterogeneous volume perception reasoning module performs the chain-thinking reasoning process including: Based on the environmental visual information, the geometric properties of obstacles in front are identified, and the ground clearance of the bottom of the obstacle is calculated; The three-dimensional geometric contour and ground clearance information of the obstacle are compared with the spatial position relationship of the preset virtual human body collision box. The verification includes determining whether the geometric attributes intrude into the horizontal boundary of the virtual human body collision box and whether the ground clearance is within the vertical height range of the virtual human body collision box. If the verification result shows that the obstacle simultaneously intrudes into the horizontal boundary and vertical height range of the virtual human body collision box, then an asymmetric collision risk is determined, a detour sign instruction is generated, and the semantic feedback output terminal is triggered to generate a warning message; if the verification result shows that the obstacle does not intrude into the virtual human body collision box and the robot body can pass through, then a straight-through sign instruction is generated.

[0010] As an optional implementation, the low-level cooperative motion unit is the core execution unit for robot motion control, and is constructed based on an asymmetric actor-critic reinforcement learning architecture; wherein: The critic network is only enabled during the training phase. Its input data includes the robot's proprioceptive information and real-time interactive force signals, which are used to output motion control commands for the robot's joints, thereby achieving real-time closed-loop motion control of the robot. The executor network is activated during the deployment phase. Its input data includes the robot's body perception information and real-time interactive force signals, which are used to output control commands for the robot's joints.

[0011] As an optional implementation, the low-level collaborative motion unit is trained based on a reward function that includes human-machine coupled dynamic constraints, the reward function including at least: A safety reward item is configured to apply a negative reward when the virtual human collision box collides with an environmental obstacle. An interactive smoothing reward is configured to penalize the second derivative of the interactive force signal with respect to time in order to suppress the sense of unease during the guide process. The speed synchronization bonus is configured to penalize the speed vector deviation between the robot and the accompanying personnel, so that the robot's guiding speed converges to the comfortable speed of the accompanying personnel.

[0012] As an optional implementation, the system further includes a hardware-level tactile reflection module, whose operating frequency is higher than the second frequency of the lower-level collaborative motion unit; the hardware-level tactile reflection module is configured with independent safety monitoring logic, and when the amplitude of the interactive force signal is detected to exceed a preset safety threshold, the hardware-level tactile reflection module takes priority over the higher-level planning unit and the lower-level collaborative motion unit to forcibly trigger the robot to enter an emergency stop state or a passive compliance state.

[0013] As an optional implementation, the semantic feedback output terminal of the dual-stream parallel output module is also configured with causal explanation generation logic; When the heterogeneous volume perception reasoning module generates a detour sign instruction due to the detection of an asymmetric spatial obstacle, the causal explanation generation logic extracts the verification results of the ground clearance and the virtual human body collision box in the chain thinking reasoning, generates a natural language description containing the obstacle type and the reason for detour, and plays it through a speech synthesis device.

[0014] As an optional implementation, in the human-computer collaborative navigation system based on the visual-touch-language-action model, the multimodal perception alignment module further includes a visual encoder, which processes RGB images using a visual transformer architecture; The feature mapping layer includes at least a linear projection layer, which projects the visual features output by the visual encoder and the tactile features output by the tactile feature extraction unit onto the same dimension, and then concatenates them with the text prompt word embedding vectors before inputting them into the preset large language model.

[0015] As an optional implementation, the human-machine collaborative navigation system based on the visual-tactile-language-action model is applied to a quadruped robot platform, and the guide device is a force feedback handle with single or multiple degrees of freedom installed on the back of the quadruped robot. The size parameters of the accompanying personnel include at least height, shoulder width, and stride length, and the preset large language model is a visual-language-action large model that has been fine-tuned by instructions.

[0016] A second aspect of this invention provides a human-machine collaborative navigation method based on a visual-tactile-language-action model. This method is a human-machine collaborative closed-loop motion control method for a guide robot, applied to the system described in the first aspect of this invention, and includes the following steps: Physical Intent Extraction Step: The multimodal perception alignment module collects environmental visual information and real-time interactive force signals from the guide device; it uses a one-dimensional convolutional neural network to extract features from the interactive force signals and maps them into tactile tokens that are semantically aligned with the large language model to represent the physical interaction intent of the accompanying person, providing human-computer interaction intent input for robot motion control; Heterogeneous Volume Reasoning Step: The heterogeneous volume perception reasoning module extracts the geometric attributes and ground clearance of environmental obstacles based on the environmental visual information, and performs a vertical spatial comparison of the geometric attributes and ground clearance with the accompanying person's height parameter in the system prompts; it drives the preset large language model to perform chain-like reasoning to identify and / or avoid asymmetric spatial obstacles, providing path decision-making basis for robot safe motion control; Dual-stream collaborative decision-making steps: Based on the tactile token and the results of heterogeneous volume inference, the dual-stream parallel output module generates in parallel a navigation landmark sequence containing spatial coordinates and semantic voice feedback for accompanying personnel; the navigation landmark sequence is used to drive the robot to perform trajectory tracking and closed-loop motion control; Dynamic closed-loop control steps: As the core execution step of the method, the hierarchical asynchronous control module inputs the navigation landmark sequence into the low-level collaborative motion unit. The low-level collaborative motion unit combines the real-time sensed interactive force signals and outputs motion control commands for the robot joints to dynamically adjust the robot's motion state in order to maintain the dynamic synchronization between the robot and the accompanying personnel, thereby realizing human-machine collaborative closed-loop motion control of the guide robot.

[0017] This invention differs from traditional navigation technologies that solely focus on spatial position measurement, path mapping, and positioning accuracy optimization. Its core improvement does not optimize the measurement methods for navigation parameters such as position and orientation. Instead, it optimizes the human-machine collaborative motion control strategy, hierarchical asynchronous closed-loop control architecture, and human-machine coupled dynamic synchronous control method for guide robots based on information acquired through multimodal perception. All navigation-related perception and reasoning processes provide decision-making basis for the robot's safe and precise motion control, ultimately achieving autonomous motion control and human-machine collaborative motion management. Therefore, compared to existing technologies, this invention has at least the following technical advantages: This invention extracts temporal features from real-time human-computer interaction force signals collected by a guide robot and converts them into tactile tokens aligned with the semantic space of a pre-set large language model through a feature mapping layer. This establishes a mapping relationship between physical interaction signals and control semantic space, solving the technical defects of traditional guide robots' one-way open-loop control and inability to distinguish between environmental resistance and subjective control intentions from the source of control input. It transforms continuous temporal tactile interaction signals into standardized control instruction semantic units that can be recognized by the control decision system, enabling the robot to make feedforward adjustments to the navigation control strategy in advance based on the recognized interaction intentions of the human, such as sudden stops and hesitations. This improves the compliance and response accuracy of human-computer interaction control during guide robot operation. This invention constructs a joint geometric envelope of the human-machine coupling system during path planning, injects the size parameters of the accompanying personnel as core control constraints into the system, and drives a pre-set large language model to perform chain-like reasoning with heterogeneous volume constraints through three-dimensional spatial comparison of obstacle geometric properties, ground clearance and accompanying personnel size parameters. This completes the safety verification and planning control of the navigation path, realizes the collaborative control verification of the passability of the robot body and the accompanying personnel, and can actively identify and avoid asymmetric spatial obstacles that the robot can pass through but pose a collision risk to the accompanying personnel, thus improving the safety level of guide control. This invention employs a parallel output architecture combining action commands and semantic feedback. While outputting navigation control landmark sequences, it simultaneously generates semantic and speech feedback matching the control decisions, establishing an interpretable feedback loop for control decisions. Through a high-frequency, low-frequency, hierarchical asynchronous control mechanism between a high-level planning unit and a low-level collaborative motion unit, it uses a large language model at low frequency for long-range navigation planning and control, and a low-frequency motion control unit for real-time motion closed-loop adjustment. This resolves the contradiction between the inference latency of the large language model and the timing adaptation of the robot's high-frequency servo control. Simultaneously, the low-level collaborative motion unit dynamically adjusts the robot's motion control parameters based on real-time interactive force signals, achieving synchronous closed-loop control of the robot and accompanying personnel's dynamics. This ensures smooth motion control, consistent dynamic following, and real-time response performance during blind navigation. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a block diagram of a human-computer collaborative navigation system based on a visual-touch-language-action model, according to a specific embodiment of the present invention.

[0020] Figure 2 This is an architecture diagram of a human-computer collaborative navigation system based on a visual-touch-language-action model, according to a specific embodiment of the present invention.

[0021] Figure 3 This is a schematic diagram of a multimodal input alignment and network structure according to a specific embodiment of the present invention.

[0022] Figure 4 This is a flowchart of a heterogeneous volume perception reasoning logic according to a specific embodiment of the present invention.

[0023] Figure 5 This is a diagram of a low-level collaborative reinforcement learning training architecture according to a specific embodiment of the present invention.

[0024] Figure 6 This is a schematic diagram of an emergency tactile reflex principle according to a specific embodiment of the present invention.

[0025] Figure 7 This is a schematic diagram of a human-machine coupling dynamics model according to a specific embodiment of the present invention.

[0026] Figure 8 This is a semantic audio feedback logic diagram based on dual-stream output according to a specific embodiment of the present invention.

[0027] Figure 9 This is a flowchart of a human-computer collaborative navigation method based on a visual-touch-language-action model, according to a specific embodiment of the present invention. Detailed Implementation

[0028] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Furthermore, it should be understood that the specific embodiments described herein are only for illustration and explanation of this application and are not intended to limit this application.

[0029] It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments of this application. Furthermore, the descriptions of each embodiment in the following embodiments have their own emphasis; for parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0030] like Figures 1-8 As shown, the first aspect of the present invention provides a human-machine collaborative navigation system based on a visual-tactile-language-action model. The system is a human-machine collaborative motion closed-loop control system for a guide robot, and includes at least a multimodal perception alignment module, a heterogeneous volume perception inference module, a dual-stream parallel output module, and a hierarchical asynchronous control module. The main contents of each module are as follows.

[0031] The multimodal perception alignment module is configured to collect environmental visual information and real-time human-computer interaction force signals on the guide device, providing human-computer interaction intention input for robot motion control. Specifically, the position of the guide device is, for example, the real-time human-computer interaction signal at the end of the guide device. Those skilled in the art can also collect real-time human-computer interaction force signals according to the actual situation.

[0032] Specifically, the environmental visual information is acquired through a visual sensor; the guide device can be a guide rope, a traction rope, or a force feedback handle installed on the robot, used to transmit interactive force signals in human-machine collaborative navigation; the real-time human-machine interactive force signal can be, for example, a guide rope interactive force signal, or simply a guide rope interactive force, which can be acquired through a force sensor.

[0033] The multimodal perception alignment module includes at least a multimodal encoder containing a tactile feature extraction unit. The tactile feature extraction unit is used to extract temporal features from the interaction force signal and convert the extracted temporal features into tactile tokens aligned with the semantic space of a pre-set large language model through a feature mapping layer, so as to characterize or identify the physical interaction intentions of the accompanying personnel.

[0034] The heterogeneous volume perception and reasoning module, integrated into the pre-set large language model, is configured to construct the joint geometric envelope of the human-machine coupling system during path planning, providing obstacle avoidance decision-making basis for robot safe motion control. The heterogeneous volume perception and reasoning module extracts the geometric attributes and ground clearance of environmental obstacles based on the environmental visual information, and compares the geometric attributes and ground clearance with the accompanying personnel size parameters in the system prompts in three-dimensional space. This drives the pre-set large language model to perform chain-of-thought reasoning with heterogeneous volume constraints when generating navigation decisions, in order to identify and avoid asymmetric spatial obstacles. Specifically, heterogeneous volume refers to the geometric volume asymmetry caused by the difference in size between the robot body and the human body. This asymmetry directly results in the robot body's traversable path failing to meet the safety passage requirements of accompanying personnel, and is the core constraint condition for the motion control of guide robots.

[0035] Specifically, human-machine coupling dynamic synchronization refers to the robot's motion state remaining consistent with the walking state of a blind companion, without dragging or lagging behind. The human-machine coupling system is the core physical foundation for the design of control strategies and the construction of reward functions for low-level collaborative motion units, providing theoretical support for achieving synchronized human-machine dynamic control.

[0036] The human-machine coupling system is characterized or analyzed by establishing a human-machine coupling dynamics model or system; the joint geometric envelope surface specifically refers to the virtual human collision volume, for example, a cylinder is established based on the height and shoulder width of the human body, which is defined as the occupied space; the virtual human collision box refers to the three-dimensional geometric boundary established based on the body parameters (such as height and shoulder width) of the accompanying person, which is used to simulate the occupied space of the blind user.

[0037] Specifically, such as Figure 7 As shown, the human-machine coupled dynamics model consists of three modules: a robot end 301, a blind user 302, and a guide harness 303. This model models the guide robot 301, the guide harness 303, and the blind user 302 as a coupled whole: the robot end defines the motion state through its mass, position, and torque, while the blind user end defines the motion state through the blind user's mass, position, and active intention force. The two interact through the guide harness, which has spring coefficient and damping coefficient characteristics, thus constituting the physical source of tactile information in the multimodal perception architecture, providing underlying mechanical support for the large model to recognize the user's intention and provide smooth traction.

[0038] The dual-stream parallel output module, connected to the heterogeneous volume perception inference module, is configured with an action command output terminal and a semantic feedback output terminal. The action command output terminal, also known as the action command head or action head, is at least used to output a navigation landmark sequence containing spatial coordinates. The navigation landmark sequence is used to drive the robot to perform trajectory tracking and closed-loop motion control. The semantic feedback output terminal, also known as the audio description head or audio head, is at least used to generate voice description text or semantic voice feedback for accompanying personnel in real time based on the current environmental state.

[0039] Please continue reading. Figure 8 After passing through the PG-VLA inference engine, the action head outputs landmark points, i.e., navigation landmark sequences, and the audio head outputs relevant voice information feedback, such as "A beam has been detected above, detour".

[0040] In this approach, the model not only outputs navigation signs to the underlying controller, but also outputs egocentric voice descriptions to the user in parallel, bridging the perception gap.

[0041] The hierarchical asynchronous control module is the core motion control execution module of the system. It includes a high-level planning unit operating at a first frequency and a low-level cooperative motion unit operating at a second frequency, wherein the second frequency is higher than the first frequency, that is, the second frequency corresponds to a high frequency and the first frequency corresponds to a low frequency. The high-level planning unit runs the preset large language model to generate the navigation sign sequence, and the low-level cooperative motion unit receives the navigation sign sequence and adjusts the robot's motion state based on real-time perceived interactive force signals to maintain human-robot dynamic synchronization.

[0042] Specifically, the control priority of the hierarchical asynchronous control module is lower than that of the hardware-level tactile reflex module, forming a three-level closed-loop control architecture of "emergency safety control - real-time motion control - long-term planning control".

[0043] In this way, the present invention can understand the tactile feedback reflected by the guide device, and has the ability to perceive volume in a human-machine heterogeneous manner and provide semantic feedback, thereby realizing human-machine collaborative closed-loop motion control of the guide robot.

[0044] like Figure 3 As shown, in one embodiment of the present invention, the tactile feature extraction unit adopts a one-dimensional convolutional neural network (1D-CNN) or a temporal attention network architecture to process the historical sequence of the interactive force signal in the form of a sliding window; The tactile feature extraction unit is pre-trained through contrastive learning to map the pulling, relaxing, or high-frequency shaking features in the interactive force signal into corresponding text semantic embedding vectors, so that the pre-set large language model can recognize the tactile token as the hesitation, fear, and / or sudden stop intention of the accompanying person.

[0045] Specifically, the tactile feature extraction unit is the tactile feature extraction network. This network processes the interactive force signals acquired by the force sensor in the form of a sliding window. The network is pre-trained through contrastive learning to align physical signal features representing "pulling", "relaxation", "shaking", etc., with the corresponding text semantic embeddings.

[0046] Please continue reading. Figure 3This invention first extracts multidimensional features from visual images, tactile information, and text prompts using a SigLIP visual encoder, a 1D-CNN tactile encoder, and a text tokenizer, respectively. Then, a linear projection layer is used to project tactile features, visual features, and prompts onto a unified representation space, achieving intermodal alignment and sequence splicing. The spliced ​​multimodal tokens are then input into a LoRA-tuned Vicuna-7B large model skeleton for deep fusion inference. Finally, a dual-stream output head generates action commands for physical control and environmental description text for voice interaction in parallel, achieving closed-loop integration of perception, decision-making, and interpretation.

[0047] In one application scenario of this invention, the feature mapping layer maps the tactile features extracted by 1D-CNN to the embedding space of a Large Language Model (LLM) using a pre-trained projection matrix. During the multimodal alignment stage, a contrastive loss is employed to maximize the similarity between the embedding vectors of tactile events (e.g., "pull backward") and their corresponding semantic texts (e.g., "stop / hesitate"). The mapped tactile feature dimensions are unified to the hidden layer dimensions of the LLM (e.g., 4096 dimensions) and concatenated sequentially with visual tokens and prompts from SigLIP, serving as the unified input to the large model.

[0048] like Figure 4 As shown, in visual encoding, the SigLIP-So400M visual transformer is used to process RGB images, which are then projected onto the LLM embedding space through two layers of MLP to generate visual tokens. In haptic encoding, a sliding window (e.g., 1 second) of interactive force signals, i.e., historical tension data (including tensile amplitude and yaw angle), is maintained. A one-dimensional convolutional neural network (1D-CNN) is used as the haptic encoder to extract temporal features of force (such as mutation rate, sustained tensile force, etc.). Through contrastive learning pre-training, the physical signals are aligned to the semantic space. For example, rapid backward tensile force is encoded and mapped to a similar...<PULL_BACK> or <hesitation>Semantic token.

[0049] This approach enables large models to understand "pulling" like they understand words, thus distinguishing between "terrain resistance" and "human refusal." This method can discretize the continuous tension signal of guide devices such as guide ropes into tactile tokens that large models can understand, giving robots "physical empathy" and enabling them to recognize the hesitation, fear, or sudden stop intentions of blind people.

[0050] In one application scenario of this invention, the input of the 1D-CNN is a sequence of tension vectors within a sliding window (e.g., 50×2-dimensional, representing the amplitude and angle within 1 second); the first convolutional layer has a kernel size of k=5, a stride of s=1, and an output dimension of 46×64. After the downsampling layer, two convolutional layers with kernels of k=3 and a stride of s=2 are followed, progressively compressing the temporal features to 10×256. The loss function uses the InfoNCE loss function, first calculating the cosine similarity between the tactile intent and its matching text (positive sample pair) and amplifying it exponentially, then comparing it with the sum of amplified similarities of all non-matching texts (negative sample pairs) to identify the probability of a correct pairing; subsequently, this probability is converted into a loss value through logarithmic transformation and a negative sign. Finally, by minimizing the loss, the correlation between positive sample pairs is brought closer, and the interference of negative sample pairs is pushed away, thereby improving the matching accuracy of cross-modal features.

[0051] like Figure 4 As shown, in one embodiment of the present invention, the process of the heterogeneous volume perception reasoning module executing the chain-like reasoning includes: Based on the environmental visual information, the geometric properties of obstacles in front are identified, and the ground clearance of the bottom of the obstacle is calculated; The geometric attributes and the ground clearance are compared with a preset virtual human body collision box to verify their spatial relationship. The verification includes determining whether the geometric attributes intrude into the horizontal boundary of the virtual human body collision box and whether the ground clearance is within the vertical height range of the virtual human body collision box. If the verification result shows that the obstacle simultaneously intrudes into the horizontal boundary and vertical height range of the virtual human body collision box, then an asymmetric collision risk is determined, a detour sign instruction is generated, and the semantic feedback output terminal is triggered to generate a warning message; if the verification result shows that the obstacle does not intrude into the virtual human body collision box and the robot body can pass through, then a straight-through sign instruction is generated.

[0052] Specifically, the construction process of the virtual human collision box involves using the reasoning ability of a large model to explicitly project a virtual human collision box during the planning phase. By injecting "human height and width constraints" into the text prompt, the system can take into account the passability of accompanying personnel when generating paths, effectively avoiding top traps.

[0053] Specifically, while robots can easily navigate over suspended obstacles, vertical obstacles, and low-lying traps, these pose a fatal risk of head collision for accompanying blind individuals, i.e., asymmetric collision risk.

[0054] In one embodiment of the present invention, by using heterogeneous volume-aware inference, the volume parameters of the person (such as height H=1.75, width W=0.6) are explicitly injected into the Prompt, constraining the inference steps of the large model: Reasoning Step 1: Analyze the height of obstacles (such as tree branches, door frames) above the ground and the passage width within the field of vision; Reasoning step 2: Compare the environmental geometry with the preset personnel volume constraints (virtual collision box); Reasoning step 3: Combine the user tension intentions detected by the tactile token (such as walking normally or hesitating) to determine the collision risk of the current path; Output generation: The corresponding navigation landmarks and voice explanations (such as "There is an obstacle overhead ahead, please follow me to the right to detour") are only generated when the logic chain determines that it is safe.

[0055] In one application scenario of this invention, in order to solve the problem of geometric asymmetry, specific prompt word engineering and thought chain are introduced in heterogeneous volume perception reasoning, and the specific process is as follows.

[0056] System prompt injection: Explicitly define human-machine geometry parameters in the system prompt: "You are a guidedog. Robot Size=[0.5m, 0.3m], Human Size=[1.75m, 0.6m]. Safety Constraint: Avoid obstacles intersecting with Human Volume."

[0057] Chain-of-Thought (CoT): The model does not directly output actions, but instead generates reasoning steps first. For example, when a beam is detected 1.4m high in front: 1) Perception: "Visual system detects horizontal bar at height 1.4m."; 2) Volume verification: "My height is 0.5m (Passable). Human height is 1.75m (Collision Risk)." 3) Decision: "Action: Stop and Replan. Generate Audio Warning.".

[0058] Dual-stream output: 1) Flow 1 (Action): Outputs the navigation sign sequence to the underlying policy; 2) Stream 2 (Speech): Outputs the text "Watch out, low hanging branch ahead" and plays it through the TTS module (Text-to-Speech module, the technical component for text-to-speech function).

[0059] like Figure 5 As shown, in one embodiment of the present invention, the low-level cooperative motion unit is the core execution unit for robot motion control, and is constructed based on an asymmetric actor-critic reinforcement learning architecture; wherein: The Critic Network is enabled only during the training phase. Its input data includes privileged information from the simulation environment, which includes at least the real position coordinates of the accompanying personnel, the coefficient of terrain friction, and the physical stiffness parameters of the guide device. The Actor Network is activated during the deployment phase. Its input data only includes the robot's body perception information and real-time interactive force signals, which are used to output control commands for the robot's joints.

[0060] Privileged information is important information that is difficult to measure with sensors in the real environment but exists in reality, such as the true location of pedestrians, terrain parameters (including friction coefficient and ground recovery capacity), stiffness and damping parameters of guide harnesses, robot parameters (load changes, center of mass shift, joint shift), and motor stiffness and damping gain. By using privileged information, value estimation can be made more accurate.

[0061] Specifically, privileged information serves to provide more accurate value estimation feedback. This type of information refers to key parameters that actually exist in the environment but are difficult to obtain directly due to limitations in sensor accuracy or physical conditions. These include, but are not limited to: external environmental characteristics, such as the precise coordinates of pedestrians and terrain parameters (such as the coefficient of friction and the ground recovery coefficient); interaction and load characteristics, such as the stiffness and damping parameters of guide harnesses and the real-time load changes of the robot; and hardware physical deviations, such as the displacement of the center of mass and joints, motor stiffness, and damping gain.

[0062] Specifically, the robot's proprioceptive information includes IMU information and joint state information. The actor network implicitly determines the status of accompanying personnel, such as whether they have fallen behind, by sensing forces.

[0063] Here, this flexible coupling control strategy enables a compliant response to traction and speed synchronization.

[0064] In one embodiment of the present invention, the low-level collaborative motion unit is trained based on a reward function that includes human-machine coupled dynamic constraints, the reward function including at least: A safety reward item is configured to apply a negative reward when the virtual human collision box collides with an environmental obstacle. The interaction smoothing reward is configured to penalize the second derivative of the interaction force signal with respect to time, that is, to penalize the rate of change of the guide rope tension, in order to suppress or reduce the sense of jerking during the guidance process and ensure a smooth traction process. The speed synchronization reward is configured to penalize the speed vector deviation between the robot and the accompanying personnel, so that the robot's guiding speed converges to the comfortable speed of the accompanying personnel, thus achieving synchronization between the robot and the accompanying personnel.

[0065] For example, this invention employs reinforcement learning training based on the Proximal Policy Optimization (PPO) algorithm. This algorithm can quickly learn stable and efficient control strategies through real-time interaction with the physical environment (such as force feedback and joint motion data). In the simulation environment (Isaac Sim), the robot is connected to a "Passive Humanoid Proxy" via a spring-damper system. The human proxy has mass and a damping coefficient, which varies with human intent; for example, the damping coefficient increases when the human is fearful, thus enabling coupled dynamics modeling.

[0066] Specifically, in the training of the underlying reinforcement learning strategy, the total reward is the weighted sum of the task, reduction, and collaboration. The task corresponds to the safety reward, the reduction represents the interaction smoothness reward, and the collaboration represents the speed synchronization reward.

[0067] Specifically, the physical meaning of the safety reward is a collision penalty, which applies a very large negative value when the envelope of the person's volume overlaps with an obstacle, forcing the strategy to avoid it; the physical meaning of the interaction smoothness reward is to penalize the second derivative of the guide rope tension, which prevents the robot from generating a sudden pulling sensation by constraining the rate of change of the guide rope tension; the physical meaning of the speed synchronization reward is a speed following reward, which prompts the robot to automatically align with the user's walking speed and to limit it when it exceeds the comfort threshold.

[0068] like Figure 6 As shown, in one embodiment of the present invention, the human-computer collaborative navigation system based on the visual-tactile-language-action model further includes a hardware-level tactile reflection module, which runs on a microcontroller and has an operating frequency higher than the second frequency of the lower-level collaborative motion unit. The hardware-level tactile reflection module is equipped with independent safety monitoring logic. When the amplitude of the interactive force signal exceeds the preset safety threshold, the hardware-level tactile reflection module takes priority over the high-level planning unit and the low-level cooperative motion unit to forcibly trigger the robot to enter an emergency stop state or a passive compliance state.

[0069] In one application scenario of this invention, the safety detection logic is an independent logic running on the MCU. If instantaneous tension is detected, then... Figure 6 If the force sensor signal exceeds the threshold, such as a threshold set to 50N, exceeding the safety threshold indicates a potential fall or collision. In this case, the system will ignore all upper-level commands and directly trigger emergency braking or emergency stop and lock the joint to ensure safety.

[0070] In one embodiment of the present invention, the semantic feedback output terminal of the dual-stream parallel output module is further configured with causal explanation generation logic; When the heterogeneous volume perception reasoning module generates a detour sign instruction due to the detection of an asymmetric spatial obstacle, the causal explanation generation logic extracts the verification results of the ground clearance and the virtual human body collision box in the chain thinking reasoning, generates a natural language description containing the obstacle type and the reason for detour, and plays it through a speech synthesis device.

[0071] In one embodiment of the present invention, the multimodal perception alignment module further includes a visual encoder, which uses a Vision Transformer (ViT) to process RGB images; The feature mapping layer includes at least a linear projection layer, which projects the visual features output by the visual encoder and the tactile features output by the tactile feature extraction unit onto the same dimension, and then concatenates them with the text prompt word embedding vectors before inputting them into the preset large language model.

[0072] In one embodiment of the present invention, the multimodal perception alignment module includes a visual encoder, which is based on the ViT architecture and is used to extract global and local visual features of RGB images.

[0073] The feature mapping layer includes at least a linear projection layer, which maps the visual features output by the visual encoder and the tactile features output by the tactile feature extraction unit to a unified modal representation space. The aligned visual and tactile features are then concatenated with the embedding vectors of the text prompts and input into the pre-set large language model to achieve deep fusion and collaborative reasoning of multimodal information.

[0074] In one embodiment of the present invention, the human-machine collaborative navigation system based on the visual-tactile-language-action model is applied to a quadruped robot platform, and the guide device is a force feedback handle or rope handle with single or multiple degrees of freedom installed on the back of the quadruped robot. The size parameters of the accompanying personnel include at least height, shoulder width, and stride length. The preset large language model is a visual-language-action large model that has been fine-tuned by instructions, such as the Vicuna-7B large model.

[0075] For example, the system described in this invention is deployed on the Unitree Go2 quadruped robot platform, and the core computing unit is an NVIDIA Jetson Orin NX.

[0076] The hardware components include the following.

[0077] Perception layer: A RealSense D435i depth camera is installed in the head to acquire RGB and depth images; a custom 1-DoF force sensor handle is installed in the back to collect the tension and direction of the guide rope at a frequency of 100Hz.

[0078] Computation layer: Adopts a layered asynchronous architecture, including: 1) High-level cognitive thread or strategy (Thread A, ~3Hz): runs the PG-VLA large model (based on Vicuna-7B fine-tuning) corresponding to the multimodal perception alignment module, and is responsible for environmental understanding, intent recognition and path planning; 2) Low-level control thread or policy (Thread B, 50Hz): Runs the reinforcement learning policy network, responsible for full-body motion control (Joint Positions / Velocities). 3) Reflection layer or strategy (MCU, 500Hz): Runs hard real-time haptic reflection logic.

[0079] In one application scenario of this invention, in order to address the contradiction between the high latency (e.g., ~300ms) of large model inference and the high frequency requirements of robot control (e.g., 20ms), the following mechanism is adopted.

[0080] Asynchronous communication mechanism: The higher-level PG-VLA updates road signs and writes them to shared memory at a frequency of 3Hz, while the lower-level control policy reads the latest road signs and executes them at a frequency of 50Hz.

[0081] Active waiting mechanism: If the underlying policy detects continuous backward tension (the large model has not yet reacted), the reinforcement learning policy will automatically reduce its speed or enter a stationary mode, using force feedback to achieve initial compliance.

[0082] Hardware-level tactile reflex mechanism: If the instantaneous tension exceeds the safety threshold (e.g., 50N, indicating a fall or serious collision), the system will ignore all upper-level instructions, directly trigger an emergency stop and lock the joint to ensure absolute safety.

[0083] like Figure 9 As shown, a second aspect of the present invention provides a human-machine collaborative navigation method based on a visual-tactile-language-action model. This method is a human-machine collaborative closed-loop motion control method for a guide robot, and is applied to a system as described in any of the above embodiments, comprising the following steps: S1. Physical Intent Extraction Step: The multimodal perception alignment module collects environmental visual information and real-time interactive force signals from the guide device (such as the end effector); the interactive force signals are used to extract features using a one-dimensional convolutional neural network and mapped to tactile tokens that are semantically aligned with the large language model to represent the physical interaction intent of the accompanying person and provide human-computer interaction intent input for robot motion control; S2. Heterogeneous volume reasoning steps: The heterogeneous volume perception reasoning module extracts the geometric attributes and ground clearance of environmental obstacles based on the environmental visual information, and performs a vertical spatial comparison between the geometric attributes and ground clearance and the height parameters of the accompanying personnel in the system prompts; drives the pre-set large language model to perform chain-like thinking reasoning to identify and / or avoid asymmetric spatial obstacles; S3. Dual-stream collaborative decision-making steps: Based on the tactile token and the result of heterogeneous volume inference, the dual-stream parallel output module generates in parallel a navigation landmark sequence containing spatial coordinates and semantic voice feedback for accompanying personnel; the navigation landmark sequence is used to drive the robot to perform trajectory tracking and closed-loop motion control; S4. Dynamic closed-loop control steps: The hierarchical asynchronous control module inputs the navigation landmark sequence into the low-level collaborative motion unit. The low-level collaborative motion unit dynamically adjusts the robot's motion state in combination with the real-time sensed interactive force signals to maintain the dynamic synchronization between the robot and the accompanying personnel, thereby realizing human-machine collaborative closed-loop motion control of the guide robot.

[0084] For example, by real-time acquisition of force sensor signals at the end of the guide rope, dynamic features are extracted and converted into a tactile token sequence, which, along with visual image features, is input into a large language model. Combining visual depth information with the prior knowledge of the large language model, a voxel analysis is performed on the space in front. The voxel analysis process involves discretizing the continuous three-dimensional space into regular cubic grids (voxels) and constructing a probability map in real time that includes the probability of obstacle occupancy and spatial attributes. Subsequently, based on the injected personnel volume constraints (i.e., considering the actual physical contour of the accompanying personnel rather than their center of mass), a refined collision risk assessment is performed in the voxel space. Based on the current tactile tokens (indicating the personnel's intentions) and the volume assessment results, navigation landmarks and voice explanation text are generated in parallel. The low-level control strategy, based on the landmark instructions and real-time tension feedback, achieves smooth traction for the accompanying personnel through variable impedance control or speed adjustment.

[0085] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.

[0086] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a system including a processing module or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

[0087] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.< / hesitation>

Claims

1. A human-computer collaborative navigation system based on a visual-tactile-language-action model, characterized in that, The system is a human-machine collaborative motion closed-loop control system for guide mobile robots, including: A multimodal perception alignment module is configured to collect environmental visual information and real-time human-computer interaction force signals from guide devices for the visually impaired, providing human-computer interaction intention input for robot motion control; the multimodal perception alignment module includes at least a tactile feature extraction unit, which is used to extract temporal features from the interaction force signals and convert the extracted temporal features into tactile tokens aligned with the semantic space of a pre-set large language model through a feature mapping layer, so as to represent the physical interaction intention of the accompanying person; A heterogeneous volume perception and reasoning module, integrated into the pre-set large language model, is configured to construct a joint geometric envelope of the human-machine coupling system during path planning, providing obstacle avoidance decision-making basis for robot safe motion control. The heterogeneous volume perception and reasoning module extracts the geometric attributes and ground clearance of environmental obstacles based on the environmental visual information, and compares the geometric attributes and ground clearance with the accompanying personnel size parameters in the system prompts in three-dimensional space. This drives the pre-set large language model to perform chain-like reasoning with heterogeneous volume constraints when generating navigation decisions, in order to identify and / or avoid asymmetric spatial obstacles. A dual-stream parallel output module, connected to the heterogeneous volume perception inference module, is configured with an action command output terminal and a semantic feedback output terminal; the action command output terminal is at least used to output a navigation landmark sequence containing spatial coordinates, the navigation landmark sequence is used to drive the robot to perform trajectory tracking and closed-loop motion control; the semantic feedback output terminal is at least used to generate semantic voice feedback for accompanying personnel in real time based on the current environmental state and inference results. The hierarchical asynchronous control module is the core motion control execution module of the system. It includes a high-level planning unit operating at a first frequency and a low-level collaborative motion unit operating at a second frequency, wherein the second frequency is higher than the first frequency. The high-level planning unit runs the preset large language model to generate the navigation sign sequence. The low-level collaborative motion unit receives the navigation sign sequence and outputs motion control commands for the robot joints based on real-time perceived interactive force signals, dynamically adjusting the robot's motion state to maintain human-machine dynamic synchronization between the robot and the accompanying person, thereby realizing human-machine collaborative closed-loop motion control of the guide robot.

2. The human-computer collaborative navigation system based on a visual-tactile-language-action model according to claim 1, characterized in that, The tactile feature extraction unit adopts a one-dimensional convolutional neural network or temporal attention network architecture and processes the historical sequence of the interactive force signal in the form of a sliding window; The tactile feature extraction unit is pre-trained through contrastive learning to map the pulling, relaxing, or high-frequency shaking features in the interactive force signal into corresponding text semantic embedding vectors, so that the pre-set large language model can recognize the tactile token as the hesitation, fear, and / or sudden stop intention of the accompanying person.

3. The human-computer collaborative navigation system based on a visual-tactile-language-action model according to claim 1, characterized in that, The heterogeneous volume perception reasoning module performs the chain-thinking reasoning process as follows: Based on the environmental visual information, the geometric properties of obstacles in front are identified, and the ground clearance of the bottom of the obstacle is calculated; The three-dimensional geometric contour and ground clearance information of the obstacle are compared with the spatial position relationship of the preset virtual human body collision box. The verification includes determining whether the geometric attributes intrude into the horizontal boundary of the virtual human body collision box and whether the ground clearance is within the vertical height range of the virtual human body collision box. If the verification result shows that the obstacle simultaneously intrudes into the horizontal boundary and vertical height range of the virtual human body collision box, then an asymmetric collision risk is determined, a detour sign instruction is generated, and the semantic feedback output terminal is triggered to generate a warning message; if the verification result shows that the obstacle does not intrude into the virtual human body collision box and the robot body can pass through, then a straight-through sign instruction is generated.

4. The human-computer collaborative navigation system based on a visual-tactile-language-action model according to claim 3, characterized in that, The low-level cooperative motion unit is the core execution unit for robot motion control, and it is constructed based on an asymmetric actor-critic reinforcement learning architecture; wherein: The critic network is enabled only during the training phase. Its input data includes privileged information from the simulation environment, which is used to output a value function to complete policy evaluation. The executor network is activated during the deployment phase. Its input data only includes the robot's body perception information and real-time interactive force signals, which are used to output motion control commands for the robot's joints, thereby realizing real-time closed-loop motion control of the robot.

5. The human-computer collaborative navigation system based on a visual-tactile-language-action model according to claim 4, characterized in that, The low-level collaborative motion unit is trained based on a reward function that includes human-machine coupled dynamic constraints, and the reward function includes at least: Safety reward items are configured to apply negative safety reward items when the virtual human collision box collides with environmental obstacles. An interactive smoothing reward is configured to penalize the second derivative of the interactive force signal with respect to time in order to suppress the sense of unease during the guide process. The speed synchronization bonus is configured to penalize the speed vector deviation between the robot and the accompanying personnel, so that the robot's guiding speed converges to the comfortable speed of the accompanying personnel.

6. The human-computer collaborative navigation system based on a visual-tactile-language-action model according to claim 1, characterized in that, It also includes a hardware-level tactile reflection module, and its operating frequency is higher than the second frequency of the lower-level cooperative motion unit; The hardware-level tactile reflection module is equipped with independent safety monitoring logic. When the amplitude of the interactive force signal exceeds the preset safety threshold, the hardware-level tactile reflection module takes priority over the high-level planning unit and the low-level cooperative motion unit to forcibly trigger the robot to enter an emergency stop state or a passive compliance state.

7. The human-computer collaborative navigation system based on a visual-tactile-language-action model according to claim 3, characterized in that, The semantic feedback output terminal of the dual-stream parallel output module is also configured with causal explanation generation logic. When the heterogeneous volume perception reasoning module generates a detour sign instruction due to the detection of an asymmetric spatial obstacle, the causal explanation generation logic extracts the verification results of the ground clearance and the virtual human body collision box in the chain thinking reasoning, generates a natural language description containing the obstacle type and the reason for detour, and plays it through a speech synthesis device.

8. The human-computer collaborative navigation system based on a visual-tactile-language-action model according to claim 1, characterized in that: The multimodal perception alignment module also includes a visual encoder, which uses a visual transformer architecture to process RGB images; The feature mapping layer includes at least a linear projection layer, which projects the visual features output by the visual encoder and the tactile features output by the tactile feature extraction unit onto the same dimension, and then concatenates them with the text prompt word embedding vectors before inputting them into the preset large language model.

9. The human-computer collaborative navigation system based on a visual-tactile-language-action model according to claim 1, characterized in that, include: The system is applied to a quadruped robot platform, and the guide device is a force feedback handle with one or more degrees of freedom installed on the back of the quadruped robot. The size parameters of the accompanying personnel include at least height, shoulder width, and stride length, and the preset large language model is a visual-language-motor large model that has been fine-tuned by instructions.

10. A human-computer collaborative navigation method based on a visual-tactile-language-action model, characterized in that, The method is a human-machine collaborative closed-loop motion control method for guide mobile robots, applied to the system described in any one of claims 1-9, and includes the following steps: Physical intent extraction steps: The multimodal perception alignment module collects environmental visual information and real-time interactive force signals from the guide device; a one-dimensional convolutional neural network is used to extract features from the interactive force signals and map them into tactile tokens that are semantically aligned with the large language model to represent the physical interaction intent of the accompanying person and provide human-computer interaction intent input for robot motion control; Heterogeneous volume reasoning steps: The heterogeneous volume perception reasoning module extracts the geometric attributes and ground clearance of environmental obstacles based on the environmental visual information, and compares the geometric attributes and ground clearance with the height parameters of the accompanying personnel in the system prompts in a vertical space; drives the pre-set large language model to perform chain-like thinking reasoning to identify and / or avoid asymmetric spatial obstacles, providing a path decision basis for the robot's safe motion control; Dual-stream collaborative decision-making steps: Based on the tactile token and the results of heterogeneous volume inference, the dual-stream parallel output module generates in parallel a navigation landmark sequence containing spatial coordinates and semantic voice feedback for accompanying personnel; the navigation landmark sequence is used to drive the robot to perform trajectory tracking and closed-loop motion control; Dynamic closed-loop control steps: As the core execution step of the method, the hierarchical asynchronous control module inputs the navigation landmark sequence into the low-level collaborative motion unit. The low-level collaborative motion unit combines the real-time sensed interactive force signals and outputs motion control commands for the robot joints to dynamically adjust the robot's motion state in order to maintain the dynamic synchronization between the robot and the accompanying personnel, thereby realizing human-machine collaborative closed-loop motion control of the guide robot.