Visual-tactile cross-modal mapping wearable path guiding system

Through the wearable path guidance system with visual-tactile cross-modal mapping, combined with the depth camera and haptic feedback device, the navigation complexity and delay problems of visual inconvenience users are solved, and high-precision and low-latency navigation guidance is achieved, improving the robustness and user experience of the system.

CN120385342APending Publication Date: 2025-07-29BEIHANG UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510491375.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing wearable navigation system has complex path planning, poor robustness, high response delay, and poor human-computer interaction among users with visual inconvenience, and has failed to effectively combine user gait and environmental information for navigation guidance.

Method used

A wearable path guidance system with visual-tactile cross-modal mapping is adopted to build an incremental scene map through a depth phase mechanism, combining IMU sensors and streamlined path planning modules to generate a simplified path suitable for people, and providing directional guidance and obstacle avoidance prompts through a tactile feedback device, and combining a voice assist module to provide navigation guidance when necessary.

Benefits of technology

It realizes high-precision and low-latency navigation in complex environments, improves the robustness of the navigation system and user operation convenience, reduces the dependence on hearing, and enhances the universality and safety of the navigation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120385342A_ABST
    Figure CN120385342A_ABST
Patent Text Reader

Abstract

The invention discloses a wearable path guiding system based on visual-tactile cross-modal mapping. The wearable path guiding system adopts a three-module architecture, namely a visual environment sensing module, a simplified path planning module and a tactile guiding module. The environment perception module constructs an incremental scene map with semantic annotations based on a depth camera. And the simplified path planning module is used for generating a steering vector containing gait information and a steering angle according to the advancing direction and simplified path pointing and performing real-time iterative optimization to realize real-time matching of the advancing direction of the user and the target direction. And the tactile guidance module is used for realizing synchronization of gait and tactile stimulation, generating priority gradient intensity field tactile feedback with a direction indication characteristic according to path deviation information, correcting the advancing direction, and starting the voice auxiliary module when multiple times of tactile correction do not reach the expectation. The method is suitable for a visual signal limited scene, the visual signal is mapped into the tactile signal to realize self-adaptive synchronization with the motion rhythm, and path guidance is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - field of intelligent wearable devices and autonomous navigation technologies, and particularly relates to a wearable path guidance system with visual - tactile cross - modal mapping and its implementation method. Background Art

[0002] With the rapid development of intelligent perception technologies and wearable devices, navigation systems are extending from traditional vehicle - mounted / air - borne platforms to human - centered applications. In the civilian assistance field, the travel needs of approximately 300 million visually impaired people globally have given rise to products such as guide canes and smart glasses; in industrial scenarios, applications such as underground pipe network inspection and high - risk scenario rescue have promoted the migration of AGV navigation technologies to wearable devices; in special and military fields, the demand for individual soldier navigation in complex environments is increasing continuously, and the cross - integration of fields such as augmented reality (AR), virtual reality (VR), and medical remote robots has further accelerated the integration of multi - modal perception and wearable devices.

[0003] However, existing wearable navigation systems face significant technical bottlenecks in cross - field applications:

[0004] (1) At the path planning level, existing path planning algorithms are mainly applied to robot and vehicle navigation. In wearable navigation systems for visually impaired assistive devices, there are still significant deficiencies in path guidance optimization. Traditional path planning methods usually construct global or local paths based on static or dynamic maps, emphasizing obstacle avoidance and optimal path calculation, but do not fully consider the biomechanical characteristics and cognitive load of human walking. For example, existing algorithms often assume that the navigation entity has good visual perception ability and can directly rely on visual feedback for path correction, but for visually impaired users, this assumption does not hold. Therefore, how to generate a simplified path suitable for human walking in a complex environment and reduce path features unfriendly to visually impaired users, such as sharp turns and narrow passages, is one of the core challenges in optimizing path planning algorithms. In addition, path planning should also combine the user's gait inertial data to dynamically adjust the guidance strategy to ensure the natural connection and smooth execution of navigation information.

[0005] (2) In the dimension of human - machine interaction, the tactile feedback channel fails to establish an information mapping mechanism that is clear in spatial guidance and conforms to human motion intuition; existing tactile feedback devices mainly transmit the azimuth information of obstacles, but lack clear path guidance, making it difficult for users to intuitively perceive the direction to move forward, resulting in fuzzy spatial perception; in addition, the synchronization of tactile feedback instructions with the user's actual movement is also a key issue affecting the navigation experience. Traditional tactile coding methods usually rely on fixed - pattern vibration prompts, but do not fully consider the user's gait rhythm and reminder timing, which can lead to misunderstandings and delays in navigation signals and increase the navigation failure rate.

[0006] The problems existing in the above methods jointly lead to common problems faced by the existing systems in cross - domain applications, such as complex path planning, poor robustness, high response latency, and poor human - machine interaction. It is urgent to achieve breakthroughs through technological innovation. Summary of the Invention

[0007] In view of the above problems, the present invention proposes a wearable path guidance system with visual - tactile cross - modal mapping. This system is applicable to complex scenarios in multiple fields such as civilian assistance, industrial operations, special applications, and military tactics. Through a closed - loop architecture of environmental perception - path planning - tactile feedback, it realizes cross - domain applications including outdoor navigation for visually impaired people, anti - wandering guardianship for Alzheimer's patients, three - dimensional space inspection of underground pipe networks, guidance in smoke environments for fire rescue, path planning for mine emergency escape, and individual soldier tactical collaborative navigation.

[0008] A wearable path guidance system with visual - tactile cross - modal mapping according to the present invention includes a visual - tactile cross - modal mapping wearable device, a visual environment perception module, a simplified path planning module, a tactile guidance module, and a voice assistance module.

[0009] The visual - tactile cross - modal mapping wearable device includes an adjustable waist belt tied to the human waist, lower - limb fixing rings symmetrically tied to the left and right thighs of the human body, and bone - conduction headphones mounted on the human ears.

[0010] Among them, a depth camera and a control chamber are installed on the adjustable waist belt. The control chamber is internally provided with an embedded processor, a differential GPS module, and a power supply. The embedded processor integrates an environment perception module, a simplified path planning module, and a tactile guidance module; the lower - limb fixing rings are two annular parts, on which an array of spatial vibrators is installed. At the same time, an IMU sensor is installed at the biceps femoris on the opposite side of the array of spatial vibrators; the array of spatial vibrators is 4 micro linear vibrators installed along the thigh ring side on the annular part, used to provide vibration stimulation for the thigh.

[0011] The visual environment perception module constructs a scene model based on the depth camera and synchronously constructs an incremental scene map with semantic annotations.

[0012] The simplified path planning module obtains an initial path based on the incremental scene map constructed by the visual environment perception module. Based on an adaptive path sparse segmentation model, the initial path is discretized into a sparse set of waypoints to generate a simplified path. Further, the embedded processor generates a steering angle Δθ in real - time according to the angle deviation between the user's current traveling direction measured by the IMU sensor and the direction of the simplified path; at the same time, the IMU sensor acquires the current speed v, angular velocity ω, and acceleration α of the user's legs and transmits them to the embedded processor. After low - pass filtering and coordinate alignment for data pre - processing, they are uniformly packaged into a steering vector as an intermediate transmission signal.

[0013] The tactile guidance module is designed with a gait phase prediction unit, which is used to accurately predict the gait phase change of the user and extract the stance phase part in the gait cycle; and it is set that the vibration stimulation trigger starts from the middle of the stance phase of a single leg, and the duration is dynamically adjusted according to the swing phase duration of the other leg predicted in real time, ensuring that it does not exceed 60% of the swing phase time of the gait cycle. Further, the embedded processor generates a steering vector containing gait information and steering angle according to the forward direction and the trimmed path direction, and iteratively optimizes it in real time to achieve real-time matching of the user's traveling direction and the target direction.

[0014] At the same time, the tactile guidance module also uses a dual-mode vibration coding strategy to achieve direction guidance and emergency obstacle avoidance according to the obstacle distance information collected by the depth camera.

[0015] When the embedded processor detects that the obstacle distance is greater than the safety threshold, the direction guidance mode is enabled. The tactile guidance module generates a spatial activation sequence of the linear vibrator array, dynamically selects and activates the corresponding vibration units according to the user's steering angle deviation, and periodically triggers in a vibration feedback manner from weak to strong to guide the user to adjust the traveling direction or maintain the current path.

[0016] The embedded processor also judges the current user's movement speed v at the same time. When the user's movement speed v user ≥2m / s, the tactile guidance module activates the prospective guidance mode. In this mode, the vibration signal is triggered in advance before the vibration stimulation trigger moment in the middle of the phase according to the aforementioned spatial activation sequence.

[0017] When the embedded processor detects that the obstacle is approaching and the distance is less than the safety threshold, the tactile guidance module activates the emergency obstacle avoidance mode. In this mode, the high-frequency vibration of the entire array is immediately excited.

[0018] The voice assistance module is used to provide voice prompts when the user fails to successfully correct the traveling direction through tactile feedback, and is transmitted by the embedded processor to the bone conduction headphones through the Bluetooth module to assist the user in path correction.

[0019] The advantages of the present invention are as follows:

[0020] 1. The visual-tactile cross-modal mapping wearable path guidance system of the present invention is applicable to a variety of application scenarios, covering fields such as civilian assistance, industrial inspection, and military tactics, and can meet the guidance needs in different scenarios, significantly improving the application universality and practical value of the guidance system.

[0021] 2. The wearable path guidance system for visual-tactile cross-modal mapping of the present invention, through a closed-loop architecture of environmental perception, simplified path planning, and tactile feedback, can monitor the user's motion state in real time, dynamically adjust the navigation strategy, provide a simplified path planning method for complex environments, reduce path complexity, ensure a high degree of matching between the user's traveling direction and the predetermined path, and thus improve the safety and accuracy of the overall guidance.

[0022] 3. The wearable path guidance system for visual-tactile cross-modal mapping of the present invention reduces the dependence on hearing through multi-sensor data fusion, realizes a multi-sensory cooperation mode such as vision, touch, and voice, forms a complementary feedback mechanism, and improves the robustness of the system in multiple scenarios and the convenience of user operation.

[0023] 4. The wearable path guidance system for visual-tactile cross-modal mapping of the present invention adopts a hybrid network architecture model of Transformer and CNN based on multi-level spatio-temporal context modeling. It dynamically locates the vibration trigger timing (midstance) through the support phase probability curve, and limits the vibration duration within 60% of the swing phase according to the predicted swing phase duration by regression. This mechanism ensures the natural synchronization of tactile stimuli with the user's gait, eliminates the cognitive misalignment caused by feedback delay, significantly improves the navigation intuitiveness, and at the same time avoids excessive vibration interfering with gait stability and improves the interaction experience.

[0024] 5. The wearable path guidance system for visual-tactile cross-modal mapping of the present invention can perceive the distance of obstacles in real time through a depth camera and dynamically switch between dual-mode tactile coding: when the distance of the obstacle is greater than the safety threshold, it adopts the direction guidance mode and provides progressive path guidance through the periodic spatial activation sequence of the linear vibration array; when the obstacle approaches the dangerous range, it immediately triggers the emergency obstacle avoidance mode, and the full array of leg vibrators synchronously vibrates at a high frequency and strongly. The system dynamically optimizes the safety threshold by integrating the user's motion state and environmental data, taking into account the navigation accuracy and the reliability of emergency obstacle avoidance, and breaks through the response lag defect of the traditional single vibration mode.

[0025] 6. The wearable path guidance system for visual-tactile cross-modal mapping of the present invention adopts a dynamic hierarchical tactile feedback mechanism: when the steering angle deviation exceeds the threshold, it triggers unilateral progressive vibration (left / right leg linear array) based on the principle of "taking the larger value" to simulate intuitive traction with a strength gradient from weak to strong; when the deviation is within the tolerance range, it activates bilateral centripetal contraction vibration (triggered sequentially from the outside to the inside) to provide a balance hint. This strategy combines human motion intuition and active error suppression, significantly improves the navigation accuracy and reduces the user's cognitive load. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0027] Figure 1 Schematic diagram of the architecture and working process of the wearable path guidance system for visual-tactile cross-modal mapping of the present invention;

[0028] Figure 2 Schematic diagram of the structure of the flexible support component in the wearable path guidance system for visual-tactile cross-modal mapping of the present invention;

[0029] Figure 3 Schematic diagram of the designed adaptive path sparse segmentation model in the wearable path guidance system for visual-tactile cross-modal mapping of the present invention;

[0030] Figure 4 Schematic diagram of the designed vibration stimulation trigger timing in the wearable path guidance system for visual-tactile cross-modal mapping of the present invention;

[0031] Figure 5 Schematic diagram of the designed dual-mode vibration coding strategy in the wearable path guidance system for visual-tactile cross-modal mapping of the present invention;

[0032] Figure 6 Schematic diagram of the designed vibration priority gradient intensity field algorithm in the wearable path guidance system for visual-tactile cross-modal mapping of the present invention;

[0033] In the figure:

[0034] 1 - Adjustable waist belt, 2 - Bone conduction headphones, 3 - Lower limb fixing ring

[0035] 101 - Camera fixing cabin, 102 - Depth camera, 103 - Rear-end control cabin

[0036] 104 - Embedded processor, 105 - Differential GPS module, 106 - Distributed power system

[0037] 107 - Spatial vibrator array, 108 - IMU sensor, 109 - Signal line Detailed implementation manners

[0038] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts fall within the scope of protection of this application.

[0039] The wearable path guidance system for visual-tactile cross-modal mapping of the present invention includes a visual-tactile cross-modal mapping wearable device, a visual environment perception module, a simplified path planning module, a tactile guidance module, and a voice assistance module, as Figure 1 shown.

[0040] The visual-tactile cross-modal mapping wearable device includes an adjustable waist belt 1, symmetric lower limb fixing rings 2, and a bone conduction headset 3, as Figure 2 shown.

[0041] The adjustable waist belt 1 is worn around the human waist and can adjust its fastening degree around the waist according to the human waist circumference. A camera fixing cabin 101 and a rear-end control cabin 102 are integrated on the adjustable waist belt 1. After the adjustable waist belt 1 is worn, the camera fixing cabin 101 and the rear-end control cabin 102 are located in the front and rear of the human body respectively.

[0042] Among them, a depth camera 103 is built into the camera fixing cabin 101. The lens of the depth camera 103 faces forward of the human body and is used to collect point cloud data and depth information in the environment, construct a sparse three-dimensional point cloud map based on ORB-SLAM3, and perform three-dimensional modeling of the scene.

[0043] The rear-end control cabin 102 is built with an embedded processor 104, a differential GPS module 105, and a distributed power system 106. Among them, the embedded processor 104 integrates an environment perception module, a simplified path planning module, and a tactile guidance module, and is used to receive and process the data collected by the depth camera 103 and each module. The differential GPS module 105 provides high-precision positioning data, updates the user's position information in real time, and combines with the environmental information collected by the depth camera 103 to ensure that the simplified path planning module obtains accurate spatial coordinates. The distributed power system 106 is placed symmetrically in a lightweight manner to provide stable power supply for all sensors, processors, and execution modules, and support the system to work continuously for a long time.

[0044] The bone conduction headset 3 is mounted on the human head and is located in the temporal bone area, and is used to provide voice reminders when tactile guidance correction fails.

[0045] The symmetric lower limb fixing ring 2 includes two annular members, which are fixed to the corresponding positions of the left and right thighs of the human body by a strap and can adjust the fastening degree on the thighs according to the thigh circumference of the human body. A spatial vibrator array 107 is installed on the annular member; at the same time, an IMU sensor 108 is installed at the biceps femoris on the opposite side of the spatial vibrator array 107. After the lower limb fixing ring is worn on the thigh, the spatial vibrator array 107 is located on the front side of the thigh, and the IMU sensor 108 is located on the back side of the thigh. The above-mentioned spatial vibrator array 107 and IMU sensor 108 are both connected to the embedded processor 104 through a signal line 109

[0046] Among them, the IMU sensor 108 collects the posture information and acceleration information of the user's legs in real time. The spatial vibrator array 107 is 4 micro linear vibrators installed along the thigh ring side on the annular member, and is evenly arranged from the rectus femoris to the vastus lateralis to form a gradient tactile feedback. Since in the human thigh, the quadriceps femoris, especially the rectus femoris, is more sensitive to the vibration feedback in the front direction, has higher stability during walking, and is not affected by the large flexion and extension of the knee joint, it can be used as a tactile feedback point for the front path reminder. The vastus lateralis can effectively sense the lateral vibration signal, enabling the user to more intuitively judge the direction; in contrast, although the inner thigh is sensitive to vibration, when used as a direction guide, the signal may be interfered by clothing or the friction between the legs. Therefore, in the present invention, for the right leg and the left leg, the 4 micro linear vibrators are evenly distributed from the rectus femoris to the vastus lateralis, and each micro linear vibrator is spaced 30°. The corresponding angles of the four micro linear vibrators J1 to J4 are respectively:

[0047] J1 = 0°, J2 = 30°, J3 = 60°, J4 = 90°.

[0048] The above-mentioned depth camera 103, differential GPS module 105, and IMU sensor 108 are all connected to the embedded processor 104 through a signal line 108. The micro linear vibrator is connected to the embedded processor 104 through the PCA9685 PWM drive module via the I2C bus; at the same time, the embedded processor 104 and the bone conduction earphone 1 perform wireless communication through a Bluetooth module. Furthermore, the embedded processor 104 receives the signals transmitted by each part, and through the built-in environment perception module, refined path planning module, and tactile guidance module, maps the visual signal to a tactile signal in a visual signal limited scenario to achieve adaptive synchronization with the movement rhythm, and completes obstacle avoidance guidance; and provides a voice prompt through the voice assistance module when the user fails to successfully correct the traveling direction through tactile feedback.

[0049] The visual environment perception module is based on the scene model constructed by the depth camera 103, and uses the YOLO algorithm to detect common objects in front of the human body, ensuring accuracy and real-time performance, and synchronously constructing an incremental scene map with semantic annotations.

[0050] The refined path planning module constructs an incremental scene map based on the visual environment perception module, obtains an initial path through a path planning algorithm, and discretizes the initial path into a sparse set of waypoints based on an adaptive path sparse segmentation model to generate a refined path. Specifically:

[0051] First, use a global path planning algorithm to obtain an original discrete path;

[0052] Subsequently, for each intermediate node p i (i = 1, …, n - 1) in the initial path segment, calculate its turning angle θ i ;

[0053]

[0054] Furthermore, an improved RDP algorithm is used for recursive simplification. Specifically, for the starting point p i and the ending point p j (i = 0, j = n) of the current path segment, first calculate the average turning angle θ avg of all intermediate nodes in the path, and then determine the local tolerance ε local .

[0055]

[0056] In the formula, ε min and ε max represent the minimum and maximum allowable errors in high-curvature and low-curvature regions respectively; θ max is a preset maximum turning angle reference value.

[0057] Within the path segment, traverse each intermediate node p m (i < m < j, m = 1, …, n - 1), and calculate the perpendicular distance m from the intermediate node p (starting point and ending point). That is, the deviation distance of each intermediate node p m from the straight line . The calculation formula is:

[0058]

[0059] Find the maximum deviation distance d max and its corresponding index point If d max ≤ ε local , it is considered that the straight line is sufficient to approximate this section of the path, and all intermediate points can be discarded. Otherwise, the path is segmented at , and the same simplification process is recursively executed for the segmented parts respectively.

[0060] Finally, by merging the endpoints of each segment, a simplified polyline path is obtained.

[0061] After the aforementioned path planning, the embedded processor 104 generates a steering angle Δθ in real time according to the angle deviation between the current traveling direction of the user measured by the IMU sensor 108 and the direction of the simplified path; the direction is defined as pointing from the current direction to the target path direction, and the positive direction is set as the counterclockwise direction; at the same time, the IMU sensor 108 obtains the current speed v, angular velocity ω, and acceleration α of the user's legs and transmits them to the embedded processor 104. After low-pass filtering and coordinate alignment for data preprocessing, they are uniformly packed into a steering vector as the intermediate transmission signal of the guidance system of the present invention.

[0062] In the tactile guidance module, the gait cycle is predicted through the gait phase prediction part, including accurately predicting the gait phase change of the user, dynamically positioning the single-leg support phase detection and the regression prediction of the swing phase duration of the other leg. The gait phase prediction part adopts a hybrid network architecture model of Transformer and CNN based on multi-level spatio-temporal context modeling.

[0063] The input X of the CNN part is the preprocessed angular velocity ω and acceleration α in the steering vector.

[0064]

[0065] Among them, is the angular velocity of the left leg in the x, y, and z axes; is the angular velocity of the right leg in the x, y, and z axes; similarly and are the accelerations of the left leg and the right leg in the x, y, and z axes respectively; T = 400 represents a time window of a fixed length, 12 represents the 3-axis angular velocity and 3-axis acceleration of each of the left and right legs, and the data constitutes a multi-channel time series at a fixed frequency. After sliding window segmentation and normalization processing to eliminate noise and dimensional differences, the multi-scale time series feature extraction unit extracts signal features from a multi-scale perspective by using the method of multi-branch parallel 1D convolution.

[0066] In the multi-scale temporal feature extraction unit, three parallel branches are designed to capture motion features at different scales. The high-frequency detail capture branch uses stacked standard 1D convolutions with kernel sizes of 5, 3, 3 in sequence and channel numbers of 64, 128, 256 in sequence to extract local instantaneous motion mutation information; the mid-range dependence modeling branch uses dilated convolution with a kernel size of 3 and a dilation rate set to 2 and a channel number of 128 to capture the periodic pattern of the stride cycle; the global context awareness branch uses deformable convolution with a kernel size set to 5 and a channel number of 64, which can adaptively adjust the receptive field to adapt to the diversity between individual gaits. The outputs of each branch are dynamically weighted and fused through a gated attention mechanism to form a unified feature representation. Provide a basis for subsequent spatio-temporal modeling.

[0067] For the Transformer part based on multi-level spatio-temporal context modeling, the fused feature F CNN is first linearly mapped to the model dimension D model = 512 and added to the learnable positional encoding to provide a representation that contains both motion information and position information for each time step, serving as the input to the encoder in the Transformer. This encoder uses a spatio-temporal separable multi-head attention mechanism (ST-SeparableAttention). Among them, the temporal attention branch F time captures the long-range dependence of gait events through 8-head attention, while the spatial attention branch F space extracts the collaborative relationship between each sensor channel through 4-head attention. The fusion of the two attention outputs is weighted by the learnable parameter α to obtain the joint attention F attn , and its mathematical expression is:

[0068] F attn = α·F time + (1 - α)·F space

[0069] Furthermore, F attn performs local and global feature interaction modeling through an enhanced feed-forward network (FFN) composed of a gated linear unit (GLU) and a depthwise separable convolution. Its calculation process is:

[0070]

[0071] Among them, represents the element-wise multiplication operation; W1 and W2 are learnable weights. After multiple layers of stacking (including residual connections and layer normalization), the context feature

[0072] F Trans = LayerNorm(F attn + FFN(F attn ))

[0073] In the task-driven dynamic two-stream decoder of the Transformer part, two parallel decoding streams are designed, which are used for stance phase detection and swing phase duration regression prediction respectively.

[0074] The stance phase detection stream first takes the output F of the Transformer Trans and inputs it into a bidirectional GRU to obtain smooth temporal features Subsequently, it forms a stance phase probability curve P support (t) ∈ [0, 1] T through a fully connected layer and Sigmoid activation, where t is time.

[0075] Meanwhile, on the basis of the spatio-temporal feature F, the swing phase duration prediction stream Trans uses the short-time Fourier transform (STFT) to extract the frequency domain energy distribution feature where F represents the dimension in the frequency domain. After fusing the temporal and frequency domain features, the swing phase duration T is gradually regressed through a Temporal Convolutional Network (TCN). swing .

[0076] To enhance the consistency between the two tasks, a cross-task interaction gating mechanism is designed to correct the stance phase probability curve P support (t) with the swing phase duration information to obtain Its expression is:

[0077]

[0078] where W g is a learnable weight matrix, and σ is an activation function, ensuring the effective transfer and fusion of cross-task information.

[0079] To achieve accurate phase boundary detection, an adaptive differentiable dynamic threshold segmentation method is adopted, and its threshold calculation formula is:

[0080]

[0081] where μ and are learnable parameters, W is the local window length, and the threshold τ(t) is used to detect the start and end points of the stance phase.

[0082] To meet the actual application requirements of real-time performance and resource constraints, the model is lightweight processed. Structured pruning is performed on the convolutional kernels and attention heads, and at the same time, dynamic 8-bit quantization technology is used for model compression and deployment, so as to achieve low latency in end-to-end system inference. The overall hybrid neural network adopts an end-to-end joint training strategy, and the joint loss function is:

[0083]

[0084] Among them, it includes the support phase detection loss (Combining BCEWithLogitsLoss and TverskyLoss, the parameter settings are α = 0.7, β = 0.3), the swing phase regression loss (HuberLoss, δ is set to 0.1), and the regularization term (including L1 sparse constraint and temporal smoothing constraint) of the multi-objective loss function to ensure that the model can still maintain a high prediction accuracy in the case of data skew.

[0085] The gait phase prediction part of the entire hybrid neural network realizes the dynamic prediction of the gait phase through a dual-task decoding mechanism. On the one hand, the fully connected layer and the Sigmoid function are used to generate the support phase probability curve, and the start and end time points of the support phase are determined through dynamic threshold segmentation, and the vibration stimulation trigger moment in the middle of the support phase is located at the probability peak; on the other hand, the swing phase duration is predicted through a parallel regression branch to provide a time constraint for the vibration stimulation. The triggering logic follows that the vibration stimulation starts from the middle of the single-leg support phase to ensure the maximization of the user's perception effect, and the duration is dynamically adjusted according to the swing phase duration of the other leg predicted in real time to ensure that it does not exceed 60% of the gait cycle swing phase (that is, the stage when the other foot leaves the ground and is ready to step), as Figure 4 shown.

[0086] The embedded processor 104 generates a steering vector containing gait information and steering angle according to the forward direction and the refined path pointing, and iteratively optimizes it in real time to achieve real-time matching of the user's walking direction and the target direction

[0087] At the same time, the tactile guidance module also adopts a dual-mode vibration coding strategy to achieve direction guidance and emergency obstacle avoidance according to the obstacle distance information collected by the depth camera 103, as Figure 5 shown.

[0088] Among them, when the embedded processor 104 detects that the obstacle distance is greater than the safety threshold, the direction guidance mode is enabled. According to the spatial activation sequence of the linear vibrator array 107 generated by the tactile guidance module, a PWM control signal is generated to control the vibration of the linear vibrator array 107, which is triggered periodically to provide vibration stimulation to the leg to represent the target direction.

[0089] The spatial activation sequence is generated based on the priority gradient intensity field haptic feedback algorithm, as Figure 6 shown, specifically:

[0090] The embedded processor 104 determines the vibration mode to be adopted according to the current steering angle deviation Δθ of the user, specifically:

[0091] When |Δθ| > 10°, it enters the direction adjustment mode for further judgment: when 10° < Δθ < 90°, the linear vibrator array 107 on the left leg is triggered to guide the user to adjust the traveling direction to the left; when -90° < Δθ < -10°, the linear vibrator vibration 107 array on the right leg is triggered to guide the user to adjust the traveling direction to the right. In the direction adjustment mode, according to the "take the larger" principle, the haptic guidance module defines the vibration sequence as:

[0092] k = min{i ∈ {1, 2, 3, 4}: J i ≥ |Δθ|}

[0093] and activates the vibration units J1, J2, …, J k . For example, when Δθ = 50°, since J3 = 60° ≥ 50°, k = 3 is taken, and the activated vibration units are J1, J2, and J3. The activated vibration units vibrate in ascending order of strength from J1 to J k , and the vibration intensity I i is distributed according to the following formula:

[0094]

[0095] where I base is the basic vibration intensity and ΔI is the maximum intensity increment.

[0096] When |Δθ| ≤ 10°, the embedded processor 104 determines that the user deviation is within the normal error range and triggers the concentric contraction mode. At this time, the linear vibrator arrays 107 on both legs are activated simultaneously, and the vibration units are activated in sequence from the outside to the inside, that is, the activation order is

[0097] J4 → J3 → J2 → J1

[0098] The vibration intensity is distributed according to the following formula:

[0099]

[0100] where order(1) = J4, order(2) = J3, order(3) = J2, order(4) = J1.

[0101] Through the above spatial activation sequence, the haptic guidance module can dynamically select and activate the corresponding vibration units according to the user's turning angle deviation, and guide the user to adjust the traveling direction or maintain the current path in a vibration feedback manner from weak to strong.

[0102] The embedded processor 104 also judges the user's movement speed v at the same time. When the user's movement speed v user ≥ 2m / s, the haptic guidance module activates the forward-looking guidance mode. In this mode, before the vibration stimulation trigger moment in the middle phase according to the aforementioned spatial activation sequence, the vibration signal is triggered in advance to help the user adjust the movement path in advance. The vibration trigger timing is:

[0103] Δt lead =0.3·v user +0.1·a user

[0104] Where, Δt lead is the time of advance trigger, v user is the user's movement speed, and a user is the user's acceleration. Ensure that when the user moves quickly, the haptic guidance module can react in advance to help the user obtain navigation information earlier and avoid potential obstacle threats.

[0105] When the embedded processor 104 detects that an obstacle is approaching and the distance is less than the safety threshold, the haptic guidance module activates the emergency obstacle avoidance mode. In this mode, the full array of high-frequency vibrations is immediately excited to attract the user's attention as quickly as possible and prompt him to take emergency braking measures.

[0106] The voice assistance module is used to provide voice prompts such as "You have deviated from the course. Please correct to the left", "You have deviated from the course. Please correct to the right" to assist the user in path correction when the user fails to successfully correct the traveling direction through haptic feedback. This module judges whether to activate the voice prompt based on the turning angle Δθ. When the expected correction threshold is exceeded, that is, |Δθ|>90°, the system will start the voice prompt, which is transmitted by the embedded processor 104 to the bone conduction earphone 1 through the Bluetooth module to ensure that the user can obtain effective navigation assistance.

[0107] Various modifications of the above embodiments (such as increasing or decreasing the number of vibration motors, replacing the IMU model, etc.) are obvious to those skilled in the art. Any equivalent transformation based on the technical solution of this application falls within the protection scope of the claims of this application.

Claims

1. A wearable path guidance system for visual-tactile cross-modal mapping, characterized in that: It includes visual-tactile cross-modal mapping wearable devices and visual environment perception modules, streamlined path planning modules, tactile guidance modules and voice assistance modules; The visual-tactile cross-modal mapping wearable device includes an adjustable waist belt tied around the waist of the human body, lower limb fixing rings symmetrically tied to the left and right thighs of the human body, and bone conduction headphones mounted on the ears of the human body; The adjustable waist strap is equipped with a depth camera and a control pod, which contains an embedded processor, a differential GPS module, and a power supply. The embedded processor integrates a visual environment perception module, a streamlined path planning module, and a tactile guidance module. The lower limb fixation ring consists of two rings, on which a spatial vibrator array is installed. An IMU sensor is also installed on the biceps femoris on the opposite side of the spatial vibrator array. The spatial vibrator array consists of four miniature linear vibrators installed on the ring along the thigh ring side to provide vibration stimulation to the thigh. The visual environment perception module builds a scene model based on the depth camera and simultaneously builds an incremental scene map with semantic annotations; The streamlined path planning module, based on the incremental scene map constructed by the visual environment perception module, obtains an initial path through a path planning algorithm. Based on the adaptive path sparse segmentation model, the initial path is discretized into a sparse set of waypoints to generate a streamlined path. The embedded processor further generates a steering angle Δθ in real time based on the angular deviation of the user's current direction of travel relative to the streamlined path direction measured by the IMU sensor. At the same time, the IMU sensor obtains the user's current leg velocity v, angular velocity ω, and acceleration α, which are transmitted to the embedded processor. After low-pass filtering and coordinate alignment for data preprocessing, the data is uniformly packaged into a steering vector as an intermediate transmission signal. The tactile guidance module is designed with a gait phase prediction unit, which uses a hybrid Transformer and CNN network architecture model based on multi-level spatiotemporal context modeling to accurately predict the user's gait cycle, including accurately predicting the user's gait phase changes, dynamically locating the single-leg stance phase detection and the other leg's swing phase duration regression prediction. The vibration stimulation trigger is set to start from the middle of the stance phase, and the duration is dynamically adjusted according to the real-time predicted swing phase duration to ensure that it does not exceed 60% of the swing phase time of the gait cycle. Furthermore, the embedded processor generates a steering vector containing gait information and steering angle based on the forward direction and the streamlined path direction, and iteratively optimizes it in real time to achieve real-time matching of the user's travel direction with the target direction. At the same time, the tactile guidance module also uses a dual-mode vibration encoding strategy to achieve directional guidance and emergency obstacle avoidance based on the obstacle distance information collected by the depth camera; When the embedded processor detects that the distance to an obstacle is greater than the safety threshold, the direction guidance mode is activated. The tactile guidance module generates a spatial activation sequence for the linear vibrator array. Based on the user's steering angle deviation, the corresponding vibration unit is dynamically activated and periodically triggered in a vibration feedback mode from weak to strong, guiding the user to adjust the direction of travel or maintain the current path. The embedded processor also judges the current user's movement speed v. When the user's movement speed v user ≥ 2 m / s, the tactile guidance module activates the prospective guidance mode. In this mode, before the vibration stimulation trigger moment in the middle phase according to the aforementioned spatial activation sequence, the vibration signal is triggered in advance; When the embedded processor detects that an obstacle is approaching and the distance is less than the safety threshold, the tactile guidance module activates the emergency obstacle avoidance mode, which immediately stimulates high-frequency vibration of the entire array; The voice assistance module is used to provide voice prompts when the user fails to successfully correct the direction of travel through tactile feedback, which are transmitted by the embedded processor to the bone conduction headset via the Bluetooth module to assist the user in path correction.

2. The wearable path guidance system for visual-tactile cross-modal mapping according to claim 1, characterized in that: Four miniature linear vibrators are evenly arranged from the rectus femoris to the vastus lateralis, with each miniature linear vibrator spaced 30° apart; at the same time, the IMU sensor is located on the biceps femoris on the opposite side of the spatial vibrator array.

3. The wearable path guidance system for visual-tactile cross-modal mapping according to claim 1, characterized in that: The path planning method of the simplified path planning module is: First, the original discrete path is obtained using the global path planning algorithm; Then, for each intermediate node p in the initial path segment i (i=1,…,n-1), calculate its rotation angle θ i ; Further, an improved RDP algorithm is used for recursive simplification. For the starting point p of the current path segment i and the ending point p j (i = 0, j = n), first calculate the average turning angle θ of all intermediate nodes in the path avg , and then determine the local tolerance ε local ; Where, ε min With ε max are the minimum and maximum allowable errors in high curvature and low curvature regions, respectively; θ max is the preset maximum turning angle reference value; Within the path segment, traverse each intermediate node p m (where i < m < j, m = 1, …, n - 1), calculate the intermediate node p m to the straight line vertical distance which is the deviation distance of each intermediate node p m to the straight line The calculation formula is as follows: Find the maximum deviation distance d max and its corresponding index point If d max ≤ε local , it is considered that the straight line is sufficient to approximate this section of the path, and all intermediate points are discarded; otherwise, the path is segmented at and the same simplification process is recursively executed for each segmented part; Finally, a simplified polyline path is obtained by merging the endpoints of each segment.

4. The wearable path guidance system for visual-tactile cross-modal mapping according to claim 1, characterized in that: The spatial activation sequence is generated based on the priority gradient intensity field tactile feedback algorithm, specifically: The embedded processor determines the vibration mode to be used based on the user's current steering angle deviation Δθ: When |Δθ|>10°, the system enters the direction adjustment mode and makes further judgments: when 10°<Δθ<90°, the linear vibrator array on the left leg is triggered to guide the user to adjust the direction of travel to the left; when -90°<Δθ<-10°, the linear vibrator array on the right leg is triggered to guide the user to adjust the direction of travel to the right. In the direction adjustment mode, the tactile guidance module defines the vibration mode as follows: k = min{i ∈ {1, 2, 3, 4}: J i ≥ |Δθ|} and activate the vibration units J1, J2, …, J k ; the activated vibration units vibrate from J1 to J k in the order from weak to strong, and their vibration intensity I i is distributed according to the following formula: where I base is the basic vibration intensity and ΔI is the maximum intensity increment; When |Δθ|≤10°, the embedded processor determines that the user's deviation is within the normal error range and triggers the concentric contraction mode. At this time, the linear vibrator arrays on both legs are activated simultaneously, and the vibration units are activated in sequence from the outside to the inside. That is, the activation order is: J4→J3→J2→J1 The vibration intensity is distributed as follows: Among them, order(1)=J4, order(2)=J3, order(3)=J2, order(4)=J1.

5. The wearable path guidance system with visual-tactile cross-modal mapping according to claim 1, characterized in that: In the gait phase prediction part, the input X of the CNN part is the preprocessed angular velocity ω and acceleration α in the steering vector; the data is formed into a multi-channel time series at a fixed frequency, and sliding window segmentation and normalization are performed to eliminate noise and dimensional differences. The multi-scale time series feature extraction unit uses a multi-branch parallel 1D convolution method to extract signal features from multi-scale perspectives; The multi-scale temporal feature extraction unit is designed with three parallel branches to capture motion features at different scales. The high-frequency detail capture branch uses stacked standard 1D convolutions to extract local instantaneous motion mutation information. The mid-range dependency modeling branch uses dilated convolutions to capture periodic patterns across gait cycles. The global context perception branch uses deformable convolutions to adaptively adjust the receptive field to accommodate the diversity of individual gaits. The outputs of each branch are dynamically weighted and fused through a gated attention mechanism to form a unified feature representation F. CNN ; For the Transformer part based on multi-level spatiotemporal context modeling, the fusion feature F CNN First, it is linearly mapped to the model dimension D model And add it to the learnable position code P to provide a representation containing both motion information and position information for each time step as the input of the encoder in the Transformer; the encoder adopts a spatiotemporal separation multi-head attention mechanism, in which the time attention branch F time Capturing the long-range dependencies of gait events through 8-head attention; spatial attention branch F space The synergistic relationship between sensor channels is extracted through four attention heads; the fusion of the two attention outputs is weighted by the learnable parameter α to obtain the joint attention F attn ; Furthermore, F attn The enhanced feed-forward network FFN composed of gated linear units and depthwise separable convolutions is used to model the interaction between local and global features. After multiple layers of stacking, the context feature F is finally generated Trans ; In the task-driven dynamic two-stream decoder of the Transformer part, two parallel decoding streams are designed, which are respectively used for stance phase detection and swing phase duration regression prediction; among them, the stance phase detection stream first takes the output F of the Transformer Trans as the input of a bidirectional GRU to obtain smooth temporal features F GRU , and then forms the stance phase probability curve P support (t) through a fully connected layer and Sigmoid activation; Meanwhile, the swing phase duration prediction flow extracts the frequency domain energy distribution feature F based on the spatio-temporal feature F Trans using the short-time Fourier transform, and after fusing the time domain and frequency domain features, gradually regresses through the TCN to obtain the swing phase duration T freq . swing .

6. The wearable path guidance system for visual-tactile cross-modal mapping according to claim 5, characterized in that: A cross-task interactive gating mechanism is designed in the task-driven dynamic dual-stream decoder to enhance the consistency between the two tasks. The support phase probability curve P is adjusted by the swing phase duration information. support (t) is corrected to obtain Its expression is: Among them, W g is a learnable weight matrix, and σ is an activation function, ensuring the effective transmission and fusion of cross-task information.

7. The wearable path guidance system for visual-tactile cross-modal mapping according to claim 5, wherein: In the Transformer part, an adaptive differentiable dynamic threshold segmentation method is used to achieve accurate phase boundary detection; the threshold calculation formula is: where μ and are learnable parameters, W is the local window length, and the threshold τ(t) is used to detect the start and end points of the stance phase.

8. The wearable path guidance system with visual-tactile cross-modal mapping according to claim 5, characterized in that: The Transformer and CNN hybrid network architecture model based on multi-level spatiotemporal context modeling adopts an end-to-end joint training strategy, and the joint loss function is: Among them, it includes the support phase detection loss swing phase regression loss and a regularization term The multi-objective loss function ensures that the model maintains a high prediction accuracy in the case of data skew.

9. The wearable path guidance system with visual-tactile cross-modal mapping as claimed in claim 1, characterized in that: In the forward-looking guidance mode, the vibration triggering timing is: Δt lead = 0.3·v user + 0.1·a user Where Δt lead is the time of advance triggering, v user is the user's movement speed, a user is the user's acceleration.

10. The wearable path guidance system with visual-tactile cross-modal mapping according to claim 1, characterized in that: The voice assistance module determines whether to activate voice prompts based on the steering angle Δθ. When it exceeds the expected correction threshold, that is, |Δθ|>90°, the system will start the voice prompt.

Citation Information

Cited By

  • High-speed refreshing positioning algorithm and system based on wireless signal processing

    CN120568465A

  • Auxiliary eye movement spatial directional prompt vibration annular equipment system

    CN121255029A

  • Stride frequency guiding method and wearable device

    CN122460931A