Robot visual navigation system and method based on Transform and dynamic feature optimization

Through the Transformer-based dynamic feature optimization method, the problems of limited generalization ability and high computing resource consumption of visual navigation technology in complex scenes are solved, and high-precision path planning and stable motion control are achieved.

CN120668144APending Publication Date: 2025-09-19CHONGQING RES INST OF HARBIN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510911547.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing visual navigation technology has limited generalization capabilities in complex and changeable real-world scenarios, consumes a lot of computing resources, and the model has different inference times when processing observation images of different qualities, resulting in poor navigation performance, delay accumulation, and incorrect waypoint selection.

Method used

A Transformer-based dynamic feature optimization method is adopted. Through image acquisition, feature matching and path point optimization modules, combined with non-local mean denoising, improved ORB feature extraction, and RANSAC algorithm to filter mismatched points, an image-path prediction handshake synchronization mechanism is introduced, and the path point optimization module is used for outlier detection and cluster analysis to ensure the accuracy and real-time performance of path point selection.

Benefits of technology

It significantly improves the positioning accuracy and navigation performance of visual navigation, solves the problems of large errors, accumulated delays and incorrect path point selection, and achieves more efficient path planning and stable motion control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120668144A_ABST
    Figure CN120668144A_ABST
Patent Text Reader

Abstract

The invention discloses a robot visual navigation system and method based on Transform and dynamic feature optimization. The robot visual navigation system comprises an image acquisition module, an image feature matching module, a path point prediction module, a path point optimization module and a motion control calculation module. According to the invention, by introducing Transform and a dynamic feature optimization mechanism, high-quality processing and accurate modeling are carried out on image data; non-local mean denoising is adopted, ORB feature extraction and matching are improved, and mismatching points are filtered in combination with an RANSAC algorithm; a handshake synchronization mechanism between an image and path prediction is designed, and it is ensured that an image used for path point prediction is the latest collected image all the time by comparing timestamps; through outlier detection, clustering and geometric center calculation, the system can optimize path point selection in real time according to environmental changes, invalid paths are automatically avoided, and the continuity, rationality and safety of navigation paths are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a visual navigation system, and in particular to a robot visual navigation system and method based on Transformer and dynamic feature optimization. Background Art

[0002] Visual navigation technology is an environmental perception and autonomous navigation technology based on visual sensors (such as cameras). It is widely used in robotics, autonomous driving, drones, and other fields. Its core is to acquire environmental images through visual sensors and process them using computer vision algorithms to achieve functions such as positioning, mapping, and path planning. Compared with traditional lidar or ultrasonic navigation systems, pure visual navigation technology offers advantages such as low cost, simple hardware requirements, and rich information. It can capture the texture and color of the environment through monocular, binocular, or multi-camera systems, providing navigation systems with more comprehensive environmental perception capabilities.

[0003] The development of visual navigation technology relies primarily on advances in computer vision and artificial intelligence. A common visual navigation algorithm involves mapping point cloud information before localizing the robot. Visual SLAM (Simultaneous Localization and Mapping) is one of the core technologies of visual navigation. Visual SLAM uses a camera to capture environmental images and, using feature extraction, matching, and optimization algorithms, estimates the robot's position and constructs a map of the environment in real time. Furthermore, deep learning-based visual navigation technology has also been widely used in recent years. Deep learning models, particularly convolutional neural networks (CNNs), can automatically extract important visual features from images, helping with tasks such as object recognition, obstacle detection, and target tracking. Semantic segmentation and target detection technologies further enhance the environmental understanding capabilities of visual navigation systems, enabling them to identify specific objects and terrain in complex scenes, thereby assisting in decision-making and path planning.

[0004] Deep learning-based visual navigation technology utilizes models such as convolutional neural networks (CNNs) to automatically extract visual features from images, enabling functions such as object recognition, obstacle detection, and target tracking. However, it is highly data-dependent and requires a large amount of labeled data for training. In practice, obtaining high-quality labeled data is expensive. Furthermore, its generalization capabilities are limited, and deep learning models can significantly degrade in environments outside the distribution of their training data, making them difficult to adapt to complex and changing real-world scenarios.

[0005] Semantic segmentation technology classifies each pixel in an image into different semantic categories (such as roads, obstacles, and buildings), while object detection technology identifies specific objects in an image (such as doors, windows, and tables). While these technologies can enhance the navigation system's understanding of the environment and assist with path planning and decision-making, they consume significant computing resources. Semantic segmentation and object detection typically require high computing resources, making them difficult to run in real time on low-power devices. Furthermore, their adaptability to complex scenes is limited. In scenes with drastic lighting changes, severe occlusion, or diverse object shapes, the accuracy of semantic segmentation and object detection can significantly decrease.

[0006] like Figure 1 As shown in the figure, NoMaD is a new architecture for robot navigation in unknown environments. It uses a unified diffusion strategy to achieve exploratory tasks and goal-oriented behavior tasks. First, it is necessary to record the topological path. The images observed during the robot's movement are saved as a topological node every 1 second to obtain the topological map under the path. As the robot moves in the environment, it continuously obtains RGB visual observations of the current position. These observation images are encoded using the visual encoder in NoMaD and converted into feature vectors. The newly obtained image feature vectors are matched with the feature vectors of existing nodes in the topological map through a pre-trained Transformer model. If the similarity exceeds a certain threshold, the current observation is considered to be similar to the scene represented by an existing node, and the position information of the node can be used as the reference position of the robot. Through the pre-trained diffusion strategy model, combined with the estimated position of the nearest topological node and the target node position, multiple sets of path points are output, and the best path point is selected from them by a hyperparameter to guide the next movement of the robot. This method has the following problems: (1) When selecting path points obtained by model inference, an uncontrollable hyperparameter is used for selection, which may result in the selection of incorrect path points.

[0007] (2) The model recognition effect is poor, and during the navigation process, errors may occur when distinguishing similar scenes.

[0008] (3) Since the model has different inference times when processing observation images of different qualities, delay accumulation occurs during long-term operation, which in turn leads to poor navigation performance and inability to navigate based on the currently observed images. Summary of the Invention

[0009] The present invention provides a robot vision navigation system and method based on Transformer and dynamic feature optimization, which can predict path points based on images transmitted by a visual camera for navigation.

[0010] The purpose of the present invention is achieved through the following technical solutions: A robot visual navigation system based on Transformer and dynamic feature optimization includes an image acquisition module, an image feature matching module, a path point prediction module, a path point optimization module, and a motion control calculation module, wherein: The image acquisition module is responsible for acquiring and preprocessing high-quality images to provide clear, low-noise image input for subsequent visual calculations; The image feature matching module is responsible for extracting and matching features from the images collected by the image acquisition module, and combining the feature filtering mechanism to achieve accurate estimation of the robot's current position; The waypoint prediction module is responsible for generating a set of waypoint candidates based on the latest observation images collected by the image acquisition module through a pre-trained Transformer model, and predicting the target waypoints in combination with the topological matching results; The path point optimization module is responsible for performing outlier detection, smoothing and cluster analysis on the candidate path point set to extract the most representative optimized path points; The motion control calculation module is responsible for calculating the linear velocity and angular velocity based on the target path point and the current state, and performing limiting processing to achieve stable and accurate motion control.

[0011] A robot visual navigation method based on Transformer and dynamic feature optimization using the above system includes the following steps: Step 1: Image acquisition: Step 1-1, image acquisition: The image acquisition module obtains raw image data through the RGB camera; Step 1-2, image processing: The image acquisition module uses the non-local means algorithm to denoise the collected images; Step 2: Image feature matching: Step 2-1, image queue management: the image feature matching module receives the image generated by the image acquisition module and saves it in an image queue; Step 2-2, grayscale conversion and noise removal: Convert the color image to a grayscale image, use Gaussian filtering to remove noise in the grayscale image, and retain important edge information; Step 2-3, key point and feature extraction: Use the ORB algorithm to extract the key points of the image and their corresponding feature descriptors. The specific steps are as follows: Step 2-3-1: Use the improved FAST-9 detector to detect key points in each scale space of the image pyramid, and select key points with significant features through non-maximum suppression; Step 2-3-2, use the intensity centroid method to calculate the main direction of the key point and give the feature point rotation invariance; Step 2-3-3: Construct a rotation-invariant rBRIEF descriptor and generate a 256-bit binary feature vector through the predefined 512 pairs of position sampling points; Step 2-3-4: Optimize the feature descriptor, including sampling point correlation optimization and variance constraint processing; Step 2-4, feature matching: Use the FLANN algorithm to perform feature matching and obtain the feature matching matrix. The specific steps are as follows: Step 2-4-1, build a multi-layer KD-tree index structure to perform spatial division of feature descriptors; Step 2-4-2: Use the approximate nearest neighbor search algorithm to quickly match in the feature space; Step 2-4-3: Calculate the Hamming distance of the feature point pairs and establish the initial matching relationship; Step 2-4-4, generate a feature matching matrix containing the coordinate information of the matching point pairs; Step 2-5, false matching filtering: Use the RANSAC algorithm to filter out false matching features and retain the correct features; Step 2-6, optimal topological point selection: select the topological point that is most similar to the current image based on the matching results as the best estimate of the current position; Step 3: Path point prediction: Step 3-1: The path point prediction module receives images collected in real time by the image acquisition module and saves these images in a time sequence in an image queue; Step 3-2: The handshake module ensures that the images in the image queue are consistent with the currently observed images through real-time timestamp comparison and process control mechanism. The specific steps are as follows: Step 3-2-1, the handshake module monitors the timestamp of the image acquisition module in real time and compares it with the timestamp of the latest image in the image queue; Step 3-2-2: If the two are consistent, it indicates that the image in the image queue is the latest observed image. The handshake module allows the waypoint prediction module to extract the image and related data from the image queue and input it into the pre-trained Transformer model for processing. If the two are inconsistent, the handshake module suspends the operation of the waypoint prediction module and waits for the image queue to be updated to the latest observed image before restarting the prediction process. Step 3-3: Input the images and data in the image queue into the pre-trained Transformer model to obtain multiple path points corresponding to each image; Step 3-4: Use the most similar topological points obtained by the image feature matching module to determine the target point based on the number of features and the set threshold. The specific steps are as follows: Step 3-4-1: The image feature matching module extracts features from the current scene and matches them with the pre-stored topological map to find the topological point that is most similar to the current scene; Step 3-4-2: Determine the selection rule for the next target point based on the number of features of the topological point and the preset threshold: If the number of features of the topological point is less than the set threshold, it indicates that the reliability of the currently matched topological point is low, and the path point corresponding to the current minimum distance is selected as the next target point, and the topological point is updated to the current minimum distance node; If the number of features of the topological point is greater than or equal to the set threshold, it indicates that the matching result is reliable, and the path point of the next node is directly selected as the next target point, and the current node is updated to the next node; Step 4: Path point optimization: Step 4-1: The path point optimization module uses the Z-score method to detect and remove outliers from the multiple path points output by the path point prediction module. The specific steps are as follows: Step 4-1-1: The path point optimization module performs multi-level optimization processing on the candidate path point set output by the path point prediction module, and detects and filters outliers; Step 4-1-2: Analyze the distribution characteristics of the path point set based on statistical principles, establish a data distribution model, use a dynamic threshold mechanism to automatically identify abnormal path points that deviate from the main distribution area, perform iterative outlier removal operations, and retain a subset of path points that conform to the distribution law; Step 4-2: Preprocess the path points: Use the moving average method to smooth them to remove noise; Step 4-3: Use the K-means clustering algorithm to cluster the preprocessed path points and take the center cluster as the best path point. The specific steps are as follows: Step 4-3-1: By analyzing the spatial distribution density of path points, the silhouette coefficient method is used to automatically determine the optimal number of clusters for K-means clustering; Step 4-3-2: After determining the number of clusters, use the K-means clustering algorithm to group the waypoints and divide them into different clusters according to their characteristics, so as to identify potential path patterns; Step 4-3-3: Evaluate the quality of each path cluster based on the density of points within the cluster and select the most representative path cluster; Step 4-3-4: Calculate the geometric center of the optimal path cluster and use it as the final optimized path point; Step 5: Motion control calculation: Step 5-1: The motion control calculation module calculates the distance error and heading error between the current position and the target path point by analyzing the optimal path point; Step 5-2: Use the PID control algorithm to calculate the linear velocity and angular velocity respectively. The linear velocity is adjusted by the weighted sum of the proportional, integral, and differential terms. The angular velocity is adjusted based on the predicted heading information when approaching the target point. Otherwise, the PID control algorithm is also used to adjust the angular velocity to ensure that the vehicle can travel along the ideal trajectory. Step 5-3: Limit the calculated velocity value to ensure that the linear velocity and angular velocity are within the threshold range; Step 5-4: After processing, the adjusted linear velocity and angular velocity are returned to the execution module as control instructions, driving the motion control system to travel along the desired trajectory.

[0012] Compared with the prior art, the present invention has the following advantages: 1. This invention utilizes the Transformer and dynamic feature optimization mechanisms to achieve high-quality image data processing and precise modeling throughout the entire process, from image acquisition, feature extraction, matching, to path prediction. In particular, the use of non-local means denoising, improved ORB feature extraction and matching, and the RANSAC algorithm for filtering mismatches significantly enhance the stability and robustness of visual features, effectively addressing the large errors in traditional RGB visual navigation and achieving superior positioning and navigation results.

[0013] 2. This invention innovatively designs a handshake synchronization mechanism between image and path prediction. By comparing timestamps, it ensures that the image used for pathpoint prediction is always the most recently acquired image, effectively avoiding navigation errors caused by image prediction being out of sync with current observations. This mechanism not only enhances the real-time performance and accuracy of pathpoint predictions, but also significantly improves the overall timing robustness of the system.

[0014] 3. Through outlier detection, clustering and geometric center calculation in the path point optimization module of the present invention, the system can optimize path point selection in real time according to environmental changes, automatically avoid invalid paths, and improve the continuity, rationality and safety of the navigation path. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a schematic diagram of the NoMad architecture; Figure 2 This is a diagram showing the relationship between the modules of the robot vision navigation system based on Transformer and dynamic feature optimization; Figure 3 It is the topological node 1 in the recorded topological map; Figure 4 It is the topological node 2 in the recorded topological map; Figure 5 It is the topological node 3 in the recorded topological map; Figure 6 Schematic diagram of predicted path point 1 output during navigation (yellow dots represent the predicted path point group); Figure 7 Schematic diagram of predicted path point 2 output during navigation (yellow dots represent the predicted path point group); Figure 8 Schematic diagram of predicted path points 3 output during navigation (yellow dots represent the predicted path point group); Figure 9 The test path for navigation (navigate from the starting point to the end point in the direction of the red arrow). DETAILED DESCRIPTION

[0016] The technical solution of the present invention is further described below with reference to the accompanying drawings, but is not limited thereto. Any modification or equivalent replacement of the technical solution of the present invention that does not depart from the spirit and scope of the technical solution of the present invention should be included in the scope of protection of the present invention.

[0017] The present invention provides a robot vision navigation system based on Transformer and dynamic feature optimization, such as Figure 2 As shown, the system includes an image acquisition module, an image feature matching module, a path point prediction module, a path point optimization module, and a motion control calculation module. The detailed process of each module is as follows: (1) Image acquisition module To improve image accuracy and the precision of subsequent image recognition, the present invention designs an image acquisition module for acquiring high-quality images. This module can be divided into two parts: image acquisition and image processing. The image acquisition module acquires raw image data using an RGB camera, which is capable of capturing high-resolution color images, providing high-quality input for subsequent processing. The image processing module uses a non-local means algorithm to denoise the acquired images. This algorithm effectively removes noise while preserving image edges and details, thereby improving the overall image quality. The processed image is saved for use in subsequent steps.

[0018] (2) Image feature matching module To achieve high-precision image feature matching and position estimation, the present invention designs an image feature matching module. This module receives images generated by the image acquisition module and, through a series of processing steps, extracts, matches, and filters image features to ultimately determine the best estimate of the current position. Specifically, this module can be divided into the following main steps: image queue management, grayscale conversion and noise removal, key point and feature extraction, feature matching, false match filtering, and optimal topological point selection.

[0019] The image feature matching module first receives the image generated by the image acquisition module and saves it in an image queue. This queue is used to store image data of multiple consecutive frames to ensure that the system can handle dynamically changing scenes.

[0020] To reduce computational complexity and improve the accuracy of feature extraction, the input color image is first converted to a grayscale image. Then, a Gaussian filter is applied to smooth the grayscale image to remove high-frequency noise and retain important edge information.

[0021] The ORB algorithm is then used to extract keypoints and their corresponding feature descriptors from the image. Keypoint detection is performed at all scales of the image pyramid using an improved FAST-9 detector. Keypoints with significant features are selected through non-maximum suppression. The main orientations of the keypoints are calculated using the intensity centroid method, making them rotationally invariant. A rotationally invariant rBRIEF descriptor is constructed, generating a 256-bit binary feature vector from 512 predefined pairs of positional sampling points. The feature descriptors are optimized, including correlation optimization and variance constraint processing, to improve feature discrimination. The FLANN algorithm is used for feature matching to find corresponding keypoint pairs between images and generate a feature matching matrix. A multi-layer kd-tree index structure is constructed to spatially partition the feature descriptors. An approximate nearest neighbor search (ANN) algorithm is used for fast matching in the feature space. The Hamming distance between feature point pairs is calculated to establish initial matching relationships. A feature matching matrix containing the coordinate information of the matching point pairs is generated. The RANSAC algorithm is used to filter out mismatched features and select a set of consistent inliers from the feature matching results, representing correct matching relationships and retaining the correct features. Finally, the topological point that is most similar to the current image is selected based on the matching results as the best estimate of the current position. Using this module, we can obtain more accurate topological points that are most similar to the current image, thereby better estimating the current position.

[0022] (3) Waypoint prediction module The path point prediction module first receives the images collected in real time by the image acquisition module and saves these images in a chronological order in an image queue; then, the handshake module ensures that the images in the image queue are consistent with the currently observed images through real-time timestamp comparison and process control mechanism, thereby avoiding the problem of inconsistency between the image based on which the prediction is made and the currently observed image due to too long prediction time.

[0023] The specific implementation method of the handshake module includes the following steps: first, the handshake module monitors the timestamp of the image acquisition module in real time and compares it with the timestamp of the latest image in the image queue; if the two are consistent, it indicates that the image in the image queue is the latest observed image, and the handshake module allows the path point prediction module to extract images and related data from the image queue and input them into the pre-trained Transformer model for processing; if the two are inconsistent, the handshake module suspends the operation of the path point prediction module and waits for the image queue to be updated to the latest observed image before restarting the prediction process.

[0024] After receiving the image and data, the pre-trained Transformer model generates multiple path points corresponding to each image. Next, the image feature matching module extracts features from the current scene and matches them with the pre-stored topological map to find the topological point that is most similar to the current scene. Based on the number of features of the topological point and the preset threshold, the system determines the selection rule for the next target point: if the number of features of the topological point is less than the set threshold, it indicates that the reliability of the currently matched topological point is low. The system selects the path point corresponding to the current minimum distance as the next target point and updates the topological point to the current minimum distance node; if the number of features of the topological point is greater than or equal to the set threshold, it indicates that the matching result is reliable. The system directly selects the path point of the next node as the next target point and updates the current node to the next node. Through the synchronization mechanism of the handshake module, the system can ensure that the path point prediction is always based on the latest observed image, thereby effectively avoiding the image inconsistency problem caused by prediction delay, and significantly improving the real-time performance, accuracy and overall robustness of the path point prediction.

[0025] (4) Path point optimization module The waypoint optimization module performs multi-level optimization processing on the candidate waypoint set output by the waypoint prediction module, detects and filters outliers, analyzes the distribution characteristics of the waypoint set based on statistical principles, establishes a data distribution model, uses a dynamic threshold mechanism to automatically identify abnormal waypoints that deviate from the main distribution area, performs iterative outlier removal operations, and retains a subset of waypoints that conform to the distribution law.

[0026] After removing outliers, in order to further improve the continuity and stability of the path point sequence, the moving average method is used to smooth the retained path points to reduce the impact of local fluctuations on subsequent clustering results.

[0027] By analyzing the spatial distribution density of waypoints, the silhouette coefficient method is used to automatically determine the optimal number of clusters for K-means clustering. After determining the number of clusters, the K-means clustering algorithm is used to group waypoints and divide them into clusters based on their characteristics, thereby identifying potential path patterns. The quality of each path cluster is evaluated based on the density of points within the cluster (such as the sum of squared distances within the cluster), and the most representative path cluster is selected. The geometric center of the optimal path cluster is calculated and used as the final optimized path point. This module can eliminate the interference of erroneous path points.

[0028] (5) Motion control calculation module The motion control calculation module first analyzes the optimal pathpoints and calculates the distance error and heading error between the current position and the target pathpoint. The distance error represents the straight-line distance between the current position and the target point, while the heading error represents the deviation between the current heading and the desired heading at the target pathpoint.

[0029] Next, the PID control algorithm is used to calculate the linear velocity and angular velocity respectively. When calculating the linear velocity, the PID controller adjusts the velocity by weighted sum of the proportional term, integral term, and differential term. The proportional term is used to deal with the immediate difference between the current position and the target path point, the integral term is responsible for eliminating long-term system deviations, and the differential term is used to predict future error trends, thereby adjusting the linear velocity more smoothly. For the control of angular velocity, when the target path point is close and the error is small, the motion control calculation module will make fine adjustments based on the predicted heading information to ensure that the vehicle can approach the target point smoothly and accurately. If the target point is far away, or the current heading error is large, the motion control calculation module will still use the PID control algorithm to make adjustments to ensure that the vehicle can travel along the ideal trajectory.

[0030] After calculating the linear and angular velocities, the motion control calculation module limits them to ensure they remain within preset thresholds to prevent control system instability or equipment damage caused by excessive control signals. This limiting effectively prevents overspeeding and oversteering, ensuring system safety and stability during path tracking.

[0031] Finally, the processed and adjusted linear and angular velocities are returned as control instructions to the execution module, driving the motion control system to drive along the desired trajectory. This module can effectively solve the adaptation problem of deploying on quadruped robots.

[0032] Example: The system development platform is the Linux operating system, the GPU is an NVIDIA Jetson Orin NX, the program is written in Python 3.8, and uses the PyTorch 1.12 framework; the algorithm is deployed on the Unitree GO2 quadruped robot of Yushu Technology.

[0033] The visual navigation and motion control implementation process of the quadruped robot GO2 is as follows: First, the image acquisition module acquires the environment image through the RGB camera, and uses the non-local means algorithm to remove noise while retaining the image edges and details, outputting high-quality images; these images are passed to the image feature matching module, which stores the images in a queue, converts the color image into a grayscale image and further removes noise using Gaussian filtering, then uses the ORB algorithm to extract key points and feature descriptors, and then uses the FLANN algorithm to perform feature matching to generate a matching matrix, and uses the RANSAC algorithm to filter out mismatched points, and finally selects the topological point that is most similar to the current image as the best estimate of the current position. Topological point image is as follows Figures 3 to 5 shown.

[0034] Subsequently, the path point prediction module receives images from the image queue and inputs them into the pre-trained model to predict multiple path points corresponding to each image. The target point is selected based on the relationship between the number of features of the most similar topological point obtained by the image feature matching module and the set threshold: if the number of features is less than the threshold, the path point corresponding to the current minimum distance is selected as the target point and the topological point is updated; otherwise, the path point of the next node is selected as the target point and the node is updated.

[0035] Next, the waypoint optimization module optimizes the predicted waypoints: first, the Z-score method is used to detect and remove outliers, then the moving average method is used to smooth the waypoints to remove noise, and finally the K-means clustering algorithm is used to cluster the waypoints and select the center cluster as the optimal waypoint. Figure 3 Corresponding nodes (the predicted path points are as follows Figure 6 As shown), when the robot approaches Figure 4 Corresponding nodes (the predicted path points are as follows Figure 7 As shown), when the robot approaches Figure 5 Corresponding nodes (the predicted path points are as follows Figure 8 As shown), predicted path points and test path (navigation test path as shown Figure 9 The trend is consistent with that shown in Figure 2, which verifies the effectiveness of the matching mechanism.

[0036] Finally, the motion control calculation module calculates the distance error and heading error based on the optimal path point, and uses the PID control algorithm to calculate the linear velocity and angular velocity, respectively. The linear velocity is adjusted through the weighted sum of the proportional, integral, and differential terms. The angular velocity is adjusted based on the predicted heading information when approaching the target point, otherwise the PID control algorithm is also used. The calculated velocity value is clipped to ensure it is within the threshold range, and the adjusted linear and angular velocities are finally returned to control the movement of the GO2 robot. Through the collaborative operation of these modules, the GO2 robot achieves vision-based autonomous navigation and stable motion.

Claims

1. A robot vision navigation system based on Transformer and dynamic feature optimization, characterized by The robot visual navigation system includes an image acquisition module, an image feature matching module, a path point prediction module, a path point optimization module and a motion control calculation module, wherein: The image acquisition module is responsible for acquiring and preprocessing high-quality images to provide clear, low-noise image input for subsequent visual calculations; The image feature matching module is responsible for extracting and matching features from the images collected by the image acquisition module, and combining the feature filtering mechanism to achieve accurate estimation of the robot's current position; The waypoint prediction module is responsible for generating a set of waypoint candidates based on the latest observation images collected by the image acquisition module through a pre-trained Transformer model, and predicting the target waypoints in combination with the topological matching results; The path point optimization module is responsible for performing outlier detection, smoothing and cluster analysis on the candidate path point set to extract the most representative optimized path points; The motion control calculation module is responsible for calculating the linear velocity and angular velocity based on the target path point and the current state, and performing limiting processing to achieve stable and accurate motion control.

2. A robot visual navigation method based on Transformer and dynamic feature optimization using the robot visual navigation system according to claim 1, characterized in that The method comprises the following steps: Step 1: Image acquisition: Step 1-1, image acquisition: The image acquisition module obtains raw image data through the RGB camera; Step 1-2, image processing: The image acquisition module uses the non-local means algorithm to denoise the collected images; Step 2: Image feature matching: Step 2-1, image queue management: the image feature matching module receives the image generated by the image acquisition module and saves it in an image queue; Step 2-2, grayscale conversion and noise removal: Convert the color image to a grayscale image, use Gaussian filtering to remove noise in the grayscale image, and retain important edge information; Step 2-3, key point and feature extraction: Use the ORB algorithm to extract the key points of the image and their corresponding feature descriptors; Step 2-4, feature matching: Use the FLANN algorithm to perform feature matching and obtain the feature matching matrix; Step 2-5, false matching filtering: Use the RANSAC algorithm to filter out false matching features and retain the correct features; Step 2-6, optimal topological point selection: select the topological point that is most similar to the current image based on the matching results as the best estimate of the current position; Step 3: Path point prediction: Step 3-1: The path point prediction module receives images collected in real time by the image acquisition module and saves these images in a time sequence in an image queue; Step 3-2: The handshake module ensures that the images in the image queue are consistent with the currently observed images through real-time timestamp comparison and process control mechanism; Step 3-3: Input the images and data in the image queue into the pre-trained Transformer model to obtain multiple path points corresponding to each image; Step 3-4: Use the most similar topological points obtained by the image feature matching module to determine the target point based on the number of features and the set threshold; Step 4: Optimize path points: Step 4-1: The path point optimization module detects and removes outliers using the Z-score method on the multiple path points output by the path point prediction module; Step 4-2: Preprocess the path points: Use the moving average method to smooth them to remove noise; Step 4-3: Use the K-means clustering algorithm to cluster the preprocessed path points and take the center cluster as the best path point; Step 5: Motion control calculation: Step 5-1: The motion control calculation module calculates the distance error and heading error between the current position and the target path point by analyzing the optimal path point; Step 5-2: Use the PID control algorithm to calculate the linear velocity and angular velocity respectively. The linear velocity is adjusted by the weighted sum of the proportional, integral, and differential terms. The angular velocity is adjusted based on the predicted heading information when approaching the target point. Otherwise, the PID control algorithm is also used to adjust the angular velocity to ensure that the vehicle can travel along the ideal trajectory. Step 5-3: Limit the calculated velocity value to ensure that the linear velocity and angular velocity are within the threshold range; Step 5-4: After processing, the adjusted linear velocity and angular velocity are returned to the execution module as control instructions, driving the motion control system to travel along the desired trajectory.

3. The robot vision navigation method based on Transformer and dynamic feature optimization according to claim 2 is characterized in that The specific steps of steps 2-3 are as follows: Step 2-3-1: Use the improved FAST-9 detector to detect key points in each scale space of the image pyramid, and select key points with significant features through non-maximum suppression; Step 2-3-2, use the intensity centroid method to calculate the main direction of the key point and give the feature point rotation invariance; Step 2-3-3: Construct a rotation-invariant rBRIEF descriptor and generate a 256-bit binary feature vector through the predefined 512 pairs of position sampling points; Step 2-3-4: Optimize the feature descriptor, including sampling point pair correlation optimization and variance constraint processing.

4. The robot vision navigation method based on Transformer and dynamic feature optimization according to claim 2 is characterized in that The specific steps of steps 2-4 are as follows: Step 2-4-1: Construct a multi-layer KD-tree index structure to perform spatial division on the feature descriptors; Step 2-4-2: Use the approximate nearest neighbor search algorithm to quickly match in the feature space; Step 2-4-3: Calculate the Hamming distance of the feature point pairs and establish the initial matching relationship; Step 2-4-4: Generate a feature matching matrix containing the coordinate information of the matching point pairs.

5. The robot vision navigation method based on Transformer and dynamic feature optimization according to claim 2 is characterized in that The specific steps of step 3-2 are as follows: Step 3-2-1, the handshake module monitors the timestamp of the image acquisition module in real time and compares it with the timestamp of the latest image in the image queue; Step 3-2-2: If the two are consistent, it indicates that the image in the image queue is the latest observed image. The handshake module allows the path point prediction module to extract the image and related data from the image queue and input it into the pre-trained Transformer model for processing. If the two are inconsistent, the handshake module suspends the operation of the path point prediction module and waits for the image queue to be updated to the latest observed image before restarting the prediction process.

6. The robot vision navigation method based on Transformer and dynamic feature optimization according to claim 2 is characterized in that The specific steps of steps 3-4 are as follows: Step 3-4-1: The image feature matching module extracts features from the current scene and matches them with the pre-stored topological map to find the topological point that is most similar to the current scene; Step 3-4-2: Determine the selection rule for the next target point based on the number of features of the topological point and the preset threshold: If the number of features of the topological point is less than the set threshold, it indicates that the reliability of the currently matched topological point is low, and the path point corresponding to the current minimum distance is selected as the next target point, and the topological point is updated to the current minimum distance node; if the number of features of the topological point is greater than or equal to the set threshold, it indicates that the matching result is reliable, and the path point of the next node is directly selected as the next target point, and the current node is updated to the next node.

7. The robot vision navigation method based on Transformer and dynamic feature optimization according to claim 2 is characterized in that The specific steps of step 4-1 are as follows: Step 4-1-1: The path point optimization module performs multi-level optimization processing on the candidate path point set output by the path point prediction module, and detects and filters outliers; Step 4-1-2: Analyze the distribution characteristics of the path point set based on statistical principles, establish a data distribution model, use a dynamic threshold mechanism to automatically identify abnormal path points that deviate from the main distribution area, perform iterative outlier removal operations, and retain a subset of path points that conform to the distribution law.

8. The robot vision navigation method based on Transformer and dynamic feature optimization according to claim 2 is characterized in that The specific steps of step 4-3 are as follows: Step 4-3-1: By analyzing the spatial distribution density of path points, the silhouette coefficient method is used to automatically determine the optimal number of clusters for K-means clustering; Step 4-3-2: After determining the number of clusters, use the K-means clustering algorithm to group the waypoints and divide them into different clusters according to their characteristics, so as to identify potential path patterns; Step 4-3-3: Evaluate the quality of each path cluster based on the density of points within the cluster and select the most representative path cluster; Step 4-3-4: Calculate the geometric center of the optimal path cluster and use it as the final optimized path point.