A robot vision navigation method and system based on ontology self-sensing

CN122813883APending Publication Date: 2026-09-25ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611332053.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-31
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]针对现有机器人视觉导航在跨本体部署时动作预测易产生歧义、缺少独立轨迹风险评估以及高风险训练样本不足的问题,本发明提出一种基于本体自感知的机器人视觉导航方法及系统

Benefits of technology

[0017]1)本体几何作为显式条件参与动作预测、距离感知和轨迹纠偏,可降低相同视觉观测下不同本体动作的歧义;2)空间感知与动作预测顺序解耦,能够独立输出逐航点最近障碍物距离并按阈值触发纠偏;3)以多候选角度可行性预测代替单一修正角回归,能够表示同一高风险轨迹存在多个可行修正方向的情形;4)离线轨迹增广从有限真实数据生成高风险轨迹及其监督标签,无需在部署时在线构建占据栅格。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122813883A_ABST
    Figure CN122813883A_ABST
Patent Text Reader

Abstract

The application discloses a kind of robot visual navigation method and system based on ontology self-perception.The method obtains RGB image sequence, depth observation, relative navigation target and robot ontology geometric parameter, and is encoded into shared context Token;Action prediction module reads the context and outputs future action sequence and is converted into local waypoint trajectory;Space perception module combines shared context, trajectory feature and ontology geometric parameter, predicts the distance of each waypoint to the nearest obstacle, and triggers risk perception correction module when the minimum predicted distance is lower than the safety threshold;Risk perception correction module selects the angle of minimum deflection from the original trajectory from the feasible correction angle and rotates the trajectory as a whole.In the training stage, the continuous trajectory scaling and rotation scanning on the occupancy grid generated by the ontology perception are used to generate training data for decoupling the training of space perception and risk perception correction module, thereby improving the safety and adaptability of visual navigation under different ontology sizes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of robot visual navigation, embodied intelligence and cross-ontology learning technology, and particularly relates to a robot visual navigation method and system based on ontology self-perception. Background Technology

[0002] Existing visual navigation methods typically map visual observations directly to future actions or waypoints. These methods can reduce reliance on traditional localization, mapping, and pipeline planning. However, most methods rely primarily on visual information for decision-making and do not explicitly model the differences in length, width, height, and obstacle-crossing capabilities among different robot bodies.

[0003] For the same visual observation, ontologies of different sizes and clearances may correspond to different safety actions. For example, a narrow area that a small ontology can pass through may pose a collision risk to a large ontology. If the training data only covers a small number of fixed platforms, the mapping of observed actions can easily be implicitly bound to the training platform, which may lead to action ambiguity and safety risks when deployed across ontologies.

[0004] Furthermore, real-world navigation data contains limited high-risk trajectories and feasible correction directions, and discrete waypoints themselves struggle to fully represent the geometric relationships between continuous trajectories and environmental obstacles. Therefore, a visual navigation method is needed that can explicitly inject ontological geometric information, independently assess trajectory risk after action prediction, and learn feasible correction directions using offline-generated high-risk samples. Summary of the Invention

[0005] To address the problems of ambiguous motion prediction, lack of independent trajectory risk assessment, and insufficient high-risk training samples in existing robot visual navigation systems deployed across multiple entities, this invention proposes a robot visual navigation method and system based on entity self-awareness. This invention encodes the robot's length, width, height, and maximum obstacle-crossing height as entity geometry tokens, forming a shared context with RGB image sequences, current depth observation, and relative targets. The motion prediction module outputs future actions and local waypoint trajectories, the spatial perception module predicts the distance from each waypoint to the nearest obstacle, and the risk perception and correction module predicts the feasibility of multiple discrete global yaw correction angles when a risk is triggered and selects the feasible angle with the smallest yaw. During the training phase, high-risk trajectories and feasible angle labels are generated by offline entity perception occupying a grid.

[0006] This invention is achieved through the following technical solution: A first aspect of the present invention: a robot visual navigation method based on self-awareness, wherein the robot is equipped with a visual sensor and a depth sensor, and the method is executed in real time by a processor mounted on the mobile robot or a computing device communicating with the mobile robot, comprising the following steps: S1. Obtain the current and historical RGB image sequences of the robot, depth observation, relative navigation target in the robot's body coordinate system, and body geometric parameters; S2. The RGB image sequence, current depth observation, relative navigation target, and ontological geometry parameters are encoded as RGB Token, depth Token, target Token, and ontological geometry Token, respectively, to form a shared context Token; S3. The action prediction module reads the shared context token through one-way cross-attention and decodes it into an action sequence, and then converts the action sequence into local waypoint trajectories and trajectory features. S4. The spatial perception module predicts the distance from each waypoint in the local waypoint trajectory to the nearest obstacle based on the shared context token and trajectory features; when the minimum predicted distance is lower than a preset safety threshold, it is determined to be a high-risk trajectory. S5. For high-risk trajectories, the risk perception and correction module predicts the feasibility probability of each discrete candidate angle within a preset global yaw correction angle interval; selects the candidate angle with the smallest absolute value among the feasible candidates as the correction angle, and rotates the high-risk trajectory as a whole. S6. Convert the local waypoint trajectories that are not judged as high-risk or have been corrected by overall rotation into executable control commands for the robot and output them.

[0007] Specifically, the ontological geometry parameters include at least the robot's length, width, height, and maximum obstacle-crossing height; the ontological geometry token is generated by a mapping network from a four-dimensional vector composed of length, width, height, and maximum obstacle-crossing height, and is used for motion prediction, distance prediction, and risk perception correction.

[0008] Furthermore, the unidirectional cross-attention is configured to allow the action query token to be updated based on the shared context token, while preventing the action query token from updating the shared context token in the reverse direction; the action prediction module outputs the action sequence within the preset prediction time domain in parallel.

[0009] Furthermore, in step S5, the candidate angle with the smallest absolute value among the feasible candidates is selected as the correction angle. Specifically, the yaw correction angle interval is divided into multiple discrete candidate angle intervals, and the feasibility of each candidate angle is monitored by multiple hot tags. During inference, the candidate angle with the smallest deflection relative to the original high-risk trajectory is selected from the set of feasible candidate angles.

[0010] Furthermore, the method also includes a training phase, which comprises a pre-training phase and a fine-tuning phase: During the pre-training phase, data constructed from cross-ontology first-person video is used to train the action prediction branch with behavior cloning loss. In the fine-tuning stage, depth observations, trajectories, and camera intrinsic parameters from multiple consecutive future moments are used to offline fuse local point clouds in the robot's central coordinate system. Ground regions are identified through ground fitting, and obstacle regions are expanded based on ontological geometric parameters to obtain an ontological perception occupancy grid. Training is performed using behavioral cloning loss, waypoint nearest obstacle distance supervision loss, and candidate correction angle feasibility supervision loss from real robot navigation data. The risk trajectory offline augmentation in the fine-tuning stage is performed on the offline constructed ontological perception occupancy grid.

[0011] Specifically, the risk trajectory broadening during the fine-tuning phase includes: Spline smoothing is applied to the original waypoint trajectory; a safety factor is calculated based on the minimum obstacle spacing and trajectory displacement scale; and a uniform scale transformation is performed on all waypoints around the safety factor. Rotate the trajectory after scale transformation at preset angle intervals within a preset angle range; detect whether each candidate trajectory collides on the grid occupied by the body perception, identify the collision candidate as a high-risk trajectory, and mark the corresponding non-collision rotation angle as a multi-hot feasible angle label.

[0012] Specifically, the ontology-aware occupancy grid is constructed in the following way: using depth observations, demonstration trajectories, and camera intrinsic parameters from multiple consecutive future moments, local point clouds in the robot's central coordinate system are fused offline; ground regions are identified through ground fitting, and obstacle regions are expanded according to ontology geometric parameters to obtain the ontology-aware occupancy grid; the point cloud and occupancy grid are not used as online strategy inputs during the deployment phase.

[0013] A second aspect of the present invention: a robot visual navigation system based on ontology self-perception, comprising: The multimodal input encoding module acquires the robot's current and historical RGB image sequences, depth observation, relative navigation target in the robot's body coordinate system, and body geometric parameters; The ontology geometry encoding module encodes the RGB image sequence, current depth observation, relative navigation target, and ontology geometry parameters into RGB Token, depth Token, target Token, and ontology geometry Token, respectively, forming a shared context Token; The action prediction module reads the shared context token through one-way cross-attention and decodes it into an action sequence, then converts the action sequence into local waypoint trajectories and trajectory features. The spatial perception module predicts the distance from each waypoint to the nearest obstacle in the local waypoint trajectory based on the shared context token and trajectory features; when the minimum predicted distance is lower than a preset safety threshold, it is determined to be a high-risk trajectory. The risk perception and correction module predicts the feasibility probability of each discrete candidate angle within a preset global yaw correction angle range for high-risk trajectories; selects the candidate angle with the smallest absolute value among the feasible candidates as the correction angle, and rotates the high-risk trajectory as a whole. The underlying control output module converts local waypoint trajectories that are not judged as high-risk or have undergone overall rotation correction into executable control commands for the robot and outputs them.

[0014] A third aspect of the present invention: an electronic device comprising one or more processors and a memory; the memory being used to store one or more programs, which, when executed by the processor, cause the processor to implement the robot visual navigation method.

[0015] A fourth aspect of the present invention: a computer-readable storage medium having computer instructions stored thereon, wherein the computer instructions, when executed by a processor, implement the robot visual navigation method.

[0016] The beneficial effects of this invention are as follows:

[0017] 1) Ontology geometry, as an explicit condition, participates in action prediction, distance perception, and trajectory correction, which can reduce ambiguity of different ontology actions under the same visual observation; 2) Spatial perception and action prediction are decoupled sequentially, enabling independent output of the nearest obstacle distance at each waypoint and triggering correction according to a threshold; 3) Replacing single correction angle regression with multi-candidate angle feasibility prediction can represent the situation where there are multiple feasible correction directions for the same high-risk trajectory; 4) Offline trajectory augmentation generates high-risk trajectories and their supervision labels from limited real data, without the need to build an occupying grid online during deployment. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the system architecture of the robot visual navigation method based on ontology self-perception of the present invention; Figure 2 This is a schematic diagram of the training data augmentation and trajectory correction process of the present invention. Detailed Implementation

[0019] The present invention will be described below with reference to the accompanying drawings and specific embodiments. These embodiments are used to explain the technical solutions of the present invention and should not be construed as limiting the scope of protection.

[0020] Example 1: like Figure 1As shown, this invention first provides a system architecture for a robot visual navigation method based on self-awareness. In one embodiment, this invention can be deployed on wheeled robots, quadruped robots, or other mobile platforms with RGB visual input, depth visual input, and a low-level speed control interface, where the robot's shape and obstacle-crossing capabilities can be described using geometric parameters. The system inputs the relative target position in the self-awareness coordinate system; this relative target can be updated via an onboard odometry system, and the robot's pose itself is not directly used as input to the policy network.

[0021] 1. Multimodal heterogeneous input encoder The navigation strategy receives multimodal input at time t, including current and historical RGB image sequences. Current depth observation The inputs are relative to the target g and the ontological geometric parameters m. Each input is independently encoded and mapped to a token space of a unified dimension, and then concatenated to form a shared context for reuse by the action prediction, spatial awareness, and risk awareness correction modules. The RGB image sequence is extracted with image patch features by a pre-trained visual encoder. The visual encoder uses a visual Transformer with frozen parameters to form RGB tokens that preserve local spatial information.

[0022] (1) RGB image encoder An RGB image encoder is used to extract environmental appearance and spatial semantic features from consecutive first-person view images. In this embodiment, a pre-trained and frozen DINOv3-S is used as the backbone network of the RGB encoder, which scales the input image to 224×224 and outputs an RGB token sequence while preserving local spatial information; the selection of the backbone network and the input resolution are alternative implementation methods.

[0023] (2) Depth map encoder The depth map encoder receives the current depth observation at time t. This is used to extract depth tokens related to local geometry. In this embodiment, a ConvNeXt-T network initialized with pre-trained weights is used as the backbone network, and the input depth map is also scaled to 224×224. The physical scale of the depth map has different requirements at different training stages: the modalities (depth, trajectory, and camera intrinsics) in the pre-training data only need to maintain consistent scales; the fine-tuning data uses a depth consistent with the real physical scale. ; Camera intrinsics are not part of the above deployment phase strategy status. This is an essential component for automatic cross-ontology video annotation and point cloud backprojection and multi-frame fusion during offline risk trajectory augmentation. In this embodiment, the RGB branch and depth branch each output 64 384-dimensional tokens.

[0024] (3) Body geometric encoder The body geometry parameter m is a four-dimensional continuous vector, whose components represent the robot's length. ,width ,high and maximum obstacle-crossing height : The vector is encoded into an ontology geometry token by a mapping network and used as conditional information in the three modules of action prediction, spatial perception, and risk perception correction.

[0025] (4) Relative target encoder The relative target g is the two-dimensional target position in the robot's body coordinate system, which is defined as follows: The target encoder maps g to a target token, which, together with the RGB token, depth token, and ontological geometry token, forms a shared context. The target position is updated as the robot moves based on the onboard odometry.

[0026] 2. Decoupling process of prediction, perception and correction.

[0027] The multimodal heterogeneous input encoder projects multimodal tokens onto a unified feature space, forming a shared context. In this embodiment, the RGB Token is 64×384, the depth Token is 64×384, and the target Token and ontology geometry Token are each 1×384, thus the shared context contains a total of 130 384-dimensional tokens. Action queries read the shared context through unidirectional cross-attention; this directional constraint is used to prevent intermediate states of action branches from changing the shared context in reverse, enabling subsequent spatial awareness and risk awareness correction modules to reuse relatively independent environment, target, and ontology information.

[0028] Phase 1: Motion Prediction In this embodiment, the multimodal heterogeneous input encoder employs a 6-layer, 384-dimensional Transformer and sets an action query token corresponding to the prediction time domain H. A one-way attention mask allows the action query token to be updated based on the shared context, but does not allow the action query token to update the shared context; the action prediction module then generates action representations in the prediction time domain in parallel.

[0029] The updated action representation is decoded by the action prediction head into an action sequence for the next H time steps. Each of the actions Including linear velocity and angular velocity In this embodiment, H=8, and the action prediction head uses a two-layer MLP. The action sequence is converted into a two-dimensional waypoint trajectory in the body's local coordinate system according to the robot's kinematics, and further encoded into trajectory features. This is for use by subsequent modules.

[0030] Phase Two: Spatial Perception The spatial awareness module is positioned after the action prediction module and is decoupled from the action prediction output sequence. This module receives the shared context updated with shared coding parameters from the action prediction branch. (i.e., the result obtained after the shared context output by the encoder is updated with encoding parameters by the action prediction module) and trajectory features It learns the relationship between ontological geometry, the current environment, and the predicted trajectory, and outputs the nearest obstacle distance for each waypoint in the prediction time domain: ; ;in, Features of spatial perception Indicates the prediction of the first time domain The predicted distance from each waypoint to the nearest obstacle To predict the temporal length. Since this task requires finer-grained geometric information, its training loss is allowed to update the encoding parameters of the shared context. It should be noted that the loss function of the spatial awareness module is allowed to backpropagate along the gradient to the encoding parameters of the shared context to adjust the representation of the shared context to better serve the fine-grained geometric awareness task; while the unidirectional cross-attention mechanism of the action prediction branch only prevents the action query token from back-updating the shared context, and the two do not conflict.

[0031] Phase Three: Correcting Risk Perception When predicting the distance vector The minimum value is lower than the safety threshold. At this time, the current local waypoint trajectory is identified as a high-risk trajectory. The risk perception and correction module then receives the shared context. Spatial perception features (i.e., intermediate features of stage two output) and trajectory features .

[0032] This module does not directly regress a single continuous angle, but instead predicts the feasibility probability of each discrete candidate angle within a preset global yaw angle range. As a non-restrictive example, [-45°, 45°] can be discretized at 5° intervals, and the feasibility probability of each candidate angle can be output using the Sigmoid function: ; ; ; in, This represents the feasible probability vector for each discrete candidate yaw correction angle. Let B represent the Sigmoid function, and let B represent the preset global yaw correction angle candidate set. express One of the candidate angles, This represents the feasible probability threshold. This represents the set of feasible candidate angles that exceed the threshold. This indicates the final selected yaw correction angle.

[0033] When feasible set When the trajectory is not empty, the system selects the candidate angle with the smallest deflection relative to the original trajectory. And rotate the robot's current local coordinate system origin as the rotation center along a high-risk trajectory. If feasible set If the value is empty, the system will not output the high-risk trajectory directly as an executable trajectory and will trigger a preset safety rollback strategy. The safety rollback strategy may include stopping and waiting, requesting the upstream strategy to re-predict, or handing over the downstream local planning or control safety module to regenerate the trajectory. The specific rollback method can be determined according to the robot platform configuration.

[0034] The above-described stages one through three outline the complete process of the deployment and inference phase in this embodiment. The training process of this embodiment is further explained below, including two stages: pre-training and fine-tuning, as well as the training objectives and data augmentation strategies designed to achieve spatial awareness and risk awareness correction capabilities.

[0035] 3. Phased training objectives for pre-training and fine-tuning

[0036] The training process is divided into two stages: pre-training and fine-tuning. The pre-training stage is used to learn basic motion prediction capabilities based on ontology geometry from heterogeneous cross-ontology data, optimizing only the motion prediction branch; the fine-tuning stage uses real robot data with physical scale consistency to train the spatial perception and risk perception correction modules.

[0037] (1) Pre-training stage: The pre-training stage is used to learn the basic action prediction ability based on ontology geometry from heterogeneous cross-ontology data. Only the action prediction branch is optimized. The spatial perception module and the risk perception correction module do not participate in the training for the time being.

[0038] The pre-training data is constructed from first-person perspective internet videos of various ontology categories, covering different ontology types such as animals, people, and vehicles. The data construction process is as follows: First, the depth, trajectory, and camera intrinsics of video clips are estimated using an automatic annotation model, and ontology geometric parameters are manually assigned according to category attributes. The pre-training data does not need to be strictly consistent with the real physical scale, but the observations, trajectories, and camera intrinsics within the same sample should maintain scale consistency. The estimated trajectories are checked for smoothness, and video clips with abrupt viewpoint changes, severe occlusion, or unreliable annotations are removed. Due to the high geometric reconstruction noise of such data, the grid-based supervision spatial perception and correction module is not used during the pre-training stage.

[0039] Behavioral cloning loss is used to ensure that the future action sequence output by the policy network is consistent with the demonstrated action sequence. Let the training samples be ( The policy network is... Then the behavioral cloning loss is defined as: ; in Representational Policy Network The trainable parameters, where 𝒟 represents the training dataset. [.] represents the expectation of the state-action sequence samples in the training dataset. This represents the policy state at time t. This demonstrates a sequence of future actions. The policy network is based on The output is the predicted action sequence. This represents the square of the L2 norm.

[0040] (2) Fine-tuning stage: The fine-tuning stage uses real robot navigation data to jointly optimize the three components. The data in this stage has physical scale consistency, specifically including: real robot data provides metric trajectory and camera intrinsic parameters to generate scale-consistent depth observations; the body geometry parameters are taken from the actual parameters of the corresponding robot platform.

[0041] The action prediction branch is supervised by real action labels, while the spatial perception and risk perception correction modules are supervised by offline augmented trajectories. The total loss in the fine-tuning phase consists of the behavior cloning loss. Spatial perception distance supervision loss and candidate correction angle feasibility monitoring loss Weighted composition: ; The specific values ​​of each loss weight can be determined based on the training data and validation results. The distance to the nearest obstacle at each waypoint is used to constrain the output of the spatial perception module, and its supervised target is calculated by the ontology perception occupying the grid. It is used to constrain the feasibility probability of each discrete candidate correction angle, and its supervision objective is whether each candidate angle is a multi-hot tag without collision.

[0042] (3) Sequential Decoupling Structure: Both the spatial perception module and the risk perception and correction module are sequentially decoupled from the action prediction branch. During training, this structure allows the intermediate trajectory generated by action prediction to be replaced with an augmented high-risk trajectory, while continuing to reuse the shared context consistent with the current observation. This structure is also combined with a unidirectional attention mask to prevent the gradient update of high-risk trajectory features from backpropagating to the shared context. Thus, while spatial perception and correction capabilities are enhanced during training, the integrity of the shared context's representation of information such as environmental observation, navigation targets, and ontological geometry is maintained.

[0043] 4. Offline augmentation and joint training strategy for risk trajectories

[0044] Risk trajectory augmentation is only used for fine-tuning data construction and supervision label generation, and is not executed online during inference deployment. For example... Figure 2 As shown, the process sequentially includes multi-frame depth back-projection and post-processing, local point cloud fusion, ground plane fitting, ontology-aware morphological expansion and grid construction, and generation of high-risk trajectories and feasible correction directions. Given depth observations, demonstration trajectories, and camera intrinsic parameters for H future timeframes, the system first back-projects multi-frame depth pixels into 3D points according to the camera intrinsic parameters, and then fuses them into a local point cloud in the robot's central coordinate system based on the pose or trajectory relationship of adjacent frames. Subsequently, effective depth points are selected from the central field of view, and the dominant ground plane is fitted using RANSAC; then, based on the robot length... ,width ,high and maximum obstacle-crossing height Obstacle points are identified, and morphological expansion is performed on the obstacle area according to the body shape and safety margin to obtain a body-aware occupancy grid for collision detection.

[0045] Calculate the safety factor for the original waypoint trajectory ,in The minimum distance from the trajectory to the nearest obstacle. and These are the starting waypoint and the ending waypoint, To prevent small constants with a denominator of zero.

[0046] Perform a uniform scaling transformation on the spline-smoothed reference trajectory: ,in Surrounded by safety factor The reference scale is determined for sampling; the transformation preserves the overall shape of the trajectory and generates candidate trajectories with different obstacle spacing.

[0047] In this embodiment, the scale-transformed trajectory is further rotated at 5° intervals within the range of [-45°, 45°] to form a candidate trajectory pool. Each candidate trajectory is projected onto the ontology-aware occupied grid for continuous collision checks: candidates that collide are designated as high-risk trajectories; the non-collision rotating candidates corresponding to the same high-risk trajectory determine the feasible correction angle and form multi-hot tags.

[0048] As an example of unrestricted training, pre-training can use the Adam optimizer, batch size of 64, and weight decay of 1×10⁻⁶. -4 and from 1×10 -4 Attenuation to 1×10 -5 A cosine learning rate plan is used, with 500k training steps; fine-tuning is performed from pre-trained weights, with a batch size of 32, and the visual encoder is frozen. The above parameters are for feasibility purposes only and do not constitute a limitation on the training configuration.

[0049] The deployment process of this invention is illustrated below using an example of a wheeled or quadruped robot performing point-to-target navigation in an indoor environment containing narrow passages and fixed obstacles. This example is only used to illustrate the input, inter-module data transfer, and risk correction processes; specifically, it includes the following three stages:

[0050] Phase 1: Ontology geometry initialization and multimodal feature acquisition The system reads the robot's length, width, height, and maximum obstacle-crossing height, forming a four-dimensional body geometry vector and encoding it as a body geometry token. Simultaneously, it acquires current and historical RGB images and current depth observations, encoding the two-dimensional target position in the robot's body coordinate system as a target token. These modal tokens are combined to form a shared context.

[0051] Phase Two: Generation of Future Actions and Local Trajectories

[0052] The system uses a historical RGB image sequence of length 4 and reads the shared context through unidirectional cross-attention using 8 action query tokens. The action prediction head outputs the linear velocity and angular velocity sequences for the next 8 time steps in parallel, and then converts the action sequence into a 2D waypoint trajectory and trajectory features in the local coordinate system of the ontology.

[0053] Phase Three: Space Risk Perception and Trajectory Correction

[0054] Online trajectory correction during this stage and Figure 2The offline augmentation process is as follows: In the offline stage, supervision labels are generated through depth backprojection and post-processing, ground fitting, obstacle morphology dilation, and candidate trajectory collision checking. In the online deployment stage, backprojection, post-processing, and morphology dilation are not re-executed. Instead, the trained spatial awareness module predicts waypoint distances, and the risk awareness correction module predicts the feasibility probability of each discrete yaw angle. The spatial awareness module predicts the nearest obstacle distance for each waypoint based on the shared context and trajectory features. If all predicted distances are not lower than the safety threshold, the original trajectory is directly output. If any predicted distance is lower than the safety threshold, the trajectory is marked as high-risk, and the risk awareness correction module is triggered.

[0055] The risk perception and correction module combines shared context, spatial perception features, and high-risk trajectory features to output the feasibility probability of each discrete candidate global yaw angle; this probability represents the likelihood that rotating the entire high-risk trajectory according to the corresponding candidate angle will satisfy the collision-free condition.

[0056] The system filters feasible candidate angles based on probability thresholds and selects the feasible angle with the smallest absolute value to preserve the original motion intent as much as possible. Then, it rotates the high-risk trajectory around the robot's current local coordinate system origin, converts the corrected trajectory into the speed or other control commands required by the underlying control interface, and executes them.

[0057] Furthermore, the present invention also provides a robot visual navigation system based on self-awareness, comprising the following modules: The multimodal input encoding module acquires the robot's current and historical RGB image sequences, depth observation, relative navigation target in the robot's body coordinate system, and body geometric parameters; The ontology geometry encoding module encodes the RGB image sequence, current depth observation, relative navigation target, and ontology geometry parameters into RGB Token, depth Token, target Token, and ontology geometry Token, respectively, forming a shared context Token; The action prediction module reads the shared context token through one-way cross-attention and decodes it into an action sequence, then converts the action sequence into local waypoint trajectories and trajectory features. The spatial perception module predicts the distance from each waypoint to the nearest obstacle in the local waypoint trajectory based on the shared context token and trajectory features; when the minimum predicted distance is lower than a preset safety threshold, it is determined to be a high-risk trajectory. The risk perception and correction module predicts the feasibility probability of each discrete candidate angle within a preset global yaw correction angle range for high-risk trajectories; selects the candidate angle with the smallest absolute value among the feasible candidates as the correction angle, and rotates the high-risk trajectory as a whole. The underlying control output module converts local waypoint trajectories that are not judged as high-risk or have undergone overall rotation correction into executable control commands for the robot and outputs them.

[0058] Accordingly, this application also provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the above-described robot visual navigation method based on proprioception. In the hardware structure of any data-processing device in which the robot visual navigation system based on proprioception provided in this embodiment of the invention is located, in addition to the processor, memory, and network interface, the data-processing device in the embodiment may also include other hardware depending on the actual function of the data-processing device, which will not be elaborated further.

[0059] Accordingly, this application also provides a computer-readable storage medium storing computer instructions thereon, which, when executed by a processor, implement the above-described robot visual navigation method based on ontology self-perception. The computer-readable storage medium can be an internal storage unit of any data-processing device as described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units of any data-processing device and external storage devices. The computer-readable storage medium is used to store the computer program and other programs and data required by the data-processing device, and can also be used to temporarily store data that has been output or will be output.

[0060] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

Claims

1. A robot visual navigation method based on self-awareness, wherein the robot is equipped with a visual sensor and a depth sensor, and the method is executed in real time by a processor mounted on the mobile robot or a computing device communicating with the mobile robot, characterized in that... Includes the following steps: S1. Obtain the current and historical RGB image sequences of the robot, depth observation, relative navigation target in the robot's body coordinate system, and body geometric parameters; S2. Encode the RGB image sequence, current depth observation, relative navigation target and body geometry parameters respectively to obtain RGB Token, depth Token, target Token and body geometry Token, forming a shared context Token; S3. The action prediction module reads the shared context token through one-way cross-attention and decodes it into an action sequence, and then converts the action sequence into local waypoint trajectories and trajectory features. S4. The spatial perception module predicts the distance from each waypoint in the local waypoint trajectory to the nearest obstacle based on the shared context token and trajectory features; when the minimum predicted distance is lower than a preset safety threshold, it is determined to be a high-risk trajectory. S5. For high-risk trajectories, the risk perception and correction module predicts the feasibility probability of each discrete candidate angle within a preset global yaw correction angle interval; selects the candidate angle with the smallest absolute value among the feasible candidates as the correction angle, and rotates the high-risk trajectory as a whole. S6. Convert the local waypoint trajectories that are not judged as high-risk or have been corrected by overall rotation into executable control commands for the robot and output them.

2. The robot visual navigation method based on ontology self-perception according to claim 1, characterized in that, The ontological geometry parameters include at least the robot's length, width, height, and maximum obstacle-crossing height; the ontological geometry token is generated by a mapping network from a four-dimensional vector composed of length, width, height, and maximum obstacle-crossing height, and is used for motion prediction, distance prediction, and risk perception correction.

3. The robot visual navigation method based on ontology self-perception according to claim 1, characterized in that, The unidirectional cross-attention is configured to allow the action query token to be updated based on the shared context token, while preventing the action query token from updating the shared context token in the reverse direction; the action prediction module outputs the action sequence within the preset prediction time domain in parallel.

4. The robot visual navigation method based on ontology self-perception according to claim 1, characterized in that, In step S5, the candidate angle with the smallest absolute value among the feasible candidates is selected as the correction angle. Specifically, the yaw correction angle interval is divided into multiple discrete candidate angle intervals, and the feasibility of each candidate angle is monitored by multiple hot tags. The multiple hot tags represent the tag vectors indicating whether each candidate angle is collision-free. During inference, the candidate angle with the smallest deflection relative to the original high-risk trajectory is selected from the set of feasible candidate angles.

5. The robot visual navigation method based on ontology self-perception according to claim 1, characterized in that, The method also includes a training phase, which comprises a pre-training phase and a fine-tuning phase: During the pre-training phase, data constructed from cross-ontology first-person video is used to train the action prediction branch with behavior cloning loss. In the fine-tuning stage, depth observations, trajectories, and camera intrinsic parameters from multiple consecutive future moments are used to offline fuse local point clouds in the robot's central coordinate system. Ground regions are identified through ground fitting, and obstacle regions are expanded based on ontological geometric parameters to obtain an ontological perception occupancy grid. Training is performed using behavioral cloning loss, waypoint nearest obstacle distance supervision loss, and candidate correction angle feasibility supervision loss from real robot navigation data. The risk trajectory offline augmentation in the fine-tuning stage is performed on the offline constructed ontological perception occupancy grid.

6. The robot visual navigation method based on ontology self-perception according to claim 5, characterized in that, The risk trajectory broadening during the fine-tuning phase includes: Spline smoothing is applied to the original waypoint trajectory; a safety factor is calculated based on the minimum obstacle spacing and trajectory displacement scale; and a uniform scale transformation is performed on all waypoints around the safety factor. Rotate the trajectory after scale transformation at preset angle intervals within a preset angle range; detect whether each candidate trajectory collides on the grid occupied by the body perception, identify the collision candidate as a high-risk trajectory, and mark the corresponding non-collision rotation angle as a multi-hot feasible angle label.

7. The robot visual navigation method based on ontology self-perception according to claim 6, characterized in that, The ontology-aware occupancy grid is constructed as follows: using depth observations, demonstration trajectories, and camera intrinsics from multiple consecutive future moments, local point clouds in the robot's central coordinate system are fused offline; ground regions are identified through ground fitting, and obstacle regions are expanded based on ontology geometric parameters to obtain the ontology-aware occupancy grid; the point clouds and occupancy grid are not used as online strategy inputs during the deployment phase.

8. A system based on the robot visual navigation method based on ontology self-perception as described in any one of claims 1-7, characterized in that, include: The multimodal input encoding module acquires the robot's current and historical RGB image sequences, depth observation, relative navigation target in the robot's body coordinate system, and body geometric parameters; The ontology geometry encoding module encodes the RGB image sequence, the current depth observation, the relative navigation target, and the ontology geometry parameters respectively, and obtains the RGB Token, depth Token, target Token, and ontology geometry Token to form a shared context Token. The action prediction module reads the shared context token through one-way cross-attention and decodes it into an action sequence, then converts the action sequence into local waypoint trajectories and trajectory features. The spatial perception module predicts the distance from each waypoint to the nearest obstacle in the local waypoint trajectory based on the shared context token and trajectory features; when the minimum predicted distance is lower than a preset safety threshold, it is determined to be a high-risk trajectory. The risk perception and correction module predicts the feasibility probability of each discrete candidate angle within a preset global yaw correction angle range for high-risk trajectories; selects the candidate angle with the smallest absolute value among the feasible candidates as the correction angle, and rotates the high-risk trajectory as a whole. The underlying control output module converts local waypoint trajectories that are not judged as high-risk or have undergone overall rotation correction into executable control commands for the robot and outputs them.

9. An electronic device, characterized in that, It includes one or more processors and a memory; the memory is used to store one or more programs, which, when executed by the processor, cause the processor to implement the robot vision navigation method as described in any one of claims 1-7.

10. A computer-readable storage medium storing computer instructions thereon, characterized in that, When the computer instructions are executed by the processor, they implement the robot visual navigation method as described in any one of claims 1-7.