Visual navigation method and system fusing diffusion model and explicit constraint guidance
Through the visual navigation method that integrates the diffusion model and explicit constraint guidance, the problems of insufficient path security and poor dynamic adaptability in traditional visual navigation are solved, and a safe and accessible path is generated, thereby achieving an efficient navigation system.
Patent Information
- Application Number
- CN202510417998.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional visual navigation methods lack adaptability to environmental changes and cannot respond to unknown obstacles in real time, resulting in insufficient path security and poor dynamic adaptability. The generated paths are prone to collision with obstacles and have weak generalization ability.
The visual navigation method that integrates the diffusion model and explicit constraint-guided visual navigation method, obtains visual observation data and navigation target data, performs feature extraction and fusion, generates candidate paths, and optimizes the path based on collision constraints and target distance constraints, and uses the conditional diffusion model and path selection strategy to determine the target navigation path.
It improves the safety and accessibility of the path, enhances the reliability and practicality of the navigation system, realizes zero sample migration to the real world, and does not require retraining for new scenarios, improving the efficiency of path navigation.
Smart Images

Figure CN120333440A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a visual navigation method and system that integrates a diffusion model and explicit constraint guidance. Background Art
[0002] Currently, visual navigation is one of the core challenges in the field of mobile robots and is widely used in scenarios such as unmanned delivery and service robots. Its core goal is to obtain environmental information through sensors and plan a safe and efficient path in real time to reach the target location. However, traditional visual navigation methods often lack sufficient adaptability to environmental changes, rely on static environment assumptions, and cannot respond to unknown obstacles in real time, resulting in limited obstacle avoidance ability for dynamic obstacles, leading to a risk of collision between the generated path and obstacles, and weak generalization ability for unknown scenarios, and being unable to adapt to diverse environments.
[0003] In summary, the technical problems existing in the related technologies need to be improved. Summary of the Invention
[0004] The embodiments of this application aim to at least solve one of the technical problems in the related technologies to some extent. For this reason, the main purpose of the embodiments of this application is to propose a visual navigation method and system that integrates a diffusion model and explicit constraint guidance, which can solve the problems of insufficient path safety and poor dynamic adaptability, significantly improve the safety and reachability of the path, and further improve the path navigation efficiency.
[0005] To achieve the above object, on the one hand, an embodiment of this application proposes a visual navigation method that integrates a diffusion model and explicit constraint guidance, and the method includes the following steps:
[0006] Obtain initial image data; the initial image data includes visual observation data and navigation target data;
[0007] Perform feature extraction processing on the visual observation data to obtain target observation semantic features;
[0008] Perform feature extraction processing on the navigation target data to obtain target image features;
[0009] Perform feature fusion processing on the target observation semantic features and the target image features to generate a conditional vector;
[0010] Input the conditional vector into a conditional diffusion model for diffusion training to generate a number of initial candidate paths;
[0011] Construct explicit constraints based on the initial image data, and perform path optimization on each of the initial candidate paths according to the explicit constraints to obtain a number of candidate optimized paths; the explicit constraints include collision constraints and target distance constraints;
[0012] Determine a target navigation path from each of the candidate optimization paths according to a path selection strategy.
[0013] In some embodiments, the feature extraction process on the visual observation data to obtain target observation semantic features includes:
[0014] Input the visual observation data into a visual encoder;
[0015] Perform a feature extraction process on the visual observation data through the visual encoder to obtain the target observation semantic features.
[0016] In some embodiments, the navigation target data includes image target data and point target data. The feature extraction process on the navigation target data to obtain target image features includes:
[0017] Input the image target data into a target encoder, and perform a feature extraction process on the image target data through the target encoder to obtain the target image features;
[0018] Or,
[0019] Input the point target data into the target encoder, and perform a feature extraction process on the point target data through the target encoder to obtain the target image features.
[0020] In some embodiments, the feature fusion process on the target observation semantic features and the target image features to generate a conditional vector includes:
[0021] Perform a feature fusion process on the target observation semantic features and the target image features through a cross-attention mechanism to generate the conditional vector.
[0022] In some embodiments, the inputting the conditional vector into a conditional diffusion model for diffusion training to generate a number of initial candidate paths includes:
[0023] Input the conditional vector into the conditional diffusion model, and perform forward noise addition training on the conditional vector and the real path point sequence through the conditional diffusion model to generate a noise path;
[0024] Perform reverse denoising inference on the noise path through the conditional diffusion model to generate a number of the initial candidate paths.
[0025] In some embodiments, the constructing an explicit constraint based on the initial image data and optimizing each of the initial candidate paths according to the explicit constraint to obtain a number of candidate optimization paths includes:
[0026] Construct the target distance constraint based on the navigation target data and the initial candidate path;
[0027] Construct the collision constraint based on the visual observation data;
[0028] Calculate the total constraint cost gradient of the explicit constraint according to the target distance constraint and the collision constraint;
[0029] Perform path optimization on each of the initial candidate paths according to the total constraint cost gradient to obtain a number of candidate optimized paths.
[0030] In some embodiments, the constructing the target distance constraint based on the navigation target data and the initial candidate path includes:
[0031] Perform data extraction processing on the navigation target data to obtain target point coordinates;
[0032] Perform data extraction processing on the initial candidate path to obtain the end point coordinates of the initial candidate path;
[0033] Construct the target distance constraint according to the target point coordinates and the end point coordinates of the initial candidate path.
[0034] In some embodiments, the constructing the collision constraint based on the visual observation data includes:
[0035] Input the visual observation data into a depth estimation model to generate a depth map;
[0036] Construct a local TSDF map according to the depth map;
[0037] Determine the collision constraint according to the local TSDF map.
[0038] In some embodiments, the path selection strategy includes a consistency screening strategy and a continuity optimization strategy. The determining the target navigation path from each of the candidate optimized paths according to the path selection strategy includes:
[0039] According to the consistency screening strategy, screen the current paths that are consistent with the path direction of the historical path from a number of the candidate optimized paths;
[0040] According to the continuity optimization strategy, perform weighted average filtering processing on the current path and the historical path to obtain the target navigation path.
[0041] To achieve the above object, another aspect of the embodiments of the present application proposes a visual navigation system that fuses a diffusion model and explicit constraint guidance. The system includes the following modules:
[0042] An initial image data acquisition module for acquiring initial image data; the initial image data includes visual observation data and navigation target data;
[0043] A first feature extraction and processing module for performing feature extraction and processing on the visual observation data to obtain target observation semantic features;
[0044] A second feature extraction and processing module for performing feature extraction and processing on the navigation target data to obtain target image features;
[0045] A feature fusion processing module for performing feature fusion processing on the target observation semantic features and the target image features to generate a conditional vector;
[0046] A conditional diffusion model training module for inputting the conditional vector into a conditional diffusion model for diffusion training to generate a number of initial candidate paths;
[0047] A candidate path optimization module for constructing explicit constraints based on the initial image data and optimizing the paths of each of the initial candidate paths according to the explicit constraints to obtain a number of candidate optimized paths; the explicit constraints include collision constraints and target distance constraints;
[0048] A target navigation path determination module for determining a target navigation path from each of the candidate optimized paths according to a path selection strategy.
[0049] The embodiments of the present application at least include the following beneficial effects: The present application provides a visual navigation method and system that combines a diffusion model and explicit constraint guidance. The solution includes obtaining initial image data, which includes visual observation data and navigation target data; performing feature extraction processing on the visual observation data to obtain target observation semantic features; performing feature extraction processing on the navigation target data to obtain target image features; performing feature fusion processing on the target observation semantic features and the target image features to generate a conditional vector; inputting the conditional vector into a conditional diffusion model for diffusion training to generate a number of initial candidate paths; constructing explicit constraints based on the initial image data, and optimizing the paths of each initial candidate path according to the explicit constraints to obtain a number of candidate optimized paths. The explicit constraints include collision constraints and target distance constraints; determining the target navigation path from each candidate optimized path according to a path selection strategy. The embodiments of the present application generate multi-modal paths through a conditional diffusion model, and introduce an explicit cost gradient guidance strategy and a path selection strategy to determine the target navigation path, effectively solving the problems of insufficient path safety and poor dynamic adaptability in traditional visual navigation methods, and significantly improving the reliability and practicality of the navigation system. Among them, the introduction of collision constraints makes the generated path unlikely to collide with obstacles, improving the safety of the path; the introduction of target distance constraints guides the path towards the target point, enhancing the reachability of the path. The embodiments of the present application learn the multi-scene path distribution through pre-training the diffusion model, and at the same time introduce explicit cost constraints, realizing zero-shot transfer to the real world without retraining for a new scene, further improving the path safety and path navigation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 FIG. is a flowchart of the steps of a visual navigation method that combines a diffusion model and explicit constraint guidance provided by the embodiments of the present application;
[0051] Figure 2 FIG. is an overall flowchart of a visual navigation method that combines a diffusion model and explicit constraint guidance provided by the embodiments of the present application;
[0052] Figure 3 FIG. is a structural diagram of a visual navigation system that combines a diffusion model and explicit constraint guidance provided by the embodiments of the present application;
[0053] Figure 4 FIG. is a hardware structural diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] To make the objectives, technical solutions, and advantages of this application more clearly understood, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not intended to limit this application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of this application. They are merely examples of systems and methods that are consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0055] It can be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "while...", or "in response to determining".
[0056] The terms "at least one", "multiple", "each", "any one", etc. used in this application, at least one includes one, two, or more than two, multiple includes two or more than two, each refers to each of the corresponding multiple, and any one refers to any one of the multiple.
[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0058] Before elaborating in detail on the embodiments of this application, first, some nouns and terms involved in the embodiments of this application are described. The nouns and terms involved in the embodiments of this application are applicable to the following explanations.
[0059] 1) Embodied intelligence, an artificial intelligence system combining software and hardware that perceives and acts based on a physical body. It obtains information, understands problems, makes decisions, and realizes actions through the interaction between the agent and the environment, thereby generating intelligent behaviors and adaptability.
[0060] 2) Diffusion model, a type of latent variable model, is a Markov chain trained with variational estimation and can generate data similar to its training data. The goal of the diffusion model is to learn the latent structure of the dataset by modeling the way data points diffuse in the latent space. It is applied to various tasks, such as image denoising, image restoration, super-resolution imaging, image generation, etc.
[0061] 3) Cost-guided diffusion model, a framework that combines a diffusion model with explicit cost constraints, optimizes path generation through gradient guidance (such as collision cost, target distance cost) to ensure the path is safe and reachable.
[0062] 4) Robot path planning refers to designing a motion trajectory from the starting point to the target point for the robot through algorithms, ensuring the path is safe (avoiding obstacles), efficient (short path length), and compliant with the robot's motion constraints (such as turning radius, speed limit).
[0063] Currently, visual navigation is one of the core challenges in the field of mobile robots and is widely used in scenarios such as unmanned delivery and service robots. Its core goal is to obtain environmental information through sensors and plan a safe and efficient path in real time to reach the target location. However, traditional visual navigation methods often lack sufficient adaptability to environmental changes, rely on static environmental assumptions, and cannot respond to unknown obstacles in real time, resulting in limited obstacle avoidance ability for dynamic obstacles, thus leading to a risk of collision with obstacles in the generated path and weak generalization ability for unknown scenarios and inability to adapt to diverse environments. Traditional methods rely on geometric modeling and modular design, while in recent years, deep learning-based methods have achieved significant performance improvements through end-to-end learning but still face problems such as insufficient generalization and strong data dependence. The current related technologies are mainly divided into two categories: classical methods and learning-based methods, each with its own advantages and disadvantages. Among them, traditional navigation systems adopt modular design and rely on manually designed cost functions (such as obstacle distance, path length), with the characteristics of strong interpretability and good scene adaptability, but the modular process is prone to information loss and it is difficult to meet the real-time obstacle avoidance requirements in dynamic environments; learning-based methods perform excellently in scenarios covered by training data, but have limited generalization ability for unknown environments and require a large amount of labeled data and high training costs.
[0064] Exemplarily, methods such as ViNT (Visual Navigation Transformer) propose a visual navigation framework based on diffusion models, which generate sub-goal images through topological maps to guide the robot to approach the goal step by step. Its advantage lies in its long-range planning ability, but it has insufficient real-time obstacle avoidance ability for local dynamic obstacles, resulting in a decrease in success rate in unknown obstacle scenarios; methods such as NoMaD (Goal Masked Diffusion Policies) use diffusion models to generate multi-modal action sequences, directly mapping visual observations to robot control commands. However, this method is prone to collisions in complex obstacle environments and has poor path continuity; diffusion policies, such as Diffusion Policy, are responsible for applying diffusion models to robot control strategy learning and achieving complex tasks through action sequence generation. Although it demonstrates the potential of multi-modal action generation, it is difficult to balance path diversity and safety. All in all, the generated paths of related visual navigation technologies may violate physical rules (such as colliding with obstacles), especially in dynamic or complex scenarios, where the planning failure rate is relatively high. There are also deficiencies such as insufficient path continuity, poor adaptability to dynamic environments, and difficulty in balancing generalization and reliability.
[0065] In view of this, in the embodiments of the present application, a visual navigation method and system that integrates a diffusion model and explicit constraint guidance are provided. The solution includes obtaining initial image data, which includes visual observation data and navigation target data; performing feature extraction processing on the visual observation data to obtain target observation semantic features; performing feature extraction processing on the navigation target data to obtain target image features; performing feature fusion processing on the target observation semantic features and the target image features to generate a conditional vector; inputting the conditional vector into a conditional diffusion model for diffusion training to generate a number of initial candidate paths; constructing explicit constraints based on the initial image data, and optimizing the paths of each initial candidate path according to the explicit constraints to obtain a number of candidate optimized paths. The explicit constraints include collision constraints and target distance constraints; determining a target navigation path from each candidate optimized path according to a path selection strategy. In the embodiments of the present application, multi-modal paths are generated through a conditional diffusion model, and an explicit cost gradient guidance strategy and a path selection strategy are introduced to determine the target navigation path, effectively solving the problems of insufficient path safety and poor dynamic adaptability in traditional visual navigation methods, and significantly improving the reliability and practicality of the navigation system. Among them, the introduction of collision constraints makes the generated paths unlikely to collide with obstacles, improving the safety of the paths; the introduction of target distance constraints guides the paths to approach the target point, enhancing the reachability of the paths. In the embodiments of the present application, the pre-trained diffusion model is used to learn the path distributions of multiple scenarios, and explicit cost constraints are introduced to achieve zero-shot transfer to the real world without retraining for new scenarios, further improving path safety and path navigation efficiency.
[0066] A visual navigation method integrating a diffusion model and explicit constraint guidance provided by an embodiment of the present application relates to the field of computer technology. The visual navigation method integrating a diffusion model and explicit constraint guidance provided by an embodiment of the present application can be applied to a terminal, a server, or software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing a visual navigation method integrating a diffusion model and explicit constraint guidance, etc., but is not limited to the above forms.
[0067] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs (Personal Computers), minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0068] Please refer to Figure 1 , Figure 1 which is an optional step flowchart of a visual navigation method integrating a diffusion model and explicit constraint guidance provided by an embodiment of the present application. Figure 1 The method in
[0069] Step S101, obtain initial image data; the initial image data includes visual observation data and navigation target data.
[0070] Among them, the initial image data refers to the original image without any processing, and the initial image data includes visual observation data and navigation target data.
[0071] For the visual observation data, it is the original image information about the observation scene or object directly obtained through visual means (such as cameras, sensors, etc.), which reflects the visual characteristics such as the appearance, position, and posture of objects in the actual scene, and is an intuitive visual record of the real world. In the embodiments of the present application, the visual observation data is the continuous moment images observed by the robot, that is, the RGB (Red Green Blue Color Mode) observation sequence.
[0072] For the navigation target data (also known as sub-goals), there are two optional forms of sub-goals, namely: image target data and point target data. The image target data (also known as the target image) refers to a single pre-given RGB-formatted picture, which is used to guide the conditional diffusion model to output an action to move towards the target image; similarly, the point target data (also known as the target coordinate) is also a pre-given coordinate point, which is used to guide the diffusion model to output an action to move towards the target coordinate.
[0073] Step S102, perform feature extraction processing on the visual observation data to obtain target observation semantic features;
[0074] In some embodiments, step S102 may include: inputting the visual observation data into a visual encoder; performing feature extraction processing on the visual observation data through the visual encoder to obtain target observation semantic features.
[0075] Among them, the visual encoder is a neural network model or algorithm module specifically used to process visual data (such as images, video frames, etc.). Its main function is to convert the input visual observation data into a feature vector with a specific representation form, that is, to perform feature extraction and encoding. In the embodiments of the present application, the visual encoder is mainly used to extract the semantic features of the RGB observation sequence.
[0076] In specific implementation, the visual encoder employs a pre-trained convolutional neural network (Convolutional Neural Network, abbreviated as CNN) or a vision transformer (Vision Transformer, abbreviated as ViT) to extract the spatio-temporal features of the RGB observation sequence. Exemplarily, input the historical frame images {I{t - T},..., I{t}} into the visual encoder by the visual encoder output the feature vector h <obs>, that is, the target observation semantic feature.
[0077] Step S103: Perform feature extraction processing on the navigation target data to obtain target image features;
[0078] Optionally, the navigation target data includes two optional forms: image target data and point target data. Among them, the image target data refers to RGB picture targets, and the point target data refers to point targets (i.e., GPS (Global Positioning System) position coordinates).
[0079] In some embodiments, step S103 may include: inputting the image target data into a target encoder, and performing feature extraction processing on the image target data through the target encoder to obtain target image features; or, inputting the point target data into the target encoder, and performing feature extraction processing on the point target data through the target encoder to obtain target image features.
[0080] Among them, for the image target data, the target image feature h of the target image is extracted using the same network model as the visual encoder <goal>。For the feature extraction process of point target data, specifically, the point target coordinates (x, y) are mapped into a feature vector h through the fully connected layer (FC) of the neural network <goal>。
[0081] Step S104: perform feature fusion processing on the target observed semantic feature and the target image feature to generate a conditional vector;
[0082] In some embodiments, step S104 may include: performing feature fusion processing on the target observed semantic feature and the target image feature through a cross-attention mechanism to generate a conditional vector.
[0083] Among them, the Goal Encoder is used to process the features of image targets or point targets.
[0084] In a specific implementation, the target observed semantic feature h is processed through a cross-attention mechanism (Cross-Attention) <obs>With the target image feature h <goal>Perform feature fusion processing to generate a conditional vector C, and input the conditional vector C into the conditional diffusion model.
[0085] Step S105: Input the conditional vector into the conditional diffusion model for diffusion training to generate a number of initial candidate paths.
[0086] In some embodiments, step S105 may include: inputting the conditional vector into the conditional diffusion model, and performing forward noise addition training by the conditional diffusion model according to the conditional vector and the real path point sequence to generate a noise path; performing reverse denoising inference on the noise path by the conditional diffusion model to generate a number of initial candidate paths.
[0087] Among them, the conditional diffusion model (Conditional Diffusion Model) mainly generates multi-modal candidate paths based on the conditional vector C.
[0088] In the embodiments of the present application, the conditional diffusion model adopts the denoising diffusion probabilistic model (Denoising Diffusion Probabilistic Model, abbreviated as DDPM), destroys the path data by gradually adding noise, and then generates a path by reverse denoising. The specific diffusion process includes a forward training stage and a reverse inference stage, and the specific implementation process is as follows:
[0089] (1) Forward training stage: First, input the path point sequence P0 (real path) and the conditional vector C into the denoising diffusion probabilistic model; then, add Gaussian noise ε to the path point sequence P0 according to the time step t by the denoising diffusion probabilistic model t , to generate a noise path P t . In the forward training stage, a loss function also needs to be designed to minimize the noise prediction error, that is, the loss function of the denoising diffusion probabilistic model in the forward training stage is used to quantify and minimize the error between the noise predicted by the model and the real noise, so as to optimize the model parameters and improve the path generation quality. Among them, the calculation formula of the loss function is as follows:
[0090]
[0091] Among them, L represents the loss function, φ θ represents the noise prediction network, MSE (Mean squared error) is the mean squared error, ε t represents the Gaussian noise, represents the noise prediction value at each step, P t represents the noise path, t represents the time step, and O represents the observed data (RGB image sequence).
[0092] (2) Reverse inference stage: mainly starts from the noise path P t Start by gradually denoising to generate path P0. The specific denoising calculation formula is as follows:
[0093]
[0094] Among them, P t-1 represents the path after denoising at each step, represents the noise prediction value at each step, and α, γ, σ 2 respectively represent the hyperparameters of this function. N(0, σ 2 I) represents a normal distribution with variance σ 2 , N represents the normal distribution, and I represents the identity matrix.
[0095] In the specific implementation, multiple groups of candidate paths can be generated through multiple diffusion samplings.
[0096] Step S106: Construct explicit constraints based on the initial image data, and optimize the paths of each of the initial candidate paths according to the explicit constraints to obtain several candidate optimized paths; the explicit constraints include collision constraints and target distance constraints;
[0097] In some embodiments, step S106 may include: constructing a target distance constraint based on the navigation target data and the initial candidate path; constructing a collision constraint based on the visual observation data; calculating the total constraint cost gradient of the explicit constraint according to the target distance constraint and the collision constraint; optimizing the paths of each of the initial candidate paths according to the total constraint cost gradient to obtain several candidate optimized paths.
[0098] In some specific embodiments, the step of constructing a target distance constraint based on the navigation target data and the initial candidate path may include: performing data extraction processing on the navigation target data to obtain the target point coordinates; performing data extraction processing on the initial candidate path to obtain the end coordinates of the initial candidate path; constructing a target distance constraint according to the target point coordinates and the end coordinates of the initial candidate path.
[0099] In some specific embodiments, the step of constructing a collision constraint based on the visual observation data may include: inputting the visual observation data into a depth estimation model to generate a depth map; constructing a local TSDF map according to the depth map; determining the collision constraint according to the local TSDF map.
[0100] Specifically, in the inference stage of the conditional diffusion model, by introducing a CostGradient Guidance module, multiple groups of candidate paths generated by the denoising diffusion probability model are optimized based on the gradients of the target distance constraint cost and the collision constraint cost to ensure that the multiple groups of candidate paths are safe and reachable.
[0101] Among them, the target distance constraint cost F g Used to measure the distance between the end point of the candidate path and the target point, and the target distance constraint cost F g The constraint function formula of
[0102] F g(P) = ||W0 - G p || 2
[0103] Where, F g(P) represents the target distance constraint cost function, W0 represents the end point of the candidate path, and G p represents the target point coordinates.
[0104] For the collision constraint cost F c , it is used to calculate the distance between the path and the obstacle based on the real-time constructed TSDF (Truncated Signed Distance Function) map. Where, the collision constraint cost F c The constraint function formula of
[0105]
[0106] Where, F c(P) represents the collision constraint cost function, k t represents the influence coefficient of each path point, W t represents all path points in the path, σ R represents the robot width radius (half width of the robot), and C represents the loss function.
[0107] Where, σ R is the half width of the robot. In the embodiments of the present application, the detection points are extended along the vertical direction of the path to ensure obstacle avoidance safety. Specifically, extending the detection points along the vertical direction of the path can be understood as increasing the safety threshold of obstacle avoidance. When the robot actually moves, its trajectory will occupy a certain spatial range. By extending the detection points along the vertical direction, the positions of both sides of the robot can be simulated, simulating an increase in the width of the robot, and keeping the robot away from obstacles.
[0108] For the construction of the TSDF map, mainly use the monocular depth estimation model (Depth Anything V2) to generate a depth map, and then convert the depth map into a local TSDF map, marking the distance of obstacles at each position. Where, TSDF is a truncated signed distance function, which constructs a distance field from each voxel in the environment to the obstacle through sensor data (such as a depth camera), indirectly helping the robot perceive the position of the obstacle. Where, the collision constraint cost is the constraint cost updated based on the TSDF map, and based on the local TSDF map, the collision cost (collision constraint) can be dynamically updated.
[0109] During the inference stage of the conditional diffusion model, by introducing a cost gradient guidance module, the gradients of the target distance constraint cost and the collision constraint cost are used to optimize multiple sets of candidate paths generated by the denoising diffusion probability model, ensuring that the multiple sets of candidate paths are safe and reachable. That is, at each step of the reverse diffusion, the gradient of the total cost is superimposed on the noise prediction result, where is the derivative of denotes the total constraint cost loss function, denotes the loss based on the coordinate target, denotes the loss based on the obstacle distance. The noise prediction result means that at each step of denoising in the reverse inference of the diffusion model, the derivative of the cost function (target distance constraint cost function and collision constraint cost function) designed in this embodiment of the application is added as a guiding quantity. Finally, after multiple steps of denoising, the final candidate path is obtained. The formula for superimposing the gradient of the total cost on the noise prediction result is as follows:
[0110]
[0111] where P t-1 represents the path after each step of noise reduction; N represents the normal distribution; s t represents the gradient scaling coefficient, which balances diversity and satisfaction of constraints. P represents the path; O represents the observed data (RGB image sequence); Σ represents the covariance matrix.
[0112] Step S107, determine the target navigation path from each of the candidate optimized paths according to the path selection strategy.
[0113] Among them, the path selection strategy (Path Selection Policy) is used to screen the optimal path and execute it smoothly. The path selection strategy includes a consistency screening strategy and a continuity optimization strategy.
[0114] In some embodiments, step S107 may include: according to the consistency screening strategy, screen the current path that is consistent with the path direction of the historical path from several candidate optimized paths; according to the continuity optimization strategy, perform weighted average filtering on the current path and the historical path to obtain the target navigation path.
[0115] In specific implementation, the optimal path is selected from the multi-modal candidate paths through the path selection strategy to ensure motion coherence.
[0116] For the consistency screening strategy, a path direction difference δ is predefined in advance, and then candidate paths that are consistent with the path direction of the historical path (based on the path direction executed by the robot at the previous moment) (δ < ε) are screened. The formula for consistency screening is as follows:
[0117] V = {P|δ(P t , P h ) < ε)
[0118] Wherein, V represents the set of paths selected, P represents a path, δ(·) represents the angle of deviation of the directions of two paths, ε is the set deviation threshold, and P h represents the historical path.
[0119] For the continuous optimization strategy, it is used to perform weighted average filtering on the current path (the only candidate path selected after the consistency screening operation) and the historical path to obtain the target path to be executed. The calculation formula of the target path is as follows:
[0120] P final = λP current + (1 - λ)P history
[0121] Wherein, P final represents the target path; λ is the smoothing coefficient to reduce motion jitter; P current represents the current path, and P history represents the historical path.
[0122] Steps S101 to S107 shown in the embodiments of the present application include obtaining initial image data; the initial image data includes visual observation data and navigation target data; performing feature extraction processing on the visual observation data to obtain target observation semantic features; performing feature extraction processing on the navigation target data to obtain target image features; performing feature fusion processing on the target observation semantic features and the target image features to generate a conditional vector; inputting the conditional vector into a conditional diffusion model for diffusion training to generate a number of initial candidate paths; constructing explicit constraints based on the initial image data, and performing path optimization on each initial candidate path according to the explicit constraints to obtain a number of candidate optimized paths; the explicit constraints include collision constraints and target distance constraints; determining the target navigation path from each candidate optimized path according to the path selection strategy. The embodiments of the present application generate multi-modal paths through a conditional diffusion model, and introduce an explicit cost gradient guidance strategy and a path selection strategy to determine the target navigation path, effectively solving the problems of insufficient path safety and poor dynamic adaptability in traditional visual navigation methods, significantly improving the reliability and practicality of the navigation system. Among them, the introduction of collision constraints makes the generated path unlikely to collide with obstacles, improving the safety of the path; the introduction of target distance constraints guides the path towards the target point, enhancing the reachability of the path; the embodiments of the present application learn the multi-scene path distribution through a pre-trained diffusion model, and at the same time introduce explicit cost constraints, realizing zero-shot transfer to the real world without retraining for new scenarios, further improving path safety and path navigation efficiency.
[0123] To explain the principle of the technical solution of the present invention in detail, the overall process of the present invention will be described below in conjunction with some specific embodiments. It is easy to understand that the following is an explanation of the technical principle of the present invention and should not be regarded as a limitation of the present invention.
[0124] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the overall process of a visual navigation method that combines a diffusion model and explicit constraint guidance provided by an embodiment of the present application; as Figure 2 shown, the visual navigation method that combines a diffusion model and explicit constraint guidance provided by an embodiment of the present application is implemented based on a cost-guided diffusion model, and its core architecture is as Figure 2 shown. The system mainly consists of five modules: a visual encoder, a target encoder, a conditional diffusion model, a cost gradient guidance module, and a path selection strategy. Specifically, the overall implementation process of the visual navigation method that combines a diffusion model and explicit constraint guidance is as follows: First, obtain an RGB observation sequence and target information (target image or point coordinates); then, input the RGB observation sequence into the visual encoder for feature extraction processing to obtain target observation semantic features, and input the target information into the target encoder for feature extraction processing to obtain target image features; then, use the cross-attention mechanism of the Transformer model to perform feature fusion processing on the target observation semantic features and target image features to generate a conditional vector; input the conditional vector into the conditional diffusion model for diffusion training to generate a number of initial candidate paths; further, construct explicit constraints (collision constraint and target distance constraint) based on the RGB observation sequence and target information. Among them, the target distance constraint is used to measure the distance between the end point of the candidate path and the target point, and the collision constraint is mainly updated in real time based on the TSDF map. For the construction of the TSDF map, mainly use a monocular depth estimation model (Depth Anything V2) to process the RGB observation sequence to generate a depth map, and then convert the depth map into a local TSDF map to mark the obstacle distance at each position. Based on the local TSDF map, the collision cost (collision constraint) can be dynamically updated; furthermore, according to the explicit constraints, optimize each initial candidate path to obtain a number of candidate optimized paths. Specifically, in the inference stage of the conditional diffusion model, by introducing a cost gradient guidance module, optimize the multi-group candidate paths generated by the denoising diffusion probability model based on the gradients of the target distance constraint cost and the collision constraint cost to ensure that the multi-group candidate paths are safe and reachable. That is, as Figure 2 shown, at each step of the reverse diffusion, the gradient of the total cost ( is The derivative) is superimposed on the noise prediction result to optimize multiple groups of candidate paths; finally, the target navigation path is determined from each candidate optimized path according to the path selection strategy.
[0125] Among them, under the guidance of the gradient of the designed cost function, this gradient information is incorporated into each denoising process of the diffusion model to guide the generation of local paths. For long-distance navigation tasks, high-level strategies, such as topological maps, are used to provide sub-goals, and this method supports both image goals and point goals.
[0126] It should be noted that this embodiment only briefly illustrates the general process of a visual navigation method that combines a diffusion model and explicit constraint guidance. For the detailed description of each step, reference can be made to the relevant content in the foregoing embodiments, and details are not described here. It can be understood that the present invention does not limit this.
[0127] In the embodiment of the present application, initial image data is obtained; the initial image data includes visual observation data and navigation target data; the visual observation data is subjected to feature extraction processing to obtain target observation semantic features; the navigation target data is subjected to feature extraction processing to obtain target image features; the target observation semantic features and target image features are subjected to feature fusion processing to generate a conditional vector; the conditional vector is input into a conditional diffusion model for diffusion training to generate a number of initial candidate paths; an explicit constraint is constructed based on the initial image data, and each initial candidate path is optimized according to the explicit constraint to obtain a number of candidate optimized paths; the explicit constraint includes a collision constraint and a target distance constraint; the target navigation path is determined from each candidate optimized path according to the path selection strategy. In the embodiment of the present application, a multi-modal path is generated through a conditional diffusion model, and an explicit cost gradient guidance strategy and a path selection strategy are introduced to determine the target navigation path, effectively solving the problems of insufficient path safety and poor dynamic adaptability in traditional visual navigation methods, and significantly improving the reliability and practicability of the navigation system. Among them, the introduction of the collision constraint makes the generated path unlikely to collide with obstacles, improving the safety of the path; the introduction of the target distance constraint guides the path to approach the target point, enhancing the reachability of the path; in the embodiment of the present application, the pre-trained diffusion model is used to learn the multi-scene path distribution, and at the same time, an explicit cost constraint is introduced, realizing zero-shot transfer to the real world without retraining for a new scene, further improving the path safety and path navigation efficiency.
[0128] A visual navigation method that combines a diffusion model and explicit constraint guidance provided in the embodiment of the present application can be applied to application scenarios such as RGB visual navigation devices, service robots, and autonomous driving, providing an efficient and reliable solution for visual navigation in complex environments.
[0129] Exemplarily, first, a vision sensor (such as an RGB camera) is responsible for collecting environmental images (RGB sequences) in real time, and the environmental images (RGB sequences) are used as the original input of the visual navigation system. The visual navigation system includes a visual encoder, a diffusion model, and a display cost constraint guidance template, and is used to process vision sensor data to generate a safe path. After the vision sensor collects vision sensor data, the vision sensor data (image sequence) and the target instruction (such as a target image or point coordinates) are input into the visual navigation system, and are processed by each module of the visual navigation system to output an optimized path point sequence. Finally, the generated path point sequence is converted into a robot motion instruction (such as speed, steering angle), and the robot is made to execute it, that is, the robot receives the target image or target coordinates to be navigated to, and outputs path points according to the real-time observed RGB image sequence. It can be understood that in practical applications, after the robot receives the target image or target coordinates to be navigated to, it can output path points according to the real-time observed RGB image sequence.
[0130] In summary, the key points of a visual navigation method that fuses a diffusion model and explicit constraint guidance provided by the embodiments of the present application are as follows:
[0131] (1) Cost-guided diffusion model framework: Combine an explicit cost function (including collision cost, target distance cost) with the diffusion model, and dynamically adjust the path generation direction by introducing a differentiable cost gradient (ΔF) in the reverse diffusion stage of the diffusion model.
[0132] (2) Real-time local TSDF map construction and collision cost calculation: Based on monocular RGB input, generate a depth map through a depth estimation model (such as Depth Anything V2), and construct a local TSDF (truncated signed distance function) map in real time to quantify the distance between path points and obstacles.
[0133] (3) Multi-modal path consistency screening and smoothing strategy: Screen candidate paths through historical path direction consistency constraints, and achieve path temporal smoothing by combining weighted average filtering.
[0134] A visual navigation method that fuses a diffusion model and explicit constraint guidance provided by the embodiments of the present application has the following advantages:
[0135] (1) Adaptability to dynamic environments and real-time obstacle avoidance: Currently, the navigation methods based on diffusion models in related technologies (such as NoMaD, ViNT) are difficult to handle unknown obstacles or dynamic scenarios. However, the embodiments of the present application can estimate depth in real time through monocular RGB input, construct a local TSDF map, dynamically update the collision cost, and use gradient guidance to adjust path generation, so that the obstacle avoidance effect of the generated path is greatly improved.
[0136] (2) Co - optimization of Generalization and Reliability: The existing end - to - end methods (such as diffusion strategies) in related technologies rely on the training data distribution and have weak generalization ability for unknown scenarios; although the classical methods are reliable, they cannot adapt to diverse environments. However, the embodiments of the present application can learn the multi - scenario path distribution through pre - training the diffusion model (data - driven generalization), and at the same time introduce explicit cost constraints (such as target distance cost), achieving zero - shot transfer to the real world without retraining for new scenarios, and taking into account both path safety and path generation efficiency, thereby improving the overall visual navigation efficiency.
[0137] (3) Balance between Path Diversity and Continuity: The path generated by the existing diffusion - model - based navigation methods (such as NoMaD) in related technologies is prone to sudden direction changes or jitters, resulting in unstable movement. However, the embodiments of the present application introduce a consistency screening and weighted smoothing strategy to select a temporally coherent trajectory from multi - modal candidate paths, reducing the jitter caused by path selection during the robot's movement.
[0138] The embodiments of the present application generate multi - modal paths through a conditional diffusion model and introduce explicit cost gradient guidance and path selection strategies, solving the problems of insufficient path safety, poor dynamic adaptability, and low continuity in related technologies. Specifically, by designing the collision cost (based on the real - time TSDF map) and the target distance cost as differentiable functions, path generation is guided by gradients during the diffusion model inference stage to ensure that the path meets safety (reducing collisions) and target reachability. Among them, a local TSDF map is constructed in real - time based on monocular RGB input, and the collision cost is updated by fusing real - time obstacle information to support dynamic obstacle avoidance; by pre - training the diffusion model to learn the multi - scenario path distribution (data - driven generalization), and at the same time using an explicit cost function to constrain path generation (rule - driven reliability), zero - shot transfer to unknown scenarios (such as a significant increase in the success rate of real - world experiments) is achieved without retraining for new scenarios; and, by introducing a consistency path selection strategy, the current path is weighted and smoothed by combining the movement trend of the historical path to avoid sudden direction changes, which can improve the movement coherence of the robot.
[0139] Please refer to Figure 3 , the embodiments of the present application also provide a visual navigation system 300 that combines a diffusion model and explicit constraint guidance, which can implement the above - mentioned visual navigation method that combines a diffusion model and explicit constraint guidance. The system includes the following modules:
[0140] An initial image data acquisition module 301, configured to acquire initial image data; the initial image data includes visual observation data and navigation target data;
[0141] The first feature extraction processing module 302 is configured to perform feature extraction processing on the visual observation data to obtain target observation semantic features;
[0142] The second feature extraction processing module 303 is configured to perform feature extraction processing on the navigation target data to obtain target image features;
[0143] The feature fusion processing module 304 is configured to perform feature fusion processing on the target observation semantic features and the target image features to generate a conditional vector;
[0144] The conditional diffusion model training module 305 is configured to input the conditional vector into a conditional diffusion model for diffusion training to generate a number of initial candidate paths;
[0145] The candidate path optimization module 306 is configured to construct explicit constraints based on the initial image data, and perform path optimization on each of the initial candidate paths according to the explicit constraints to obtain a number of candidate optimized paths; the explicit constraints include a collision constraint and a target distance constraint;
[0146] The target navigation path determination module 307 is configured to determine a target navigation path from each of the candidate optimized paths according to a path selection strategy.
[0147] It can be understood that the content in the above method embodiments is applicable to the system embodiments of the present application. The functions specifically implemented in the system embodiments of the present application are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0148] An embodiment of the present application further provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned visual navigation method guided by a fusion diffusion model and explicit constraints. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.
[0149] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented in the device embodiments of the present application are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0150] Please refer to Figure 4 , Figure 4 which schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0151] The processor 401 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0152] The memory 402 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 402 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 402 and are called by the processor 401 to execute a visual navigation method combining a diffusion model and explicit constraint guidance provided in the embodiments of the present application;
[0153] The input / output interface 403 is used to implement information input and output;
[0154] The communication interface 404 is used to implement communication interaction between this device and other devices, and can communicate through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0155] The bus 405 transmits information between various components of the device (such as the processor 401, the memory 402, the input / output interface 403, and the communication interface 404);
[0156] Among them, the processor 401, the memory 402, the input / output interface 403, and the communication interface 404 are communicatively connected to each other inside the device through the bus 405.
[0157] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned visual navigation method combining a diffusion model and explicit constraint guidance is implemented.
[0158] It can be understood that the content in the above method embodiments is applicable to the storage medium embodiments. The functions specifically implemented by the storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0159] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0160] A visual navigation method integrating a diffusion model and explicit constraint guidance, and a visual navigation system integrating a diffusion model and explicit constraint guidance provided by embodiments of the present application. The method includes obtaining initial image data; the initial image data includes visual observation data and navigation target data; performing feature extraction processing on the visual observation data to obtain target observation semantic features; performing feature extraction processing on the navigation target data to obtain target image features; performing feature fusion processing on the target observation semantic features and the target image features to generate a conditional vector; inputting the conditional vector into a conditional diffusion model for diffusion training to generate a number of initial candidate paths; constructing explicit constraints based on the initial image data, and optimizing the paths of each initial candidate path according to the explicit constraints to obtain a number of candidate optimized paths; the explicit constraints include a collision constraint and a target distance constraint; determining a target navigation path from each candidate optimized path according to a path selection strategy. Embodiments of the present application generate multi-modal paths through a conditional diffusion model, and introduce an explicit cost gradient guidance strategy and a path selection strategy to determine the target navigation path, effectively solving the problems of insufficient path safety and poor dynamic adaptability in traditional visual navigation methods, and significantly improving the reliability and practicality of the navigation system. Among them, the introduction of the collision constraint makes the generated path unlikely to collide with obstacles, improving the safety of the path; the introduction of the target distance constraint guides the path towards the target point, enhancing the reachability of the path; embodiments of the present application learn the multi-scene path distribution through a pre-trained diffusion model, and at the same time introduce explicit cost constraints, realizing zero-shot transfer to the real world without retraining for a new scene, further improving the path safety and path navigation efficiency.
[0161] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0162] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0163] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0164] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware and their appropriate combinations.
[0165] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0166] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0167] In several embodiments provided by the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of systems or units can be in electrical, mechanical or other forms.
[0168] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0169] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0170] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs.
[0171] The preferred embodiments of the embodiments of the present application have been described above with reference to the drawings, and thus do not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.< / goal> < / obs> < / goal> < / goal> < / obs>
Claims
1. A visual navigation method integrating diffusion models and explicit constraint guidance, characterized in that, The method includes the following steps: Obtain initial image data; the initial image data includes visual observation data and navigation target data; Perform feature extraction processing on the visual observation data to obtain target observation semantic features; Perform feature extraction processing on the navigation target data to obtain target image features; Perform feature fusion processing on the target observation semantic features and the target image features to generate a conditional vector; Input the conditional vector into a conditional diffusion model for diffusion training to generate a number of initial candidate paths; Construct explicit constraints based on the initial image data, and optimize the paths of each of the initial candidate paths according to the explicit constraints to obtain a number of candidate optimized paths; the explicit constraints include collision constraints and target distance constraints; Determine a target navigation path from each of the candidate optimized paths according to a path selection strategy.
2. The method according to claim 1, wherein The performing feature extraction processing on the visual observation data to obtain target observation semantic features includes: Input the visual observation data into a visual encoder; Perform feature extraction processing on the visual observation data through the visual encoder to obtain the target observation semantic features.
3. The method according to claim 1, wherein, The navigation target data includes image target data and point target data. The performing feature extraction processing on the navigation target data to obtain target image features includes: Input the image target data into a target encoder, and perform feature extraction processing on the image target data through the target encoder to obtain the target image features; Or, Input the point target data into the target encoder, and perform feature extraction processing on the point target data through the target encoder to obtain the target image features.
4. The method according to claim 1, characterized in that The performing feature fusion processing on the target observation semantic features and the target image features to generate a conditional vector includes: Perform feature fusion processing on the target observation semantic features and the target image features through a cross-attention mechanism to generate the conditional vector.
5. The method according to claim 1, wherein The inputting the conditional vector into a conditional diffusion model for diffusion training to generate a number of initial candidate paths includes: Input the conditional vector into the conditional diffusion model, and perform forward noise addition training on the conditional vector and a real path point sequence through the conditional diffusion model to generate a noise path; Perform reverse denoising inference on the noise path through the conditional diffusion model to generate a number of the initial candidate paths.
6. The method according to claim 1, wherein The constructing explicit constraints based on the initial image data, and optimizing the paths of each of the initial candidate paths according to the explicit constraints to obtain a number of candidate optimized paths includes: Construct the target distance constraint based on the navigation target data and the initial candidate paths; Construct the collision constraint based on the visual observation data; Calculate the total constraint cost gradient of the explicit constraints according to the target distance constraint and the collision constraint; Optimize the paths of each of the initial candidate paths according to the total constraint cost gradient to obtain a number of the candidate optimized paths.
7. The method according to claim 6, wherein The constructing the target distance constraint based on the navigation target data and the initial candidate paths includes: Perform data extraction processing on the navigation target data to obtain the target point coordinates; Perform data extraction processing on the initial candidate path to obtain the end coordinates of the initial candidate path; Construct the target distance constraint according to the target point coordinates and the end coordinates of the initial candidate path.
8. The method according to claim 6, wherein The constructing of the collision constraint based on the visual observation data includes: Input the visual observation data into a depth estimation model to generate a depth map; Construct a local TSDF map according to the depth map; Determine the collision constraint according to the local TSDF map.
9. The method according to claim 1, characterized in that, The path selection strategy includes a consistency screening strategy and a continuity optimization strategy. The determining of the target navigation path from each of the candidate optimized paths according to the path selection strategy includes: According to the consistency screening strategy, screen the current path with the same path direction as the historical path from several candidate optimized paths; According to the continuity optimization strategy, perform weighted average filtering on the current path and the historical path to obtain the target navigation path.
10. A visual navigation system integrating a diffusion model and explicit constraint guidance, characterized in that The system includes the following modules: An initial image data acquisition module for acquiring initial image data; the initial image data includes visual observation data and navigation target data; A first feature extraction processing module for performing feature extraction processing on the visual observation data to obtain target observation semantic features; A second feature extraction processing module for performing feature extraction processing on the navigation target data to obtain target image features; A feature fusion processing module for performing feature fusion processing on the target observation semantic features and the target image features to generate a conditional vector; A conditional diffusion model training module for inputting the conditional vector into a conditional diffusion model for diffusion training to generate several initial candidate paths; A candidate path optimization module for constructing explicit constraints based on the initial image data and performing path optimization on each of the initial candidate paths according to the explicit constraints to obtain several candidate optimized paths; the explicit constraints include collision constraints and target distance constraints; A target navigation path determination module for determining the target navigation path from each of the candidate optimized paths according to the path selection strategy.
Citation Information
Cited By
Port multi-intelligent-terminal cooperative trajectory generation method, device and equipment
CN120972956A
Self-driving automobile diffusion navigation method based on visual language model guidance
CN121026178A
Quadruped robot parkour navigation method and system based on multi-modal feature fusion
CN121384037A
Action planning method and device based on visual language guidance and differentiation diffusion
CN122067159A