Visual navigation method and system based on de-noising diffusion bridge model
Through the visual navigation method of the denoising diffusion bridge model, the reverse diffusion is used to use prior actions as the initial state to perform reverse diffusion, which solves the problem of low adaptability and efficiency of visual navigation in complex environments, and achieves a fast and stable navigation strategy.
Patent Information
- Application Number
- CN202510417988.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing visual navigation technology has poor adaptability in unknown or dynamic environments. Traditional methods rely on high-precision sensors to update maps in real time. Deep learning methods have high training costs and limited generalization capabilities. The initial noise and target action distribution of the method based on diffusion model are large, resulting in high computational complexity and sparse generated actions.
The denoising diffusion bridge model is used to obtain the initial image data for feature encoding, generate context vectors and perform linear modulation, and generate prior actions in combination with preset motion rules. The prior actions are used as the initial state of the diffusion bridge model for reverse denoising processing to generate target navigation actions.
Effectively reduce redundant iterations in the diffusion process, reduce cumulative errors, improve the efficiency and stability of action generation, adapt to complex dynamic environments, and realize fast and stable visual navigation strategies.
Smart Images

Figure CN120333439A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a visual navigation method and system based on a denoising diffusion bridge model. Background Art
[0002] Visual navigation is an important technology for autonomous mobile robots to achieve goal-oriented movement in complex environments. The current visual navigation technologies mainly include methods based on mapping and path planning or methods based on deep learning. Traditional path planning-based methods are stable, but highly dependent on accurate maps and environmental perception, with poor adaptability in unknown or dynamic environments and difficulty in real-time adjustment of navigation strategies; deep reinforcement learning-based methods can adapt to complex environments, but have high training costs, low sample efficiency, and limited generalization ability, requiring a large amount of training data to adapt to new environments; imitation learning methods can learn navigation strategies through expert demonstration data, but the data quality has a great impact on the final navigation performance, and it is also difficult to be generalized in unknown environments. In recent years, visual navigation methods based on diffusion models have become a research hotspot. However, such methods rely on Gaussian noise as the initial input, resulting in a large deviation between the distribution of target actions and the actual navigation requirements, thereby increasing the computational complexity of the denoising step. In addition, the target actions generated by the diffusion model are relatively sparse and difficult to directly apply to complex navigation tasks, especially in application scenarios that require precise control and immediate response.
[0003] In summary, the technical problems existing in the related technologies need to be improved. Summary of the Invention
[0004] The embodiments of this application aim to at least solve one of the technical problems in the related technologies to some extent. For this reason, the main purpose of the embodiments of this application is to propose a visual navigation method and system based on a denoising diffusion bridge model, which can effectively reduce redundant iterations in the diffusion process, reduce cumulative errors, greatly improve the efficiency and stability of action generation, and at the same time can adapt to complex dynamic environments.
[0005] To achieve the above object, on the one hand, an embodiment of this application proposes a visual navigation method based on a denoising diffusion bridge model, and the method includes the following steps:
[0006] Obtain initial image data;
[0007] Perform feature encoding processing on the initial image data to generate a context vector;
[0008] Perform linear modulation processing on the context vector to generate a conditional variable;
[0009] Generate a prior action according to the context vector and a preset motion rule;
[0010] Input the prior action and the conditional variable into a denoising diffusion bridge model for diffusion training to generate a target navigation action; wherein, the prior action is used as the initial state when the denoising diffusion bridge model performs reverse denoising processing.
[0011] In some embodiments, the initial image data includes an initial observation image and an initial target image. The feature encoding process for the initial image data to generate a context vector includes:
[0012] Perform feature extraction on the initial observation image to obtain observation image features;
[0013] Perform feature extraction on the initial target image to obtain target image features;
[0014] Perform feature fusion on the observation image features and the target image features to generate the context vector.
[0015] In some embodiments, the linear modulation process for the context vector to generate a conditional variable includes:
[0016] Input the context vector into a feature linear modulation module;
[0017] Generate modulation parameters according to the context vector through the feature linear modulation module;
[0018] Determine the conditional variable according to the modulation parameters through the feature linear modulation module.
[0019] In some embodiments, the generation of the prior action according to the context vector and a preset motion rule includes:
[0020] Map the context vector to an action space through a fully connected layer of a neural network model and generate low-dimensional action features corresponding to the context vector;
[0021] Generate an initial candidate path according to the low-dimensional action features and the preset motion rule;
[0022] Determine the prior action according to the initial candidate path.
[0023] In some embodiments, the method further includes:
[0024] Based on the Gaussian prior rule, sample data from a Gaussian white noise distribution to obtain the prior action;
[0025] Or,
[0026] Input the initial observation image in the initial image data into a conditional variational autoencoder model for training to obtain the prior action.
[0027] In some embodiments, inputting the prior action and the conditional variable into the denoising diffusion bridge model for diffusion training to generate a target navigation action includes:
[0028] Inputting the prior action and the conditional variable into the denoising diffusion bridge model;
[0029] Performing forward noise addition training according to the real path point sequence through the denoising diffusion bridge model to generate a noise action;
[0030] Performing backward denoising processing on the noise action according to the prior action and the conditional variable through the denoising diffusion bridge model to generate the target navigation action.
[0031] In some embodiments, after inputting the prior action and the conditional variable into the denoising diffusion bridge model for diffusion training to generate a target navigation action, the method further includes:
[0032] Generating a navigation action instruction according to the target navigation action;
[0033] Sending the target navigation action and the navigation action instruction to the motion control module of the robot, and controlling the robot to execute the target navigation action according to the navigation action instruction through the motion control module.
[0034] To achieve the above object, on the other hand, an embodiment of the present application proposes a visual navigation system based on a denoising diffusion bridge model, and the system includes the following modules:
[0035] An initial image data acquisition module, configured to acquire initial image data;
[0036] A feature encoding processing module, configured to perform feature encoding processing on the initial image data to generate a context vector;
[0037] A linear modulation processing module, configured to perform linear modulation processing on the context vector to generate a conditional variable;
[0038] A prior action generation module, configured to generate a prior action according to the context vector and a preset motion rule;
[0039] A target navigation action generation module, configured to input the prior action and the conditional variable into the denoising diffusion bridge model for diffusion training to generate a target navigation action; wherein, the prior action is used as an initial state when the denoising diffusion bridge model performs backward denoising processing.
[0040] To achieve the above object, on the other hand, an embodiment of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the foregoing method is implemented.
[0041] To achieve the above object, on the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the foregoing method is implemented.
[0042] The embodiments of the present application at least include the following beneficial effects: The present application provides a visual navigation method and system based on a denoising diffusion bridge model. The solution includes obtaining initial image data; performing feature encoding processing on the initial image data to generate a context vector; performing linear modulation processing on the context vector to generate a conditional variable; generating a prior action according to the context vector and a preset motion rule; inputting the prior action and the conditional variable into the denoising diffusion bridge model for diffusion training to generate a target navigation action; wherein, the prior action is used as the initial state when the denoising diffusion bridge model performs reverse denoising processing. By using the denoising diffusion bridge model, the embodiments of the present application perform a reverse diffusion process with the prior action distribution as the initial state of the diffusion bridge model, so that the initial state can be efficiently corrected to the target action distribution, thereby realizing a fast and stable visual navigation strategy. That is, the embodiments of the present application convert the diffusion denoising process starting from Gaussian noise in the traditional visual navigation method into a diffusion denoising process starting from an informative prior action distribution, which can effectively reduce redundant iterations in the diffusion process, reduce cumulative errors, greatly improve the efficiency and stability of action generation, and can adapt to complex dynamic environments at the same time. Description of the Drawings
[0043] Figure 1 is a flowchart of the steps of a visual navigation method based on a denoising diffusion bridge model provided by an embodiment of the present application;
[0044] Figure 2 is an overall flow schematic diagram of a visual navigation method based on a denoising diffusion bridge model provided by an embodiment of the present application;
[0045] Figure 3 is a structural schematic diagram of a visual navigation system based on a denoising diffusion bridge model provided by an embodiment of the present application;
[0046] Figure 4 is a hardware structural schematic diagram of the electronic device provided by an embodiment of the present application. Detailed Embodiments
[0047] In order to make the objectives, technical solutions, and advantages of this application more clearly understood, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not intended to limit this application. When the following description involves the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of this application. They are merely examples of systems and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0048] It can be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "while...", or "in response to determining".
[0049] The terms "at least one", "multiple", "each", "any one", etc. used in this application, at least one includes one, two, or more than two, multiple includes two or more than two, each refers to each of the corresponding multiple, and any one refers to any one of the multiple.
[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0051] Before elaborating in detail on the embodiments of this application, first explain some of the nouns and terms involved in the embodiments of this application. The nouns and terms involved in the embodiments of this application are subject to the following explanations.
[0052] 1) Embodied Intelligence refers to the ability of an intelligent agent to interact with the environment in real time through a perception-action closed loop to achieve autonomous decision-making and task execution. In the embodiments of this application, it specifically refers to an intelligent system that obtains environmental information through a visual sensor and combines physical movement to achieve a navigation goal.
[0053] 2) Visual Navigation, a technology that obtains environmental information through visual sensors (such as cameras, RGB-D (RGB-Depth) sensors), combines path planning and motion control algorithms, and guides an intelligent agent from a starting point to a target position. In the embodiments of this application, it specifically refers to visual navigation based on pure RGB images.
[0054] 3) Diffusion Model, a deep learning framework based on probabilistic generative models, which learns the data distribution by gradually adding noise (forward process) and denoising (reverse process). In the embodiments of this application, the Denoising Diffusion Bridge Model (DDBM for short) is used to generate the optimal action sequence from the current state to the target state.
[0055] 4) Imitation Learning, a method of training the behavior of an intelligent agent by imitating expert demonstration data (such as human operations or predefined policies). In the embodiments of this application, combined with the diffusion model, the noise distribution of the expert trajectory is modeled as the basis for generating the navigation policy.
[0056] 5) Action Policy, the rule or probability distribution by which an intelligent agent generates the next action according to the current state. In the embodiments of this application, the action policy is generated by the denoising diffusion bridge model, and the long-term cumulative reward is optimized with the Markov Decision Process (MDP) framework.
[0057] 6) Denoising Process, the reverse operation in the diffusion model that gradually restores the noisy data to the original data. In the embodiments of this application, this process is used to iteratively optimize from the initially noisy trajectory to generate a smooth and physically constrained navigation path.
[0058] Visual navigation is an important technology for autonomous mobile robots to achieve goal-oriented movement in complex environments. Current visual navigation technologies mainly include methods based on mapping and path planning or methods based on deep learning. Traditional path planning-based methods are stable, but highly dependent on accurate maps and environmental perception, with poor adaptability in unknown or dynamic environments and difficulty in real-time adjusting navigation strategies; deep reinforcement learning-based methods can adapt to complex environments, but have high training costs, low sample efficiency, and limited generalization ability, requiring a large amount of training data to adapt to new environments; imitation learning methods can learn navigation strategies through expert demonstration data, but the data quality has a great impact on the final navigation performance, and it is also difficult to generalize in unknown environments. In recent years, visual navigation methods based on diffusion models have become a research hotspot. However, such methods rely on Gaussian noise as the initial input, resulting in a large deviation between the distribution of target actions and the actual navigation requirements, thus increasing the computational complexity of the denoising step. In addition, the target actions generated by diffusion models are relatively sparse and difficult to directly apply to complex navigation tasks, especially in application scenarios that require precise control and instant response.
[0059] Exemplarily, traditional mapping and path planning technologies usually rely on high-precision sensors (such as lidar) for accurate ranging and positioning. However, in unknown or dynamic environments, it is difficult to update the map in real time, with high computational complexity, and it is also sensitive to noise and lighting conditions. Deep Reinforcement Learning (DRL) methods optimize navigation strategies through the trial-and-error learning of agents in simulation environments, but have high training costs, low sample efficiency, and difficulty in generalizing in different environments. Imitation Learning (IL) methods can quickly adapt to target tasks by learning expert demonstration data, but are highly dependent on the quality of expert data. In the framework of imitation learning, in recent years, visual navigation methods based on Diffusion Model (DM) have become a new research direction. These methods use the diffusion process to model data, adopting the Denoising Diffusion Probabilistic Model (DDPM). During the training stage, Gaussian noise is added to the data to make its distribution tend to the standard Gaussian distribution. During the inference stage, target navigation actions are generated through the denoising process. However, traditional diffusion models start from random Gaussian noise and require a large number of steps during the generation process to "correct" the noise, resulting in an inefficient generation process; the generated actions often lack sufficient guidance and are difficult to achieve stable navigation in complex scenarios; due to the large difference between the initial noise and the target actions, the model needs to continuously make a large number of corrections during the step-by-step denoising process, increasing the inference time and computational overhead.
[0060] In view of this, an embodiment of the present application provides a visual navigation method and system based on a denoising diffusion bridge model. The solution includes obtaining initial image data; performing feature encoding processing on the initial image data to generate a context vector; performing linear modulation processing on the context vector to generate a conditional variable; generating a prior action according to the context vector and a preset motion rule; inputting the prior action and the conditional variable into the denoising diffusion bridge model for diffusion training to generate a target navigation action; wherein the prior action is used as the initial state when the denoising diffusion bridge model performs reverse denoising processing. By using the denoising diffusion bridge model, the embodiment of the present application takes the prior action distribution as the initial state of the diffusion bridge model to perform the reverse diffusion process, so that the initial state can be efficiently corrected to the target action distribution, thereby realizing a fast and stable visual navigation strategy. That is, the embodiment of the present application converts the diffusion denoising process starting from Gaussian noise in the traditional visual navigation method into a diffusion denoising process starting from an informative prior action distribution, which can effectively reduce redundant iterations in the diffusion process, reduce cumulative errors, greatly improve the efficiency and stability of action generation, and at the same time can adapt to complex dynamic environments.
[0061] A visual navigation method based on a denoising diffusion bridge model provided by an embodiment of the present application relates to the field of computer technology. The visual navigation method based on a denoising diffusion bridge model provided by an embodiment of the present application can be applied to a terminal, can also be applied to a server, or can be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements a visual navigation method based on a denoising diffusion bridge model, etc., but is not limited to the above forms.
[0062] This application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs (Personal Computers), minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0063] Please refer to Figure 1 , Figure 1 which is an optional step flowchart of a visual navigation method based on a denoising diffusion bridge model provided by an embodiment of this application. Figure 1 The method in
[0064] Step S101, obtain initial image data;
[0065] Among them, the initial image data refers to the original image without any processing. The initial image data includes an initial observation image and an initial target image.
[0066] For the initial observation image, it is the original image information about the observation scene or object directly obtained through visual means (such as cameras, sensors, etc.), which reflects the visual characteristics such as the appearance, position, and posture of objects in the actual scene and is an intuitive visual record of the real world. In the embodiment of this application, the initial observation image is the image of consecutive moments observed by the robot, that is, the RGB (Red Green Blue Color Mode) observation sequence.
[0067] For the initial target image, it refers to a single RGB-formatted picture given in advance, which is used to guide the denoising diffusion bridge model to output actions moving towards the target image.
[0068] Step S102, perform feature encoding processing on the initial image data to generate a context vector;
[0069] Among them, the initial image data includes an initial observation image and an initial target image.
[0070] In some embodiments, step S102 may include: performing feature extraction processing on the initial observation image to obtain observation image features; performing feature extraction processing on the initial target image to obtain target image features; performing feature fusion processing on the observation image features and the target image features to generate a context vector.
[0071] In a specific implementation, the robot collects a continuous RGB observation sequence and a target image, extracts visual features of the RGB observation sequence and the target image via a convolutional network or a Transformer encoder respectively, and performs feature fusion processing on the visual features corresponding to the RGB observation sequence and the target image to generate a context vector C that describes the current scene and target information. t 。
[0072] Step S103, performing linear modulation processing on the context vector to generate a conditional variable;
[0073] In some embodiments, step S103 may include: inputting the context vector into a feature linear modulation module; generating modulation parameters according to the context vector by the feature linear modulation module; determining the conditional variable according to the modulation parameters by the feature linear modulation module.
[0074] In a specific implementation, the context vector C t is passed to the feature linear modulation module, and the feature linear modulation module generates modulation parameters a and b according to the input context vector C t and then forms a conditional variable through feature linear modulation according to the modulation parameters a and b. The conditional variable is used to adjust the behavior of the generation strategy, and the conditional variable obtained after modulation is subsequently input into the diffusion bridge module to guide the generation of the final navigation action.
[0075] Step S104, generating a prior action according to the context vector and a preset motion rule;
[0076] In some embodiments, step S104 may include: mapping the context vector to an action space through a fully connected layer of a neural network model and generating a low-dimensional action feature corresponding to the context vector; generating an initial candidate path according to the low-dimensional action feature and the preset motion rule; determining the prior action according to the initial candidate path.
[0077] In some embodiments, it may further include: sampling data from a Gaussian white noise distribution based on a Gaussian prior rule to obtain a prior action; or inputting the initial observation image in the initial image data into a conditional variational autoencoder model for training to obtain a prior action.
[0078] In a specific implementation, the prior action generation module provides three prior generation strategies, namely, the Gaussian prior generation strategy, the rule-based prior generation strategy, and the learning-based prior generation strategy.
[0079] Among them, the Gaussian prior generation strategy refers to directly sampling from the Gaussian white noise distribution as a comparison benchmark for traditional diffusion models. The rule-based prior generation strategy refers to mapping the context vector C t to the action space and combining preset rules (for example, using the predicted path length ρ and the confidence θ of motion direction classification, such as going straight, turning left, turning right, U-turn, etc.) to generate a parabolic path as the initial action. The learning-based prior generation strategy refers to adopting a lightweight Conditional Variational Autoencoder (CVAE) model to learn the relationship between observations and actions from expert data to obtain prior actions that better meet the requirements of actual tasks.
[0080] Among them, the rule-based prior generation strategy maps C t , combines preset motion rules (such as the parabolic generation formula) to generate the initial action. Exemplarily, the parabola can solve for parameters according to the standard form y = ax 2 + bx + c or the vertex form y = a(x - h) 2 + k to ensure that the path is smooth and meets the kinematic requirements.
[0081] In a specific implementation, the specific generation principle and generation process of the rule-based prior generation strategy are as follows:
[0082] 1) Design principle: Use a Fully Connected Layer (FC) to map environmental perception information (such as observed image features) to the action space, extract low-dimensional features to generate preliminary prior actions; the prior actions follow preset trajectory constraints, which are usually set as parabolic trajectories to simulate natural motion trends and improve motion smoothness; determine the preliminary path based on the predicted path length ρ and motion behavior classification (for example: going straight, turning left, turning right, right U-turn, left U-turn).
[0083] 2) Trajectory modeling: Use the parabolic function y = a(x - h) 2 + k to define the path so that it passes through the current robot position and the target point. Among them, the direction and shape of the parabola are affected by the path length and angle perturbation, thereby generating different types of trajectories.
[0084] Among them, h represents the axis of symmetry of the set path, and k represents the vertex height of the set path. The constraint a < 0 is ensured to make the trajectory bend downward, so as to simulate a reasonable path for the robot to move towards the target point.
[0085] 3) Noise modeling: To improve the generalization ability, an adaptive noise mechanism based on prediction confidence is designed. Specifically: when the model has a high prediction confidence in the current state, the noise perturbation is reduced to maintain accuracy; when the prediction confidence is low, the noise perturbation is increased to encourage exploration of different paths.
[0086] Among them, the noise standard deviation σ is calculated according to the following formula:
[0087] σ = min_std + (max_std - min_std)·(1 - confidence)
[0088] Among them, max_std is the pre - defined maximum value of the standard deviation, min_std is the pre - defined minimum value of the standard deviation, and confidence is the confidence of the behavior classification output by the fully - connected layer.
[0089] The rule - based prior generation strategy is applicable to the noise adjustment of the path length and angle direction, ensuring that the path is both constrained and adaptable to different scenarios.
[0090] In the embodiments of the present application, by providing a variety of prior generation strategies (including Gaussian, rule - based, learning - based, etc.), the system can obtain better performance in different scenarios and adapt to complex dynamic environments.
[0091] Step S105: Input the prior action and the conditional variable into the denoising diffusion bridge model for diffusion training to generate the target navigation action; among them, the prior action is used as the initial state when the denoising diffusion bridge model performs reverse denoising processing.
[0092] In some embodiments, step S105 may include: inputting the prior action and the conditional variable into the denoising diffusion bridge model; through the denoising diffusion bridge model, performing forward noise - adding training according to the real path point sequence to generate a noise action; through the denoising diffusion bridge model, performing reverse denoising processing on the noise action according to the prior action and the conditional variable to generate the target navigation action.
[0093] In some embodiments, after step S105, it may further include: generating a navigation action instruction according to the target navigation action; sending the target navigation action and the navigation action instruction to the motion control module of the robot, and controlling the robot to execute the target navigation action according to the navigation action instruction through the motion control module.
[0094] In the embodiments of the present application, the prior action a TAs the initial state, reverse diffusion is performed through the DDBM framework. In each step of reverse diffusion, the neural network calculates the scoring function s(a t , time step t, considering a t ,t,a T ,T), and at the same time combines h(a t ,t,a T ,T) (the h-transform term of Doob) to adjust the action update direction. After k steps of iteration (usually fewer steps than required by traditional diffusion models), the target action a0 is output. The generated target action sequence a0 will be directly used for robot control to achieve the navigation task from the current state to the target position.
[0095] It should be noted that only in the first step of reverse diffusion, the initial action is used as a constraint term for the network, that is, the current state is equal to the initial action (prior action). During training, first, based on the h-transform term of Doob, the multi-step intermediate states between the target action and the initial action are obtained, and then the neural network is input with the current state, time step, and the change amount between two adjacent intermediate states corresponding to the learning time step.
[0096] It should be noted that the forward operation is to add noise (diffusion) to the ground truth in the dataset during model training to obtain intermediate states between the ground truth and the prior, enabling the model to learn the denoising process of gradually changing from the prior action to the ground truth. Since this is the default operation of the diffusion model, it will not be elaborated in this embodiment of the present application.
[0097] In specific implementation, the process of diffusion training for the denoising diffusion bridge model includes a forward diffusion process and a reverse denoising process. Specifically, the general forward diffusion process and reverse denoising process are as follows:
[0098] (1) Forward diffusion process: The data starts from the original distribution p0 and is gradually transformed into a simple distribution (usually a Gaussian distribution P T =N(0,I), where P T represents the initial data distribution in the diffusion model, N represents the normal distribution, and I represents the identity matrix. This forward diffusion process can be represented by a stochastic differential equation (SDE), and the specific expression is as follows:
[0099] dx t =f(x t ,t)dt + g(t)dw t
[0100] where x t represents the system state or sample variable at time t, which will change as time evolves; dx t represents the change in state x between time t and t + dtt Minor changes in; f(x t , t) is expressed as the drift term, which is a function of the current state and time and represents the trend of deterministic change; g(t) is expressed as the diffusion coefficient, which controls the intensity of random perturbation and is a function of time; dw t is expressed as a Wiener process, representing noise perturbation, and satisfies That is, a normal distribution with zero mean and variance dt, where d represents differentiation.
[0101] (2) Reverse denoising process: To recover the original data from the noise, a reverse SDE is used for representation, and the specific expression is as follows:
[0102]
[0103] Among them, represents the logarithmic probability density gradient with respect to x t , also known as the score function, which is used to indicate the direction in which the current state point rises fastest in the probability distribution; represents the Wiener process (Brownian motion) in the reverse time, which is different from w t in the forward process and is used for reverse simulation.
[0104] In the embodiments of the present application, in order to make the initial action distribution (prior distribution) closer to the target action distribution, the Doob's h-transform is introduced. The forward diffusion process and the reverse denoising process of the diffusion bridge in the embodiments of the present application are specifically as follows:
[0105] (1) The forward process of the diffusion bridge in the embodiments of the present application can be written as the following expression:
[0106] da t =[f(a t , t)+g(t) 2 h(a t , t, a0, T)]dt+g(t)dw t
[0107] Among them, is used to force the process to a fixed end point y; a t represents the action state at time t, which is an action variable rather than a state variable (different from the aforementioned x t ); a0 represents the target end point action of the diffusion process, usually the action in the data that is real or provided by an expert; T represents the total time or end time of the diffusion process.
[0108] (2) In the imitation learning scenario, the reverse diffusion bridge process can be represented by the following expression:
[0109]
[0110] Among them, s(a t , t, a T , T) is a conditional scoring function (which can be approximated as ), and in the embodiment of the present application, it is a convolutional neural network; h(a t , t, a T , T) is also used to guide the smooth transition of the distribution and serves as the target for the convolutional neural network to fit during the reverse denoising process.
[0111] In the embodiment of the present application, to quantify the difference between the prior distribution π s (a T ) and the target distribution π(a0), the KL (Kullback-Leibler Divergence) divergence is introduced, and the specific calculation formula is as follows:
[0112]
[0113] Among them, when the prior distribution is closer to the target distribution, the correction required for each step in the diffusion process is smaller, so the lower bound of the error is lower. That is:
[0114]
[0115] Among them, C represents a constant related to the noise schedule g(t) and the time step; represents the expected value when a0 is sampled from a certain distribution (usually the initial distribution π(a0)); D KL (π s (a T )||π(a0)) represents the Kullback–Leibler divergence (KL divergence), also known as relative entropy, which is an index to measure the difference between two probability distributions.
[0116] The embodiment of the present application constructs an initial state closer to the target action distribution, reduces the error caused by the KL divergence during the diffusion process, and thus generates an action sequence that better meets the task requirements.
[0117] In addition, the embodiment of the present application also designs a training loss function for the network of the entire visual navigation system. The training loss of the entire network consists of the following three parts (diffusion bridge loss L b , prior generation loss L p , and temporal distance loss L d ):
[0118] (1) Diffusion bridge loss L b ;
[0119] The weighted mean squared error (MSE) loss is adopted to measure the difference between the target action a0 and the output D of the network after denoising and correction. θ (a t ,t,a T ):
[0120]
[0121] where w(t) is the time weight scheduling function, and a0 represents the true value in the dataset.
[0122] (2) Prior generation loss L p ;
[0123] For the regular prior, when using the fully connected layer to generate the preliminary action, the MSE loss and the cross-entropy classification loss are used for calculation. The specific expressions are as follows:
[0124]
[0125] where p c is the one-hot encoding of the true motion category; is the predicted probability; represents the difference between the predicted value and the true value of the i-th sample; represents the predicted value of the i-th sample; a i represents the true value of the i-th sample; λ c and λ a are the weight coefficients of the classification loss and the regression loss respectively, both used to balance the loss effects of these two different tasks.
[0126] For the learned prior, when using the CVAE to generate the prior, the loss consists of the reconstruction error and the KL divergence. The specific expressions are as follows:
[0127]
[0128] where N represents the number of samples; O represents the conditional variable, which refers to the features of the image in the embodiments of this application; a represents the known action information; z represents the latent variable; KL(·) is a divergence calculation used to calculate the difference between two distributions. The KL divergence is used to force the posterior distribution to be close to the prior distribution to ensure that the model does not depend on the known action information a during generation; is the posterior distribution estimated by the CVAE inference model; π(z|O) is the conditional prior distribution of the CVAE, which only depends on the conditional variable O.
[0129] (3) Temporal distance loss L d ;
[0130] To better capture the temporal relationship between the target image and the current observation, the MSE loss between the predicted time distance and the true distance is introduced, and the specific expression is as follows:
[0131]
[0132] where f(c t,i ) is the temporal distance predicted by the context vector, and d i is the corresponding true distance.
[0133] Finally, the overall loss of the entire network is:
[0134] L = λ b L b + λ p L p + λ d L d
[0135] where the weights λ b , λ p and λ d are used to control the balance between modules.
[0136] By comprehensively using the diffusion bridge loss, the prior generation loss, and the temporal distance loss in the embodiments of this application, the collaborative optimization of each module is achieved, thereby improving the stability and accuracy of the overall navigation.
[0137] A visual navigation method based on the denoising diffusion bridge model provided by the embodiments of this application uses the denoising diffusion bridge model to transform the action generation problem in the visual navigation task into a continuous denoising process starting from an arbitrary informative prior (instead of traditional Gaussian noise). By introducing a variety of prior generation strategies, designing the corresponding denoising diffusion process, and jointly optimizing the loss function, not only the policy inference is accelerated (reducing the denoising steps), but also the smoothness of the navigation path and the task success rate are significantly improved. At the same time, it has high robustness and efficiency in a variety of environments (including indoor and outdoor, zero-shot, and adaptive tasks).
[0138] Steps S101 to S105 shown in the embodiments of the present application include obtaining initial image data; performing feature encoding processing on the initial image data to generate a context vector; performing linear modulation processing on the context vector to generate a conditional variable; generating a prior action according to the context vector and a preset motion rule; and inputting the prior action and the conditional variable into a denoising diffusion bridge model for diffusion training to generate a target navigation action. Among them, the prior action is used as the initial state when the denoising diffusion bridge model performs reverse denoising processing. By using the denoising diffusion bridge model, the embodiments of the present application perform a reverse diffusion process with the prior action distribution as the initial state of the diffusion bridge model, so that the initial state can be efficiently corrected to the target action distribution, thereby realizing a fast and stable visual navigation strategy. That is, the embodiments of the present application convert the diffusion denoising process starting from Gaussian noise in traditional visual navigation methods into a diffusion denoising process starting from an informative prior action distribution, which can effectively reduce redundant iterations in the diffusion process, reduce cumulative errors, greatly improve the efficiency and stability of action generation, and can adapt to complex dynamic environments at the same time.
[0139] To explain the principle of the technical solution of the present invention in detail, the overall process of the present invention will be described below in conjunction with some specific embodiments. It is easy to understand that the following is an explanation of the technical principle of the present invention and should not be regarded as a limitation of the present invention.
[0140] The core idea of the visual navigation method based on the denoising diffusion bridge model is to use the denoising diffusion bridge model (DDBM) to convert the traditional denoising process starting from Gaussian noise into a continuous denoising process starting from an informative prior action, so as to generate an action sequence that more efficiently and accurately meets the requirements of the target task. Please refer to Figure 2 , Figure 2 is a schematic diagram of the overall process of a visual navigation method based on the denoising diffusion bridge model provided by the embodiments of the present application; as Figure 2 shown, the entire visual navigation system mainly includes a feature extraction module, a feature linear modulation module, a prior generation module, and a diffusion bridge module. Among them, Figure 2 the orange square columns in the feature extraction module represent the features of all images after passing through the feature encoder. The red vertical bars of the prior generation strategy module based on learning represent the encoders of the features, and the blue-green vertical bars represent the decoders of the features respectively. z represents the latent space of the features. The feature extraction module uses RGB observations and target images, and performs feature extraction on the RGB observations and target images through a transformation encoder to generate a context vector; the feature linear modulation module applies the learned conditions to improve the path accuracy; and the prior generation module ensures that the policy output is better consistent with the target action, starting from the alternative prior actions generated using an optional prior policy; finally, in the policy output module, the system takes the current action and the features extracted from the image as inputs, and after denoising processing by the network, outputs the final action sequence.
[0141] In a specific implementation, such as Figure 2 shown, first, the feature extraction module receives the RGB observation image and the target image, extracts the key features of the RGB observation image and the target image through a visual encoder, and then aggregates these key features to form an embedding vector (context vector) that can describe the current state and the target state, and transmits this context vector as an input to the feature linear modulation module; then, in the feature linear modulation module, the system generates a conditional variable according to the input state feature (context vector), and this conditional variable adjusts the path generation process in the subsequent diffusion bridge module through feature linear modulation. The feature linear modulation module makes the generated path more precisely point to the target by adjusting the conditions of the policy, improving the accuracy and efficiency of the task; subsequently, the prior generation module generates alternative action trajectories according to the current state, and these trajectories are provided by the prior policy generator to provide a reference for the optimization of the policy; finally, in the policy output module, the system takes the current action and the features extracted from the image as inputs, and after denoising processing by the network, outputs the final action sequence, and this action sequence can effectively guide the agent to execute the task and achieve the expected goal in the environment.
[0142] Specifically, the specific implementation processes of the feature extraction module, the feature linear modulation module, the prior generation module, and the diffusion bridge module are as follows:
[0143] 1) The feature extraction module and the feature linear modulation module;
[0144] In a specific implementation, the robot collects a continuous RGB observation sequence and the target image, extracts the visual features of the RGB observation sequence and the target image through a convolutional network or a Transformer encoder respectively, and performs feature fusion processing on the visual features corresponding to the RGB observation sequence and the target image to generate a context vector C that describes the current scene and target information t ; subsequently, the context vector C t is transmitted to the feature linear modulation module, and the feature linear modulation module generates modulation parameters a and b according to the input context vector C t and then forms a conditional variable through feature linear modulation according to the modulation parameters a and b. The conditional variable is used to adjust the behavior of the generation policy, and the conditional variable obtained after modulation is then input into the diffusion bridge module to guide the generation of the final navigation action.
[0145] 2) The prior action generation module;
[0146] Utilize the context vector C t to generate a preliminary action estimate a T (prior action, that is Figure 2(source actions in). The prior action generation module provides three prior generation strategies, namely, Gaussian prior generation strategy, rule-based prior generation strategy, and learning-based prior generation strategy.
[0147] Among them, the Gaussian prior generation strategy refers to directly sampling from the Gaussian white noise distribution as a comparison benchmark for traditional diffusion models. The rule-based prior generation strategy refers to mapping the context vector C t to the action space and combining preset rules (for example, using the predicted path length ρ and the confidence θ of motion direction classification, such as going straight, turning left, turning right, U-turn, etc.) to generate a parabolic path as the initial action. The learning-based prior generation strategy refers to adopting a lightweight conditional variational autoencoder model to learn the relationship between observations and actions from expert data to obtain prior actions that better meet the actual task requirements. Among them, Figure 2 in the learning-based prior generation strategy, the input image is the initial RGB observation sequence.
[0148] 3) Denoising diffusion bridge module;
[0149] Taking the prior action a T as the starting point of the diffusion process, it is gradually corrected through multiple steps of reverse diffusion (using reverse SDE (Stochastic Differential Equation) or probability flow ODE (Ordinary Differential Equation)) to finally obtain the target action a0. In this process, the network continuously updates the action in an iterative manner to make it gradually converge to the target action distribution, thereby generating a high-quality action sequence for the robot to execute.
[0150] As Figure 2 shown, the overall implementation process of a visual navigation method based on the denoising diffusion bridge model includes the following four steps:
[0151] (1) Input processing;
[0152] Specifically, first, collect the RGB image sequence and the target image I g during the current time period; then, use a convolutional network or a Transformer encoder to extract the features of each frame of the RGB image sequence and the target image I g respectively; next, through an aggregation technique, aggregate the features extracted from the RGB image sequence and the target image I g to generate the context vector C t .
[0153] (2) Prior generation (generating a preliminary action a through different strategies) T ):
[0154] One is the Gaussian prior generation strategy: directly sample a T ~ N(0, I).
[0155] The other is the rule-based prior generation strategy: the fully connected layer maps C t , combined with a preset motion rule (such as a parabola generation formula) to generate an initial action.
[0156] Exemplarily, the parabola can solve for parameters according to the standard form y = ax 2 + bx + c or the vertex form y = a(x - h) 2 + k to ensure that the path is smooth and meets the kinematic requirements.
[0157] In specific implementation, the specific generation principle and generation process of the rule-based prior generation strategy are as follows:
[0158] 1) Design principle: Use a fully connected layer to map environmental perception information (such as observed image features) to the action space, extract low-dimensional features to generate a preliminary prior action; the prior action follows a preset trajectory constraint, which is usually set as a parabola trajectory to simulate the natural motion trend and improve the motion smoothness; determine the preliminary path according to the predicted path length d and motion behavior classification (for example: straight, left turn, right turn, right U-turn, left U-turn).
[0159] 2) Trajectory modeling: Use the parabola function y = a(x - h) 2 + k to define the path so that it passes through the current robot position and the target point. Among them, the direction and shape of the parabola are affected by the path length and angle perturbation, thereby generating different types of trajectories.
[0160] Among them, h represents the axis of symmetry of the set path, k represents the vertex height of the set path, and the constraint a < 0 is used to ensure that the trajectory bends downward, thereby simulating a reasonable path for the robot to move towards the target point.
[0161] 3) Noise modeling: To improve the generalization ability, an adaptive noise mechanism based on prediction confidence is designed. Specifically: when the model has a high prediction confidence in the current state, reduce the noise perturbation to maintain accuracy; when the prediction confidence is low, increase the noise perturbation to encourage exploration of different paths.
[0162] Among them, the noise standard deviation σ is calculated according to the following formula:
[0163] σ = min_std + (max_std - min_std) · (1 - confidence)
[0164] Among them, max_std is the predefined maximum standard deviation, min_std is the predefined minimum standard deviation, and confidence is the confidence of the behavior classification output by the fully connected layer.
[0165] The rule-based prior generation strategy is applicable to the noise adjustment of the path length and angle direction, ensuring that the path is both restrictive and adaptable to different scenarios.
[0166] Thirdly, the learning-based prior generation strategy: Using the trained CVAE model, generate prior actions that are more in line with the expert demonstrations according to the input observation data.
[0167] (3) Denoising diffusion bridge process: Take the prior action a T as the initial state, perform reverse diffusion through the DDBM framework. In each step of reverse diffusion, the neural network calculates the scoring function s(a t , t, a t , t, a T , T) according to the current state a t , t, a T , T), and at the same time combine h(a t , t, a T , T) (Doob's h-transform term) to adjust the action update direction. After k steps of iteration (usually fewer steps than required by traditional diffusion models), output the target action a0.
[0168] (4) Output and decision-making: The generated target action sequence a0 will be directly used for robot control to achieve the navigation task from the current state to the target position.
[0169] It should be noted that this embodiment only briefly illustrates the general process of a visual navigation method based on the denoising diffusion bridge model. For the detailed description of each step, reference can be made to the relevant content in the foregoing embodiments, which will not be elaborated here. It can be understood that the present invention places no restrictions on this.
[0170] In the embodiments of the present application, initial image data is obtained; feature encoding processing is performed on the initial image data to generate a context vector; linear modulation processing is performed on the context vector to generate a conditional variable; a prior action is generated according to the context vector and a preset motion rule; the prior action and the conditional variable are input into a denoising diffusion bridge model for diffusion training to generate a target navigation action; wherein, the prior action is used as the initial state when the denoising diffusion bridge model performs reverse denoising processing. In the embodiments of the present application, by using the denoising diffusion bridge model and taking the prior action distribution as the initial state of the diffusion bridge model for the reverse diffusion process, the initial state can be efficiently corrected to the target action distribution, thereby realizing a fast and stable visual navigation strategy. That is, the embodiments of the present application convert the diffusion denoising process starting from Gaussian noise in the traditional visual navigation method into a diffusion denoising process starting from an informative prior action distribution, which can effectively reduce redundant iterations in the diffusion process, reduce cumulative errors, greatly improve the efficiency and stability of action generation, and at the same time can adapt to complex dynamic environments.
[0171] A visual navigation method based on a denoising diffusion bridge model provided by the embodiments of the present application can be applied to application scenarios such as mobile robots and autonomous driving systems. Exemplarily, the visual navigation method based on the denoising diffusion bridge model provided in the embodiments of the present application is integrated into a mobile robot system, and the mobile robot system includes a visual sensor, a navigation strategy module, and a motion control module. Specifically, the visual sensor (such as an RGB camera) is responsible for collecting the current environment image and the target image; the navigation strategy module uses the denoising diffusion bridge model to process the image features and generate a navigation action instruction; after receiving the navigation action instruction, the motion control module drives the robot to execute the corresponding action, thereby realizing the autonomous navigation and target positioning of the robot in the environment. The entire system forms a closed-loop process of perception - decision - execution, improving the real-time performance and robustness of navigation.
[0172] In the embodiments of the present application, by introducing prior action information, the target action is generated in a more efficient manner, improving the computational efficiency and generalization ability of the model. Compared with traditional methods, the embodiments of the present application reduce the denoising steps, optimize the generation quality of the target action, and enhance the adaptability to different environments.
[0173] In addition, for the visual navigation method based on the denoising diffusion bridge model (DDBM) proposed in the embodiments of the present application, its core lies in using an informative prior action as the initial state of the diffusion process and guiding the prior to the target action distribution through Doob's h-transform, thereby realizing more efficient and accurate action generation. Although there are some alternative solutions currently, they all have obvious deficiencies, and the specific analysis is as follows:
[0174] 1) For traditional diffusion model solutions: methods such as NoMaD (Goal-Masked Diffusion Policies for Navigation and Exploration). The NoMaD method directly starts the diffusion process from Gaussian noise as the initial state. Since the initial state has a large difference from the target action distribution, such methods require a large number of denoising iteration steps, resulting in long inference time and poor real-time performance. At the same time, the generation accuracy is also difficult to guarantee.
[0175] 2) For generative adversarial networks (GANs) or flow-based models, these methods can also be used to generate continuous actions. However, the training process of GANs is unstable, and there are often mode collapse problems in high-dimensional action spaces. Although flow-based models are reversible, they have high computational complexity and weak real-time response capabilities, and are difficult to meet the requirements of visual navigation in complex dynamic environments.
[0176] In summary, the visual navigation method based on the denoising diffusion bridge model provided by the embodiments of this application can simultaneously have the following technical contents: using prior information (such as rule-based or learning-based prior generation) to significantly reduce the distribution difference between the initial state and the target state; effectively reducing the denoising iteration steps through the diffusion bridge technology to improve the generation efficiency and action accuracy; meeting the high requirements of visual navigation in complex environments in terms of real-time performance and robustness. Therefore, the denoising diffusion bridge model guided by informative priors in the embodiments of this application can achieve faster, more stable, and higher-precision action generation in visual navigation, and can adapt to complex dynamic environments at the same time.
[0177] The key point of the visual navigation method based on the denoising diffusion bridge model (DDBM) proposed by the embodiments of this application is:
[0178] 1) Using the denoising diffusion bridge model, the traditional diffusion denoising process starting from Gaussian noise is converted to start from the informative prior action distribution, that is, using the prior action distribution as the initial state of the diffusion bridge model, and guiding the reverse diffusion process through Doob's $h$-transform, so that the initial state is efficiently corrected to the target action distribution, realizing a fast and stable visual navigation strategy, thereby greatly improving the efficiency and stability of action generation.
[0179] 2) Multi-strategy prior generation method. The multi-strategy prior generation method mainly includes Gaussian prior generation method, rule-based prior generation method, and learning-based prior generation method. Among them, the rule-based prior generation (using fully connected mapping combined with parabolic path generation, analyzing scene information in real time and adding appropriate noise to ensure action diversity) and the learning-based prior generation (using a lightweight CVAE model to learn expert data) are the two main prior generation mechanisms.
[0180] 3) Joint optimization training scheme: By designing diffusion bridge loss, prior generation loss (including action prediction error and motion direction classification loss), and temporal distance loss, the collaborative optimization of visual feature extraction, prior generation, and denoising diffusion bridge module is realized, thereby improving the consistency between the generated action and the target action distribution.
[0181] 4) Introduce KL divergence as a measure of the closeness between the prior and the target distribution, and prove the relationship between error bound and noise scheduling, providing theoretical support for method design. It is theoretically proved that the KL divergence between the prior distribution and the target action distribution affects the error of the diffusion process, ensuring that when the prior is closer to the target distribution, the fewer the number of denoising correction steps required, thus improving the accuracy of the generated action.
[0182] Based on these four key points, the visual navigation method based on the denoising diffusion bridge model (DDBM) proposed in the embodiment of this application can improve the adaptability and robustness of visual navigation in complex dynamic environments. Even in the presence of noise and environmental changes, it can effectively reduce the redundant iterations in the diffusion process, reduce the cumulative error, and ensure the smoothness of the navigation path and the task success rate.
[0183] Therefore, the visual navigation method based on the denoising diffusion bridge model (DDBM) proposed in the embodiment of this application effectively avoids the redundant denoising steps brought by the traditional diffusion model starting from Gaussian noise by introducing informative prior action distributions (including rule-based and learning-based priors), thereby greatly improving the efficiency and accuracy of target action generation. In addition, by jointly optimizing the diffusion bridge loss, prior generation loss, and temporal distance loss, the model's fitting ability to the target action distribution and its adaptability to different environments are further enhanced, achieving faster and more stable real-time visual navigation as a whole. Compared with traditional visual navigation methods, the embodiment of this application has obvious advantages in terms of real-time performance, accuracy, and adaptability to complex dynamic environments.
[0184] Please refer to Figure 3 , the embodiment of this application also provides a visual navigation system 300 based on the denoising diffusion bridge model, which can implement the above-mentioned visual navigation method based on the denoising diffusion bridge model. The system includes the following modules:
[0185] The initial image data acquisition module 301 is configured to acquire initial image data;
[0186] The feature encoding processing module 302 is configured to perform feature encoding processing on the initial image data to generate a context vector;
[0187] The linear modulation processing module 303 is configured to perform linear modulation processing on the context vector to generate a conditional variable;
[0188] The prior action generation module 304 is configured to generate a prior action according to the context vector and a preset motion rule;
[0189] The target navigation action generation module 305 is configured to input the prior action and the conditional variable into a denoising diffusion bridge model for diffusion training to generate a target navigation action; wherein, the prior action is used as an initial state when the denoising diffusion bridge model performs reverse denoising processing.
[0190] It can be understood that the content in the above method embodiments is applicable to the system embodiments of the present application. The functions specifically implemented by the system embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0191] An embodiment of the present application further provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned visual navigation method based on a denoising diffusion bridge model. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0192] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0193] Please refer to Figure 4 , Figure 4 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0194] The processor 401 can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, abbreviated as ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0195] The memory 402 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 402 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 402 and are called by the processor 401 to execute a visual navigation method based on a denoising diffusion bridge model according to an embodiment of the present application;
[0196] The input / output interface 403 is used to implement information input and output;
[0197] The communication interface 404 is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0198] The bus 405 transmits information between various components of the device (such as the processor 401, the memory 402, the input / output interface 403, and the communication interface 404);
[0199] Among them, the processor 401, the memory 402, the input / output interface 403, and the communication interface 404 are communicatively connected to each other inside the device through the bus 405.
[0200] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned visual navigation method based on a denoising diffusion bridge model.
[0201] It can be understood that the content in the above method embodiments is applicable to this storage medium embodiment. The functions specifically implemented by this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0202] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely provided relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0203] A visual navigation method and a visual navigation system based on a denoising diffusion bridge model provided by an embodiment of the present application. The method includes obtaining initial image data; performing feature encoding processing on the initial image data to generate a context vector; performing linear modulation processing on the context vector to generate a conditional variable; generating a prior action according to the context vector and a preset motion rule; and inputting the prior action and the conditional variable into the denoising diffusion bridge model for diffusion training to generate a target navigation action. The prior action is used as the initial state when the denoising diffusion bridge model performs reverse denoising processing. By using the denoising diffusion bridge model, the embodiment of the present application performs a reverse diffusion process with the prior action distribution as the initial state of the diffusion bridge model, so that the initial state can be efficiently corrected to the target action distribution, thereby realizing a fast and stable visual navigation strategy. That is, the embodiment of the present application converts the diffusion denoising process starting from Gaussian noise in the traditional visual navigation method into a diffusion denoising process starting from an informative prior action distribution, which can effectively reduce redundant iterations in the diffusion process, reduce cumulative errors, greatly improve the efficiency and stability of action generation, and at the same time adapt to complex dynamic environments.
[0204] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0205] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0206] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0207] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware and their appropriate combinations.
[0208] In the description of the present application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0209] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0210] In several embodiments provided by the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of systems or units can be in electrical, mechanical or other forms.
[0211] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0212] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0213] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0214] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. This does not limit the scope of the rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.
Claims
1. A visual navigation method based on the denoising diffusion bridge model, characterized in that, The method includes the following steps: Obtain initial image data; Perform feature encoding processing on the initial image data to generate a context vector; Perform linear modulation processing on the context vector to generate a conditional variable; Generate a prior action according to the context vector and a preset motion rule; Input the prior action and the conditional variable into a denoising diffusion bridge model for diffusion training to generate a target navigation action; wherein, the prior action is used as the initial state when the denoising diffusion bridge model performs reverse denoising processing.
2. The method according to claim 1, characterized in that The initial image data includes an initial observation image and an initial target image. The performing feature encoding processing on the initial image data to generate a context vector includes: Perform feature extraction processing on the initial observation image to obtain observation image features; Perform feature extraction processing on the initial target image to obtain target image features; Perform feature fusion processing on the observation image features and the target image features to generate the context vector.
3. The method according to claim 1, wherein The performing linear modulation processing on the context vector to generate a conditional variable includes: Input the context vector into a feature linear modulation module; Generate modulation parameters according to the context vector through the feature linear modulation module; Determine the conditional variable according to the modulation parameters through the feature linear modulation module.
4. The method according to claim 1, wherein The generating a prior action according to the context vector and a preset motion rule includes: Map the context vector to an action space through a fully connected layer of a neural network model and generate low-dimensional action features corresponding to the context vector; Generate an initial candidate path according to the low-dimensional action features and the preset motion rule; Determine the prior action according to the initial candidate path.
5. The method according to claim 1, wherein The method further includes: Perform data sampling from a Gaussian white noise distribution based on a Gaussian prior rule to obtain the prior action; Or, Input the initial observation image in the initial image data into a conditional variational autoencoder model for training to obtain the prior action.
6. The method according to claim 1, wherein The inputting the prior action and the conditional variable into a denoising diffusion bridge model for diffusion training to generate a target navigation action includes: Input the prior action and the conditional variable into the denoising diffusion bridge model; Perform forward noise addition training according to a real path point sequence through the denoising diffusion bridge model to generate a noise action; Perform reverse denoising processing on the noise action according to the prior action and the conditional variable through the denoising diffusion bridge model to generate the target navigation action.
7. The method according to claim 1, wherein After the inputting the prior action and the conditional variable into a denoising diffusion bridge model for diffusion training to generate a target navigation action, the method further includes: Generate a navigation action instruction according to the target navigation action; Send the target navigation action and the navigation action instruction to a motion control module of a robot, and control the robot to execute the target navigation action according to the navigation action instruction through the motion control module.
8. A visual navigation system based on a denoising diffusion bridge model, characterized in that, The system includes the following modules: An initial image data acquisition module for acquiring initial image data; A feature encoding processing module for performing feature encoding processing on the initial image data to generate a context vector; A linear modulation processing module for performing linear modulation processing on the context vector to generate a conditional variable; A prior action generation module for generating a prior action according to the context vector and a preset motion rule; A target navigation action generation module for inputting the prior action and the conditional variable into a denoising diffusion bridge model for diffusion training to generate a target navigation action; wherein, the prior action is used as an initial state when the denoising diffusion bridge model performs reverse denoising processing.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Diffusion model wind power prediction method based on prior knowledge
CN117239730A
Magnetic particle image denoising method, system and equipment based on two-stage diffusion model
CN117635479A
Diffusion model power load prediction method based on continuous standardized flow optimization
CN117788204A
Video deblurring method based on latent variable prior knowledge guidance, computer equipment, readable storage medium and program product
CN117830154A
Method and system for controlling distributions of attributes in language models for text generation
US20220108081A1