Unmanned vehicle navigation method based on multi-modal fusion
Through an adaptive sensor fusion strategy based on variational inference, combining information geometric measurement and Markov decision-making process, sensor weights are dynamically adjusted, and the problem of insufficient navigation accuracy and robustness of multimodal fusion in complex environments in the existing technology is solved, and efficient and adaptive navigation capabilities are achieved.
Patent Information
- Application Number
- CN202510332271.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing multimodal fusion method is difficult to achieve efficient sensor data fusion in complex environments, resulting in insufficient navigation accuracy and robustness, especially in dynamic environments and emergencies.
Adaptive sensor fusion strategy based on variational inference is adopted, and by optimizing sensor weights, combining information geometric measurements and Markov decision-making process, the weights are dynamically adjusted to adapt to environmental changes, and real-time environmental state prediction is used to use dynamic Bayesian networks.
It realizes efficient integration of multimodal data in a dynamic environment, improves navigation accuracy and robustness, can quickly adapt to complex environment changes, reduce environmental noise interference, and improves the adaptive navigation capabilities of unmanned vehicles.
Smart Images

Figure CN120141518A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and specifically to an unmanned vehicle navigation method based on multi-modal fusion. Background Art
[0002] In recent years, significant progress has been made in unmanned driving technology. As the core technology of the autonomous driving perception system, sensor fusion plays a crucial role in improving navigation accuracy, environmental perception ability, and decision-making stability. Currently, unmanned vehicles mainly rely on a variety of sensors such as lidar, millimeter-wave radar, cameras, IMU, GPS / RTK, etc. to obtain environmental information, and improve the accuracy and robustness of perception by fusing different modal data. However, in complex scenarios, existing multi-modal fusion methods still face many challenges, mainly reflected in sensor weight optimization, cross-modal information extraction and reasoning, and path planning based on environmental adaptability.
[0003] Currently, multi-modal sensor data fusion mainly relies on fixed weighted averaging and simple machine learning methods, such as weighted accumulation, Bayesian estimation, and deep learning models. The mentioned methods may be effective in static environments, but in dynamic environments, the quality of sensors will be affected by factors such as lighting, weather, and obstacle occlusion. For example, in foggy and rainy scenarios, the sparsity of lidar point clouds increases, and the imaging quality of cameras decreases. If the weights are fixed, it may lead to the navigation system relying on low-quality data, affecting decision-making accuracy.
[0004] Traditional deep learning methods usually use end-to-end networks for feature extraction and fusion, but there are certain limitations. On the one hand, deep learning models require a large amount of high-quality data for training. In the field of autonomous driving, the cost of collecting real-scene data is extremely high, and the data distribution cannot cover all possible situations. On the other hand, the black-box nature of end-to-end models cannot explain how cross-modal data affects decision-making, and it is prone to fusion failure and misjudgment when encountering scenarios not seen in the training set. For example, some neural network fusion strategies perform well in good outdoor environments, but may fail in tunnels, shadow areas, and extreme weather. Therefore, a cross-modal fusion method that can perform efficient reasoning with a small amount of data and quickly adapt to environmental changes is needed.
[0005] The path planning of driverless vehicles usually relies on high-definition maps and rule-based path generation methods, such as Dijkstra algorithm, A* algorithm, and reinforcement learning methods. However, traditional path planning methods cannot perceive environmental changes in real time. Especially in unstructured roads, complex traffic environments, and emergencies, the navigation strategy cannot be quickly adjusted. For example, in the case of construction sections, sudden traffic jams, and road damage, the predefined navigation path may fail, resulting in the vehicle being unable to adjust its driving direction in time. Existing reinforcement learning methods can optimize path planning to a certain extent, but the training cost is high, and it is impossible to achieve efficient adjustment in a dynamic environment. Therefore, how to combine environmental perception information for path planning and adaptively adjust the sensor weights so that the planning strategy can adapt to different road conditions is an urgent problem to be solved.
[0006] Therefore, those skilled in the art provide a driverless vehicle navigation method based on multi-modal fusion to solve the above-mentioned problems. Summary of the Invention
[0007] Aiming at the deficiencies of the existing technology, the present invention provides a driverless vehicle navigation method based on multi-modal fusion to solve the problems raised in the above background technology.
[0008] To achieve the above objectives, the present invention is realized through the following technical solutions: A driverless vehicle navigation method based on multi-modal fusion, including:
[0009] Step S1, sensor data acquisition and preprocessing, obtaining multi-modal sensor data and performing data alignment, filtering, and normalization processing;
[0010] Step S2, constructing an optimization framework for sensor fusion weights based on the variational inference method, using the preprocessed data as input, defining an optimization objective function, and solving the optimal solution of the fusion weights through the variational inference method;
[0011] Step S3, on the basis of the variational inference optimization model, optimizing the sensor weight update process using information geometric metrics;
[0012] Step S4, after completing the information geometric optimization acceleration, combining the decision field theory for dynamic weight adjustment, and adaptively adjusting the sensor data fusion weights according to the environmental state;
[0013] Step S5, on the basis of the dynamic weight adjustment, performing online inference and real-time fusion;
[0014] Step S6, after completing the real-time fusion, conducting scheme verification and experimental evaluation for different environments and simulation platforms.
[0015] Preferably, in the step S1, the sensor data acquisition and preprocessing further includes:
[0016] Step 1.1, it is necessary to select multi-modal sensors to obtain environmental perception data;
[0017] Step 1.2, perform time synchronization, spatial alignment and filtering on the collected data, and remove noise;
[0018] Specifically, time synchronization is carried out through the timestamp alignment algorithm, spatial alignment is processed using a specific coordinate transformation method. In terms of data filtering, voxel filtering is applied to lidar data, Kalman filtering is applied to millimeter-wave radar data, and image enhancement processing is performed on camera images;
[0019] Step 1.3, through data normalization, map all sensor data to a unified scale range;
[0020] For the data d i (t) and d j (t) between different sensors i and j, time synchronization is carried out through timestamps t i and t j as follows: sync(d i (t i ),d j (t j ));
[0021] The voxel filtering of lidar data is realized through the following formula, dividing the three-dimensional point cloud data into voxel grids and calculating the center points of each grid:
[0022] where P filtered is the center point of the grid, P i is each point in the original point cloud, and N is the number of points in each system;
[0023] The millimeter-wave radar data is denoised using Kalman filtering. The state equation of Kalman filtering is:
[0024] x k = Fx k-1 + Bu k + w k ,
[0025] where x k is the current state, F is the state transition matrix, x k-1 is the previous state, B is the control input matrix, u k is the control input, and w k is the process noise;
[0026] For the enhancement of camera images, an image filtering algorithm is used to improve the image quality. The formula of the image filtering algorithm:
[0027] Among them, h(j) is the gray-level distribution of the image, and H(i) is the gray value of the enhanced image;
[0028] All sensor data are mapped to the same scale interval [1, 2] through linear normalization:
[0029]
[0030] Among them, d is the original data, and min(d) and max(d) are the minimum and maximum values of the data.
[0031] Preferably, in step S2, the optimization of the multi-modal data fusion weight further includes:
[0032] Step 2.1, sensor weight modeling: Set the fusion weight matrix W = {w ij}, where w ij represents the weighting coefficient of the i-th group of sensors for the j-th dimensional feature, and construct a multi-modal data weighted fusion model:
[0033]
[0034] Among them, Y represents the fused data matrix, X i represents the data matrix of the i-th group of sensors, W i is the corresponding weight matrix, ∈ is the noise term, M is the number of groups of sensors, and ° represents element-wise multiplication operation;
[0035] To optimize the weight matrix W, variational inference is introduced for optimization;
[0036] Step 2.2, variational inference to optimize the weight: Assume that the posterior distribution P(W|X) of W cannot be directly calculated,
[0037] Use the variational distribution Q(W) for approximation, that is: Q(W) ≈ P(W|X),
[0038] Assume that Q(W) follows a parameterized probability distribution:
[0039]
[0040] Among them, q(w ij |λ ij ) is the variational distribution of w ij , and λ ij is the variational parameter;
[0041] Step 2.3, optimize the objective function:
[0042] The goal is to maximize the variational lower bound:
[0043] L(λ) = E Q(W)[logP(Y∣W,X)] - KL(Q(W)∥P(W)),
[0044] where [logP(Y∣W,X)] represents the expectation of the log-likelihood, and KL is the relative entropy,
[0045] KL(Q(W)∥P(W)) represents the relative entropy divergence between Q(W) and P(W), which is used for regularization optimization;
[0046] By maximizing L(λ), the variational parameter λ can be optimized, and then the optimal fusion weight can be obtained;
[0047] Step 2.4, Mean-field approximation and Stochastic Gradient Variational Inference optimization:
[0048] Assume that Q(W) can be decomposed into the product of multiple independent factors, that is, the mean-field approximation is adopted:
[0049] D is the feature dimension, and M is the number of sensor groups;
[0050] Using the gradient variational inference method, by sampling W (s) perform gradient optimization:
[0051]
[0052] where S is the number of samplings, and W (s) ~Q(W) is the sample obtained by sampling from the variational distribution,
[0053] is the gradient of the log-likelihood under the sampling weight W (s) and
[0054] is the gradient of the variational parameter λ ij with respect to the variational lower bound L(λ);
[0055] Optimize the variational parameter λ by the gradient descent method ij , and finally obtain the fusion weight matrix W.
[0056] Preferably, in the step S3, the weight optimization based on information geometry further includes:
[0057] Step 3.1, Calculate the information geometry metric: Define the metric tensor G(W) of the weight matrix W on the Riemannian manifold, which is used to describe the local geometric structure of the parameter space. The expression is as follows:
[0058]
[0059] where G(W) represents the metric tensor of the weight space, is the gradient vector of the variational distribution, and T represents the transpose operation;
[0060] By calculating G(W), the local curvature information of the manifold is obtained for optimizing the weight update strategy;
[0061] Step 3.2, calculate the natural gradient: In the standard gradient descent method, the gradient update direction is determined by the gradient in the Euclidean space, while the natural gradient method adjusts the gradient direction based on information geometry to make the optimization process stable. The calculation formula of the natural gradient is as follows:
[0062]
[0063] where, is the natural gradient direction, and G(W) -1 is the inverse matrix of the metric tensor,
[0064] is the ordinary gradient direction;
[0065] Through natural gradient optimization, the path dependence problem is avoided and the convergence efficiency is improved;
[0066] Step 3.3, optimize the weight update strategy: To accelerate the convergence speed of the SGVI algorithm, the Riemannian optimization method is used for weight update. The update formula is as follows:
[0067]
[0068] where, W (t+1) is the weight matrix after the (t + 1)-th round of update, W (t) is the current weight matrix, η is the learning rate, is the natural gradient direction.
[0069] Preferably, in the step S4, the weight dynamic optimization based on the environmental state and the Markov decision process further includes:
[0070] Step 4.1, environmental state modeling: Define the environmental state vector s t , which describes the environmental state information at time t. The environmental state in the model is modeled in the following way:
[0071] s t = f(X t , U t ),
[0072] where, s t is the environmental state vector at the current time t, X t is the sensor data matrix, U t is the external control variable. Through environmental state modeling, the relationship between the environmental state and the sensor data can be dynamically captured;
[0073] Step 4.2, Markov decision process modeling: To describe the change of sensor weights in a dynamic system, a Markov decision process is used for modeling;
[0074] Set the state space of the Markov process as S, the action space as A, and the state transition probability P(s t+1 ∣s t ,a t ), where a t is the action at time t, and the update process of the sensor weights at time t is described as: W t+1 = W t + a t
[0075] where W t+1 is the weight matrix updated at time t + 1, W t is the current weight matrix at time t, and a t is the action at time t;
[0076] The state transition of the environment P(s t+1 ∣s t ,a t ) is given by a probability distribution, indicating the probability that the state s t transits to the state s t by the action a t+1 ;
[0077] Step 4.3, weight adjustment strategy: The weight adjustment strategy is modeled by introducing an environmental impact function g(s t ), and the specific weight adjustment strategy is expressed as:
[0078] where a t is the action at time t, g(s t ) is the environmental impact function, is the gradient of the current weight matrix.
[0079] Preferably, in the step S5, the real-time weight update and fusion calculation based on the dynamic Bayesian network further includes:
[0080] Step 5.1, dynamic Bayesian network predicting the environmental state: To improve the prediction accuracy of the environmental state, a dynamic Bayesian network is used to model the time evolution process of the environmental state. Set the environmental state at time t as s t , then the state transition is described by the following probability distribution:
[0081]
[0082] where s t is the environmental state vector at time t, z tis a latent variable, X t is the sensor data matrix at the current time t, P(s t ∣z t ) represents the conditional probability of the latent variable z t to the environmental state s t ;
[0083] P(z t ∣s t-1 , X t ) represents the distribution of the latent variable jointly determined by the environmental state s t-1 at the previous time and the current sensor data X t ;
[0084] Through DBN inference, the optimal estimated value of the environmental state at the current time is obtained:
[0085]
[0086] where X 1:t represents all the sensor data collected from time 1 to t, is the optimal estimated value of the environmental state at the current time;
[0087] Step 5.2, variational inference updates the weights in real time: Based on the environmental state predicted in Step 5.1 Variational inference is used to optimize the sensor weights W t in real time;
[0088] Assume that the posterior distribution of the sensor weights is complex and difficult to solve, so a variational distribution Q(W t ) is introduced for approximation:
[0089] Set Q(W t ) to follow a parameterized Gaussian distribution:
[0090]
[0091] where q(w t,ij ∣λ t,ij ) is the variational distribution of the weight w t,ij , λ t,ij is the variational parameter, and M is the number of sensor groups;
[0092] The parameters of the variational distribution Q(W t ) are optimized by maximizing the variational lower bound:
[0093]
[0094] Update the variational parameters through gradient variational inference:
[0095]
[0096] Among them, S is the number of sampling times, is a sample obtained by sampling from the variational distribution, W t is the sensor weight matrix at time t, Q(W t ) is the variational approximation distribution of the posterior distribution, L(λ t ) is the variational lower bound function, is the gradient of the variational lower bound with respect to the variational parameter λ t,j , is the sth weight matrix sample obtained by sampling from the variational distribution Q(W t );
[0097] Through the optimization process, the optimal sensor weight matrix W t is obtained, which is dynamically adjusted according to the change of the environmental state.
[0098] Preferably, in step 5.3, the sensor data is weighted and fused: based on the optimized weight matrix W t calculated in step 5.2, the sensor data is weighted and fused to calculate the final output data:
[0099]
[0100] where Y represents the fused data matrix, X i represents the data matrix of the i-th group of sensors, W i is the corresponding weight matrix, ∈ is the noise term, M is the number of groups of sensors, and ° represents the element-wise multiplication operation.
[0101] Preferably, the multi-modal fusion-based unmanned vehicle navigation method is applied to an autonomous vehicle to achieve adaptive navigation of the vehicle under different environmental conditions.
[0102] A terminal device includes a processor, a memory, and sensors communicating with the processor. The memory stores computer-executable instructions, and when the processor executes the computer-executable instructions, it is used to implement the multi-modal fusion-based unmanned vehicle navigation method.
[0103] A storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the multi-modal fusion-based unmanned vehicle navigation method.
[0104] The present invention provides a multi-modal fusion-based unmanned vehicle navigation method, having the following beneficial effects:
[0105] 1. The present invention adopts an adaptive sensor fusion strategy based on variational inference. By optimizing the sensor weights, it enables the optimal fusion of data from different modalities in a dynamic environment. Compared with the existing fusion schemes with fixed weights and simple weighting, the present invention can adjust the sensor contribution in real time, reduce environmental noise interference, and improve the navigation accuracy.
[0106] 2. The present invention introduces a cross-modal learning and inference mechanism, and uses the information geometry optimization method to accelerate the weight update, enabling the unmanned vehicle to quickly adapt to sudden environmental changes. Different from traditional neural network models that require a large amount of training data, the method of the present invention does not require pre-training with a large-scale data set, can adaptively adjust decisions, and improve the generalization ability.
[0107] 3. The present invention combines a path planning method based on environmental adaptation, predicts the environmental state through a dynamic Bayesian network, and adjusts the sensor weights to make the path planning conform to the real-time traffic conditions. Compared with the existing static map and rule-based planning schemes, this method can make reasonable navigation decisions in complex traffic environments and enhance the survival ability of unmanned vehicles in unstructured environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0108] Figure 1 is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0109] To enable those skilled in the art to understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0110] The present invention will be described in detail below with reference to the accompanying drawings:
[0111] Embodiment:
[0112] Please refer to the appended Figure 1 , the embodiment of the present invention provides a navigation method for an unmanned vehicle based on multi-modal fusion, including:
[0113] Step S1, sensor data acquisition and preprocessing, obtaining multi-modal sensor data and performing data alignment, filtering, and normalization processing;
[0114] Step 1.1, it is necessary to select multi-modal sensors to obtain environmental perception data;
[0115] Step 1.2, perform time synchronization, spatial alignment, and filtering processing on the collected data, and remove noise;
[0116] Specifically, time synchronization is carried out through a timestamp alignment algorithm, spatial alignment is processed using a specific coordinate transformation method. In terms of data filtering, voxel filtering is applied to lidar data, Kalman filtering is applied to millimeter-wave radar data, and image enhancement processing is performed on camera images;
[0117] Step 1.3, through data normalization, map all sensor data to a unified scale range;
[0118] For the data d i (t) and d j (t) between different sensors i and j, time synchronization is carried out through timestamps t i and t j : sync(d i (t i ), d j (t j ));
[0119] The voxel filtering of lidar data is achieved through the following formula, dividing the three-dimensional point cloud data into voxel grids and calculating the center points of each grid:
[0120] where P filtered is the center point of the grid, P i is each point in the original point cloud, and N is the number of points in each system;
[0121] Kalman filtering is used for noise reduction of millimeter-wave radar data, and the state equation of Kalman filtering is:
[0122] x k = Fx k-1 + Bu k + w k ,
[0123] where x k is the current state, F is the state transition matrix, x k-1 is the previous state, B is the control input matrix, u k is the control input, and w k is the process noise;
[0124] For the enhancement of camera images, an image filtering algorithm is used to improve the image quality, and the formula of the image filtering algorithm is:
[0125] where h(j) is the gray-level distribution of the image, and H(i) is the gray value of the enhanced image;
[0126] Map all sensor data to the same scale interval [1, 2] through linear normalization:
[0127]
[0128] Among them, d is the original data, and min(d) and max(d) are the minimum and maximum values of the data;
[0129] Step S2, construct an optimization framework for sensor fusion weights based on variational inference method. Use the preprocessed data as input, define the optimization objective function, and solve the optimal solution of the fusion weights through the variational inference method;
[0130] Step 2.1, sensor weight modeling: Set the fusion weight matrix W = {w ij}), where w ij represents the weighting coefficient of the i-th group of sensors for the j-th dimensional feature, and construct a multimodal data weighted fusion model:
[0131]
[0132] Among them, Y represents the fused data matrix, X i represents the data matrix of the i-th group of sensors, W i is the corresponding weight matrix, ∈ is the noise term, M is the number of groups of sensors, and ° represents the element-wise multiplication operation;
[0133] To optimize the weight matrix W, introduce variational inference for optimization;
[0134] Step 2.2, variational inference to optimize weights: Assume that the posterior distribution P(W|X) of W cannot be directly calculated,
[0135] Use the variational distribution Q(W) for approximation, that is: Q(W) ≈ P(W|X),
[0136] Assume that Q(W) follows a parameterized probability distribution:
[0137]
[0138] Among them, q(w ij |λ ij ) is the variational distribution of w ij , and λ ij is the variational parameter;
[0139] Step 2.3, optimization objective function:
[0140] The goal is to maximize the variational lower bound:
[0141] L(λ) = E Q(W) [logP(Y|W,X)] - KL(Q(W)∥P(W)),
[0142] Among them, [logP(Y|W,X)] represents the expectation of the log-likelihood, and KL is the relative entropy,
[0143] KL(Q(W)∥P(W)) represents the relative entropy divergence between Q(W) and P(W), which is used for regularization optimization;
[0144] By maximizing L(λ), the variational parameter λ can be optimized, and then the optimal fusion weight can be obtained;
[0145] Step 2.4, mean field approximation and stochastic gradient variational inference optimization:
[0146] Assume that Q(W) can be decomposed into the product of multiple independent factors, that is, the mean field approximation is adopted:
[0147] D is the feature dimension, and M is the number of sensor groups;
[0148] Using the gradient variational inference method, by sampling W (s) Perform gradient optimization:
[0149]
[0150] where S is the number of samplings, and W (s) ~Q(W) is the sample obtained by sampling from the variational distribution,
[0151] is the logarithmic likelihood gradient under the sampling weight W (s) ,
[0152] is the variational parameter λ ij for the gradient of the variational lower bound L(λ);
[0153] Optimize the variational parameter λ by the gradient descent method ij , and finally obtain the fusion weight matrix W;
[0154] Step S3, based on the variational inference optimization model, use information geometric metrics to optimize the sensor weight update process;
[0155] Step 3.1, calculate the information geometric metric: Define the metric tensor G(W) of the weight matrix W on the Riemannian manifold, which is used to describe the local geometric structure of the parameter space. The expression is as follows:
[0156]
[0157] where G(W) represents the metric tensor of the weight space, is the gradient vector of the variational distribution, and T represents the transpose operation;
[0158] By calculating G(W), obtain the local curvature information of the manifold, which is used to optimize the weight update strategy;
[0159] Step 3.2, Calculate the natural gradient: In the standard gradient descent method, the gradient update direction is determined by the gradient in the Euclidean space, while the natural gradient method adjusts the gradient direction based on information geometry to make the optimization process stable. The calculation formula of the natural gradient is as follows:
[0160]
[0161] where, is the natural gradient direction, and G(W) -1 is the inverse matrix of the metric tensor,
[0162] is the ordinary gradient direction;
[0163] Through natural gradient optimization, the path dependence problem is avoided and the convergence efficiency is improved;
[0164] Step 3.3, Optimize the weight update strategy: To accelerate the convergence speed of the SGVI algorithm, the Riemannian optimization method is used for weight update. The update formula is as follows:
[0165]
[0166] where, W (t+1) is the weight matrix after the (t + 1)-th round of update, W (t) is the current weight matrix, η is the learning rate, is the natural gradient direction;
[0167] Step S4, After completing the information geometry optimization acceleration, combine the decision field theory to perform dynamic weight adjustment, and adaptively adjust the sensor data fusion weight according to the environmental state;
[0168] Step 4.1, Environmental state modeling: Define the environmental state vector s t , which describes the environmental state information at time t. The environmental state in the model is modeled in the following way:
[0169] s t = f(X t , U t ),
[0170] where, s t is the environmental state vector at the current time t, X t is the sensor data matrix, U t is the external control variable. Through environmental state modeling, the relationship between the environmental state and sensor data can be dynamically captured;
[0171] Step 4.2, Markov decision process modeling: To describe the change of sensor weights in a dynamic system, the Markov decision process is used for modeling;
[0172] Set the state space of the Markov process as \(S\), the action space as \(A\), and the state transition probability \(P(s t+1 |s t ,a t ), where \(a t is the action at time \(t\). The update process of the sensor weight at time \(t\) is described as: \(W t+1 = W t + a t
[0173] where \(W t+1 is the weight matrix updated at time \(t + 1\), \(W t is the current weight matrix at time \(t\), and \(a t is the action at time \(t\);
[0174] The state transition of the environment \(P(s t+1 |s t ,a t ) is given by a probability distribution, indicating the probability that state \(s t is transferred to state \(s t by action \(a t+1 ;
[0175] Step 4.3, weight adjustment strategy: The weight adjustment strategy is modeled by introducing an environmental impact function \(g(s t ). The specific weight adjustment strategy is expressed as:
[0176]
[0177] where \(a t is the action at time \(t\), \(g(s t ) is the environmental impact function, is the gradient of the current weight matrix;
[0178] Step S5, based on the dynamic weight adjustment, perform online inference and real-time fusion;
[0179] Step 5.1, dynamic Bayesian network to predict the environmental state: To improve the prediction accuracy of the environmental state, a dynamic Bayesian network is used to model the time evolution process of the environmental state. Set the environmental state at time \(t\) as \(s t , then the state transition is described by the following probability distribution:
[0180]
[0181] where \(s t is the environmental state vector at time \(t\), \(z t is the latent variable, \(X t is the sensor data matrix at the current time \(t\), \(P(s t|z t ) represents the latent variable z t to the environmental state s t 's conditional probability,
[0182] P(z t |s t-1 ,X t ) represents the environmental state s at the previous moment t-1 and the current sensor data X t jointly determine the latent variable distribution;
[0183] Through DBN inference, the optimal estimated value of the environmental state at the current moment is obtained:
[0184]
[0185] where X 1:t represents all sensor data collected from time 1 to t, is the optimal estimated value of the environmental state at the current moment;
[0186] Step 5.2, variational inference updates the weights in real time: Based on the environmental state predicted in Step 5.1 Use variational inference to optimize the sensor weights W t in real time;
[0187] Assume that the posterior distribution of the sensor weights is complex and difficult to solve, so a variational distribution Q(W t ) is introduced for approximation:
[0188] Set Q(W t ) to follow a parameterized Gaussian distribution:
[0189]
[0190] where q(w t,ij |λ t,ij ) is the variational distribution of the weight w t,ij , λ t,ij is the variational parameter, and M is the number of sensor groups;
[0191] The parameters of the variational distribution Q(W t ) are optimized by maximizing the variational lower bound:
[0192]
[0193] Update the variational parameters through gradient variational inference:
[0194]
[0195] where S is the number of samplings, is a sample obtained by sampling from a variational distribution, and W t is the sensor weight matrix at time t, and Q(W t ) is the variational approximation distribution of the posterior distribution, and L(λ t ) is the variational lower bound function, is the gradient of the variational lower bound with respect to the variational parameter λ t,j ; is the sth weight matrix sample obtained by sampling from the variational distribution Q(W t );
[0196] Through the optimization process, the optimal sensor weight matrix W t ;
[0197] Step 5.3, weighted fusion of sensor data: Based on the optimized weight matrix W t calculated in Step 5.2, perform weighted fusion on the sensor data to calculate the final output data Y:
[0198]
[0199] where Y represents the fused data matrix, X i represents the data matrix of the i-th group of sensors, W i is the corresponding weight matrix, ∈ is the noise term, M is the number of groups of sensors, and ° represents the element-wise multiplication operation;
[0200] Step S6, after completing real-time fusion, conduct scheme verification and experimental evaluation for different environments and simulation platforms.
[0201] In Step S1, select lidar, millimeter-wave radar, and camera to obtain environmental data, and perform time synchronization, spatial alignment, denoising, and normalization of the data to the same scale to ensure the consistency of different sensor data. By removing noise, the perception quality is enhanced and the data accuracy is improved; timestamp synchronization and spatial coordinate unification improve the effect of multi-sensor fusion; after normalization, different sensor data can be fused in the same computing framework;
[0202] In Step S2, variational inference can dynamically optimize the weights to adapt to different environments; stochastic gradient variational inference can avoid the complexity of directly calculating the posterior distribution; the optimized fusion weights reduce the influence of noise and outliers;
[0203] The advantage of Step S3 is more efficient than the ordinary gradient descent method; it avoids path dependence and improves the accuracy of weight adjustment;
[0204] The advantage of step S4 is that the weights can be automatically adjusted according to the changes in the environmental state; the Markov decision process combined with the environmental impact function ensures the stable operation of the navigation system in complex environments;
[0205] The advantages of step S5 are that the Markov decision process can quickly predict the environmental state, enabling the system to respond quickly; variational inference enables the weights to be adaptively adjusted according to environmental changes, improving the navigation robustness; after weighted fusion, the system can maintain high-precision perception in complex environments;
[0206] The advantages of step S6 are verified in multiple scenarios, enabling the method to be applied to real unmanned vehicle systems; by experimentally evaluating the effects of different parameters, the system performance is optimized.
[0207] The unmanned vehicle navigation method based on multimodal fusion is applied to autonomous vehicles to achieve adaptive navigation of the vehicle under different environmental conditions.
[0208] A terminal device includes a processor, a memory, and a sensor communicating with the processor. The memory stores computer-executable instructions, and when the processor executes the computer-executable instructions, it is used to implement the unmanned vehicle navigation method based on multimodal fusion.
[0209] A storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the unmanned vehicle navigation method based on multimodal fusion.
[0210] By fusing multimodal data (lidar, millimeter-wave radar, camera), the environmental perception accuracy of the unmanned vehicle is improved. Combining variational inference, information geometric optimization, Markov decision process, and dynamic Bayesian network, an adaptive, real-time, and efficient navigation ability is achieved. Gradient variational inference + natural gradient optimization is used to improve the calculation efficiency, and the Markov decision process + dynamic Bayesian network is used for dynamic weight adjustment, enabling the unmanned vehicle to adapt to complex environmental changes and make intelligent decisions. This solution is applicable to the fields of autonomous driving, intelligent robots, intelligent logistics, and intelligent transportation, and ultimately achieves high-precision perception, intelligent decision-making, and real-time optimization, improving the safety, stability, and navigation efficiency of the unmanned vehicle.
[0211] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A navigation method for an unmanned vehicle based on multimodal fusion, characterized in that: include: Step S1, sensor data acquisition and preprocessing, obtaining multimodal sensor data and performing data alignment, filtering and normalization processing; Step S2, constructing an optimization framework for sensor fusion weights based on a variational reasoning method, using the preprocessed data as input, defining an optimization objective function, and solving the optimal solution for the fusion weights through a variational reasoning method; Step S3, based on the variational inference optimization model, using information geometry metrics to optimize the sensor weight update process; Step S4, after completing the information geometry optimization acceleration, dynamically adjust the weights in combination with the decision field theory, and adaptively adjust the sensor data fusion weights according to the environmental state; Step S5, performing online reasoning and real-time fusion based on dynamic weight adjustment; Step S6, after the real-time fusion is completed, the solution verification and experimental evaluation are carried out for different environments and simulation platforms.
2. The unmanned vehicle navigation method based on multimodal fusion according to claim 1, characterized in that: In step S1, sensor data collection and preprocessing further include: Step 1.1, it is necessary to select a multimodal sensor to obtain environmental perception data; Step 1.2, performing time synchronization, spatial alignment and filtering on the collected data, and removing noise; Specifically, time synchronization is performed through a timestamp alignment algorithm, and spatial alignment is processed using a specific coordinate transformation method. In terms of data filtering, voxel filtering is applied to lidar data, Kalman filtering is applied to millimeter-wave radar data, and image enhancement processing is performed on camera images. Step 1.3, map all sensor data to a uniform scale range through data normalization; For the data d between different sensors i and j i (t) and d j (t), time synchronization is achieved through the timestamp t i and t j Perform: sync(d i (t i ),d j (t j )), Voxel filtering of LiDAR data is achieved by dividing the 3D point cloud data into voxel grids and calculating the center point of each grid: Among them, P filtered is the center point of the grid, P i are the points in the original point cloud, and N is the number of points in each voxel; The millimeter-wave radar data is denoised using Kalman filtering, and the state equation of the Kalman filter is: x k =Fx k-1 +This k +w k , Among them, x k is the current state, F is the state transfer matrix, x k-1 is the state at the previous moment, B is the control input matrix, u k is the control input, w k is the process noise; To enhance the camera image, an image filtering algorithm is used to improve the image quality. The image filtering algorithm formula is: Among them, h(j) is the grayscale distribution of the image, and H(i) is the grayscale value of the enhanced image; All sensor data are mapped to the same scale interval through linear normalization [1,2]: Where d is the original data, min(d) and max(d) are the minimum and maximum values of the data.
3. The unmanned vehicle navigation method based on multimodal fusion according to claim 1, characterized in that: In step S2, the multimodal data fusion weight optimization further includes: Step 2.1, sensor weight modeling: Set the fusion weight matrix W = {w ij }, where w ij Represents the weighted coefficient of the i-th group of sensors on the j-th dimension feature, and constructs a multimodal data weighted fusion model: Among them, Y represents the fused data matrix, X i represents the data matrix of the i-th group of sensors, W i is the corresponding weight matrix, ∈ is the noise term, M is the number of sensor groups, and ° represents the element-by-element multiplication operation; In order to optimize the weight matrix W, variational inference is introduced for optimization; Step 2.2, variational inference weight optimization: Assume that the posterior distribution P(W|X) of W cannot be calculated directly, and use the variational distribution Q(W) for approximation, that is: Q(W)≈P(W|X), Assume that Q(W) obeys a parameterized probability distribution: Among them, q(w ij ∣λ ij ) is w ij The variational distribution of ij is the variational parameter; Step 2.3, optimize the objective function: The goal is to maximize the variational lower bound: L(λ)=E Q(W) [logP(Y∣W,X)]-KL(Q(W)∥P(W)), Among them, [logP(Y|W,X)] represents the expectation of log likelihood, KL is the relative entropy, KL(Q(W)∥P(W)) represents the relative entropy divergence between Q(W) and P(W), which is used for regularization optimization; By maximizing L(λ), the variational parameter λ can be optimized, and the optimal fusion weight can be obtained; Step 2.4, mean field approximation and stochastic gradient variational inference optimization: Assume that Q(W) can be decomposed into the product of multiple independent factors, that is, the mean field approximation is adopted: D is the feature dimension, M is the number of sensor groups; Using the gradient variational inference method, by sampling W (s) Perform gradient optimization: Among them, S is the number of sampling times, W (s) ~Q(W) is the sample obtained by sampling the variational distribution, is the sampling weight W (s) The log-likelihood gradient under , is the variational parameter λ ij The gradient of the variational lower bound L(λ); Optimize the variational parameter λ by gradient descent ij , and finally the fusion weight matrix W is obtained.
4. The unmanned vehicle navigation method based on multimodal fusion according to claim 1, characterized in that: In step S3, the weight optimization based on information geometry further includes: Step 3.1, calculate the information geometry metric: define the metric tensor G(W) of the weight matrix W on the Riemann manifold to describe the local geometric structure of the parameter space. The expression is as follows: Among them, G(W) represents the metric tensor of the weight space, is the gradient vector of the variational distribution, T represents the transpose operation; By calculating G(W), we can obtain the local curvature information of the manifold, which is used to optimize the weight update strategy. Step 3.2, calculate the natural gradient: In the standard gradient descent method, the gradient update direction is determined by the gradient of the Euclidean space, while the natural gradient method adjusts the gradient direction based on information geometry to stabilize the optimization process. The calculation formula of the natural gradient is as follows: in, is the natural gradient direction, G(W) -1 is the inverse matrix of the metric tensor, is the normal gradient direction; Through natural gradient optimization, path dependency problems can be avoided and convergence efficiency can be improved; Step 3.3, optimize the weight update strategy: In order to accelerate the convergence speed of the SGVI algorithm, the Riemann optimization method is used to update the weights. The update formula is as follows: Among them, W (t+1) is the weight matrix after the update in the t+1th round, W (t) is the current weight matrix, η is the learning rate, is the natural gradient direction.
5. The unmanned vehicle navigation method based on multimodal fusion according to claim 1, characterized in that: In step S4, the weight dynamic optimization based on the environment state and the Markov decision process further includes: Step 4.1, Environmental state modeling: Define the environmental state vector s t , describes the environmental state information at time t. The environmental state in the model is modeled in the following way: s t =f(X t ,U t ), Among them, s t is the environmental state vector at the current time t, X t is the sensor data matrix, U t It is an external control variable, which can dynamically capture the relationship between environmental state and sensor data through environmental state modeling; Step 4.2, Markov decision process modeling: To describe the changes of sensor weights in a dynamic system, a Markov decision process is used for modeling; Assume that the state space of the Markov process is S, the action space is A, and the state transition probability P(s t+1 ∣s t ,a t ), where a t is the action at time t, and the update process of the sensor weight at time t is described as: W t+1 =W t +a t Among them, W t+1 is the weight matrix updated at time t+1, W t is the current weight matrix at time t, a t is the action at time t; The state transition of the environment P(s t+1 ∣s t ,a t ) is given by the probability distribution, indicating that the state s t By action a t Transfer to state s t+1 probability; Step 4.3, weight adjustment strategy: The weight adjustment strategy is to introduce the environmental impact function g(s t ) is used for modeling, and the specific weight adjustment strategy is expressed as: Among them, a t is the action at time t, g(s t ) is the environmental impact function, is the gradient of the current weight matrix.
6. The unmanned vehicle navigation method based on multimodal fusion according to claim 1, characterized in that: In step S5, the real-time weight update and fusion calculation based on the dynamic Bayesian network further includes: Step 5.1, dynamic Bayesian network predicts environmental state: In order to improve the prediction accuracy of environmental state, a dynamic Bayesian network is used to model the time evolution of environmental state, and the environmental state at time t is set to s t , then the state transition is described by the following probability distribution: Among them, s t is the environmental state vector at time t, z t is a hidden variable, X t is the sensor data matrix at the current time t, P(s t ∣z t ) represents the latent variable z t To the environment state s t The conditional probability of P(z t ∣s t-1 ,X t ) represents the environmental state s at the previous moment t-1 With the current sensor data X t Jointly determined latent variable distribution; Through DBN reasoning, we can get the optimal estimate of the current state of the environment: Among them, X 1:t Represents all sensor data collected from time 1 to t, is the optimal estimate of the current state of the environment; Step 5.2: Variational inference updates weights in real time: based on the environment state predicted in step 5.1 The sensor weight W is trained using variational reasoning t Perform real-time optimization; Assume the posterior distribution of sensor weights It is complex and difficult to solve, so we introduce the variational distribution Q(W t ) to approximate: Set Q(W t ) obeys a parameterized Gaussian distribution: Among them, q(w t,ij ∣λ t,ij ) is the weight w t,ij The variational distribution of t,ij is the variational parameter, M is the number of sensor groups; Variational distribution Q(W t ) is optimized by maximizing the variational lower bound: Update the variational parameters via gradient variational inference: Where S is the number of sampling times, is the sample obtained by sampling from the variational distribution, W t is the sensor weight matrix at time t, Q(W t ) is the variational approximation of the posterior distribution, L(λ t ) is the variational lower bound function, is the variational parameter λ t,j Find the gradient of the variational lower bound, is the variational distribution Q(W t ) obtained by sampling the s-th weight matrix sample; Through the optimization process, the optimal sensor weight matrix W that is dynamically adjusted by the environmental state changes is obtained. t .
7. The unmanned vehicle navigation method based on multimodal fusion according to claim 6, characterized in that: Step 5.3, weighted fusion sensor data: based on the optimized weight matrix W calculated in step 5.2 t , perform weighted fusion on the sensor data and calculate the final output data Y: Among them, Y represents the fused data matrix, X i represents the data matrix of the i-th group of sensors, W i is the corresponding weight matrix, ∈ is the noise term, M is the number of sensor groups, and ° represents element-by-element multiplication operation.
8. The unmanned vehicle navigation method based on multimodal fusion according to claim 1, characterized in that: The unmanned vehicle navigation method based on multimodal fusion is applied to an autonomous driving vehicle to achieve adaptive navigation of the vehicle under different environmental conditions.
9. A terminal device, characterized in that: It includes a processor, a memory and a sensor communicating with the processor, the memory stores computer-executable instructions, and the processor is used to implement the unmanned vehicle navigation method based on multimodal fusion as described in any one of claims 1 to 8 when executing the computer-executable instructions.
10. A storage medium, characterized in that: Computer executable instructions are stored, and the computer executable instructions are used to enable a computer to execute the unmanned vehicle navigation method based on multimodal fusion as described in any one of claims 1 to 8.
Citation Information
Cited By
Fusion positioning method and system for autonomous vehicle
CN120427014A
Vehicle bottom detection robot local path planning method based on cross-modal collaborative learning
CN120800384A
Vehicle weather environment self-adaption method based on online learning
CN121492940A
Environment-adaptive data acquisition unmanned vehicle and control method thereof
CN121613712A