Agent-enhanced path planning algorithm based on multimodal information fusion
Through the multimodal information fusion method, sensor capture and network processing of multiple environmental information, combined with reinforcement learning algorithms, the problem of reinforcement learning obtaining single information in complex environments is solved, and more accurate and efficient path planning is achieved.
Patent Information
- Application Number
- CN202311670287.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-12-07
AI Technical Summary
When faced with complex environments, reinforcement learning is too single to obtain information and cannot make full use of multiple types of information, resulting in poor path planning in complex scenarios.
The multimodal information fusion method is adopted to capture multiple environmental information through sensors, and information fusion is used to express the correlation between state spaces and path planning is performed in combination with reinforcement learning algorithms.
It realizes the full utilization of complex environment information, improves the accuracy and efficiency of path planning, avoids the one-sidedness and inaccuracy of human considerations, and is suitable for multi-task development in complex scenarios.
Smart Images

Figure CN117784776B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a path planning algorithm, and in particular to a multi-modal information fusion method for fully utilizing environmental information in a reinforcement field. Background Art
[0002] Path planning is the process of finding the optimal path from the starting point to the target point or a path that meets specific conditions in a given environment. Traditional path planning algorithms include sampling optimization algorithms, genetic algorithms, etc., and applying reinforcement learning to path planning to allow the intelligent agent to explore the path that meets the conditions by itself is also very popular.
[0003] Reinforcement learning algorithm can be simply understood as a Markov process, which is a process in which intelligent agents interact with the environment by taking actions and constantly obtaining feedback to update their decision-making methods. Reinforcement learning has a very impressive performance in the fields of machine learning and games, thanks to strong expert experience and rich game data in games. However, in the face of the high complexity of the environment, including complex terrain, communication sensor limitations, environmental dynamics, biological barriers, energy endurance and many other problems, reinforcement learning has not yet shown sufficient advantages. The main problem is that information acquisition is too single, and rich and diverse information is not fully utilized. At the same time, the relationship between each feature state variable and its influence on the output action are also unknown, so it can only be limited to specific task scenarios, resulting in limited technology development. Summary of the invention
[0004] The purpose of the present invention is to avoid the insufficient utilization of environmental information in reinforcement learning and propose a multimodal information fusion method. An intelligent agent reinforcement path planning algorithm based on multimodal information fusion includes multimodal information fusion technology, reinforcement learning algorithm and high-complexity and high-accuracy environmental model simulation. The technical solution is as follows:
[0005] The first step is to build complex environment and intelligent agent models.
[0006] (1) Download and store information within a certain longitude and latitude range, including longitude and latitude range, depth, density, terrain, and environmental dynamics. At the same time, in order to ensure the refinement of the simulation, interpolate the existing environmental dynamics data to fill in the missing information;
[0007] (2) Using the Gazebo platform, GPS sensors, infrared sensors, attitude sensors, and speed sensors are installed on the intelligent body. At the same time, the collision properties, gravity, and buoyancy are defined, and the corresponding environmental dynamic coefficients are calculated. These are loaded into the Gazebo platform together with the environmental information, and are called in real time through the Python algorithm and ROS control mechanism.
[0008] The second step is to design the multimodal information fusion network structure.
[0009] In view of the complex and diverse environmental information, the many sensors on the intelligent agent are used to capture it, which can make full use of the environmental information and avoid certain limitations in information utilization. At the same time, the changes in various attributes in the environment are continuous and there is a certain degree of correlation. After the sensor collects a lot of information, the infrared image information is pooled through the convolutional neural network, which reduces the spatial dimension while retaining the main features in the feature image, improving its robustness. Then, the output data is combined with the nine-dimensional information of posture angle, speed, and position obtained by the intelligent agent to form a multidimensional vector, and then input into a fully connected network to obtain the final output. The former pooling is to reduce the dimension of the image data, so that it is convenient to make full use of other sensor information, and the latter network structure is to extract the internal connection of the environmental information. The design of the state space is not a simple permutation and combination. There is also a certain correlation within the environmental information. It is difficult to achieve the ideal effect by artificial guessing in the face of a complex environment, so the black box characteristics of the network are used to achieve this effect. In the subsequent training results, the parameters in the network structure are reversely optimized, so that the relationship in the state space is also part of the reinforcement learning strategy selection.
[0010] The third step is to strengthen the learning algorithm architecture.
[0011] (1) Reinforcement learning can be summarized as a process in which an intelligent agent takes actions to affect the environment and gives feedback in order to maximize rewards. In the A2C algorithm, there are decision makers and evaluators, corresponding to the policy network and value network respectively. The policy network guides the intelligent agent to perform actions, and the value network gives evaluations. The action space dimension of the intelligent agent in the environment is 6, corresponding to roll, yaw, pitch, floating, diving, and moving forward. The state input is generated by the multimodal information fusion network in the second step. With the help of the network, the optimal action corresponding to the current environment is output, and feedback is received from the environment to obtain reward accumulation. The residual network structure is applied in this network to solve the gradient vanishing problem and improve the network performance.
[0012] (2) Design of the experience pool. After each interaction between the agent and the environment, the experience data is stored and reused during the training process, which improves sample efficiency. At the same time, high-reward action-value pairs are marked to increase the probability of each sampling.
[0013] The fourth step is training and hyperparameter adjustment.
[0014] Through the alternating process of network state input and action output until the task is completed or the number of training steps is completed, the training results are weighed to see whether they meet the established task conditions. The strategy is evaluated by comprehensively considering the reward value at the end of the round, the number of steps in a single round to complete the task, and the actual route effect, to characterize the quality of the current strategy. At the same time, the hyperparameters in the network structure are adjusted to achieve effect optimization.
[0015] The beneficial effects of the present invention are:
[0016] (1) The present invention makes full use of a lot of information about complex environments with the help of sensors, avoids information errors, and simulates the actual environment and task conditions more accurately;
[0017] (2) The present invention uses the network structure to express the correlation between state spaces, which to a certain extent avoids the one-sidedness and inaccuracy of human considerations;
[0018] (3) The present invention demonstrates the advantages of multimodal information fusion with the help of path planning tasks in the environment, which is helpful for the development and implementation of many tasks in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a multimodal information fusion network structure diagram of the present invention;
[0020] Figure 2 It is a flow chart of the intelligent agent enhanced path planning algorithm based on multimodal information fusion of the present invention. DETAILED DESCRIPTION
[0021] To make the technical solution of the present invention clearer, the present invention is further described below in conjunction with the accompanying drawings. The present invention is specifically implemented according to the following steps:
[0022] The first step is to build complex environment and intelligent agent models.
[0023] (1) Download and store information within a certain longitude and latitude range, including longitude and latitude range, depth, density, terrain, and environmental dynamics. At the same time, in order to ensure the refinement of the simulation, interpolate the existing environmental dynamics data to fill in the missing information;
[0024] (2) Using the Gazebo platform, GPS sensors, infrared sensors, attitude sensors, and speed sensors are installed on the intelligent body. At the same time, the collision properties, gravity, and buoyancy are defined, and the corresponding environmental dynamic coefficients are calculated. These are loaded into the Gazebo platform together with the environmental information, and are called in real time through the Python algorithm and ROS control mechanism.
[0025] The second step is to design the multimodal information fusion network structure.
[0026] After the sensor collects a lot of information, the infrared image information is pooled through a convolutional neural network, which reduces the spatial dimension while retaining the main features in the feature image. The output data is then combined with the nine-dimensional information of posture angle, speed, and position obtained by the intelligent agent to form a multidimensional vector, which is then input into a fully connected network to obtain the final output.
[0027] The third step is to set up the task scenario.
[0028] Define the specific tasks in the scenario, including the starting point, target point, number of obstacles, obstacle type, and task objectives (such as the shortest path, the fastest path, a single-objective task, or a multi-objective task); set the basic properties of the agent, including initial position, speed, and posture information; and initialize environmental data.
[0029] The fourth step is to set up the reinforcement learning algorithm architecture.
[0030] (1) The state space is obtained through a multimodal information fusion network;
[0031] (2) Network structure. Reinforcement learning can be summarized as a process in which an intelligent agent takes actions to affect the environment and gives feedback in order to maximize the reward. In the A2C algorithm, there are decision makers and evaluators, corresponding to the policy network and value network respectively. The policy network guides the intelligent agent to perform actions, and the value network gives evaluations. A residual-like structure is added to the network architecture, and the normalized state information and the feature vectors of each layer of the network are input into the next layer of the network, making full use of the residual learning unit, solving the gradient vanishing problem, and improving network performance.
[0032] (3) Action space. Considering the similarity with the real environment, the actions taken by the agent are divided into six dimensions: roll, yaw, pitch, float, dive, and move forward. In addition, appropriate adjustments will be made according to the task type.
[0033] (4) Design of the experience pool. Set the appropriate capacity, batch size, and update frequency of the experience pool according to the task scenario. In addition, mark the action-value pairs with high rewards to increase their probability of being sampled each time, thereby balancing the ratio of exploration and utilization.
[0034] Step 5: Training and hyperparameter adjustment.
[0035] Through the alternating process of network state input and action output until the task is completed or the number of training steps is completed, each piece of experience data is put into the experience pool, from which the strategy network and value network are sampled and updated iteratively, and the training results are weighed to see whether they meet the established task conditions. The strategy is evaluated by comprehensively utilizing the reward value at the end of the round, the number of steps in a single round to complete the task, and the actual route effect to characterize the quality of the current strategy. At the same time, the hyperparameters in the network structure are adjusted to achieve effect optimization.
Claims
1. An agent-enhanced path planning algorithm based on multimodal information fusion, characterized in that: The first step is to build a complex environment and intelligent agent model; (1) Download and store information within a certain longitude and latitude range, including longitude and latitude range, depth, density, terrain, and environmental dynamics size and direction information, perform interpolation calculations on existing environmental dynamics data, and fill in the missing information; (2) Using the Gazebo platform, GPS sensors, infrared sensors, attitude sensors, and speed sensors are installed on the intelligent agent. At the same time, the collision properties, gravity, and buoyancy are defined, and the corresponding environmental dynamic coefficients are calculated. These are loaded into the Gazebo platform together with the environmental information, and are called in real time through Python algorithms and ROS control mechanisms. The second step is to design the multimodal information fusion network structure; After the sensor collects environmental information, the infrared image information is pooled through a convolutional neural network, and then the output data is combined with the nine-dimensional information of posture angle, speed, and position obtained by the intelligent agent to form a multi-dimensional vector, which is input into a fully connected network to obtain the final output; The third step is to set up the task scene; Define the specific tasks in the scene, including the starting point, target point, number of obstacles, obstacle types, and task objectives; set the basic properties of the agent, including initial position, speed, and posture information; initialize environmental data; Step 4: Reinforce learning algorithm architecture setting; (1) The state space is obtained through a multimodal information fusion network; (2) Network structure: Reinforcement learning can be summarized as a process in which an intelligent agent takes actions to affect the environment and gives feedback in order to maximize the reward. In the A2C algorithm, it includes a decision maker and an evaluator, which correspond to the policy network and the value network respectively. The policy network guides the intelligent agent to perform actions, and the value network gives evaluations. A residual-like structure is added to the network architecture, and the normalized state information and the feature vectors of each layer of the network are input into the next layer of the network. (3) Action space: Considering that it is more similar to the real environment, the actions taken by the agent are divided into six dimensions: roll, yaw, pitch, float, dive, and move forward; (4) Design of the experience pool: Set the appropriate capacity, batch size, and update frequency of the experience pool according to the task scenario. In addition, mark the action-value pairs with high rewards to increase their probability of being sampled each time. Step 5: Training and hyperparameter adjustment; The process of network state input and action output is carried out alternately until the task is completed or the number of training steps is completed. Each piece of experience data is put into the experience pool, from which the strategy network and value network are sampled and updated iteratively. It is weighed whether the training results meet the established task conditions. The reward value at the end of the round, the number of steps in a single round to complete the task, and the actual route effect are comprehensively used to evaluate the strategy to characterize the quality of the current strategy. At the same time, the hyperparameters in the network structure are adjusted to achieve effect optimization.
Citation Information
Patent Citations
Agent unknown environment exploration method based on reinforcement learning
CN111062491A
Structure inspection agent navigation method based on damage driving and multi-mode multi-task learning
CN116824303A