Indoor automatic guided vehicle navigation obstacle avoidance method with high robustness

By using memory attention generalized network and enhanced imitation learning technology in the automatic guide vehicle navigation obstacle avoidance model, the problem of insufficient performance of deep reinforcement learning models in high-dimensional state space is solved, and more efficient training and better environmental adaptability and robustness are achieved.

CN120029277APending Publication Date: 2025-05-23SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510106479.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Deep reinforcement learning models encounter the problem of 'dimensionality curse' in high-dimensional state space, resulting in performance not meeting actual needs, low training efficiency, and inability to make accurate decisions in complex and rapidly changing environments, showing poor adaptability and robustness.

Method used

The generalized network of memory attention is adopted, combining multi-head attention mechanism, gated cycle unit and deep-density residual learning technology to enhance the action command calculation ability of the automatic guide vehicle, and through enhanced imitation learning, data augmentation and priority experience replay technology, the generalization and robustness of the model are improved.

Benefits of technology

It significantly reduces the computational burden brought by complex state space, improves the adaptability and decision-making efficiency of the model in dynamic environments, enhances the generalization ability and robustness of the model, and ensures stable performance in different environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029277A_ABST
    Figure CN120029277A_ABST
Patent Text Reader

Abstract

The invention relates to a high-robustness indoor automatic guided vehicle navigation obstacle avoidance method, which comprises the following steps that: step 1, an automatic guided vehicle acquires environment scanning data and position data; step 2, environment scanning data is processed by a multi-head attention mechanism of the memory attention generalized network, and is combined with position data to be input into a gating circulation unit and a deep dense residual learning technology integrated structure, and an action instruction is output; 3, obtaining a reward value according to the action of the action instruction; 4, updating the navigation obstacle avoidance strategy, and iterating the steps 1 to 4 to obtain an optimal low-layer navigation obstacle avoidance strategy; and 5, deploying enhanced imitation learning, a high-level navigation obstacle avoidance strategy and an optimal low-level navigation obstacle avoidance strategy, and carrying out iterative training from the step 1 to the step 4 under the guidance of the high-level navigation obstacle avoidance strategy to obtain a high-robustness indoor automatic guided vehicle navigation obstacle avoidance strategy. The calculation burden can be reduced, the training efficiency is improved, the model generalization is enhanced, and the stable performance of the model in the environment is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of automatic navigation, and in particular to a highly robust indoor automatic guided vehicle navigation and obstacle avoidance method. Background Art

[0002] In recent years, automated guided vehicles, as an important part of industrial automation and intelligent logistics, have attracted widespread attention due to their ability to autonomously complete material transportation and path planning.

[0003] The rule-based navigation method realizes the navigation and obstacle avoidance of the automated guided vehicle in the environment through preset rules and path planning. In practical applications, the rule-based method is difficult to meet the complex and changing environmental requirements faced by the automated guided vehicle, and cannot maintain efficient navigation and obstacle avoidance performance in a dynamic environment.

[0004] Traditional machine learning methods extract features and recognize patterns from environmental data, and use algorithms such as support vector machines or decision trees for navigation and obstacle avoidance. However, traditional machine learning methods often show poor generalization capabilities and lack effective response mechanisms when encountering unknown environments.

[0005] The prior art proposes a hierarchical deep reinforcement learning framework, which improves the safety and sampling efficiency of the navigation process through the collaborative work of low-level and high-level deep reinforcement learning strategies. The low-level deep reinforcement learning strategy is responsible for ensuring that the robot moves along the predetermined path to the target position and avoids obstacles in real time; the high-level deep reinforcement learning strategy further optimizes the path selection and enhances the safety of the overall navigation. In order to improve the sampling efficiency of the system and avoid the sparse reward problem in traditional reinforcement learning, the invention reduces the dimension of the state space by selecting sub-target points on the target path, thereby accelerating the training process.

[0006] It has the following technical problems: This deep reinforcement learning model usually encounters the "curse of dimensionality" problem when faced with high-dimensional state space, resulting in the model performance being unable to meet actual needs; there is still much room for improvement in the training efficiency in automated guided vehicle navigation tasks; when dealing with complex and rapidly changing environments, it is often unable to make accurate decisions, showing poor adaptability and robustness. Summary of the invention

[0007] In view of the problems existing in the prior art, the purpose of the present invention is to provide a highly robust indoor automatic guided vehicle navigation and obstacle avoidance method, which can reduce the computational burden brought by complex state space, improve training efficiency, enhance model generalization, and ensure the stable performance of the model in the environment.

[0008] In order to achieve the above object, the present invention adopts the following technical solution: A highly robust indoor automatic guided vehicle navigation and obstacle avoidance method comprises the following steps: Step 1: The automated guided vehicle obtains environmental scanning data and position data; Step 2: The environment scan data is processed by the multi-head attention mechanism of the memory attention generalized network, and then combined with the position data, it is input into the gated recurrent unit and the deep dense residual learning technology integrated structure to calculate and output the action instructions of the automatic guided vehicle; Step 3: The automated guided vehicle performs actions in the environment according to the action instructions to obtain a reward value; Step 4: After collecting a certain number of reward values, update the navigation obstacle avoidance strategy, and iterate steps 1 to 4 until the optimal low-level navigation obstacle avoidance strategy is obtained; Step 5: Deploy enhanced imitation learning, high-level navigation obstacle avoidance strategy and optimal low-level navigation obstacle avoidance strategy in the automated guided vehicle, and iterate the training from step 1 to step 4 under the guidance of the high-level navigation obstacle avoidance strategy until a highly robust indoor automated guided vehicle navigation obstacle avoidance strategy is obtained; Step 6: Deploy the trained indoor automatic guided vehicle navigation and obstacle avoidance strategy model on the automatic guided vehicle to realize the indoor navigation and obstacle avoidance function.

[0009] Furthermore, the multi-head attention mechanism processes the environmental scanning data by dividing the observation state acquired by the onboard sensor of the automated guided vehicle into multiple segments, and applying the self-attention operation on each segment in parallel while paying attention to the input features of multiple dimensions.

[0010] Furthermore, the multi-head attention mechanism adopts a scaled dot-product attention mechanism to quantify the importance of specific features to the automated guided vehicle through the interaction of queries, keys, and values. In this process, the set of queries, keys, and values ​​are determined through self-attention operations to represent different aspects of interest, feature identifiers, and related information content, respectively.

[0011] Furthermore, the gated recurrent unit manages the information flow by using an update gate and a reset gate, where the update gate determines the integration of the current hidden state with the new hidden state, and the reset gate controls the integration of historical information to update the current hidden state.

[0012] Furthermore, the deep dense residual learning technique connects the input data of each layer of the gated recurrent unit with the hidden state, allowing the input information to be directly propagated through the deep network.

[0013] Furthermore, the action space of the low-level navigation obstacle avoidance strategy includes that the angular velocity of the automatic guided vehicle is associated with the deviation angle of the target point, and the linear velocity is discretized into multiple non-negative actions.

[0014] Furthermore, the low-level navigation obstacle avoidance strategy is designed with a hybrid reward function consisting of sparse rewards and dense rewards, and the hybrid reward function includes obstacle avoidance rewards, target reaching rewards and speed rewards.

[0015] Furthermore, enhanced imitation learning adopts a hybrid strategy, using expert strategies in simple and specific boundary conditions and optimizing strategies through exploration and learning in other cases.

[0016] Furthermore, before the training phase, the enhanced imitation learning algorithm collects expert strategy data sets as the initial experience pool. In the subsequent training process, the generated data gradually replaces the original experience pool data, and through the principle of gradual approximation, the gap between the simulated data and the real data is narrowed.

[0017] Furthermore, the enhanced imitation learning algorithm combines the priority experience replay mechanism to weight the experience samples according to their importance, sampling and learning important experience data more frequently and sampling less important experience data less frequently.

[0018] In general, the present invention has the following advantages: (1) Alleviate the "curse of dimensionality" problem: The present invention proposes a memory attention generalized network, which combines a multi-head attention mechanism, a gated recurrent unit, and a deep dense residual learning technology to effectively improve the feature extraction capability of high-dimensional environmental data, thereby significantly reducing the computational burden brought by complex state space and improving the adaptability and decision-making efficiency of the model in a dynamic environment; (2) Improve training efficiency: By improving the existing hierarchical deep reinforcement learning framework, the present invention optimizes the synergy between low-level and high-level strategies, making the training process more efficient. Specifically, by combining sub-goal decomposition with path planning, the sparse reward phenomenon is reduced, thereby accelerating the training convergence of the model and improving the training efficiency and sample utilization; (3) Enhance model generalization: In order to solve the generalization problem of the model in a changing environment, the present invention adopts an enhanced imitation learning method. By integrating data enhancement technology, priority experience playback, and expert strategy guidance, the adaptability of the model in a variety of scenarios is effectively improved, while reducing the dependence of training data on a specific environment, ensuring the stable performance of the model in an unknown or dynamic environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 The flowchart is a highly robust indoor automatic guided vehicle navigation and obstacle avoidance method.

[0020] Figure 2 Schematic diagram of the multi-head attention mechanism in a generalized network with memory attention.

[0021] Figure 3 Schematic diagram of the integrated structure of gated recurrent unit and deep dense residual learning technology.

[0022] Figure 4 Schematic diagram of the update process of deep reinforcement learning model combined with enhanced imitation learning.

[0023] Figure 5 Schematic diagram of the migration of a highly robust indoor automated guided vehicle navigation and obstacle avoidance strategy method from simulation to real environment. DETAILED DESCRIPTION

[0024] In order to improve the generalization ability of automatic guided vehicle navigation and ensure fast and safe navigation in different indoor scenes, the present invention proposes an innovative and robust automatic guided vehicle navigation and obstacle avoidance method based on imitation-enhanced hierarchical deep reinforcement learning. First, the present invention proposes a memory-attention generalized network to address the "curse of dimensionality" problem in deep reinforcement learning methods. Secondly, the present invention improves the hierarchical deep reinforcement learning framework of existing methods and improves the efficiency of model training. Finally, the present invention proposes an enhanced imitation learning method to address the problem of model generalization.

[0025] The present invention will be described in further detail below.

[0026] like Figure 1 As shown, a highly robust indoor automatic guided vehicle navigation and obstacle avoidance method is characterized by comprising the following steps: Step 1: The automated guided vehicle obtains environmental scanning data and position data; Step 2: The environment scan data is processed by the multi-head attention mechanism of the memory attention generalized network, and then combined with the position data, it is input into the gated recurrent unit and the deep dense residual learning technology integrated structure to calculate and output the action instructions of the automatic guided vehicle; Step 3: The automated guided vehicle performs actions in the environment according to the action instructions to obtain a reward value; Step 4: After collecting a certain number of reward values, update the navigation obstacle avoidance strategy, and iterate steps 1 to 4 until the optimal low-level navigation obstacle avoidance strategy is obtained; Step 5: Deploy enhanced imitation learning, high-level navigation obstacle avoidance strategy and optimal low-level navigation obstacle avoidance strategy in the automated guided vehicle, and iterate the training from step 1 to step 4 under the guidance of the high-level navigation obstacle avoidance strategy until a highly robust indoor automated guided vehicle navigation obstacle avoidance strategy is obtained; Step 6: Deploy the trained indoor automatic guided vehicle navigation and obstacle avoidance strategy model on the automatic guided vehicle to realize the indoor navigation and obstacle avoidance function.

[0027] In complex navigation tasks, the AGV makes decisions based on its internal state and partial environmental information collected by sensors, enabling it to avoid obstacles and successfully reach the target location. This process can be formalized as a Markov decision process, which includes state space, action space, reward value, and next-moment state space. The state space contains the current state of the AGV, the layout of surrounding obstacles, and the relative position of the target. The action space consists of the linear velocity and angular velocity of the AGV. The next-moment state space is generated by the current action and state. The reward value represents the immediate reward obtained after performing a specific action, providing feedback on the effectiveness of the action. The core goal of the navigation task is to maximize the cumulative reward of the AGV in the process of navigation and obstacle avoidance through policy optimization. This requires the policy model to adapt to uncertain environmental conditions during training and make effective decisions in dynamic scenarios. By optimizing the long-term cumulative reward, the model ensures that the AGV completes the navigation task in a safe and efficient manner in a complex environment.

[0028] The underlying framework of the robust automatic guided vehicle navigation and obstacle avoidance method based on imitation-enhanced hierarchical deep reinforcement learning proposed in the present invention is based on a double-delay double-deep Q-network algorithm. The core principle of the double-delay double-deep Q-network is to improve the stability and efficiency of Q-learning through a dual mechanism. First, it decouples the action selection and evaluation process through the dual Q-network structure, solving the problem of overestimation of Q-values, which is a common problem in traditional Q-learning. This separation reduces the risk of over-dependence of the two tasks on the same network. Second, the double-delay double-deep Q-network algorithm decomposes the Q-value into state value and advantage function, allowing the network to better evaluate the overall value of the state. This design can more effectively learn critical state values. In addition, compared with algorithms that require additional policy networks, the computational complexity of the double-delay double-deep Q-network algorithm is relatively low, thereby reducing the demand for computing resources. The robust automatic guided vehicle navigation and obstacle avoidance method based on imitation-enhanced hierarchical deep reinforcement learning designed by the present invention includes three parts: (1) memory attention generalized network; (2) hierarchical policy structure; (3) enhanced imitation learning.

[0029] Memory-Attention Generalized Networks: The implementation scheme of the prior art only uses a simple linear layer for state feature extraction, and the performance still needs to be optimized. The memory attention generalized network proposed in the present invention enhances the automatic guided vehicle navigation and obstacle avoidance method by combining a multi-head attention mechanism, a gated recurrent unit, and a deep dense residual learning technique to create a network that captures the long-term spatiotemporal interaction features between the automatic guided vehicle and its environment. These technologies significantly improve the performance of the network in complex tasks. The network enhances the automatic guided vehicle's attention to key obstacles by implicitly inferring the motion characteristics of obstacles, and improves the automatic guided vehicle's understanding and adaptability to the environment. The framework is applied to the feature extraction layers of the high-level and low-level networks of the automatic guided vehicle, thereby improving the generalization performance of the automatic guided vehicle in unfamiliar environments.

[0030] Hierarchical policy structure: The hierarchical strategy structure of the prior art implementation scheme includes obstacle avoidance reward, target achievement reward and speed reward. The present invention improves the target reward and optimizes the collision penalty coefficient. The improved hierarchical strategy structure improves the efficiency of model training.

[0031] Enhanced Imitation Learning: The implementation schemes of the prior art do not take into account the gap between the simulation and the real environment when the method is transferred. In order to eliminate the gap between simulation and reality, the present invention proposes an innovative enhanced imitation learning technology. This technology effectively integrates data enhancement technology, priority experience playback and expert strategy, ensuring that the deep reinforcement learning model can make safe and effective decisions in clear and high-risk situations, while retaining its ability to optimize strategies through exploration and learning in other situations; and this technology reduces dependence on specific environments, while enhancing its generalization ability without affecting learning efficiency, ensuring that automated guided vehicles can make optimal decisions in different scenarios.

[0032] Specifically, in the real world, humans navigating in complex environments usually prioritize key objects and predict future dynamics based on long-term spatiotemporal features, thereby adjusting their actions accordingly. The navigation task of an automated guided vehicle is similar to that of humans in complex environments. Therefore, it is logical to enhance the attention of the automated guided vehicle to nearby key obstacles, such as those that are close, move quickly, or are on the collision course. The automated guided vehicle can then predict the behavior of these obstacles based on their recent states and adjust its own state accordingly. The memory attention generalized network proposed in this invention enhances the automated guided vehicle navigation and obstacle avoidance method by combining multi-head attention mechanism, gated recurrent unit and deep dense residual learning technology to create a network that captures the long-term spatiotemporal interaction features between the automated guided vehicle and its environment. These technologies significantly improve the performance of the network in complex tasks. The network enhances the attention of the automated guided vehicle to key obstacles by implicitly inferring the motion features of obstacles, and improves the automated guided vehicle's ability to understand and adapt to the environment. The framework is applied to the feature extraction layers of the high-level and low-level networks of the automated guided vehicle, thereby improving the generalization performance of the automated guided vehicle in unfamiliar environments.

[0033] The AGV obtains observation states through onboard sensors. In the memory attention generalized network structure, a multi-head attention mechanism is first used to process these observation states, such as Figure 2 As shown. The mechanism divides the observed state into multiple segments and applies self-attention operations on each segment in parallel. This parallel processing strategy enables the automated guided vehicle to simultaneously focus on input features of multiple dimensions, revealing the complex and rich feature relationships between the automated guided vehicle and its environment. Specifically, the present invention adopts a scaled dot product attention mechanism. It quantifies the importance of specific features to the automated guided vehicle through the interaction of queries, keys, and values. In this process, the set of queries, keys, and values ​​is determined through self-attention operations, which represent different aspects of interest, feature identifiers, and related information content, respectively. This mechanism allows the network to adaptively adjust the attention weights on environmental features, focusing on key information while ignoring less important information in the current task. This dynamic attention allocation strategy enhances the environmental perception capability of the automated guided vehicle and improves the decision-making and navigation efficiency of the automated guided vehicle in complex environments.

[0034] When solving the collision avoidance problem in AGV navigation, each decision of the AGV is time-dependent. In order to effectively utilize long-term information and enhance the learning process of the AGV, the present invention integrates gated recurrent units and deep dense residual learning techniques to train the policy model. The feature information extracted from the observed state by the attention mechanism is combined with the posture data of the AGV to form a new state, which is then passed as input to the subsequent part of the memory attention generalized network for further policy model training, such as Figure 3shown.

[0035] First, the gated recurrent unit skillfully manages the information flow after being processed by the attention mechanism by using update and reset gates, allowing the network to retain or discard historical information as needed. Specifically, the update gate determines the integration of the current hidden state with the new hidden state, while the reset gate controls the integration of historical information to update the current hidden state. This module significantly improves the efficiency and accuracy of time series data processing.

[0036] Next, deep dense residual learning technology is introduced to further optimize the performance of the model. The core idea of ​​deep dense residual learning technology is to connect the input data of each layer with the hidden state, so that the input information can be directly transmitted through the deep network. Compared with the traditional deep neural network structure, deep dense residual learning technology effectively alleviates the problems of gradient vanishing and information loss that occur as the network depth increases. By continuously passing the original input to each layer, this architecture not only retains the global features of the input, but also significantly enhances the network's ability to capture long-term dependencies.

[0037] The hierarchical strategy network structure proposed in this paper is based on a memory-attention generalized network. Considering that the AGV mainly relies on sensors such as lidar or cameras to perceive the environment, these sensors usually generate high-dimensional data. To address this problem, the raw data collected by the sensors are processed through a multi-head attention mechanism. It aims to map the high-dimensional raw scan data to a low-dimensional feature vector, so as to more accurately extract the collision information and effectively characterize the obstacle features. In this process, the absolute state information such as the absolute position of the AGV and the obstacle position is intentionally excluded to enhance the generalization ability of the method.

[0038] In the control of an automated guided vehicle, commands usually consist of linear velocity and angular velocity. In order to reduce ineffective exploration during training and improve learning efficiency, the present invention associates the angular velocity of the automated guided vehicle with the deviation angle of the target point, discretizes the linear velocity into five non-negative actions, and thus defines the action space of the low-level policy network.

[0039] In order to further optimize the low-level navigation strategy of the automated guided vehicle, the present invention designs a hybrid reward function consisting of sparse rewards and dense rewards. This function includes obstacle avoidance rewards, target reaching rewards and speed rewards.

[0040] In order to further improve the generalization ability of the automated guided vehicle and the safety of the collision avoidance strategy, the present invention also designs a high-level navigation obstacle avoidance strategy network. This network maintains consistency in structure and state space configuration with the low-level strategy network, but differs in the definition of action space. Specifically, the effective low-level strategy is regarded as an executable action of the high-level strategy, and the action with a linear velocity of 0 is listed as a separate option. This design encourages the automated guided vehicle to notice the action with a linear velocity of 0, avoid obstacles by turning when the safety distance is insufficient, and move quickly using the low-level strategy when there is no risk of collision. At the same time, when designing the reward function, the collision penalty coefficient in the high-level strategy network is increased to ensure that the automated guided vehicle maintains a sufficient safety distance from the obstacle. When the automated guided vehicle enters an open area, it is encouraged to move forward quickly by increasing the positive reward. In addition, in order to prevent the automated guided vehicle from rotating in place or stalling for a long time in a complex environment, a behavioral reward mechanism is introduced to avoid the automated guided vehicle from falling into a cycle of invalid actions. This reward mechanism ensures the safe navigation of the high-level strategy in various environments.

[0041] Extensive exploration is often required during deep reinforcement learning training, especially in environments with large state and action spaces. In order to improve learning efficiency, the present invention combines the concept of imitation learning. In the initial exploration phase, expert knowledge is used as prior information to guide the automated guided vehicle in the optimal direction. This approach reduces the time and resources required for random exploration, which is particularly important in complex learning tasks or scenarios with high simulation costs. In addition, combining expert experience helps the automated guided vehicle avoid suboptimal or dangerous behaviors during training, ensuring that decision-making behaviors are close to the best practices in practical applications before the environment is fully explored, thereby maintaining high decision quality. However, deep reinforcement learning completely relies on expert rules for decision-making, which deviates from the core framework of learning and transforms the process into a rule-based engine rather than a learning model. This approach essentially brings deep reinforcement learning closer to supervised learning. The basic principle of deep reinforcement learning is that the agent gradually learns unknown rules through interaction with the environment, rather than following a strict set of predefined rules. In deep reinforcement learning, if expert rules are used entirely for decision-making, the applicability of the model in unknown or complex scenarios will be limited because it cannot go beyond the boundaries defined by the expert rules.

[0042] Therefore, the present invention proposes a reasonable compromise: using expert policies only in simple and specific boundary conditions. This ensures that the deep reinforcement learning model can make safe and effective decisions in clear and high-risk situations, while retaining its ability to optimize policies through exploration and learning in other situations. This method combines the safety of expert knowledge with the adaptive learning and generalization capabilities of deep reinforcement learning, providing a simple and effective strategy for handling complex environments.

[0043] To achieve this goal, the present invention proposes an innovative training algorithm combined with expert strategy guidance. The algorithm adopts a hybrid strategy. In certain challenging situations, the expert strategy guides the behavior of the automated guided vehicle, while allowing the automated guided vehicle to make autonomous decisions in other situations. This hybrid strategy ensures that the automated guided vehicle can make safe and effective decisions when encountering clear, high-risk situations, while still retaining the ability to optimize strategies through exploration and learning in other situations, thereby enhancing the adaptive learning ability of the method. Before the training phase, the algorithm collects an expert strategy data set as the initial experience pool. In the subsequent training process, the generated data gradually replaces the original experience pool data, and through the principle of gradual approximation, the gap between the simulated data and the real data is narrowed. In addition, the present invention combines a priority experience replay mechanism to weight the experience samples according to their importance, such as Figure 4 As shown. Important experience data is sampled and learned more frequently, while less important experience data is sampled less frequently, allowing for more efficient learning from accumulated experience. High-value experience data generated by expert policies are given higher priority and added to the replay buffer, ensuring that the network can learn from these key experience data frequently. This not only speeds up the process of learning from expert behavior, but also enhances the generalization ability of the automated guided vehicle when dealing with complex, rare, and high-risk situations. Priority is primarily determined by the temporal difference error of experience, which is the difference between the predicted and actual rewards.

[0044] The challenges that the uncertainty and diversity in real environments bring to automated guided vehicles are often more complex than those in virtual environments. In order to bridge the gap between simulation and reality, the present invention introduces a data enhancement module. This method aims to reduce the dependence on a specific environment while enhancing its generalization ability without affecting the learning efficiency, ensuring that the automated guided vehicle can make optimal decisions in different scenarios. The present invention designs a data enhancement module using a limited parameter randomization strategy, which carefully randomizes the key parameters by adding Gaussian noise to key parameters such as the laser distance matrix, the relative distance and direction to the target, the linear velocity and angular velocity of the automated guided vehicle, and combines it with an imitation learning module to promote learning without overcomplicating the problem. In addition, although real-world data is difficult to obtain, the randomization strategy simulates the uncertainty and diversity of the real world in a simulated environment, thereby reducing the dependence on a large amount of real data. It overcomes the learning difficulties that may be caused by excessive randomization, while improving the generalization ability and performance in a real environment different from the training environment through limited and goal-oriented randomization.

[0045] By effectively integrating data augmentation technology, priority experience playback and expert strategies, the robust automated guided vehicle navigation and obstacle avoidance method based on imitation-enhanced hierarchical deep reinforcement learning not only ensures efficient learning of decision patterns from expert strategies, but also significantly reduces prediction errors and improves strategy performance. In addition, it also enhances the robustness of the strategy model, thereby optimizing the overall learning results.

[0046] The present invention first conducts simulation experiments in a physical simulator to create a complex indoor training environment in which the size, shape and position of obstacles are randomized. At the beginning of each training iteration in the virtual environment, the initial posture and final target position of the automated guided vehicle are randomly determined to explore the environment as thoroughly as possible. The automated guided vehicle receives environmental information state input through sensors. Figure 1 The imitation-enhanced hierarchical deep reinforcement learning model in . Part of the state of the environment information state is passed in Figure 2 The multi-head attention mechanism in the algorithm is used to extract features, and then combined with the position information of the automated guided vehicle, it is input together Figure 3 The integrated structure in the training is calculated, and enhanced imitation learning is combined during the training process to obtain the optimal action in the current state. Finally, the automated guided vehicle performs actions in the environment and enters the next state. The automated guided vehicle repeats the above cycle in the training environment until it reaches the destination or triggers a failure condition, and then the training is considered complete. After a certain amount of training, the model executes Figure 4 The update process in the present invention optimizes and improves the current strategy. The model of the present invention has a hierarchical strategy. According to different goals, the low-level strategy is first trained until the optimal low-level navigation and obstacle avoidance strategy is obtained. Then, the high-level navigation and obstacle avoidance strategy is trained on the basis of the low-level strategy, and finally a strong and robust indoor automatic guided vehicle navigation and obstacle avoidance strategy model is obtained. After the simulation experiment training is completed, the present invention transfers the obtained model to the indoor automatic guided vehicle in the real environment for performance verification, such as Figure 5 shown.

[0047] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.

Claims

1. A highly robust indoor automatic guided vehicle navigation and obstacle avoidance method, characterized by: The following steps are included: Step 1: The automated guided vehicle obtains environmental scanning data and position data; Step 2: The environment scan data is processed by the multi-head attention mechanism of the memory attention generalized network, and then combined with the position data, it is input into the gated recurrent unit and the deep dense residual learning technology integrated structure to calculate and output the action instructions of the automatic guided vehicle; Step 3: The automated guided vehicle performs actions in the environment according to the action instructions to obtain a reward value; Step 4: After collecting a certain number of reward values, update the navigation obstacle avoidance strategy, and iterate steps 1 to 4 until the optimal low-level navigation obstacle avoidance strategy is obtained; Step 5: Deploy enhanced imitation learning, high-level navigation obstacle avoidance strategy and optimal low-level navigation obstacle avoidance strategy in the automated guided vehicle, and iterate the training from step 1 to step 4 under the guidance of the high-level navigation obstacle avoidance strategy until a highly robust indoor automated guided vehicle navigation obstacle avoidance strategy is obtained; Step 6: Deploy the trained indoor automatic guided vehicle navigation and obstacle avoidance strategy model on the automatic guided vehicle to realize the indoor navigation and obstacle avoidance function.

2. The method for indoor automatic guided vehicle navigation and obstacle avoidance with strong robustness according to claim 1, characterized in that: The multi-head attention mechanism processes environmental scanning data by dividing the observation state obtained by the onboard sensors of the automated guided vehicle into multiple segments and applying self-attention operations on each segment in parallel, while paying attention to input features of multiple dimensions.

3. The method for indoor automatic guided vehicle navigation and obstacle avoidance with strong robustness according to claim 1, characterized in that: The multi-head attention mechanism adopts a scaled dot-product attention mechanism to quantify the importance of specific features to the automated guided vehicle through the interaction of queries, keys, and values. In this process, the set of queries, keys, and values ​​is determined through self-attention operations to represent different aspects of interest, feature identifiers, and related information content, respectively.

4. The method for indoor automatic guided vehicle navigation and obstacle avoidance with strong robustness according to claim 1, characterized in that: The gated recurrent unit manages the information flow by using an update gate that determines the integration of the current hidden state with the new hidden state and a reset gate that controls the integration of historical information to update the current hidden state.

5. The method for indoor automatic guided vehicle navigation and obstacle avoidance with strong robustness according to claim 4, characterized in that: The deep dense residual learning technique connects the input data of each layer of the gated recurrent unit with the hidden state, allowing the input information to be directly propagated through the deep network.

6. The method for indoor automatic guided vehicle navigation and obstacle avoidance with strong robustness according to claim 1, characterized in that: The action space of the low-level navigation and obstacle avoidance strategy includes the angular velocity of the automated guided vehicle associated with the deviation angle of the target point, and the linear velocity discretized into multiple non-negative actions.

7. The method for indoor automatic guided vehicle navigation and obstacle avoidance with strong robustness according to claim 1, characterized in that: The low-level navigation and obstacle avoidance strategy is designed with a hybrid reward function consisting of sparse rewards and dense rewards. The hybrid reward function includes obstacle avoidance rewards, target reaching rewards and speed rewards.

8. The method for indoor automatic guided vehicle navigation and obstacle avoidance with strong robustness according to claim 1, characterized in that: Enhanced imitation learning adopts a hybrid strategy, using an expert policy in simple and specific boundary conditions and optimizing the policy through exploration and learning in other cases.

9. The method for indoor automatic guided vehicle navigation and obstacle avoidance with strong robustness according to claim 1, characterized in that: Before the training phase, the enhanced imitation learning algorithm collects expert strategy data sets as the initial experience pool. In the subsequent training process, the generated data gradually replaces the original experience pool data, and through the principle of gradual approximation, the gap between simulated data and real data is narrowed.

10. The method for indoor automatic guided vehicle navigation and obstacle avoidance with strong robustness according to claim 1, characterized in that: The enhanced imitation learning algorithm combines the prioritized experience replay mechanism to weight the experience samples according to their importance, sampling and learning important experience data more frequently and sampling less important experience data less frequently.

Citation Information

Cited By

  • Unmanned aerial vehicle autonomous navigation system based on hierarchical reinforcement learning strategy

    CN120800385A