Intelligent path-finding algorithm and device based on multi-scale average field representation method

Through the multi-scale average field representation method combined with the attention mechanism, the problems of high computational complexity and insufficient accuracy of traditional multi-agent reinforcement learning methods are solved, and more efficient and accurate multi-agent decision-making and collaboration are achieved.

CN119940141APending Publication Date: 2025-05-06ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510118519.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional multi-agent reinforcement learning methods have high computational complexity, low computational efficiency and insufficient accuracy when processing large numbers of agents, making it difficult to effectively simulate complex multi-agent interactions.

Method used

The multi-scale average field representation method is adopted, and the average field representation of the near-field and far-field scales combined with the attention mechanism, the local interaction and global impact between the agents are captured and the decision-making process of the agent is optimized.

Benefits of technology

It significantly improves the decision-making accuracy and system collaboration capabilities of multi-agent systems, enhances the perception of complex environments, and improves the performance and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940141A_ABST
    Figure CN119940141A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent way-finding algorithm and device based on a multi-scale mean field representation method, and the method comprises the steps: selecting a reinforcement learning environment, and initializing the environment; interacting with the environment, obtaining feedback data and storing the feedback data into an experience playback pool; acquiring environment interaction data from the experience playback pool in batches to train the reinforcement learning model; defining a multi-scale average field representation formula; and defining a robot intelligent path-finding process based on multi-scale average field multi-agent reinforcement learning. According to the invention, the accuracy and decision-making efficiency of behavior prediction are obviously improved. The invention provides a powerful and flexible intelligent path-finding algorithm based on a multi-scale mean field representation method, excellent performance is shown in multi-agent reinforcement learning, more efficient and more accurate agent behavior prediction is realized especially in a large-scale complex environment, and powerful support is provided for development of multi-agent reinforcement learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of reinforcement learning and artificial intelligence, and in particular to a method for representing a mean field in reinforcement learning, and specifically to an intelligent path-finding algorithm and device for a multi-scale mean field representation method. Background Art

[0002] In recent years, Multi-Agent Reinforcement Learning (MARL) has been widely used in many fields, such as traffic dispatching, robot collaboration, resource allocation, etc., which has greatly promoted the penetration and development of intelligent technology in actual production and life. However, with the continuous advancement of technology, a severe challenge facing MARL is the surge in the number of agents. In practical applications, with the increase in the number of agents, traditional MARL methods often encounter problems such as high computational complexity, low computational efficiency, insufficient accuracy, and even inability to effectively extract key information from the environment. Especially in large-scale, multi-agent scenarios, how to ensure the scalability and efficiency of the system has become a bottleneck restricting the development of MARL. Therefore, the development of efficient reinforcement learning methods that can handle a large number of agents has become a core issue that needs to be solved urgently.

[0003] In order to solve the problem of explosive growth in interaction dimensions caused by a large number of intelligent agents, researchers first proposed to use the mean field game method to solve it, thus developing mean field reinforcement learning (MFRL). This method has largely alleviated the computational complexity problem caused by the surge in the number of intelligent agents. Specifically, mean field reinforcement learning abstracts the state characteristics or action characteristics of the global or local intelligent agent into a virtual mean intelligent agent, and transforms the interaction between intelligent agents into the interaction between intelligent agents and the virtual mean intelligent agent, thereby significantly reducing the dimension of the interaction information and successfully reducing the computational burden. However, since this method constructs a "completely balanced" mean field by taking the mean, it is difficult to truly restore the complex interactions in the actual environment. In practical applications, the decision-making of intelligent agents is often strongly influenced by neighboring intelligent agents, and this local influence is often not effectively expressed by a simple mean abstraction. To this end, researchers proposed an adaptive mean field method, which uses the attention mechanism to give higher weights to neighboring intelligent agents, thereby better simulating the local interaction relationship in reality. However, this method still has the following two shortcomings: (1) A single mean field cannot fully capture the multi-level interactions between intelligent agents. Although the adaptive mean field enhances local interactions by adjusting neighbor weights, it still relies on only one layer of the mean field and lacks comprehensive modeling of complex, hierarchical interactions between agents. (2) Lack of global vision. Although the adaptive mean field gives higher weights to neighboring agents, it ignores the potential impact of global agent information on individual agent decisions, which makes it impossible for agents to fully refer to global information when making decisions.

[0004] Traditional mean field multi-agent reinforcement learning methods show certain limitations when dealing with complex environments and agent interaction logic. There is an urgent need for more adaptive and intelligent methods to improve the accuracy and real-time performance of reinforcement learning methods. Summary of the invention

[0005] The present invention is aimed at mean field multi-agent reinforcement learning. In order to overcome the above-mentioned shortcomings of the prior art, an intelligent path-finding algorithm and device based on a multi-scale mean field representation method are provided to achieve accurate simulation of real multi-agent scenarios.

[0006] The present invention first abstracts the multi-agent system into an environment model containing multiple agents, where each agent represents an individual, and the interactive relationship between agents is modeled through a multi-scale mean field representation method. The present invention introduces a multi-scale modeling module to capture the local interactions and global influences between agents from different scale perspectives, and then represents and optimizes the states and strategies in the multi-agent system. Through this multi-scale approach, the system can analyze at both the local and global levels to more comprehensively understand the behavior patterns between agents.

[0007] Secondly, the present invention ensures that the model can learn the behavior patterns of the agent in a variety of situations by conducting simulation training in a simulation scenario and diversifying the environment. For example, the model is trained and tested locally and globally through different maps. Then, these training parameters are input into the reinforcement learning model to obtain the potential representation of the agent at different scales and gradually extract effective behavior strategies. In this process, the model learns the interaction information at different scales through multi-scale learning. This includes learning the fine-grained behavior between agents at the local scale and learning the overall behavior pattern at the global scale. Through these learning and training, the system can effectively capture local and global multi-level information, thereby optimizing the decision-making process of the agent.

[0008] Finally, the information learned at each scale is integrated, and information extraction and learning are performed through the decoder. During the training process, the model judges the performance of the agent in different environments by comparing the similarity of the agent's behavior embedding vector, and optimizes it through the reward mechanism and loss function in reinforcement learning to ensure that the agent can demonstrate efficient collaboration and decision-making capabilities in a variety of situations. Ultimately, the model outputs a Q value containing the agent's decision information, which enables efficient path planning and conflict avoidance tasks in a complex multi-agent environment, thereby improving the performance and stability of the overall system. Through multi-scale modeling and learning, the present invention provides an effective solution that can better handle complex and changeable multi-agent systems, improve the accuracy of decision-making and the system's collaborative capabilities.

[0009] The first aspect of the present invention relates to an intelligent pathfinding algorithm based on a multi-scale mean field representation method, comprising the following implementation steps:

[0010] S1: Select a reinforcement learning environment and initialize the environment;

[0011] S2: interact with the environment and obtain feedback data and store it in the experience replay pool;

[0012] S3: Obtain environmental interaction data in batches from the experience replay pool to train the reinforcement learning model;

[0013] S4: Define the multi-scale mean field representation formula;

[0014] S5: Define the robot intelligent pathfinding process based on multi-scale mean field multi-agent reinforcement learning.

[0015] S1 specifically includes the following steps:

[0016] S1.1: Selecting a suitable reinforcement learning environment is crucial for reinforcement learning model training. In order to observe the effectiveness of the method proposed in the present invention and simulate the behavior of individuals in the real world, the present invention selects a two-dimensional grid world mapped by a real-world map for path planning as the environment for reinforcement learning model training to test the effectiveness of the method. The environment comes from the two-dimensional grid map mapped by the data sorted out according to the CAD drawing and field annotation in this study. In the environment, the agent is regarded as a point on the two-dimensional grid, and the behavior of the agent is simulated by the action of this point. In the simulation environment, the agent has the behavior of moving up, down, left, right, and staying in place. This environment can train and test the accuracy and effectiveness of the intelligent pathfinding algorithm based on the multi-scale mean field representation method from many aspects.

[0017] S1.2: Data sorting. In the simulation environment, the collected data needs to be cleaned up, mainly by deleting some points with repeated annotations, correcting offset points, etc. After sorting, it is visualized in the form of a two-dimensional grid.

[0018] S2 specifically includes the following steps:

[0019] S2.1: Interact with the environment and obtain experience replay pool data. Because the data used to train the intelligent pathfinding algorithm model based on the multi-scale mean field representation method is obtained by the interaction between the method and the environment, in the simulation environment, it is necessary to initialize the state of each agent in the environment according to the created model. The model gives a strategy based on the state of the agent at the current time step and interacts with the environment to obtain the state information and reward information of the next time step. Collect and organize this data and information, and store it in the experience replay pool for model training.

[0020] S2.2: Notation. For a multi-agent environment E, which contains N agents, we use represents the position of each agent in the environment, represents the action space of each agent, Represents the characteristic information of each agent, represents the reward of each agent after executing the action, s i Represents the state of each agent, near_field i Indicates the near-field information of the current agent, far_field irepresents the far-field information of the current agent, d i Indicates whether the current agent is in a terminal state.

[0021] S2.3: The process of acquiring data and storing it in the experience replay pool. In each step, first obtain the state s of each agent in the environment i , including the field of view v i and feature f i :

[0022] s i =(v i ,f i ) (1) At the same time, get the current position pi of the intelligent agent:

[0023]

[0024] Using the current agent's position information pi and the surrounding agent's action information a i , calculate the near-field feature near_field of each agent i Information, that is, the action of selecting the nearest k neighbors based on distance:

[0025] neighbors=top-k(|pi-pj|) (3)

[0026] near_field i =get_neighbors_actions(neighbors) (4) At the same time, calculate the far-field information of each agent:

[0027]

[0028] Based on the current state and near-field and far-field feature information, the reinforcement learning Q-Learning network model is used to predict the action of each agent. The final action distribution can be expressed as:

[0029] o i =Q-Learning(v i ,f i ,near_field i ,far_field i ) (6)

[0030] Use the ε-greedy policy to select actions:

[0031] a i = arg max(o i ) (7)

[0032] After each agent selects an action, the environment performs a simulation step based on the current action to obtain a new state next_states i and reward r i At the same time, update the agent's alive status i :

[0033] next_states i =(next_views i ,next_features i ) (8)

[0034] r i ,alive i =env.step(a i ) (9)

[0035] The current state s i , select action a i , and the next state next_states i , reward value r i , near field and far field information near_field i and far_field i And the termination mark d i , stored in the experience replay pool. The format of each agent's experience storage pool is:

[0036] E i =(s i ,a i ,near_field i ,far_field i ,r i ,next_states i ,d i ) (10)

[0037] S3 specifically includes the following steps:

[0038] S3.1: Based on reinforcement learning, MARL assigns a strategy to each agent i, and the agent selects a strategy at each time step. The future discounted return is used to measure the quality of the selected strategy. For a given strategy And the initial state s0~b 0 , the future discounted return of agent i can be expressed as follows:

[0039]

[0040] The Q-value function for the multi-agent version can be given by the Bellman equation:

[0041]

[0042] In the decision-making process, each agent will try to choose the best strategy to act in order to maximize its cumulative reward. Therefore, the corresponding Q value function of agent i is can be written as:

[0043]

[0044] S3.2: The mean field idea is to approximate the interaction between many agents as the interaction between an agent and a virtual agent that integrates other agents. Assuming that the agents in the environment only interact weakly with other agents through the mean field, the mean field can be written as:

[0045]

[0046] Among them j and a j They represent the local state and actions taken by agent j respectively, and -i represents the set of other agents except i.

[0047] S4 specifically includes the following steps:

[0048] S4.1: In the present invention, unlike the traditional single-scale mean field representation method, the present invention adopts a two-scale mean field representation, where the far field represents the action distribution of all other intelligent agents, which can be expressed by the following formula:

[0049]

[0050] Where N is the number of agents in the environment, π i (a i |s i ) indicates that the i-th agent is in state s i Next select action a i The probability distribution of .

[0051] The near field represents the weighted sum of the action distributions of the k neighbors of agent j. The weights are calculated based on the interaction strength of the agents through the attention mechanism. The formula is as follows:

[0052]

[0053] Finally, the mean field distributions of the two scales are combined to obtain the final multi-scale mean field representation:

[0054]

[0055] S4.2: The present invention uses the attention mechanism to weight the action distribution of the neighboring agents and uses the far-field information to learn the global context information to make up for the limitations of the neighborhood information. and the jth key vector The attention weights of k neighbors in the neighborhood are obtained as:

[0056]

[0057] Finally, the near-field output is:

[0058]

[0059] in is the j-th value vector.

[0060] For the far field, since there is only one global value in the far field, the i-th query vector of the near field is and the global key vector The attention weight of the far field is obtained as:

[0061]

[0062] Finally, the far-field output is:

[0063]

[0064] in is the global value vector.

[0065] Finally, the average fields of the two scales are fused to obtain the final output vector of the attention module:

[0066] y i =y near,i +y far,i (twenty two)

[0067] S5 specifically includes the following steps:

[0068] S5.1: In the previous step, we got the state of the agent i , including views i and features i , and the attention mechanism integrates the near-field and far-field information y i Next, embed the state information of the agent:

[0069] ev i =embedd i ng(views i ),ef i= embedding(features i ) (twenty three)

[0070] And embed the vector ev i , ef i and mean field information y i The information vectors are integrated to form a comprehensive information vector of the intelligent agent. Finally, we only need to adjust the MLP of the model output part according to business needs and conduct personalized information learning of the intelligent agent to output the required Q value.

[0071] S5.2: In the process of intelligent pathfinding of robots, Q value is used to represent the expected return of the robot moving in different directions at a specific position, reflecting the probability value of moving in each direction. By sorting these probability values ​​and selecting the direction with the highest probability value to move, the robot can make the optimal movement direction at the current position, thereby achieving more accurate robot movement guidance and achieving the purpose of reaching the target position. Ultimately, the robot can efficiently complete intelligent navigation tasks in dynamic or complex real-world environments, providing technical support for multi-scenario intelligent navigation.

[0072] The second aspect of the present invention relates to an intelligent path-finding device based on a multi-scale mean field representation method, comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the intelligent path-finding algorithm based on the multi-scale mean field representation method of the present invention.

[0073] A third aspect of the present invention relates to a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the intelligent path-finding algorithm based on the multi-scale mean field representation method of the present invention.

[0074] The present invention first selects a suitable multi-agent reinforcement learning environment, performs environmental preprocessing on it, and constructs a simulation environment suitable for reinforcement learning, in which each agent is represented as a point in a grid to simulate the interaction between its behavior and the environment. Secondly, through the MLP module of the embedding layer, the present invention linearly converts the observation value and eigenvalue of the agent into an embedding vector, realizes efficient encoding of the agent state information, and lays the foundation for subsequent processing. Subsequently, the attention mechanism is used to calculate and fuse the near-field and far-field information to form a multi-scale mean field representation. Specifically, the attention module simulates the different effects of neighbors on the agent by dynamically allocating weights, aggregates the weighted near-field neighbor information, and fuses it with the far-field information to generate the output of multi-scale information. This method not only accurately captures local interaction information, but also takes into account global features, and effectively reflects the multi-level interaction relationship in the environment. Finally, through the MLP module of the output layer, the embedded observation value, eigenvalue and attention fusion information are jointly learned to generate a Q value output. The Q value represents the expected return of the agent performing different actions in a specific state, and guides the agent to make the best decision by sorting and selecting the action with the highest probability value.

[0075] The present invention proposes two scales of average fields, near field and far field. The near field is the neighborhood field of the agent, and the attention mechanism is mainly combined to dynamically assign weights to neighboring agents. The far field is a global field, and the traditional average field method is used to model the global action distribution. By introducing the attention mechanism at the multi-scale level, the average field model can perform information fusion in a more refined way, thereby improving the learning efficiency and decision-making quality of the reinforcement learning algorithm in a large-scale, dynamically changing environment. Especially in tasks where multi-agent competition and cooperation coexist, the introduction of the attention mechanism helps the agent to optimize the performance of the overall system by focusing on key collaborative partners or potential threats. Since the multi-scale model processes information at multiple levels, the decision-making of the agent can more comprehensively consider local and global factors. In this way, even in an environment with greater uncertainty, the decision-making of the agent is more robust and can more effectively cope with interference and changes in the environment. In addition, through multi-scale hierarchical modeling, the stability of the system is further enhanced, thereby reducing the negative feedback and unstable behavior that may occur in the system.

[0076] This technological innovation not only promotes the application of multi-agent reinforcement learning methods in the fields of intelligent transportation and robot collaboration, but also provides new solutions for decision-making problems in large-scale complex environments. By improving the intelligence level of multi-agent systems, the attention mechanism lays a solid foundation for realizing more efficient and flexible artificial intelligence systems.

[0077] The advantages of the present invention are: significantly improving the accuracy of behavior prediction and decision-making efficiency, proposing a powerful and flexible intelligent path-finding algorithm based on a multi-scale mean field representation method, showing excellent performance in multi-agent reinforcement learning, especially in large-scale complex environments, achieving more efficient and accurate agent behavior prediction, and providing strong support for the development of multi-agent reinforcement learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 It is a structural diagram of the method of the present invention;

[0079] Figure 2 is a two-dimensional grid map of the real map mapping adopted by the method of this aspect;

[0080] Figure 3 A system flow chart for implementing the method of the present invention;

[0081] Figure 4 It is a schematic diagram of the device of the present invention. DETAILED DESCRIPTION

[0082] In order to make the objectives, technical solutions and advantages of the present invention more clear, the specific implementation modes of the present invention will be further described in detail below.

[0083] Example 1

[0084] The embodiment of the present invention provides an intelligent pathfinding algorithm based on a multi-scale mean field representation method. The system flow is as follows: Figure 3 As shown, the method includes:

[0085] S1: Select a reinforcement learning environment and initialize the environment. The steps are as follows:

[0086] S1.1: The present invention selects a two-dimensional grid world based on real-world map mapping for path planning as a training and testing environment, aiming to simultaneously evaluate the effectiveness of the algorithm and its accuracy in practical applications. The environment is based on a two-dimensional grid map compiled by this study. The map data is jointly constructed by CAD drawings and field annotation information, reflecting the spatial layout in the actual scene. In the environment, the agent is abstracted as a point on a two-dimensional grid, and its behavior is simulated by the movement and state changes of the point. In the simulation environment, the behavior of the agent is limited to movement up, down, left, right, or the choice of staying in place. Through environmental settings, the present invention can comprehensively evaluate the adaptability and performance of the algorithm in different scenarios, ensuring its feasibility and robustness in practical applications.

[0087] S1.2: Data sorting. In the simulation environment, the collected data needs to be cleaned up. The main cleaning methods include: 1. Deleting the points that are repeatedly marked during the marking process; 2. Correcting the offset points to ensure that the simulation environment truly restores the original environment. The reward value settings for this environment are shown in Table 1.

[0088] Table 1

[0089]

[0090] S2: Interact with the environment and obtain feedback data and store it in the experience replay pool; the specific steps are as follows:

[0091] S2.1: Interact with the environment and obtain experience replay pool data. Because the data used to train the intelligent pathfinding algorithm model based on the multi-scale mean field representation method is obtained by the interaction between the method and the environment, when processing the simulation environment, it is necessary to initialize the state of each agent in the environment according to the created model. The model gives a strategy based on the state of the agent at the current time step and interacts with the environment to obtain the state information and reward information of the next time step. Collect and organize this data and information and store it in the experience replay pool for model training.

[0092] S2.2: Notation. For a multi-agent environment E, which contains N agents, we use represents the position of each agent in the environment, represents the action space of each agent, Represents the characteristic information of each agent, represents the reward of each agent after executing the action, s i Represents the state of each agent, near_field i Indicates the near-field information of the current agent, far_field i represents the far-field information of the current agent, d i Indicates whether the current agent is in a terminal state.

[0093] S2.3: The process of acquiring data and storing it in the experience replay pool. In each step, first obtain the state s of each agent in the environment i , including the field of view v i and feature f i :

[0094] s i =(v i ,f i ) (1)

[0095] At the same time, get the current position p of the intelligent agent i :

[0096]

[0097] Using the current agent's position information p i and the action information a of the surrounding agents i , calculate the near-field feature near_field of each agent i Information, that is, the action of selecting the nearest k neighbors based on distance:

[0098] neighbors = top-k(|p i -p j |) (3)

[0099] near_field i =get_neighbors_actions(neighbors) (4)

[0100] At the same time, calculate the far-field information of each agent:

[0101]

[0102] Based on the current state and near-field and far-field feature information, the reinforcement learning Q-Learning network model is used to predict the action of each agent. The final action distribution can be expressed as:

[0103] o i =Q-Learning(v i ,f i ,near_field i ,far_field i ) (6)

[0104] Use the ε-greedy policy to select actions:

[0105] a i = arg max(o i ) (7)

[0106] After each agent selects an action, the environment performs a simulation step based on the current action to obtain a new state next_states i and reward r i At the same time, update the agent's alive status i :

[0107] next_states i =(next_views i ,next_features i ) (8)

[0108] r i,alive i =env.step(a i ) (9)

[0109] The current state s i , select action a i , and the next state next_states i , reward value r i , near field and far field information near_field i and far_field i And the termination mark d i , stored in the experience replay pool. The format of each agent's experience storage pool is:

[0110] E i =(s i ,a i ,near_field i ,far_field i ,r i ,next_states i ,d i ) (10)

[0111] S3 specifically includes the following steps:

[0112] S3.1: In the basic term of reinforcement learning, MARL assigns a strategy to each agent i, and the agent selects a strategy at each time step. The future discounted return is used to measure the quality of the selected strategy. For a given strategy And the initial state s0~b 0 , the future discounted return of agent i can be expressed as follows:

[0113]

[0114] The Q-value function for the multi-agent version can be given by the Bellman equation:

[0115]

[0116] In the decision-making process, each agent will try to choose the best strategy to act in order to maximize its cumulative reward. Therefore, the corresponding Q value function of agent i is can be written as:

[0117]

[0118] S3.2: The mean field idea is to approximate the interaction between many agents as the interaction between an agent and a virtual agent that integrates other agents. Assuming that the agents in the environment only interact weakly with other agents through the mean field, the mean field can be written as:

[0119]

[0120] Among them j and a j They represent the local state and actions taken by agent j respectively, and -i represents the set of other agents except i.

[0121] S4 specifically includes the following steps:

[0122] S4.1: In the present invention, unlike the traditional single-scale mean field representation method, the present invention adopts a two-scale mean field representation, where the far field represents the action distribution of all other intelligent agents, which can be expressed by the following formula:

[0123]

[0124] Where N is the number of agents in the environment, π i (a i |s i ) indicates that the i-th agent is in state s i Next select action a i The probability distribution of .

[0125] The near field represents the weighted sum of the action distributions of the k neighbors of agent j. The weights are calculated based on the interaction strength of the agents through the attention mechanism. The formula is as follows:

[0126]

[0127] Finally, the mean field distributions of the two scales are combined to obtain the final multi-scale mean field representation:

[0128]

[0129] S4.2: The present invention uses the attention mechanism to weight the action distribution of the neighboring agents and uses the far-field information to learn the global context information to make up for the limitations of the neighborhood information. and the jth key vector The attention weights of k neighbors in the neighborhood are obtained as:

[0130]

[0131] Finally, the near-field output is:

[0132]

[0133] in is the j-th value vector.

[0134] For the far field, since there is only one global value in the far field, the i-th query vector of the near field is and the global key vector The attention weight of the far field is obtained as:

[0135]

[0136] Finally, the far-field output is:

[0137]

[0138] in is the global value vector.

[0139] Finally, the average fields of the two scales are fused to obtain the final output vector of the attention module:

[0140] y i =y near,i +y far,i (twenty two)

[0141] S5 specifically includes the following steps:

[0142] S5.1: In the previous step, we got the state of the agent i , including views i and features i , and the attention mechanism integrates the near-field and far-field information y i Next, embed the state information of the agent:

[0143] ev i = embedding(views i ),ef i = embedding(features i ) (twenty three)

[0144] And embed the vector ev i , ef i and mean field information y i The information vectors are integrated to form a comprehensive information vector of the intelligent agent. Finally, we only need to adjust the MLP of the model output part according to business needs and conduct personalized information learning of the intelligent agent to output the required Q value.

[0145] S5.2: In robot intelligent pathfinding, the Q-value decision method is used to help the robot efficiently plan paths and optimize decisions. First, the environment is modeled by a reinforcement learning algorithm, and the positions in the grid map are defined as states. The robot selects actions by combining exploration and utilization. Secondly, a reasonable reward function is designed, including goal achievement rewards, collision penalties, and step penalties to guide the robot to optimize the path. The Q value is gradually updated through the Bellman equation, that is, formula (13), and the robot can find the optimal path after multiple iterations. In a high-dimensional state space, the multi-scale mean field representation method is combined with technologies such as priority experience replay and target network to enhance learning efficiency and stability. Finally, the robot can efficiently complete intelligent navigation tasks in dynamic or complex real environments, providing technical support for multi-scenario intelligent navigation.

[0146] The implementation and application cases show that the intelligent pathfinding algorithm based on the multi-scale mean field representation method proposed in the present invention has significant effectiveness. Compared with other design methods, the present invention innovatively introduces the mean field representation method of two scales, near field and far field, and successfully applies it to the real-world two-dimensional reconstruction scene environment. After effective training, the superiority of this method is verified. In terms of near-field information processing, the present invention uses the attention mechanism to dynamically assign weights to the k neighbors in the neighborhood of the intelligent agent, focusing on the influence of different neighbors on the decision-making of the intelligent agent, so as to accurately capture the local environmental information of the intelligent agent. In terms of far-field information processing, by averaging the global action features, the present invention obtains the global action distribution, emphasizes the guiding role of global information on the behavior of the intelligent agent, and effectively improves the perception of the global environment. The intelligent pathfinding algorithm based on the multi-scale mean field representation method of the present invention can comprehensively consider environmental observation information, near-field information and far-field information as input, and guide the intelligent agent to make the best decision at each moment by outputting the action probability distribution of the intelligent agent. By sorting the Q value, the intelligent agent can select the action with the highest probability value, thereby achieving more accurate behavior guidance. The experimental results fully verify the feasibility, flexibility and superior performance of the model.

[0147] Example 2

[0148] Reference Figure 4 This embodiment relates to an intelligent path-finding device based on a multi-scale mean field representation method, comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the intelligent path-finding algorithm based on the multi-scale mean field representation method of Example 1.

[0149] Example 1

[0150] This embodiment relates to a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the intelligent path-finding algorithm based on the multi-scale mean field representation method of embodiment 1 is implemented.

[0151] The above description is a specific embodiment of the present invention and the technical principles used. If the changes made according to the concept of the present invention do not exceed the spirit covered by the description and drawings, they should still fall within the scope of protection of the present invention.

Claims

1. An intelligent pathfinding algorithm based on a multi-scale mean field representation method, comprising the following steps: S1: Select a suitable reinforcement learning environment and initialize the environment; S2: interact with the environment and obtain feedback data and store it in the experience replay pool; S3: Obtain environmental interaction data in batches from the experience replay pool to train the reinforcement learning model; S4: Define the multi-scale mean field representation formula; S5: Define the robot intelligent pathfinding process based on multi-scale mean field multi-agent reinforcement learning.

2. The intelligent pathfinding algorithm based on the multi-scale mean field representation method according to claim 1, characterized in that: The step S1 comprises the following steps: S1.1: Choose an appropriate reinforcement learning environment; S1.2: Data organization: After selecting a suitable environment, you first need to perform preliminary processing on the environment to ensure the quality of the data and the training effect, including preprocessing of environmental feedback data, including standardization, denoising and filling of missing value pairs of input data to improve data consistency and accuracy; it also includes setting the environmental reward value and designing the reward function according to the task objectives to ensure that the reward mechanism can guide the agent to learn effective strategies.

3. The intelligent path-finding algorithm based on the multi-scale mean field representation method according to claim 1, characterized in that: Step S2 comprises the following steps: S2.1: When training a multi-scale mean field multi-agent reinforcement learning model, data comes from the interaction with the environment. First, the state of each agent is initialized according to the created model; then, the model gives a strategy based on the current state, and the agent interacts with the environment to obtain the state and reward information for the next time step; finally, these data will be sorted and stored in the experience replay pool for subsequent model training; S2.2: Notation: For a multi-agent environment E, which contains N agents, we use represents the position of each agent in the environment, represents the action space of each agent, Represents the characteristic information of each agent, represents the reward of each agent after executing the action, s i Represents the state of each agent, near_field i Indicates the near-field information of the current agent, far_field i represents the far-field information of the current agent, d i Indicates whether the current agent is in a terminal state; S2.3: Obtain data and store it in the experience replay pool; At each step, we first obtain the state s of each agent in the environment i , including the field of view v i and feature f i : s i =(v i ,f i ) (1) At the same time, get the current position p of the intelligent agent i : Using the current agent's position information p i and the action information a of the surrounding agents i , calculate the near-field feature near_field of each agent i Information, that is, the action of selecting the nearest k neighbors based on distance: neighbors=top-k(|p i -p j |) (3) near_field i =get_neighbors_actions(neighbors) (4) At the same time, calculate the far-field information of each agent: Based on the current state and near-field and far-field feature information, the reinforcement learning Q-Learning network model is used to predict the action of each agent. The final action distribution can be expressed as: o i =Q-Learning(v i ,f i ,near_field i ,far_field i ) (6) Use the ε-greedy policy to select actions: A i =arg max(o i ) (7) After each agent selects an action, the environment performs a simulation step based on the current action to obtain a new state next_states i and reward r i At the same time, update the agent's alive status i : next_states i =(next_views i ,next_features i ) (8) r i ,alive i =env.step(a i ) (9) The current state s i , the selected action a i , and the next state next_states i , reward value r i , near field and far field information near_field i and far_ffeld i And the termination mark d i , stored in the experience replay pool. The format of each agent's experience storage pool is: E i =(s i ,a i ,near_field i ,far_field i ,r i ,next_states i ,d i ) (10) 4. The intelligent path-finding algorithm based on the multi-scale mean field representation method according to claim 1, characterized in that: Step S3 includes the following steps: S3.1: For a given strategy And the initial state s0~b 0 , the future discounted return of agent i is expressed as follows: The Q-value function for the multi-agent version is given by the Bellman equation: Agent i, corresponding Q value function Written as: Q i (s, a)=(1-α)Q i (s, a)+α[r i (s, a)+yV i (s′)] (13) S3.2: Assume that the agents in the environment interact only weakly with other agents through the mean field. Written as: Among them j and a j They represent the local state and actions taken by agent j respectively, and -i represents the set of other agents except i.

5. The intelligent path-finding algorithm based on the multi-scale mean field representation method according to claim 1, characterized in that: Step S4 comprises the following steps: S4.1: We adopt the average field representation of two scales, where the far field represents the action distribution of all other agents, expressed by the following formula: Where N is the number of agents in the environment, π i (a i |s i ) indicates that the i-th agent is in state s i Next select action a i The probability distribution of The near field represents the weighted sum of the action distributions of the k neighbors of agent j. The weights are calculated based on the interaction strength of the agents through the attention mechanism. The formula is as follows: Finally, the mean field distributions of the two scales are combined to obtain the final multi-scale mean field representation: S4.2: Use the attention mechanism to weight the action distribution of the neighboring agents and use the far-field information to learn the global context information. For the i-th query vector in the near field and the jth key vector The attention weights of k neighbors in the neighborhood are obtained as: Finally, the near-field output is: in is the i-th value vector; For the far field, since there is only one global value in the far field, the i-th query vector of the near field is and the global key vector The attention weight of the far field is obtained as: Finally, the far-field output is: in is the global value vector; Finally, the average fields of the two scales are fused to obtain the final output vector of the attention module: and i =and near,i +y far,i (22) 6. The intelligent path-finding algorithm based on the multi-scale mean field representation method according to claim 1, characterized in that: Step S5 comprises the following steps: S5.1: The state information of the agent is embedded as: ev i =embedding(views i ),ef i =embedding(features i ) (23) And embed the vector ev i , ef i and mean field information y i Fusion is performed to form a comprehensive information vector of the intelligent agent; the MLP of the model output part is adjusted according to business needs, and the intelligent agent personalized information learning outputs the required Q value; S5.2: In the process of intelligent pathfinding, the Q value is used to represent the expected return of the robot moving in different directions at a specific position, reflecting the probability value of moving in each direction. By sorting these probability values ​​and selecting the direction with the highest probability value to move, the robot can select the optimal moving direction at the current position and accurately guide the robot to move to the target position.

7. An intelligent path-finding device based on a multi-scale mean field representation method, characterized in that: It comprises a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the intelligent path-finding algorithm based on the multi-scale mean field representation method described in any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the intelligent path-finding algorithm based on the multi-scale mean field representation method described in any one of claims 1 to 6 is implemented.