A method and system for generating a robot navigation path

By constructing a dynamic relationship diagram and combining multiple algorithms to optimize the adjacency matrix, the optimal navigation path is generated, which solves the uncertainty problem of traditional navigation methods in the dynamic environment, and realizes efficient navigation of robots in the human-computer coexistence environment.

CN120008616BActive Publication Date: 2025-07-22EAST CHINA JIAOTONG UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510487508.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-22
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

Traditional robot navigation methods rely on static maps and preset rules, making it difficult to deal with uncertainty in dynamic environments, especially in environments where human-computer coexistence, resulting in increased navigation complexity.

Method used

By constructing a dynamic relationship graph, the adjacency matrix is optimized using the side channel state sequence matching strategy, and the optimal navigation path is generated by combining graph convolution networks, non-negative matrix decomposition, Transformer model and improved Monte Carlo tree search algorithm.

Benefits of technology

In-depth modeling and efficient planning of robot and pedestrian interactions in dynamic environments are achieved, and the accuracy of robot autonomous navigation is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120008616B_ABST
    Figure CN120008616B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for generating a robot navigation path. The method includes: performing non-negative matrix factorization on an adjacency matrix and a node feature matrix to obtain a first latent factor matrix and a second latent factor matrix; fusing the first latent factor matrix, the second latent factor matrix, and the node feature matrix to obtain a weighted fusion matrix; screening the weighted fusion matrix according to a preset adaptive motion structure robust screening strategy to obtain a key state matrix; linearly inputting the key state matrix into a preset Transformer model, and the Transformer model outputs an enhanced feature matrix; dynamically updating the adjacency matrix according to the enhanced feature matrix to obtain a target adjacency matrix; inputting the target adjacency matrix into a process reward model, and performing a search according to an improved Monte Carlo tree search algorithm, and the process reward model outputs an optimal navigation path. It realizes the deep modeling and efficient planning of the interaction between the robot and pedestrians in a dynamic environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of robot control, and particularly relates to a method and system for generating a navigation path of a robot. Background Art

[0002] With the rapid development of robot technology, autonomous navigation has become an important research direction in the field of robots. In a dynamic environment, a robot needs to perceive the surrounding environment in real time and make reasonable navigation decisions. Especially in an environment where humans and robots coexist (such as shopping malls, hospitals, airports, etc.), the interaction between the robot and pedestrians increases the complexity of navigation. Traditional navigation methods usually rely on static maps and preset rules, and it is difficult to cope with the uncertainties in a dynamic environment. Summary of the Invention

[0003] The present invention provides a method and system for generating a navigation path of a robot, which are used to solve the technical problem that it is usually difficult to cope with the uncertainties in a dynamic environment by relying on static maps and preset rules.

[0004] In a first aspect, the present invention provides a method for generating a navigation path of a robot, including:

[0005] Constructing a dynamic relationship graph according to the potential state of the robot and the potential state of pedestrians;

[0006] Optimizing the feature matrix in the dynamic relationship graph according to a preset side-channel state sequence matching strategy to obtain an adjacency matrix, and inputting the adjacency matrix into a graph convolutional network for learning to obtain a node feature matrix;

[0007] Performing non-negative matrix factorization according to the adjacency matrix and the node feature matrix to obtain a first latent factor matrix and a second latent factor matrix;

[0008] Fusing the first latent factor matrix, the second latent factor matrix and the node feature matrix to obtain a weighted fusion matrix, and screening the weighted fusion matrix according to a preset adaptive motion structure robust screening strategy to obtain a key state matrix;

[0009] Linearly converting the key state matrix into a query matrix, a key matrix and a value matrix, and inputting them into a preset Transformer model, and the Transformer model outputs an enhanced feature matrix;

[0010] Dynamically updating the adjacency matrix according to the enhanced feature matrix to obtain a target adjacency matrix, inputting the target adjacency matrix into a process reward model, and performing search according to an improved Monte Carlo tree search algorithm, and the process reward model outputs an optimal navigation path.

[0011] In a second aspect, the present invention provides a system for generating a robot navigation path, including:

[0012] A construction module configured to construct a dynamic relationship graph based on the potential states of the robot and the pedestrians;

[0013] An optimization module configured to optimize the feature matrix in the dynamic relationship graph according to a preset side-channel state sequence matching strategy to obtain an adjacency matrix, and input the adjacency matrix into a graph convolutional network for learning to obtain a node feature matrix;

[0014] A decomposition module configured to perform non-negative matrix decomposition on the basis of the adjacency matrix and the node feature matrix to obtain a first latent factor matrix and a second latent factor matrix;

[0015] A screening module configured to fuse the first latent factor matrix, the second latent factor matrix and the node feature matrix to obtain a weighted fusion matrix, and screen the weighted fusion matrix according to a preset adaptive motion structure robust screening strategy to obtain a key state matrix;

[0016] A first output module configured to linearly transform the key state matrix into a query matrix, a key matrix and a value matrix, and input them into a preset Transformer model, and the Transformer model outputs an enhanced feature matrix;

[0017] A second output module configured to dynamically update the adjacency matrix according to the enhanced feature matrix to obtain a target adjacency matrix, input the target adjacency matrix into a process reward model, and perform a search according to an improved Monte Carlo tree search algorithm, and the process reward model outputs an optimal navigation path.

[0018] In a third aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the method for generating a robot navigation path according to any embodiment of the present invention.

[0019] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program instructions are executed by a processor, the processor is enabled to execute the steps of the method for generating a robot navigation path according to any embodiment of the present invention.

[0020] The method and system for generating a robot navigation path of the present application optimize the feature matrix in the dynamic relationship graph according to a preset side-channel state sequence matching strategy to obtain an adjacency matrix, input the adjacency matrix into a graph convolutional network for learning to obtain a node feature matrix, perform non-negative matrix factorization on the adjacency matrix and the node feature matrix to obtain a first latent factor matrix and a second latent factor matrix, fuse the first latent factor matrix, the second latent factor matrix and the node feature matrix to obtain a weighted fusion matrix, and screen the weighted fusion matrix according to a preset adaptive motion structure robust screening strategy to obtain a key state matrix, linearly transform the key state matrix into a query matrix, a key matrix and a value matrix, and input them into a preset Transformer model. The Transformer model outputs an enhanced feature matrix, dynamically updates the adjacency matrix according to the enhanced feature matrix to obtain a target adjacency matrix, inputs the target adjacency matrix into a process reward model, and performs a search according to an improved Monte Carlo tree search algorithm. The process reward model outputs an optimal navigation path, realizing in-depth modeling and efficient planning of the interaction between the robot and pedestrians in a dynamic environment, improving the accuracy of the robot's autonomous navigation, and being beneficial to the efficient navigation and decision-making of the robot in a dynamic environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a flowchart of a method for generating a robot navigation path provided by an embodiment of the present invention;

[0023] Figure 2 It is a structural block diagram of a system for generating a robot navigation path provided by an embodiment of the present invention;

[0024] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0026] Please refer to Figure 1 , which shows a flowchart of a method for generating a robot navigation path according to the present application.

[0027] As Figure 1 shown, the method for generating a robot navigation path specifically includes the following steps:

[0028] Step S101, constructing a dynamic relationship graph according to the potential states of the robot and the pedestrians.

[0029] In this step, let be the state node of the robot, be the state node of the pedestrian, d is the state dimension, be a real number, and the weights of the edges in the dynamic relationship graph are obtained by calculating the similarity, which is used to represent the interaction intensity between the robot and the pedestrians, and a dynamic relationship graph is obtained. The expression is:

[0030]

[0031] In the formula, is the kernel matrix, representing the similarity between the potential state of the robot and the potential state of the pedestrians, is the bandwidth function, which is used to control the width of the kernel function, is the Euclidean distance.

[0032] Step S102, optimizing the feature matrix in the dynamic relationship graph according to a preset side-channel state sequence matching strategy to obtain an adjacency matrix, and inputting the adjacency matrix into a graph convolutional network for learning to obtain a node feature matrix.

[0033] In this step, to improve the accuracy of the algorithm, the SCSS matching strategy is proposed. The original features are weighted and adjusted through side-channel information, and the similarity between the predicted sequence and the actual sequence is more accurately measured through the side-channel enhanced feature vector, so as to optimize the learning of the relationship graph. Specifically, the feature matrix is extracted from the dynamic relationship graph according to the sum method. The expression is:

[0034] ,

[0035] ,

[0036] ,

[0037] In the formula, is the normalized weight of node i, is the sum of the weights of all the edges connected to node i, is the degree centrality of node k, is the total number of nodes in the relationship graph, is the node connected to node i, is the interaction intensity between nodes i and j in the dynamic relationship graph, is for the value obtained by column normalization;

[0038] Optimize the feature matrix according to the preset side-channel state sequence matching strategy to obtain the adjacency matrix. The expression is:

[0039] ,

[0040] ,

[0041] ,

[0042] ,

[0043] ,

[0044] ,

[0045] In the formula, is the element of the adjacency matrix optimized by the side-channel state sequence matching strategy, is the element of the original adjacency matrix, representing the interaction intensity between nodes i and j, is the hyperparameter for adjusting the optimization intensity, is the Gaussian kernel function, is the comprehensive matching function of the side-channel state sequence, is the predicted state sequence, is the actual state sequence, is the time decay coefficient, is the matching score at time t, is the matching score between the predicted state sequence and the actual state sequence at time t, is the matching score between the predicted state sequence and the actual state sequence at time t - 1, which is used to smooth the optimization process, is the matching score at time t - 1, d is the number of side-channel feature dimensions, is the weight coefficient of the j-th dimension feature, is the j-th dimension predicted state side-channel enhanced feature after the SCSS matching strategy, is the j-th dimension actual state side-channel enhanced feature after the SCSS matching strategy, σ is the Gaussian kernel bandwidth, is the i-th dimension side-channel enhanced feature after the SCSS matching strategy, is the i-th dimension feature at time t, m is the number of side-channel information types, Denotes element-wise multiplication, is the weight coefficient of the k-th type of side-channel information, is the k-th type of side-channel information of node i, is the weight coefficient of node i at time t-1, is the position of node i at time t-1, is the side-channel information of node i at time t-1, is the additional state of node i at time t-1, is the weight of node i at time t, is the influence weight of node l on node i at time t, is the position of node l at time t-1, is the weight coefficient of node l at time t-1, is the influence weight of node k on i at time t, is the weight coefficient of node k at time t-1, is the position of node k at time t-1, is the side-channel information of node k at time t-1, is the additional state of node k at time t-1, where k is a node, represents the input set of node k at time t, is the influence weight of node i on y at time t-1, is the weight of node y at time t-1, is the side-channel information of node y at time t-1, is the additional state of node y at time t-1, is the influence weight of node l on y at time t-1, is a node, is the output set of node i at time t, is the set of side-channel information, is the f-dimensional side-channel information, is the one-dimensional side-channel information, is the two-dimensional side-channel information. It should be noted that when the adjacency matrix is input into the graph convolutional network for learning, the expression for the node feature matrix is:

[0046] ,

[0047] In the formula, is the node feature matrix of the (m + 1)-th layer, is the activation function, is the degree matrix of, is the adjacency matrix, is the node feature matrix of the m-th layer, is the weight matrix of the m-th layer.

[0048] Step S103, perform non - negative matrix factorization based on the adjacency matrix and the node feature matrix to obtain a first latent factor matrix and a second latent factor matrix.

[0049] In this step, the expressions for obtaining the first latent factor matrix and the second latent factor matrix by performing non - negative matrix factorization based on the adjacency matrix and the node feature matrix are as follows:

[0050] ,

[0051] ,

[0052] where, is the objective function of the first latent factor matrix W, is the objective function of the second latent factor matrix P, is the first latent factor matrix, is the second latent factor matrix, , is the weight for controlling the self - correlation of features, is the node feature matrix of the m - th layer, is the transpose of the node feature matrix of the m - th layer, , is the weight for controlling the graph structure matching, is the trace of the matrix, , is the weight for controlling the global constraint, is the graph structure constraint matrix, is the identity matrix, is the transpose of the second latent factor matrix.

[0053] Step S104, fuse the first latent factor matrix, the second latent factor matrix and the node feature matrix to obtain a weighted fusion matrix, and screen the weighted fusion matrix according to a preset adaptive motion structure robust screening strategy to obtain a key state matrix.

[0054] In this step, the expression for fusing the first latent factor matrix, the second latent factor matrix and the node feature matrix to obtain a weighted fusion matrix is as follows:

[0055] ,

[0056] ,

[0057] where, is the deep interaction feature matrix, is the first latent factor matrix, is the second latent factor matrix, is the weighted fusion matrix, is the fusion weight.

[0058] It should be noted that an Adaptive Motion Structure Robust Screening Strategy (AMSRSS) is proposed for the prediction and modeling of pedestrian motion states, to improve robustness and prediction accuracy. According to the preset Adaptive Motion Structure Robust Screening Strategy, the weighted fusion matrix is screened, and the expression for the key state matrix is:

[0059] ,

[0060] ,

[0061] ,

[0062] ,

[0063] ,

[0064] In the formula, is the key state matrix, is the motion structure mask matrix, identifying the nodes to be retained in the dynamic environment, is element-wise multiplication, is the set of significance states, is the normalized adjacency matrix, is the weighted fusion matrix, is the transpose symbol, is the feature vector of node i, is the significance scoring function, is the significance threshold, is the significance scoring function of node i at time t, is the temperature coefficient, controlling the sharpness of Softmax, is the adaptive parameter at time t, is the distance function, measuring the node features and the global average feature difference, is the global average feature, is the feature of node i at time t, is the feature of node j at time t, is the adaptive parameter at time t - 1, is the derivative of the learning rate decay coefficient, is the momentum coefficient, balancing the historical gradient and the current gradient, is the loss function gradient, is the derivative of the high-order adjustment parameter, , is the system state variable, , is the derivative of the system state variable, is the adjustment parameter.

[0065] Step S105: linearly transform the key state matrix into a query matrix, a key matrix, and a value matrix, and input them into a preset Transformer model, and the Transformer model outputs an enhanced feature matrix.

[0066] In this step, the key state matrix is linearly transformed into a query matrix, a key matrix, and a value matrix, and the expression is:

[0067] ,

[0068] ,

[0069] ,

[0070] In the formula, is the query matrix, is the key matrix, is the value matrix, is the key state matrix, is the learnable weight matrix of the query matrix, is the learnable weight matrix of the value matrix, is the learnable weight matrix of the key state matrix;

[0071] Input the query matrix, the key matrix, and the value matrix into the preset Transformer model, capture the global spatial dependence relationship through multi-head attention, enhance the expression ability of node features, and output the enhanced feature matrix. The expression is:

[0072] ,

[0073] ,

[0074] ,

[0075] In the formula, is the enhanced feature matrix, is layer normalization, which normalizes the input features, is the multi-layer perceptron, is the adaptive multi-head attention mechanism, is concatenation, which combines the outputs of multiple attention heads, is the output of the first attention head, is the output of the h-th attention head, is a learnable linear projection matrix, is the output of the i-th attention head, is the query matrix of the i-th head, is the key matrix of the i-th head, is the feature dimension of each attention head, is the value matrix of the i-th head, is the Softmax function, which normalizes the attention weights, is the transpose symbol.

[0076] Step S106: Dynamically update the adjacency matrix according to the enhanced feature matrix to obtain a target adjacency matrix, input the target adjacency matrix into the process reward model, and perform a search according to the improved Monte Carlo tree search algorithm. The process reward model outputs an optimal navigation path.

[0077] In this step, the adjacency matrix is dynamically updated according to the enhanced feature matrix to obtain a target adjacency matrix. The expression is:

[0078] ,

[0079] In the formula, is the target adjacency matrix, is the activation function, , is the enhanced feature matrix is the global feature vector of node i and node j in

[0080] An improved Monte Carlo tree search algorithm is adopted, combined with the process reward model to guide the search. Starting from the newly expanded node, it is randomly simulated until the termination state is reached, and the potential value of the node is evaluated. The process reward model is used to evaluate the reward of each step during the simulation process, and the simulation results are backpropagated into the search tree to update the statistical information of all nodes on the path to obtain the optimal navigation path. The expression is:

[0081] ,

[0082] ,

[0083] ,

[0084] ,

[0085] In the formula, is the dynamic policy parameter, is the policy function, is the scaling coefficient, is the Softmax function, which normalizes the attention weights. is the enhanced feature matrix, which contains global spatial dependencies. is the learnable weight matrix. is the action-value function. is the state vector. is the action vector. is the query weight matrix. is the bias term. is the upper confidence bound, a node selection strategy that balances exploration and exploitation. is the cumulative reward of node n. is the number of visits to node n. is the exploration coefficient. is the number of visits to the parent node p(n). is the latent factor weight. is the trace of the matrix. is the first latent factor matrix. is the second latent factor matrix. is the comprehensive reward function, which combines immediate rewards, state values, and graph structure information. is the immediate reward. is the discount factor, which weighs current and future rewards. is the state-value function. is the adjacency matrix weight.

[0086] In summary, the method of this application optimizes the feature matrix in the dynamic relationship graph according to the preset side-channel state sequence matching strategy to obtain the adjacency matrix, inputs the adjacency matrix into the graph convolutional network for learning to obtain the node feature matrix, performs non-negative matrix factorization on the adjacency matrix and the node feature matrix to obtain the first latent factor matrix and the second latent factor matrix, fuses the first latent factor matrix, the second latent factor matrix, and the node feature matrix to obtain the weighted fusion matrix, screens the weighted fusion matrix according to the preset adaptive motion structure robust screening strategy to obtain the key state matrix, linearly transforms the key state matrix into the query matrix, the key matrix, and the value matrix, and inputs them into the preset Transformer model. The Transformer model outputs to obtain the enhanced feature matrix, dynamically updates the adjacency matrix according to the enhanced feature matrix to obtain the target adjacency matrix, inputs the target adjacency matrix into the process reward model, and performs search according to the improved Monte Carlo tree search algorithm. The process reward model outputs to obtain the optimal navigation path, realizing the deep modeling and efficient planning of the interaction between the robot and the pedestrian in the dynamic environment, improving the accuracy of the robot's autonomous navigation, and facilitating the efficient navigation and decision-making of the robot in the dynamic environment.

[0087] Please refer to Figure 2, which shows a structural block diagram of a system for generating a robot navigation path according to the present application.

[0088] As Figure 2 shown, the system 200 for generating a robot navigation path includes a construction module 210, an optimization module 220, a decomposition module 230, a screening module 240, a first output module 250, and a second output module 260.

[0089] Among them, the construction module 210 is configured to construct a dynamic relationship graph according to the potential states of the robot and the pedestrians; the optimization module 220 is configured to optimize the feature matrix in the dynamic relationship graph according to a preset side-channel state sequence matching strategy to obtain an adjacency matrix, and input the adjacency matrix into a graph convolutional network for learning to obtain a node feature matrix; the decomposition module 230 is configured to perform non-negative matrix decomposition according to the adjacency matrix and the node feature matrix to obtain a first latent factor matrix and a second latent factor matrix; the screening module 240 is configured to fuse the first latent factor matrix, the second latent factor matrix, and the node feature matrix to obtain a weighted fusion matrix, and screen the weighted fusion matrix according to a preset adaptive motion structure robust screening strategy to obtain a key state matrix; the first output module 250 is configured to linearly transform the key state matrix into a query matrix, a key matrix, and a value matrix, and input them into a preset Transformer model, and the Transformer model outputs an enhanced feature matrix; the second output module 260 is configured to dynamically update the adjacency matrix according to the enhanced feature matrix to obtain a target adjacency matrix, input the target adjacency matrix into a process reward model, and perform a search according to an improved Monte Carlo tree search algorithm, and the process reward model outputs an optimal navigation path.

[0090] It should be understood that Figure 2 the modules described in Figure 1 correspond to the respective steps in the method described in reference Figure 2 . Thus, the operations, features, and corresponding technical effects described above for the method also apply to the modules in

[0091] and will not be elaborated here.

[0092] As an implementation, the computer-readable storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as:

[0093] Construct a dynamic relationship graph based on the potential states of the robot and the pedestrian;

[0094] Optimize the feature matrix in the dynamic relationship graph according to a preset side-channel state sequence matching strategy to obtain an adjacency matrix, and input the adjacency matrix into a graph convolutional network for learning to obtain a node feature matrix;

[0095] Perform non-negative matrix factorization on the adjacency matrix and the node feature matrix to obtain a first latent factor matrix and a second latent factor matrix;

[0096] Fuse the first latent factor matrix, the second latent factor matrix, and the node feature matrix to obtain a weighted fusion matrix, and screen the weighted fusion matrix according to a preset adaptive motion structure robust screening strategy to obtain a key state matrix;

[0097] Linearly transform the key state matrix into a query matrix, a key matrix, and a value matrix, and input them into a preset Transformer model, and the Transformer model outputs an enhanced feature matrix;

[0098] Dynamically update the adjacency matrix according to the enhanced feature matrix to obtain a target adjacency matrix, input the target adjacency matrix into a process reward model, and perform a search according to an improved Monte Carlo tree search algorithm, and the process reward model outputs an optimal navigation path.

[0099] A computer-readable storage medium may include a storage program area and a storage data area. Among them, the storage program area may store an operating system and application programs required for at least one function; the storage data area may store data created according to the use of the robot navigation path generation system, etc. In addition, the computer-readable storage medium may include high-speed random access memory, and may also include a memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the computer-readable storage medium may optionally include a memory remotely provided with respect to the processor, and these remote memories may be connected to the robot navigation path generation system through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0100] Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention, as Figure 3 shown, the device includes: a processor 310 and a memory 320. The electronic device may further include: an input device 330 and an output device 340. The processor 310, the memory 320, the input device 330, and the output device 340 may be connected through a bus or other means, Figure 3Take the bus connection as an example. The memory 320 is the computer-readable storage medium described above. The processor 310 executes various functional applications and data processing of the server by running non-volatile software programs, instructions, and modules stored in the memory 320, that is, implements the method for generating the robot navigation path in the above method embodiments. The input device 330 can receive input digital or character information, and generate key signal inputs related to user settings and function controls of the robot navigation path generation system. The output device 340 can include display devices such as a display screen.

[0101] The above electronic device can execute the method provided by the embodiments of the present invention, and has corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided by the embodiments of the present invention.

[0102] As an implementation manner, the above electronic device is applied to the robot navigation path generation system and is used for the client, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to:

[0103] Construct a dynamic relationship graph according to the potential state of the robot and the potential state of the pedestrian;

[0104] Optimize the feature matrix in the dynamic relationship graph according to the preset side-channel state sequence matching strategy to obtain an adjacency matrix, and input the adjacency matrix into a graph convolutional network for learning to obtain a node feature matrix;

[0105] Perform non-negative matrix factorization according to the adjacency matrix and the node feature matrix to obtain a first latent factor matrix and a second latent factor matrix;

[0106] Fuse the first latent factor matrix, the second latent factor matrix, and the node feature matrix to obtain a weighted fusion matrix, and screen the weighted fusion matrix according to the preset adaptive motion structure robust screening strategy to obtain a key state matrix;

[0107] Linearly transform the key state matrix into a query matrix, a key matrix, and a value matrix, and input them into a preset Transformer model, and the Transformer model outputs an enhanced feature matrix;

[0108] Dynamically update the adjacency matrix according to the enhanced feature matrix to obtain a target adjacency matrix, input the target adjacency matrix into a process reward model, and perform search according to the improved Monte Carlo tree search algorithm, and the process reward model outputs an optimal navigation path.

[0109] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.

[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A method for generating a robot navigation path, characterized in that Including: Construct a dynamic relationship graph based on the potential state of the robot and the potential state of the pedestrian; Optimize the feature matrix in the dynamic relationship graph according to a preset side-channel state sequence matching strategy to obtain an adjacency matrix, and input the adjacency matrix into a graph convolutional network for learning to obtain a node feature matrix; Perform non-negative matrix factorization on the adjacency matrix and the node feature matrix to obtain a first latent factor matrix and a second latent factor matrix; Fuse the first latent factor matrix, the second latent factor matrix, and the node feature matrix to obtain a weighted fusion matrix, and screen the weighted fusion matrix according to a preset adaptive motion structure robust screening strategy to obtain a key state matrix; Linearly transform the key state matrix into a query matrix, a key matrix, and a value matrix, and input them into a preset Transformer model, and the Transformer model outputs an enhanced feature matrix; Dynamically update the adjacency matrix according to the enhanced feature matrix to obtain a target adjacency matrix, input the target adjacency matrix into a process reward model, and perform search according to an improved Monte Carlo tree search algorithm, and the process reward model outputs an optimal navigation path.

2. The method for generating a robot navigation path according to claim 1, wherein The constructing a dynamic relationship graph based on the potential state of the robot and the potential state of the pedestrian includes: Let be the state node of the robot, be the state node of the pedestrian, d is the state dimension, be a real number. By calculating the similarity, the weight of the edge in the dynamic relationship graph is obtained, which is used to represent the interaction intensity between the robot and the pedestrian, and the dynamic relationship graph is obtained. The expression is: , In the formula, is the kernel matrix, representing the potential state of the robot and the potential state of the pedestrian similarity, is the bandwidth function, used to control the width of the kernel function, is the Euclidean distance.

3. A method for generating a robot navigation path according to claim 1, characterized in that, The optimizing the feature matrix in the dynamic relationship graph according to a preset side-channel state sequence matching strategy to obtain an adjacency matrix includes: Extract a feature matrix from the dynamic relationship graph according to the summation method, and the expression is: , , , Wherein, is the normalized weight of node i, is the sum of the weights of all edges connected to node i, is the degree centrality of node k, is the total number of nodes in the relational graph, is the node connected to node i, is the interaction strength between nodes i and j in the dynamic relational graph, is for the value obtained by column normalization; Optimize the feature matrix according to a preset side-channel state sequence matching strategy to obtain an adjacency matrix, and the expression is: , , , , , , Wherein, is the element of the adjacency matrix optimized by the side-channel state sequence matching strategy, is the element of the original adjacency matrix, representing the interaction intensity between nodes i and j, is the hyperparameter for adjusting the optimization intensity, is the Gaussian kernel function, is the comprehensive matching function of the side-channel state sequence, is the predicted state sequence, is the actual state sequence, is the time decay coefficient, is the matching score at time t, is the matching score between the predicted state sequence and the actual state sequence at time t, is the matching score between the predicted state sequence and the actual state sequence at time t-1, used for smoothing the optimization process, is the matching score at time t-1, d is the dimension number of the side-channel feature, is the weight coefficient of the j-th dimension feature, is the j-th dimensional predicted state side-channel enhanced feature after the SCSS matching strategy, is the j-th dimensional actual state side-channel enhanced feature after the SCSS matching strategy, σ is the Gaussian kernel bandwidth, is the i-th dimensional side-channel enhanced feature after the SCSS matching strategy, is the i-th dimensional feature at time t, m is the number of side-channel information types, represents element-wise multiplication, is the weight coefficient of the k-th type of side-channel information, is the k-th type of side-channel information of node i, is the weight coefficient of node i at time t-1, is the position of node i at time t-1, is the side-channel information of node i at time t-1, is the additional state of node i at time t-1, is the weight of node i at time t, is the influence weight of node l on node i at time t, is the position of node l at time t-1, is the weight coefficient of node l at time t-1, is the influence weight of node k on i at time t, is the weight coefficient of node k at time t-1, is the position of node k at time t-1, is the side-channel information of node k at time t-1, is the additional state of node k at time t-1, k is the node, To represent the input set of node k at time t, is the influence weight of node i on y at time t-1, is the weight of node y at time t-1, is the side-channel information of node y at time t-1, is the additional state of node y at time t-1, is the influence weight of node l on y at time t-1, is a node, is the output set of node i at time t, is a set of side-channel information, is f-dimensional side-channel information, is one-dimensional side-channel information, is two-dimensional side-channel information.

4. A method for generating a robot navigation path according to claim 1, wherein Where, The expression for inputting the adjacency matrix into a graph convolutional network for learning to obtain a node feature matrix is: , Wherein, is the node feature matrix of the (m + 1)-th layer, is the activation function, is the degree matrix of, is the adjacency matrix, is the node feature matrix of the m-th layer, is the weight matrix of the m-th layer.

5. A method for generating a robot navigation path according to claim 1, characterized in that, Where, The expression for performing non-negative matrix factorization on the adjacency matrix and the node feature matrix to obtain a first latent factor matrix and a second latent factor matrix is: , , Wherein, is the objective function of the first latent factor matrix W, is the objective function of the second latent factor matrix P, is the first latent factor matrix, is the second latent factor matrix, , is the weight for controlling the autocorrelation of features, is the node feature matrix of the m-th layer, is the transpose of the node feature matrix of the m-th layer, , is the weight for controlling the graph structure matching, is the trace of the matrix, , is the weight for controlling the global constraint, is the graph structure constraint matrix, is the identity matrix, is the transpose of the second latent factor matrix.

6. The method for generating a robot navigation path according to claim 1, wherein Where, The expression for fusing the first latent factor matrix, the second latent factor matrix, and the node feature matrix to obtain a weighted fusion matrix is: , , In the formula, is the deep interaction feature matrix, is the first latent factor matrix, is the second latent factor matrix, is the weighted fusion matrix, is the fusion weight.

7. A method for generating a robot navigation path according to claim 1, wherein Where, The expression for screening the weighted fusion matrix according to a preset adaptive motion structure robust screening strategy to obtain a key state matrix is: , , , , , Wherein, is the key state matrix, is the motion structure mask matrix, identifying the nodes to be retained in the dynamic environment, is element-wise multiplication, is the set of significance states, is the normalized adjacency matrix, is the weighted fusion matrix, is the transpose symbol, is the feature vector of node i, is the significance scoring function, is the significance threshold, is the significance scoring function of node i at time t, is the temperature coefficient, controlling the sharpness of Softmax, is the adaptive parameter at time t, is the distance function, measuring the node features and the global average feature difference, is the global average feature, is the feature of node i at time t, is the feature of node j at time t, is the adaptive parameter at time t-1, is the derivative of the learning rate decay coefficient, is the momentum coefficient, balancing the historical gradient and the current gradient, is the gradient of the loss function, is the derivative of the high-order adjustment parameter, , is the system state variable, , is the derivative of the system state variable, is the adjustment parameter.

8. A method for generating a robot navigation path according to claim 1, characterized in that, The linearly transforming the key state matrix into a query matrix, a key matrix, and a value matrix, and inputting them into a preset Transformer model, and the Transformer model outputs an enhanced feature matrix includes: The expression for linearly transforming the key state matrix into a query matrix, a key matrix, and a value matrix is: , , , In the formula, is the query matrix, is the key matrix, is the value matrix, is the key status matrix, is the learnable weight matrix of the query matrix, is the learnable weight matrix of the value matrix, is the learnable weight matrix of the key status matrix; Input the query matrix, the key matrix, and the value matrix into a preset Transformer model, capture global spatial dependencies through multi-head attention, enhance the expression ability of node features, and output an enhanced feature matrix, and the expression is: , , , Wherein, is the enhanced feature matrix, is layer normalization, which normalizes the input features, is the multi-layer perceptron, is the adaptive multi-head attention mechanism, is concatenation, which combines the outputs of multiple attention heads, is the output of the first attention head, is the output of the h-th attention head, is the learnable linear projection matrix, is the output of the i-th attention head, is the query matrix of the i-th head, is the key matrix of the i-th head, is the feature dimension of each attention head, is the value matrix of the i-th head, is the Softmax function, which normalizes the attention weights, is the transpose symbol.

9. A method for generating a robot navigation path according to claim 1, characterized in that, Dynamically updating the adjacency matrix according to the enhanced feature matrix to obtain a target adjacency matrix, inputting the target adjacency matrix into a process reward model, and performing a search according to an improved Monte Carlo tree search algorithm, the process reward model outputs an optimal navigation path, including: Dynamically updating the adjacency matrix according to the enhanced feature matrix to obtain a target adjacency matrix, and the expression is: , In the formula, is the target adjacency matrix, is the activation function, , are the enhanced feature matrices are the global feature vectors of node i and node j in Using an improved Monte Carlo tree search algorithm, combining a process reward model to guide the search, starting from a newly expanded node, randomly simulating until reaching a termination state, evaluating the potential value of this node, using the process reward model to evaluate the reward of each step during the simulation process, backpropagating the simulation result into the search tree, updating the statistical information of all nodes on the path, and obtaining an optimal navigation path, and the expression is: , , , , Wherein, is a dynamic policy parameter, is a policy function, is a scaling factor, is the Softmax function, which normalizes the attention weights, is an enhanced feature matrix that contains global spatial dependencies, is a learnable weight matrix, is an action value function, is a state vector, is an action vector, is a query weight matrix, is a bias term, is the upper confidence bound, a node selection strategy that balances exploration and exploitation, is the cumulative reward of node n, is the number of visits to node n, is an exploration coefficient, is the number of visits to the parent node p(n), is a latent factor weight, is the trace of the matrix, is the first latent factor matrix, is the second latent factor matrix, is a comprehensive reward function that combines immediate reward, state value, and graph structure information, is the immediate reward, is a discount factor that weighs current and future rewards, is a state value function, is the adjacency matrix weight.

10. A system for generating a navigation path of a robot, characterized in that, Including: A construction module configured to construct a dynamic relationship graph according to the potential state of the robot and the potential state of the pedestrian; An optimization module configured to optimize the feature matrix in the dynamic relationship graph according to a preset side-channel state sequence matching strategy to obtain an adjacency matrix, and input the adjacency matrix into a graph convolutional network for learning to obtain a node feature matrix; A decomposition module configured to perform non-negative matrix decomposition according to the adjacency matrix and the node feature matrix to obtain a first latent factor matrix and a second latent factor matrix; A screening module configured to fuse the first latent factor matrix, the second latent factor matrix and the node feature matrix to obtain a weighted fusion matrix, and screen the weighted fusion matrix according to a preset adaptive motion structure robust screening strategy to obtain a key state matrix; A first output module configured to linearly transform the key state matrix into a query matrix, a key matrix and a value matrix, and input them into a preset Transformer model, and the Transformer model outputs an enhanced feature matrix; A second output module configured to dynamically update the adjacency matrix according to the enhanced feature matrix to obtain a target adjacency matrix, input the target adjacency matrix into a process reward model, and perform a search according to an improved Monte Carlo tree search algorithm, and the process reward model outputs an optimal navigation path.

Citation Information

Patent Citations

  • Multi-agent path planning method and system based on improved RND3QN network

    CN119960464A