A multi-agent cooperative positioning and navigation method and device based on quantum computing
The multi-agent cooperative localization and navigation method based on quantum computing utilizes the properties of quantum superposition and entanglement for modeling, which solves the problems of particle degradation and insufficient nonlinear processing capabilities in multi-agent systems, and achieves efficient and accurate path planning and cooperative localization.
Patent Information
- Application Number
- CN202510904139.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-07-01
AI Technical Summary
Existing positioning and navigation methods suffer from problems such as particle degeneration and impoverishment, insufficient nonlinear processing capabilities, slow policy convergence, and weak cooperative mechanisms when dealing with multi-agent cooperative positioning tasks, making it difficult to achieve efficient and accurate path planning in complex environments.
A quantum computing-based multi-agent cooperative localization and navigation method is adopted. It utilizes the properties of quantum superposition and entanglement for quantum modeling, filters target states through quantum search algorithms, and constructs a quantum network for action selection and path navigation by combining quantum encoding and factor graph models.
It improves nonlinear processing capabilities and policy search efficiency, enabling more accurate capture of complex collaborative relationships between targets and enhancing the positioning accuracy and path planning efficiency of multi-agent systems.
Smart Images

Figure CN120760715B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of quantum multi-agent system, in particular to a multi-agent cooperative positioning and navigation method and device based on quantum computing. BACKGROUND
[0002] The existing positioning and navigation methods mostly adopt design methods based on particle filtering, belief propagation and reinforcement learning. These traditional methods have many bottleneck problems that are difficult to break through in practical applications, such as strong sample dependence, example degradation and poverty, insufficient nonlinear processing capability, etc.
[0003] In the aspect of filtering technology, the existing method approximates the posterior probability density function of the target node state according to the sample observation result, which significantly improves the accuracy of inertial navigation positioning. The Bayesian estimation method based on Kalman filtering (KF) is the most widely used strategy at present. However, its applicability is limited to a certain extent by its Gaussian distribution assumption of noise and strict requirement of linear system model. Extended Kalman filtering (EKF) is used to solve nonlinear problems, but this method is mainly suitable for light nonlinear problems, because the high-order terms ignored in the model linearization process are likely to cause large truncation errors. Unscented Kalman filtering (UKF) approximates the posterior probability density function by means of deterministic sampling technology, which improves the estimation effect of the posterior probability density function of the nonlinear random model, but its application in non-Gaussian systems is still limited due to its limitation to the framework of Kalman filtering. The particle filtering (PF) method uses Monte Carlo sampling and weighted particle average, which can intuitively approximate any shape distribution, and effectively approximate the posterior distribution, and shows superiority in handling nonlinear non-Gaussian estimation problems. However, when dealing with high-dimensional problems, its computational complexity increases sharply with the increase of the number of particles, and there is a serious problem of particle degeneration.
[0004] In belief propagation, existing methods are mostly implemented in the framework of factor graph to achieve efficient propagation and fusion of information. Factor graph is a kind of directed graph structure, the node represents the variable, and the edge represents the relationship between variables (represented by factors). It can convert complex nonlinear multi-objective problems into a set of local constraint conditions, making the overall problem easier to handle. However, the computational complexity of classical factor graph and belief propagation usually increases exponentially when dealing with multi-objective cooperative localization tasks. The reason is that with the increase of the number of targets, the number of nodes and edges that need to be considered in the factor graph increases significantly, leading to rapid expansion of the size of the graph. In addition, the classical factor graph and belief propagation method usually assumes that the relationship between system states can be represented by a linear or approximately linear model. In actual multi-objective cooperative localization tasks, many cases involve strong nonlinear dynamic systems, such as mutual interference between targets, changes in target motion trajectories, and sensor errors. The classical factor graph and belief propagation method may not be able to effectively handle these complex nonlinear relationships, resulting in a decline in localization accuracy or a decrease in computational efficiency.
[0005] In reinforcement learning strategy, traditional reinforcement learning methods rely on the relatively mature Markov decision process theory framework. In simple scenarios of multi-objective cooperative path planning, agents have certain self-learning ability. They can continuously interact with the environment, try different actions, and gradually adjust their strategies based on the rewards obtained to learn a relatively reasonable path planning strategy. When facing complex environments and multi-agent interactions, the state space and action space will quickly expand, causing the dimension disaster problem, greatly increasing the learning difficulty of the algorithm, leading to very slow learning speed, and even making it difficult to converge to the optimal strategy.
[0006] Traditional reinforcement learning has defects in dealing with the mutual influence of multiple agents. Each agent usually learns only based on its own reward, easily ignoring the influence of other agents' behavior on its own reward, leading to conflicts and uncoordinated behavior among agents, and failing to achieve true cooperation.
[0007] Complex environments often have dynamic characteristics. The learning process of traditional reinforcement learning algorithms is relatively slow, and it is difficult to quickly adapt to environmental changes. Once the environment changes significantly, the previously learned strategy may no longer be applicable, and long-term learning and adjustment are needed. In a multi-agent system, communication and information sharing among agents are of great significance for cooperative path planning. However, traditional reinforcement learning algorithms usually do not fully consider the communication mechanism, making it difficult to effectively use other agent information to optimize their path planning. Even if communication is introduced, how to efficiently integrate and use this information is also a major challenge.
[0008] In the prior art, there is a lack of an efficient and accurate multi-agent cooperative localization method based on quantum computing. SUMMARY
[0009] In order to solve the technical problems of particle degradation and depletion in the conventional positioning and navigation method, limitation in high-dimensional nonlinear scene, slow strategy convergence and weak coordination mechanism in the prior art, the embodiment of the present application provides a multi-agent cooperative positioning and navigation method and device based on quantum computing. The technical solution is as follows:
[0010] In one aspect, a multi-agent cooperative positioning and navigation method based on quantum computing is provided, which is realized by a multi-agent cooperative positioning and navigation device. The method comprises:
[0011] An initial position set, a velocity space set and an observation information set of the swarm robot are obtained; based on Hilbert space, the initial position set and the velocity space set are used to perform particle quantumization using a Hadamard gate to obtain a position superposition state set;
[0012] Based on a preset weight threshold, position estimation is performed on the position superposition state set and the observation information set by using a quantum search algorithm to obtain a position estimation set of the swarm robot;
[0013] Based on the observation information set, quantum state encoding is performed using a quantum encoding circuit diagram to obtain a quantum state observation information set;
[0014] Based on a quantum variational circuit diagram, evaluation learning is performed on the quantum state observation information set using a quantum actor-critic network to obtain an action information set, an action value set and an action reward set;
[0015] Based on the quantum encoding circuit diagram, a quantum state factor graph adjacency matrix is constructed according to the velocity space set, the observation information set and the position estimation set; an experience pool is constructed according to the quantum state observation information, the action information set, the action value set, the action reward set and the quantum state factor graph adjacency matrix;
[0016] The quantum actor-critic network is parameter gradient optimized according to the data of the experience pool to obtain an optimized quantum actor-critic network;
[0017] The current position set and the current observation information set of the swarm robot are obtained; based on a preset action selection strategy, action selection is performed on the data of the experience pool, the current position set and the current observation information set using the optimized quantum actor-critic network to obtain a current action set of the swarm robot; the swarm robot performs path navigation according to the current action set.
[0018] In another aspect, a multi-agent cooperative positioning and navigation device based on quantum computing is provided, which is applied to a multi-agent cooperative positioning and navigation method based on quantum computing. The device comprises:
[0019] The particle quantumization module is configured to obtain an initial position set, a velocity space set and an observation information set of the swarm robot; based on a Hilbert space, the initial position set and the velocity space set are used to perform particle quantumization by using a Hadamard gate to obtain a position superposition state set;
[0020] The position estimation module is configured to perform position estimation by using a quantum search algorithm based on the preset weight threshold, the position superposition state set and the observation information set to obtain a position estimation set of the swarm robot.
[0021] The quantum encoding module is configured to perform quantum state encoding by using a quantum encoding circuit based on the observation information set to obtain a quantum state observation information set.
[0022] The reinforcement learning module is configured to perform evaluation learning by using a quantum actor-critic network based on a quantum variational circuit and the quantum state observation information set to obtain an action information set, an action value set and an action reward set.
[0023] The experience pool construction module is configured to construct a quantum state factor graph adjacency matrix based on the quantum encoding circuit and the velocity space set, the observation information set and the position estimation set; and construct an experience pool based on the quantum state observation information, the action information set, the action value set, the action reward set and the quantum state factor graph adjacency matrix.
[0024] The network optimization module is configured to perform parameter gradient optimization on the quantum actor-critic network based on the data of the experience pool to obtain an optimized quantum actor-critic network.
[0025] The navigation application module is configured to obtain a current position set and a current observation information set of the swarm robot; perform action selection by using the optimized quantum actor-critic network based on a preset action selection strategy, the data of the experience pool, the current position set and the current observation information set to obtain a current action set of the swarm robot; and perform path navigation of the swarm robot based on the current action set.
[0026] In another aspect, a multi-agent cooperative positioning and navigation device is provided, which includes a processor and a memory having computer readable instructions stored thereon, the computer readable instructions being executed by the processor to implement any one of the above quantum computing based multi-agent cooperative positioning and navigation methods.
[0027] In another aspect, a computer readable storage medium is provided, which stores at least one instruction, the at least one instruction being loaded and executed by a processor to implement any one of the above quantum computing based multi-agent cooperative positioning and navigation methods.
[0028] The technical scheme provided by the embodiment of the present application has at least the following beneficial effects:
[0029] The present application provides a multi-agent cooperative positioning and navigation method based on quantum computing, which uses the characteristics of quantum superposition and quantum entanglement for quantum modeling, enhances the nonlinear processing capability and strategy search efficiency, maps the velocity space into the quantum Hilbert space, realizes particle quantization, filters out the quantum state that may represent the real state of the target through the quantum search process, performs quantum transmission according to the quantum superposition theory, enhances the collapse probability of high-weight messages based on quantum unitary transformation and quantum search algorithm, maps the dependency relationship between targets into a quantum network model based on quantum entanglement and factor graph, makes it possible to fully capture the complex cooperative relationship between targets, guides the construction and evolution of quantum states in the quantum network, and improves the training performance of the quantum network and makes it easier to converge to the optimal cooperative strategy. The present application is a high-efficiency and accurate multi-agent cooperative positioning method based on quantum computing. BRIEF DESCRIPTION OF DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0031] Figure 1 is a flow chart of a multi-agent cooperative positioning and navigation method based on quantum computing provided by the embodiment of the present application;
[0032] Figure 2 is a schematic diagram of a quantum encoding circuit diagram provided by the embodiment of the present application;
[0033] Figure 3 is a schematic diagram of a quantum variation circuit diagram provided by the embodiment of the present application;
[0034] Figure 4 is a block diagram of a multi-agent cooperative positioning and navigation device based on quantum computing provided by the embodiment of the present application;
[0035] Figure 5 is a structural schematic diagram of a multi-agent cooperative positioning and navigation device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0036] The technical solutions in the present application will be described below with reference to the drawings.
[0037] In the embodiments of the present application, the words such as "example", "for example" are used to represent an example, illustration, or description. Any embodiment or design scheme described as "example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the word "example" is intended to present the concept in a specific manner. In addition, in the embodiments of the present application, the meaning expressed by "and / or" can be both, or can be one of the two.
[0038] In the embodiments of the present application, "image" and "picture" can be used interchangeably at times, and it should be pointed out that the meanings expressed are consistent when the distinction is not emphasized. "Of", "corresponding" and "corresponding" can be used interchangeably at times, and it should be pointed out that the meanings expressed are consistent when the distinction is not emphasized.
[0039] In the embodiments of the present application, sometimes the subscript such as W1 can be written in the form of non-subscript such as W1, and the meanings expressed are consistent when the distinction is not emphasized.
[0040] In order to make the technical problems, technical schemes and advantages to be solved by the present application more clear, the following will be described in detail in combination with the drawings and specific embodiments.
[0041] The embodiments of the present application provide a multi-agent cooperative positioning and navigation method based on quantum computing, which can be realized by a multi-agent cooperative positioning and navigation device. The multi-agent cooperative positioning and navigation device can be a terminal or a server. As shown in the flow chart of the multi-agent cooperative positioning and navigation method based on quantum computing, the processing flow of the method can include the following steps: Figure 1
[0042] S1, obtaining an initial position set, a velocity space set and an observation information set of a swarm robot; based on a Hilbert space, using a Hadamard gate to perform particle quantization according to the initial position set and the velocity space set, and obtaining a position superposition state set.
[0043] Optionally, based on the Hilbert space, the position superposition state set is obtained by using the Hadamard gate to perform particle quantization according to the initial position set and the velocity space set, including:
[0044] According to the velocity space set, the velocity variable of the swarm robot is discretized to obtain a finite numerical value set of the velocity variable;
[0045] Based on the Hilbert space, the Hadamard gate is used to perform equal-probability quantum state mapping according to the finite numerical value set of the velocity variable to obtain a base velocity superposition state set of the swarm robot;
[0046] According to the base velocity superposition state set and the initial position set, state transition is performed to obtain a position superposition state set of the swarm robot.
[0047] In a feasible implementation, in the classical particle filtering algorithm, the number of particles needs to be increased to cover a wider state space. For the quantum particle filtering algorithm, the classical particles are expressed in a quantumized manner before state transition, and the same particle is given the ability to simultaneously perform state transition to the next time with multiple different speeds. In order to facilitate subsequent calculation processing and theoretical analysis, the speed variable will be discretized and divided into a finite set of numerical values. Each discrete speed value represents the average speed or representative speed of a speed interval. Let the number of target speed space elements (base speed) be , wherein is the number of quantum bits required to expand the Hilbert space. Then The superposition state of the speed at time t is as follows (1):
[0048] (1);
[0049] , wherein represents The amplitude of the i-th base speed at time t, and satisfies the normalization condition .
[0050] Quantum bits are usually initialized to the ground state . In order to evolve the state to a superposition state through a unitary transformation, a special quantum logic gate, Hadamard gate (denoted as H gate), is used. The ground state or excited state of a quantum bit can be evolved to a superposition state. When the Hadamard gate acts on a single quantum bit, the ground state can be evolved to , and the excited state can be evolved to . As can be seen, when the Hadamard gate acts on the ground state or the excited state, the obtained superposition state collapses to or with a probability of 50%. The matrix form of the H gate is as follows (2):
[0051] (2);
[0052] With the help of the Hadamard gate, the corresponding tensor generation operation can be performed to evolve the ground state composite system to obtain the target speed space composite system as follows (3):
[0053] (3);
[0054] wherein, The amplitude of the corresponding base speed represents the base speed in the base speed space Each element has an equal probability The goal is At the same time, the particle filter can consider multiple state transition paths when estimating the state, which greatly enhances its state prediction ability.
[0055] S2, based on the preset weight threshold, according to the position superposition state set and the observation information set, the position estimation is carried out through the quantum search algorithm, and the position estimation set of the swarm robot is obtained.
[0056] Optionally, based on the preset weight threshold, according to the position superposition state set and the observation information set, the position estimation is carried out through the quantum search algorithm, and the position estimation set of the swarm robot is obtained, comprising:
[0057] Based on the observation information set, the position superposition state set is assigned a weight to obtain a weighted position superposition state set;
[0058] According to the weighted position superposition state set, the quantum state exceeding the preset weight threshold is marked using a reflection operator to obtain a marked superposition state set;
[0059] According to the marked superposition state set, the quantum state exceeding the preset weight threshold is amplitude enhanced using a diffusion operator to obtain an enhanced superposition state set;
[0060] According to the enhanced superposition state set, a quantum measurement node is used for projection measurement, and the expected value of the measurement result is calculated to obtain the position estimation set of the swarm robot.
[0061] In a feasible implementation, in this step, the observation of each state in the quantumized superposition state particle is compared with the true observation of the target to determine the closeness of the quantum state observation to the target observation. The higher the closeness of the quantum state observation to the target observation, the greater the weight obtained. This weight distribution is crucial because it determines the quantum state in the quantum particle superposition state that is most likely to represent the true state of the target. The weight threshold is mapped and input into the quantum circuit using auxiliary quantum bits, and the quantum state with a weight greater than the threshold will be marked. In this way, the quantum state with a higher weight in the superposition state of each quantum particle can be effectively located and extracted. Applying quantum search to the sampling process of quantum particle filtering can consciously amplify the amplitude of the quantum state with a higher weight, so that these quantum states are more likely to collapse into a classical state when observed. This enables the algorithm to more accurately sample the superposition state of the quantized particle that is beneficial to the estimation result, and thus obtain an estimation result close to the true state of the target.
[0062] Taking a particle in a swarm as an example, suppose it is in The state at time is Then, the state transition is performed through the superposition velocity as shown in equation (4):
[0063] (4);
[0064] Particles at It is in a superposition state at all times The goal is to search for states with higher weights from the superposition state with high probability and then collapse them. A weight threshold parameter is set. Using auxiliary qubits to map weight thresholds and input them into the quantum circuit, quantum states with weights greater than the threshold are labeled. Let... , Let's assume... a certain basis vector For weight greater than The quantum state element, i.e. the desired search result.
[0065] By the reflection operator and diffusion operator Unitary transformation To achieve quantum search with increased probability, where the reflection operator is... The diffusion operator is When the reflection operator is used Acting on At that time, This achieves phase reversal of the target element, aiming to mark states that meet the search conditions. On the other hand, the diffusion operator... The aim is to enhance the amplitude of phase-reversed quantum states. To expand the amplitude of states that meet the search conditions, this invention uses a unitary transformation. Application in the initial state of particles The process is as follows (5):
[0066] (5);
[0067] in, express Orthogonal vectors. initial amplitude Expand to .go through The same unitary transformation Target state The amplitude was expanded to After the particle resampling based on the quantum search algorithm, all particle superposition states are observed, and the superposition states are collapsed to high-weight classical deterministic states with high probability. The quantum particles are collapsed to high-weight classical particles for estimating the target position at the current time and participating in the algorithm iteration at the next time.
[0068] S3, according to the observation information set, using a quantum encoding circuit diagram to encode the quantum state, obtaining a quantum state observation information set.
[0069] Wherein, the quantum encoding circuit diagram is used to convert the quantum state through the quantum bit rotation gate according to the phase characteristics and radiation characteristics of the input data.
[0070] In a feasible implementation, quantum angle encoding is an efficient and commonly used strategy, especially suitable for processing high-dimensional data and continuous variables. The phase and amplitude of the quantum bit state are used to encode information. The basic idea is to encode the input classical information by means of the rotation gate in the quantum circuit, and the rotation angle of the rotation gate depends on the classical information. The first dimension
[0071] acts on the circuit of the first quantum bit. From the geometric perspective, the state of the quantum bit can be mapped to the Bloch sphere. Quantum angle encoding is actually achieved by controlling the angle of rotation around different axes (such as x, y, z axes) on the Bloch sphere (corresponding to the operation of 、 、 quantum rotation gate).
[0072] For a classical information with length , the first dimension acts on the circuit of the first quantum bit, and the process is as follows formula (6):
[0073] (6);
[0074] Wherein, is a general single quantum bit rotation gate, which can usually be gate, gate, gate or its deformation and combination. However, usually the quantum bit also contains phase information, that is, A quantum bit can not only load angle information , but also load phase information . Therefore, the present application encodes a classical data with length to The process is as follows formula (7) on the qubit:
[0075] (7);
[0076] Using this mapping relationship, classical data can be effectively expressed in quantum circuits, and subsequent calculations can be performed using the superposition and entanglement properties of quantum states.
[0077] For a classical feature data with a length of 16 bits, its quantum encoding circuit As Figure 2 shown.
[0078] S4, based on the quantum variational circuit diagram, according to the quantum state observation information set, using quantum actor-critic network for evaluation learning, obtaining action information set, action value set and action reward set.
[0079] Optionally, based on the quantum variational circuit diagram, according to the quantum state observation information set, using quantum actor-critic network for evaluation learning, obtaining action information set, action value set and action reward set, comprising:
[0080] Based on the actor quantum variational circuit diagram, according to the quantum state observation information set, using the quantum actor network to perform action probability distribution prediction, obtaining the action information set of the swarm robot;
[0081] Based on the critic quantum variational circuit diagram, according to the quantum state observation information set and the action information set, using the quantum critic network to perform value function estimation, obtaining the action value set;
[0082] Based on the preset reward function, according to the action value set and the quantum state observation information set, the action reward set is obtained.
[0083] In a feasible implementation, the quantum variational circuit diagram combines quantum circuits with classical deep neural networks, and uses quantum entanglement to reduce the operation of classical network model parameters without sacrificing accuracy.
[0084] In the actor part, each quantum actor maps the environment information to a quantum state through a state encoding module, and then uses a parameterized quantum circuit module for feature extraction and action decision. Using the characteristics of quantum superposition and entangled states, quantum parallelism is introduced in action selection and policy update, significantly improving the efficiency of policy exploration.
[0085] In the critic part, the quantum centralized critic integrates the quantum state information of multiple agents to evaluate the joint value using a deep parameterized quantum circuit module. This module uses a stacked parameterized quantum circuit structure with quantum entanglement operations to estimate the global value of multi-agent interaction. This centralized strategy helps to solve the common instability and credit allocation problems in multi-agent systems.
[0086] In addition to quantum rotation gates that act on a single qubit, there are controlled rotation gates that act on multiple qubits, which act on a qubit according to the signals of several control qubits, thereby generating quantum entanglement between these qubits. Among them, the most widely used control gate is the controlled non gate (CNOT gate), which has two input qubits, control qubit and target qubit. If the control qubit is set to 0, the target qubit remains unchanged; if the control qubit is set to 1, the target qubit is flipped.
[0087] The entanglement between qubits is generated by controlled gates, and the qubits are rotated by parameterized rotation gates. This process can be repeated in multiple layers with more parameters to improve the performance of the circuit. Both the quantum actor network and the quantum critic network use the variational circuit as shown in Figure 3 The measurement of the quantum variational circuit needs to measure the expected value of the qubit state according to the corresponding computational basis, and the process is as follows formula (8):
[0088] (8);
[0089] where, denotes the tensor product operator. denotes the output of the variational quantum circuit when the input is and the circuit parameter is , and is the set of quantum measurement bases in the variational quantum circuit.
[0090] At time , the action strategy of the th agent is made according to the given observation . This strategy is denoted as . Among them, denotes the th quantum actor model parameter.
[0091] At time , the parameterized critic model estimates the long-term return under the current strategy according to the experience in the experience pool as follows formula (9):
[0092] (9);
[0093] wherein, denotes a discount factor. denotes the length of a training episode. denotes the set of actions of all targets at time denotes a reward function.
[0094] S5, constructing a quantum state factor graph adjacency matrix according to the velocity space set, the observation information set and the position estimation set based on the quantum encoding circuit graph; and constructing an experience pool according to the quantum state observation information, the action information set, the action value set, the action reward set and the quantum state factor graph adjacency matrix.
[0095] Optionally, constructing a quantum state factor graph adjacency matrix according to the velocity space set, the observation information set and the position estimation set based on the quantum encoding circuit graph comprises:
[0096] constructing a factor graph according to the velocity space set, the observation information set and the position estimation set;
[0097] traversing the edges and nodes of the factor graph to construct a factor graph adjacency matrix;
[0098] using the quantum encoding circuit graph to encode the quantum state according to the factor graph adjacency matrix to obtain a quantum state factor graph adjacency matrix.
[0099] In an available implementation, the present application clearly expresses the dependency relationship between targets by using a factor graph. In the context of multi-target cooperative path planning, variable nodes can represent individual targets, while factor nodes are used to represent the dependency relationship between targets, which is defined by the ranging information. Specifically, each factor node connects the variable nodes related to it, reflecting the constraints or relationships between these variables. To facilitate model input and further analysis, the present application converts the factor graph into the form of an adjacency matrix, and the value of a matrix element represents whether there is a connection between the corresponding variable nodes. If there is a connection, it is 1, otherwise it is 0.
[0100] In the training phase, the quantum state observation information, the action information set, the action value set, the action reward set and the quantum state factor graph adjacency matrix of each target are collected, and these data are summarized into an experience pool. By using a centralized learning algorithm, the state and action information of all targets are comprehensively considered to optimize and train the policy network and value network of the targets. In this way, global information can be fully utilized to better learn the cooperative strategy between targets.
[0101] S6, performing parameter gradient optimization on the quantum actor-critic network according to the data of the experience pool to obtain an optimized quantum actor-critic network.
[0102] Optionally, based on the data from the experience pool, the quantum actor-critic network is optimized using parameter gradients to obtain an optimized quantum actor-critic network, including:
[0103] Based on the preset first loss function, the loss function is calculated according to the data in the experience pool to obtain the commentator network prediction loss;
[0104] Based on the loss predicted by the commentator network, the parameters of the quantum commentator network are updated using gradient descent to obtain the optimized quantum commentator network and the optimized commentator network parameters.
[0105] Based on the preset second loss function, the loss function is calculated according to the data in the experience pool and the optimized commentator network parameters to obtain the actor network prediction loss;
[0106] Based on the loss predicted by the actor network, the gradient ascent method is used to update the parameters of the quantum actor network to obtain an optimized quantum actor network.
[0107] In one feasible implementation, the present invention is based on the traditional multi-agent policy gradient loss, and alternately optimizes and updates the parameters of the quantum commenting network and the quantum actor network.
[0108] Calculate the parameter gradients of the quantum actor network and the quantum critic network: , ,in , These are the parameters of the periodically updated target quantum commentator network. and These are the parameters of the quantum actor network and the quantum critic network, respectively.
[0109] S7. Obtain the current position set and current observation information set of the cluster robots; based on the preset action selection strategy, and according to the data in the experience pool, the current position set and the current observation information set, use an optimized quantum actor-critic network to select actions and obtain the current action set of the cluster robots; based on the current action set, the cluster robots perform path navigation.
[0110] In one feasible implementation, in a multi-objective cooperative path planning task, each objective corresponds to a quantum actor model to calculate the probability of its different actions at various time points. During the execution phase, each objective operates independently, making decisions based solely on its own local observation information, without requiring real-time communication with other objectives. Each objective autonomously executes actions within its local environment according to a pre-trained strategy, achieving distributed operation of the system. This approach reduces communication overhead and improves the system's scalability and robustness.
[0111] The application provides a multi-agent cooperative positioning and navigation method based on quantum computing, which uses the characteristics of quantum superposition and quantum entanglement for quantum modeling, enhances the nonlinear processing capability and the strategy search efficiency, maps the speed space into the quantum Hilbert space, realizes particle quantization, filters out quantum states that may represent the real state of the target through a quantum search process, performs quantum transmission according to the quantum superposition theory, uses a quantum search algorithm based on quantum unitary transformation to enhance the collapse probability of high-weight messages, maps the dependency relationship between targets into a quantum network model based on quantum entanglement and a factor graph, makes it possible to comprehensively capture the complex cooperative relationship between targets, guides the construction and evolution of quantum states in the quantum network, and improves the training performance of the quantum network and makes it easier to converge to the optimal cooperative strategy. The application is a multi-agent cooperative positioning method based on quantum computing, which is efficient and accurate.
[0112] Figure 4 Fig. 1 is a block diagram of a multi-agent cooperative positioning and navigation device based on quantum computing according to an exemplary embodiment, which is used for the multi-agent cooperative positioning and navigation method based on quantum computing. Referring to Fig. 1, Figure 4 the device includes a particle quantization module 410, a position estimation module 420, a quantum encoding module 430, a reinforcement learning module 440, an experience pool construction module 450, a network optimization module 460, and a navigation application module 470. Wherein:
[0113] The particle quantization module 410 is configured to obtain an initial position set, a speed space set, and an observation information set of the swarm robots, perform particle quantization using a Hadamard gate based on the Hilbert space according to the initial position set and the speed space set, and obtain a position superposition state set.
[0114] The position estimation module 420 is configured to perform position estimation based on a preset weight threshold according to the position superposition state set and the observation information set through a quantum search algorithm, and obtain a position estimation set of the swarm robots.
[0115] The quantum encoding module 430 is configured to use a quantum encoding circuit diagram to encode quantum states according to the observation information set, and obtain a quantum state observation information set.
[0116] The reinforcement learning module 440 is configured to use a quantum actor-critic network to perform evaluation learning based on a quantum variational circuit diagram according to the quantum state observation information set, and obtain an action information set, an action value set, and an action reward set.
[0117] The experience pool construction module 450 is configured to construct a quantum state factor graph adjacency matrix according to the speed space set, the observation information set and the position estimation set based on the quantum encoding circuit diagram; and construct an experience pool according to the quantum state observation information, the action information set, the action value set, the action reward set and the quantum state factor graph adjacency matrix.
[0118] The network optimization module 460 is configured to perform parameter gradient optimization on the quantum actor-critic network according to data of the experience pool, and obtain an optimized quantum actor-critic network.
[0119] The navigation application module 470 is configured to acquire a current position set and a current observation information set of the swarm robot; perform action selection on the optimized quantum actor-critic network according to data of the experience pool, the current position set and the current observation information set based on a preset action selection strategy, and obtain a current action set of the swarm robot; and perform path navigation on the swarm robot according to the current action set.
[0120] Optionally, the particle quantization module 410 is further configured to:
[0121] discretize a speed variable of the swarm robot according to the speed space set, and obtain a finite numerical value set of the speed variable;
[0122] perform equal-probability quantum state mapping on the finite numerical value set of the speed variable based on a Hilbert space and using a Hadamard gate, and obtain a base speed superposition state set of the swarm robot;
[0123] perform state transition on the base speed superposition state set and the initial position set, and obtain a position superposition state set of the swarm robot.
[0124] Optionally, the position estimation module 420 is further configured to:
[0125] perform weight allocation on the position superposition state set based on the observation information set, and obtain a weighted position superposition state set;
[0126] perform marking on a quantum state exceeding a preset weight threshold using a reflection operator according to the weighted position superposition state set, and obtain a marked superposition state set;
[0127] perform amplitude enhancement on the quantum state exceeding the preset weight threshold using a diffusion operator according to the marked superposition state set, and obtain an enhanced superposition state set;
[0128] perform projection measurement using a quantum measurement node according to the enhanced superposition state set, and calculate an expected value of a measurement result, to obtain the position estimation set of the swarm robot.
[0129] The quantum encoding circuit diagram is used for quantum state conversion according to phase characteristics and radiation characteristics of input data through a quantum bit rotation gate.
[0130] Optionally, the reinforcement learning module 440 is further configured to:
[0131] Based on the actor quantum variational circuit diagram, action information of the swarm robot is obtained by using a quantum actor network to predict a probability distribution of an action according to the quantum state observation information set.
[0132] Based on the critic quantum variational circuit diagram, an action value set is obtained by using a quantum critic network to estimate a value function according to the quantum state observation information set and the action information set.
[0133] Based on a preset reward function, an action reward set is obtained by calculating the action value set and the quantum state observation information set.
[0134] Optionally, the experience pool construction module 450 is further configured to:
[0135] A factor graph is constructed according to the speed space set, the observation information set and the position estimation set.
[0136] A factor graph adjacency matrix is constructed by traversing edges and nodes of the factor graph.
[0137] Quantum state encoding is performed on the factor graph adjacency matrix by using a quantum encoding circuit diagram to obtain a quantum state factor graph adjacency matrix.
[0138] Optionally, the network optimization module 460 is further configured to:
[0139] Based on a preset first loss function, a critic network prediction loss is obtained by performing loss function calculation on data of the experience pool.
[0140] Based on the critic network prediction loss, the quantum critic network is updated by using a gradient descent method to obtain an optimized quantum critic network and an optimized critic network parameter.
[0141] Based on a preset second loss function, an actor network prediction loss is obtained by performing loss function calculation on data of the experience pool and the optimized critic network parameter.
[0142] Based on the actor network prediction loss, the quantum actor network is updated by using a gradient ascent method to obtain an optimized quantum actor network.
[0143] The application provides a multi-agent cooperative positioning and navigation method based on quantum computing, quantum modeling is performed by using the characteristics of quantum superposition and quantum entanglement, non-linear processing capacity and strategy search efficiency are enhanced, speed space is discretized and mapped into quantum Hilbert space, particle quantumization is realized, quantum states that may represent the real state of a target are screened out through a quantum search process, quantumization transmission is performed according to the quantum superposition theory, the collapse probability of high-weight messages is enhanced based on quantum unitary transformation and a quantum search algorithm, and the dependence relationship between targets is mapped into a quantum network model based on quantum entanglement and a factor graph, so that the complex cooperative relationship between targets can be fully captured, the construction and evolution of quantum states in the quantum network are guided, the training performance of the quantum network is improved, and the optimal cooperative strategy is more easily converged.
[0144] Figure 5 is a structural schematic diagram of a multi-agent cooperative positioning and navigation device provided by the embodiment of the application, as Figure 5 shown, the multi-agent cooperative positioning and navigation device can include the multi-agent cooperative positioning and navigation apparatus based on quantum computing shown in the above Figure 4 . Optionally, the multi-agent cooperative positioning and navigation device 510 can include the first processor 2001.
[0145] Optionally, the multi-agent cooperative positioning and navigation device 510 can further include the memory 2002 and the transceiver 2003.
[0146] The first processor 2001, the memory 2002 and the transceiver 2003 can be connected through a communication bus.
[0147] The various constituent components of the multi-agent cooperative positioning and navigation device 510 will be specifically introduced below. Figure 5
[0148] The first processor 2001 is the control center of the multi-agent cooperative positioning and navigation device 510, and can be one processor or a collective term of multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPU), can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiment of the application, such as one or more digital signal processors (DSP), or one or more field programmable gate arrays (FPGA).
[0149] Optionally, the first processor 2001 can execute various functions of the multi-agent cooperative positioning and navigation device 510 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0150] In a specific implementation, as an example, the first processor 2001 can include one or more CPUs, such as the CPU0 and the CPU1 shown in FIG. 2. Figure 5
[0151] In a specific implementation, as an example, the multi-agent cooperative positioning and navigation device 510 can also include multiple processors, such as the first processor 2001 and the second processor 2004 shown in FIG. 2. Each of these processors can be a single-CPU or a multi-CPU. The processor here can refer to one or more devices, circuit diagrams, and / or processing cores for processing data (such as computer program instructions). Figure 5
[0152] The memory 2002 is configured to store software programs for implementing the solutions of the present application, and the first processor 2001 is configured to control the execution. For specific implementation manners, refer to the above-mentioned method embodiments, which will not be repeated here.
[0153] Optionally, the memory 2002 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, an optical disk storage (including a compact disk, a laser disk, an optical disk, a digital versatile disk, a Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program codes in the form of instructions or data structures and can be accessed by a computer, but is not limited to this. The memory 2002 can be integrated with the first processor 2001 or exist independently and be coupled with the first processor 2001 through an interface circuit diagram (not shown in FIG. 2) of the multi-agent cooperative positioning and navigation device 510, and the embodiments of the present application do not make a specific limitation here. Figure 5
[0154] The transceiver 2003 is used to communicate with network devices or with terminal devices.
[0155] Alternatively, transceiver 2003 may include a receiver and a transmitter. Figure 5 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.
[0156] Optionally, the transceiver 2003 can be integrated with the first processor 2001, or it can exist independently, and its interface circuit diagram is shown in the multi-agent collaborative positioning and navigation device 510. Figure 5 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.
[0157] It should be noted that, Figure 5 The structure of the multi-agent cooperative positioning and navigation device 510 shown does not constitute a limitation on the router. Actual knowledge structure recognition devices may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0158] Furthermore, the technical effects of the multi-agent cooperative positioning and navigation device 510 can be referenced from the technical effects of the quantum computing-based multi-agent cooperative positioning and navigation method described in the above method embodiments, and will not be repeated here.
[0159] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0160] It should also be understood that the memory in the embodiments of the present application can be volatile or nonvolatile memory, or can include both volatile and nonvolatile memory. The nonvolatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically EPROM (EEPROM), or flash memory. The volatile memory can be random access memory (RAM) used as external cache. By way of example, and not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0161] The above-described embodiments can be implemented in whole or in part by software, hardware (such as a circuit diagram), firmware, or any combination thereof. When implemented by software, the above-described embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through a wired (for example, infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. containing one or more available medium collections. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid state disk.
[0162] It should be understood that the term "and / or" herein merely describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent three cases of A alone, A and B together, and B alone, where A and B can be singular or plural. In addition, the character " / " herein generally represents that the associated objects before and after are an "or" relationship, but can also represent an "and / or" relationship, which can be understood according to the context before and after.
[0163] In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0164] It should be understood that in various embodiments of the present application, the size of the sequence number of the above-described processes does not mean the order of execution, and the execution order of the processes should be determined according to their functions and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0165] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0166] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the devices, apparatuses and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0167] In several embodiments provided by the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0168] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0169] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit.
[0170] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the present application that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0171] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A multi-agent cooperative localization and navigation method based on quantum computing, characterized in that, The method includes: The initial position set, velocity space set, and observation information set of the swarm robot are obtained; based on Hilbert space, particle quantization is performed using the Hadamard gate according to the initial position set and velocity space set to obtain the set of position superposition states; Based on a preset weight threshold, and according to the set of position superposition states and the set of observation information, a quantum search algorithm is used to estimate the position of the swarm robots, resulting in a set of estimated positions, including: Based on the set of observation information, a weighted set of position superposition states is obtained by weighting the set of position superposition states. Based on the weighted position superposition state set, the reflection operator is used to mark the quantum states that exceed the preset weight threshold, and the marked superposition state set is obtained. Based on the set of superposition states after labeling, the diffusion operator is used to enhance the amplitude of quantum states that exceed a preset weight threshold, thereby obtaining an enhanced set of superposition states; Based on the enhanced superposition state set, projection measurements are performed using quantum measurement nodes, and the expected value of the measurement results is calculated to obtain the position estimation set of the swarm robot; Based on the set of observation information, quantum state encoding is performed using a quantum coding circuit diagram to obtain the set of quantum state observation information. The quantum state observation information set includes the angle information and phase information of the observation information set; Based on the quantum variational circuit diagram, and according to the quantum state observation information set, a quantum actor-critic network is used for evaluation learning to obtain the action information set, action value set, and action reward set. Based on the quantum-encoded circuit diagram, a quantum state factor graph adjacency matrix is constructed according to the velocity space set, the observation information set, and the position estimation set, including: A factor map is constructed based on the velocity space set, the observation information set, and the position estimation set; Traverse the edges and nodes of the factor graph and construct the adjacency matrix of the factor graph; Based on the factor graph adjacency matrix, quantum state encoding is performed using a quantum coding circuit diagram to obtain the quantum state factor graph adjacency matrix; An experience pool is constructed based on quantum state observation information, action information set, action value set, action reward set, and quantum state factor graph adjacency matrix; Based on the data from the experience pool, the parameters of the quantum actor-critic network are optimized using gradient optimization to obtain the optimized quantum actor-critic network. Obtain the current position set and current observation information set of the swarm robots; based on the preset action selection strategy, and according to the data in the experience pool, the current position set, and the current observation information set, use an optimized quantum actor-critic network to select actions and obtain the current action set of the swarm robots; based on the current action set, the swarm robots perform path navigation.
2. The multi-agent cooperative positioning and navigation method based on quantum computing according to claim 1, characterized in that, The process, based on Hilbert space, involves using Hadamard gates to quantize particles according to an initial set of positions and a set of velocity spaces, to obtain a set of position superposition states, including: Based on the velocity space set, the velocity variables of the swarm robot are discretized to obtain a finite set of velocity variable values; Based on Hilbert space, and according to the finite set of values of velocity variables, the Hadamard gate is used to perform equal-probability quantum state mapping to obtain the set of basic velocity superposition states of the swarm robot; The state transition is performed based on the set of superposition states of the base velocity and the set of the initial position to obtain the set of superposition states of the swarm robot.
3. The multi-agent cooperative positioning and navigation method based on quantum computing according to claim 1, characterized in that, The quantum coding circuit diagram is used to perform quantum state transformation based on the phase and radiation characteristics of the input data through a qubit rotation gate.
4. The multi-agent cooperative positioning and navigation method based on quantum computing according to claim 1, characterized in that, The method, based on quantum variational circuit diagrams and using a quantum actor-critic network for evaluation learning according to the set of quantum state observation information, obtains a set of action information, a set of action value, and a set of action reward, including: Based on the quantum variational circuit diagram of the actor, and according to the set of quantum state observation information, the quantum actor network is used to predict the probability distribution of actions and obtain the action information set of the swarm robot. Based on the commentator quantum variational circuit diagram, the value function is estimated using a quantum commentator network according to the quantum state observation information set and the action information set, to obtain the action value set; Based on a preset reward function, the action reward set is calculated according to the action value set and the quantum state observation information set.
5. The multi-agent cooperative positioning and navigation method based on quantum computing according to claim 1, characterized in that, The step of optimizing the quantum actor-critic network by performing parameter gradient optimization based on data from the experience pool to obtain an optimized quantum actor-critic network includes: Based on the preset first loss function, the loss function is calculated according to the data in the experience pool to obtain the commentator network prediction loss; Based on the loss predicted by the commentator network, the parameters of the quantum commentator network are updated using gradient descent to obtain the optimized quantum commentator network and the optimized commentator network parameters. Based on the preset second loss function, the loss function is calculated according to the data in the experience pool and the optimized commentator network parameters to obtain the actor network prediction loss; Based on the loss predicted by the actor network, the gradient ascent method is used to update the parameters of the quantum actor network to obtain an optimized quantum actor network.
6. A quantum computing-based multi-agent cooperative positioning and navigation device, wherein the quantum computing-based multi-agent cooperative positioning and navigation device is used to implement the quantum computing-based multi-agent cooperative positioning and navigation method as described in any one of claims 1-5, characterized in that, The device includes: The particle quantization module is used to obtain the initial position set, velocity space set, and observation information set of the swarm robot; based on Hilbert space, particle quantization is performed using the Hadamard gate according to the initial position set and velocity space set to obtain the set of position superposition states. The position estimation module is used to estimate the position of the cluster robot by using a quantum search algorithm based on a preset weight threshold, the set of position superposition states and the set of observation information, and to obtain the position estimation set of the cluster robot. The quantum coding module is used to encode quantum states using a quantum coding circuit diagram based on the set of observation information to obtain the set of quantum state observation information. The reinforcement learning module is used to perform evaluation learning based on quantum variational circuit diagrams and quantum state observation information set using a quantum actor-critic network to obtain action information set, action value set, and action reward set; The experience pool construction module is used to construct a quantum state factor graph adjacency matrix based on the quantum coded circuit diagram, according to the velocity space set, observation information set, and position estimation set; and to construct an experience pool based on quantum state observation information, action information set, action value set, action reward set, and quantum state factor graph adjacency matrix. The network optimization module is used to perform parameter gradient optimization on the quantum actor-critic network based on the data in the experience pool, so as to obtain an optimized quantum actor-critic network. The navigation application module is used to obtain the current position set and current observation information set of the cluster robots; based on the preset action selection strategy, and according to the data in the experience pool, the current position set and the current observation information set, the optimized quantum actor-critic network is used to select actions to obtain the current action set of the cluster robots; based on the current action set, the cluster robots perform path navigation.
7. A multi-agent cooperative positioning and navigation device, characterized in that, The multi-agent cooperative positioning and navigation device includes: processor; A memory storing computer-readable instructions that, when executed by the processor, implement the method as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1 to 5.