Cooperative positioning method based on hierarchical game and multi-agent near-end strategy optimization
Through the collaborative positioning method of hierarchical game and multi-agent proximal strategy optimization, the problem of insufficient wireless positioning accuracy in complex environments is solved, resource optimization allocation and collaborative positioning accuracy are improved to meet the positioning needs of different users.
Patent Information
- Application Number
- CN202510970787.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-15
AI Technical Summary
In complex environments, traditional wireless positioning technology has the problem of insufficient positioning accuracy, especially in signal-blocked areas such as indoors, urban canyons, and overpasses. Resource-constrained users are selfish and need to balance collaborative resource investment and positioning benefits.
A collaborative positioning method based on hierarchical game and multi-agent proximal strategy optimization is adopted. By building an integrated air-ground collaborative positioning network, the complementary information of base stations, users and drones is used for collaborative positioning, and the power and bandwidth resource allocation is optimized. The hierarchical game and multi-agent deep reinforcement learning are combined to achieve collaborative decision-making between users and drones.
It effectively reduces the system's positioning error, meets the positioning accuracy requirements of different users, and at the same time minimizes system energy consumption and bandwidth occupancy, improving positioning accuracy and efficiency.
Smart Images

Figure CN120703685A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless positioning, and in particular relates to a collaborative positioning method based on layered game and multi-agent proximal strategy optimization. Background Art
[0002] In the sixth-generation mobile communications (6G) era, the explosive growth of smart IoT devices has created an urgent need for high-precision, low-power, real-time positioning services. However, traditional wireless positioning technologies face significant challenges in complex scenarios: satellite positioning suffers from significant accuracy deficiencies in signal-blocked areas such as indoor locations, urban canyons, and overpasses. To overcome this bottleneck, collaborative positioning technology has emerged, achieving significant improvements in positioning performance by building a multi-device collaborative network. This technology divides positioning network nodes into base stations (known locations) and users (undetermined locations), innovatively leveraging complementary information between nodes for collaborative positioning. By exchanging multivariate measurement data such as distance, angle, and signal strength in real time, and leveraging advanced data fusion algorithms, the system can significantly reduce positioning errors. The collaborative nodes are diverse, including environmental sensors, vehicle-mounted terminals, mobile devices, drones, and satellites, forming a three-dimensional positioning network that integrates space and land.
[0003] However, complex environments present two key challenges: First, some users may experience significant positioning errors due to a lack of access to sufficient base station information. Indiscriminate collaboration can lead to error propagation. Second, resource-constrained users are selfish, necessitating a balance between collaborative resource investment and positioning benefits. In particular, drones, as highly maneuverable platforms, while their air-ground collaborative capabilities can enhance user positioning, require the design of multi-dimensional decision-making algorithms to coordinate ground and air-ground collaboration, as well as intelligently optimize resource allocation strategies for power, bandwidth, and other resources. Addressing these issues directly impacts the accuracy and fairness of collaborative positioning systems.
[0004] In terms of algorithms, in complex systems, multiple agents often need to collaborate or compete to achieve common or individual goals. Hierarchical game theory and multi-agent deep reinforcement learning are two important approaches to addressing these problems. Hierarchical game theory decomposes the decision-making process into multiple layers, each corresponding to a different policy granularity. For example, in a collaborative positioning system, a high-level game might determine a drone's node selection decision, while a low-level game optimizes the specific positioning algorithm. This hierarchical structure reduces computational complexity and adapts to the heterogeneity of different agents, ensuring that the system achieves a stable state during competition or collaboration. Multi-agent deep reinforcement learning combines deep neural networks with reinforcement learning, enabling multiple agents to autonomously learn optimal strategies in a dynamic environment. The combination of hierarchical game theory and multi-agent deep reinforcement learning can further enhance the adaptability of the agent system. For example, in a drone-assisted collaborative positioning network, a high-level game theory can optimize drone decision allocation, while reinforcement learning is used to dynamically adjust the positioning strategy, thereby improving overall accuracy and efficiency. Therefore, for collaborative positioning in air-ground integrated networks, it is necessary to consider the user's node selection and resource allocation, and adopt layered game and multi-agent proximal strategy optimization to jointly solve the collaborative positioning problems of users and the collaborative positioning problems of users and drones, so as to meet the positioning needs of different types of users and achieve improved system positioning accuracy. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a collaborative positioning method based on hierarchical game and multi-agent proximal strategy optimization. Hierarchical game is used to coordinate ground and ground-to-air positioning collaboration, and power allocation optimization is achieved through a semi-positive transformation method; ultimately, the positioning error of the system is minimized.
[0006] The technical solutions of the present invention are as follows:
[0007] A collaborative localization method based on hierarchical game and multi-agent proximal strategy optimization includes the following steps:
[0008] Step 1: In the UAV-assisted air-ground integrated collaborative positioning network area, build an air-ground integrated collaborative positioning network model consisting of three types of nodes: base stations, users, and UAVs;
[0009] Step 2: Establish the positioning signal model, user positioning error model, and the overall optimization problem of the collaborative positioning network model, and decouple the overall optimization problem into three sub-problems: user collaboration, user-UAV collaboration, and power allocation.
[0010] Step 3: Construct a collaborative positioning algorithm based on hierarchical game and multi-agent proximal strategy optimization to achieve alternating solutions to the three sub-problems and obtain the optimal collaborative positioning decision.
[0011] Furthermore, in step 1, the user is a mobile device to be located in the network and is located on the ground; the base station is deployed on the ground for positioning; the drone is a positioning auxiliary device hovering in the air with a known position and altitude, serving as a drone node in the model; the user and the base station undertake the positioning task, and the drone undertakes the positioning auxiliary task; users, base stations and users, and drones and users are all connected through wireless links.
[0012] Furthermore, the specific process of step 2 is as follows:
[0013] Step 2.1, construct positioning signal model and user positioning error model;
[0014] Step 2.2: Construct the overall optimization problem of the collaborative positioning network model;
[0015] Step 2.3: Decouple the overall optimization problem according to different optimization variables and decouple the overall optimization problem into three sub-problems: user collaboration, user-UAV collaboration, and user power allocation.
[0016] Furthermore, the specific process of step 2.1 is as follows:
[0017] Step 2.1.1, define the spatial geometric relationship of nodes in the collaborative positioning network model; let the kth user u k With the i-th base station a i The distance is The azimuth is The kth user u k With the jth user u j The distance is The azimuth is The kth user u k With the first drone v l The distance is The azimuth is The measured direction matrix from the drone contains the pitch angle information of the drone towards the user:
[0018]
[0019] in, For the kth user u k With the first drone v l The unit direction vector of ; For the kth user u k For the lth drone v l The pitch angle; T is the transpose sign;
[0020] Step 2.1.2: Define the communication and positioning signal resources carried by the nodes in the cooperative positioning network model; let the transmission power of the i-th base station be P a; kth user u k To the jth user u j The transmission power is P k,j ; The maximum transmission power of the user is P u,max ; kth user u k The total energy budget for the localization task is is the duration of the positioning task; the collaboration bandwidth between users is a fixed value β u , and the collaboration bandwidth between each user is equal; define the Boolean variable c k,j , represents the kth user u k Whether a positioning relationship is established with the jth user; the kth user u k To the lth drone v l The transmission power is The first drone v l To the kth user u k The allocated bandwidth is β l,k , the maximum transmission bandwidth of the UAV is β v,max ; Define drone positioning selection variables represents the positioning selection variable of the l-th UAV to the j-th user;
[0021] Step 2.1.3: Establish a collaborative positioning network model and a positioning signal model. The positioning signal model includes a receiving signal model, a geometric relationship model between signal propagation delay and target position, and a Fisher information matrix. The receiving signal model is:
[0022]
[0023] Among them, r j (t) is the waveform of the signal received by the jth user at time t; L j is the total number of paths; s(·) is the waveform of the transmitted signal; and Represents the jth user The complex gain and signal propagation delay of each path; j (t) is the additive Gaussian white noise of the jth user at time t;
[0024] The geometric relationship model between signal propagation delay and target position p is:
[0025]
[0026] Where c is a constant related to the channel environment; p j is the position of the jth user; For the jth user The delay deviation of each path;
[0027] The Fisher information matrix is:
[0028]
[0029] Where J(·) is the Fisher information matrix; N b is the number of multipath parameters; P t,j is the transmission power of the jth user at time t; u j is the unit direction vector defining the positioning error direction of the jth user;
[0030] Step 2.1.4: Construct a user positioning error model, including the Fisher information matrix received by the user from the base station, the equivalent Fisher information matrix received by the user from other users, and the equivalent Fisher information matrix received by the user from the drone.
[0031] The equivalent Fisher information matrix received by the user from the base station is:
[0032]
[0033] Where, J A (·) is the equivalent Fisher information matrix received by the user from the base station; N a is the total number of users; For the kth user u k Positioning information strength between the i-th base station; For the kth user u k With the i-th base station a i The plane direction angle between represents the i-th base station a i With the kth user u k The angle information between them;
[0034] The equivalent Fisher information matrix that a user receives from other users is:
[0035]
[0036] Among them, J U (·) is the equivalent Fisher information matrix received by a user from other users; K is the total number of users; For the kth user u k With the jth user u j The positioning information strength between For the kth user u k For the jth user u j Pitch angle; represents the kth user u k With the jth user u j Plane angle information of each received cooperative signal;
[0037] The equivalent Fisher information matrix received by the user from the drone is:
[0038]
[0039] Among them, J V (·) is the equivalent Fisher information matrix received by the user from the UAV; N v is the total number of drones; For the kth user u k With the first drone v l The positioning information strength between For the kth user u k With the first drone v l Plane angle information of each received cooperative signal;
[0040] The total equivalent Fisher information matrix is:
[0041] J(u k )=J A (u k )+J U (u k )+J V (u k );
[0042] Where J(·) is the total equivalent Fisher information matrix;
[0043] For the cooperative positioning network model, the squared position error bound is used as the measurement standard for positioning accuracy; the squared position error bound of the kth user is defined as:
[0044]
[0045] in, is the estimated location of the kth user; p k is the real location of the kth user; is the squared position error bound of the user position estimate; tr{·} is the trace of the matrix; is the total equivalent Fisher information matrix of the kth user;
[0046] The more location information a node obtains, the larger the eigenvalue of the user's total equivalent Fisher information matrix, and the smaller the node's squared location error. At this time, the user's squared location error bound is expressed as the trace of the inverse matrix of the user's own positioning information matrix:
[0047]
[0048] in, is the squared position error bound of the user.
[0049] Furthermore, in step 2.2, the overall optimization problem is:
[0050]
[0051] in, is the overall optimization problem; C1, C2, C3, C4, C5, C6, C7, C8, and C9 are different constraints; P total_max N is the maximum positioning signal transmission power that can be tolerated in the positioning network; c The maximum number of other users a user can connect to at the same time; The maximum number of users that can connect to the drone at the same time;
[0052] The specific process of step 2.3 is as follows:
[0053] Step 2.3.1: The user-drone collaboration sub-problem is modeled as
[0054]
[0055] in, is the kth user u at the current moment n k The equivalent Fisher information matrix received from other users; Step 2.3.2, the user collaboration sub-problem is modeled as
[0056]
[0057] in, is the kth user u at time n+1 k The equivalent Fisher information matrix received from the UAV;
[0058] Step 2.3.3: The user power allocation sub-problem is modeled as
[0059]
[0060] in, is the kth user u at time n+1 k The equivalent Fisher information matrix received from other users.
[0061] Furthermore, the specific process of step 3 is as follows:
[0062] Step 3.1: Transform and solve the user power allocation sub-problem. The specific process is: Converted to a semidefinite programming problem, In the transformation, by introducing auxiliary variables General Rewritten as:
[0063]
[0064] At the same time, requirements Using Schur's complement lemma, this condition can be expressed as a block matrix constraint:
[0065]
[0066] According to Schur's complement lemma, the feasible solutions of the two types of problems have a one-to-one correspondence with the objective values;
[0067] Step 3.2: Construct a hierarchical game model to solve the user-drone collaboration subproblem and the user collaboration subproblem.
[0068] Step 3.3: Apply the multi-agent proximal strategy optimization algorithm to optimize the positioning request decision in the user network formation game model, and form an efficient collaborative positioning network topology structure among users on the ground.
[0069] Furthermore, the specific process of step 3.2 is as follows:
[0070] Step 3.2.1. Construct the energy-sensitive UAV connection decision as the upper-level game to select the appropriate air-ground positioning link and the allocated bandwidth resources on the established link;
[0071] Step 3.2.2: Construct a user network formation game model as the lower-level game to achieve efficient reference node selection for ground collaborative positioning.
[0072] Furthermore, the specific process of step 3.2.1 is as follows:
[0073] A drone collection consisting of L drones As a set of auctioneers, the maximum transmission bandwidth of each drone is β v,max According to the bandwidth constraints of the UAV, the continuous bandwidth resources are divided into M orthogonal sub-channels to form the bandwidth block set of the lth UAV Among them, b l,M is the Mth bandwidth block of the lth UAV; the capacity of each bandwidth block is Δβ=β v,max / M; user collection As a bidder, the user submits a positioning assistance request to a drone within its communication range. The user participates in the auction by submitting its bid, positioning error, and remaining energy information to the drone. The drone collects this information and decides on the allocation of positioning services to the user based on the mechanism.
[0074] Each user's bidding strategy is formally represented as a triple:
[0075]
[0076] Among them, k is the bidding strategy of the kth user; Indicates the subset of drones requesting connection; w k is the bandwidth demand vector of the kth user; π k The bid vector for the unit service of the kth user;
[0077] The objective function of the drone as the auctioneer is:
[0078]
[0079] Among them, ω is the trade-off coefficient between positioning error and auction revenue, J(u k ) is the kth user u k The total equivalent Fisher information matrix of ; For the kth user u k Variables for positioning selection of the lth UAV; π k,l is the value of the k-th user's payment to the l-th drone;
[0080] An energy-sensitive nonlinear bidding cap model is proposed in the user bidding strategy:
[0081]
[0082] in, The maximum bid from the kth user to the lth user; is the γ-th power of the remaining current positioning energy of the k-th user; is the γth power of the remaining current positioning energy of the lth user; π max is the user's maximum bid; γ is the energy sensitivity coefficient.
[0083] Furthermore, the specific process of step 3.2.2 is as follows:
[0084] Step 3.2.2.1, the establishment of user network formation game; the user collection As a player in the game, each user can act as a target node to send a connection request, and can also act as a positioning reference node to receive and decide whether to accept the connection request.
[0085] The goal of each user is to reduce the positioning error of the system as much as possible through collaboration during the game. Under this goal, the user independently decides whether to accept the positioning connection request from the target user; for the kth user u k , whose positioning strategy is represented as a binary vector s k :
[0086] sk =(c k,1 ,c k,2 ,…,c k,k-1 ,c k,k+1 ,…,c k,K );
[0087] The restrictions are:
[0088]
[0089] Then u k Strategy space writing S k :
[0090]
[0091] Each j-th user u receiving the positioning request j As a reference node, it evaluates whether it agrees with the kth user u based on the feature function. k Establishing positioning collaboration;
[0092] Step 3.2.2.2. Design a feature function for each user in the game:
[0093]
[0094] in, is the characteristic function; u r is the rth user, serving as the reference node; represents the rth user u r The set of target users with established connections; ΔJ -1 (u k ) represents u k The decrease of positioning error; P k For the rth user u r for u k The power consumed in providing services; E r is the rth user u r The remaining energy;
[0095] Step 3.2.2.3: Execute the user network formation game process; specifically:
[0096] In the iterative process of forming the user network game, the connection status of all users is first initialized randomly or based on experience. And calculate the initial network G (0) The characteristic function value of , where is the initial Boolean variable; in each iteration, each user acts as the target user and sends a positioning connection request to all other users within its communication range; each other user who receives the connection request acts as a reference node and evaluates whether to agree to establish a collaborative relationship with the target user based on its characteristic function. If this link decision increases its own, it agrees to the current connection request; according to the decision of the reference node, the connection relationship {c k,j}; Recalculate the characteristic function values of all reference nodes; If in the current iteration, all users have not changed their connection decisions, the game is considered to have converged and the iteration is terminated; otherwise, continue to the next round of iteration.
[0097] Furthermore, the specific process of step 3.3 is as follows:
[0098] Step 3.3.1, obtain the user's current location estimate, positioning error and remaining energy value;
[0099] Step 3.3.2: Apply the multi-agent proximal strategy optimization algorithm to obtain the user's optimal request decision; the specific process is as follows:
[0100] First, the state space Defined as: User's remaining energy E k 、Current number of connections c k,j , squared position error bound Neighbor distance matrix, the neighbor distance matrix consists of the distances between all users, each element in the matrix is the kth user u k With the jth user u j distance
[0101] Then, define each user's action as:
[0102]
[0103] in, is the action of the kth user at time t, indicating whether the kth user accepts the connection request from other users; c' k,K is the connection indicator variable between the kth user and the Kth user; the actions of all users in the same iteration step constitute the joint action space
[0104] Finally, the reward is defined as a characteristic function of the network formation game, which encourages users to autonomously choose other users who can bring them higher benefits.
[0105]
[0106] in, is the reward for the kth user at time t.
[0107] The beneficial technical effects brought about by the present invention are as follows: The present invention considers an air-ground integrated collaborative positioning network scenario consisting of base stations, users and drones. In this scenario, a collaborative positioning method based on hierarchical game and multi-agent proximal strategy optimization is constructed to reasonably allocate the power and bandwidth resources of base stations and drones to users. From the perspective of reducing the average positioning error of the system, the present invention decouples the complex network scenario into simple sub-problems and introduces an energy-sensitive air-ground positioning mechanism and a network formation game model respectively. It also uses a multi-agent proximal strategy optimization algorithm to solve the coordination of air-ground decisions in the air-ground integrated collaborative positioning network, thereby meeting the positioning accuracy requirements of different users while minimizing system energy consumption and bandwidth occupancy. BRIEF DESCRIPTION OF THE DRAWINGS
[0108] Figure 1 This is a flow chart of the collaborative positioning method based on hierarchical game and multi-agent proximal strategy optimization of the present invention.
[0109] Figure 2 Schematic diagram of the air-ground integrated collaborative positioning network scenario in the present invention.
[0110] Figure 3 This is a flow chart of the collaborative positioning algorithm based on hierarchical game and multi-agent proximal strategy optimization in the present invention. DETAILED DESCRIPTION
[0111] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:
[0112] In the air-ground integrated collaborative positioning network, decisions such as positioning power allocation, bandwidth allocation, and positioning reference node selection between users and drones are coupled and influence each other, and users show imbalance in positioning accuracy and power allocation. In order to make more effective use of the power and bandwidth resources of users and drones, and to encourage users to participate in positioning tasks to improve the overall positioning accuracy of the system, a collaborative positioning algorithm framework based on hierarchical game is constructed, and a collaborative positioning algorithm based on hierarchical game and multi-agent proximal strategy optimization is proposed to obtain the optimal power bandwidth allocation and node selection strategy for different nodes in the network in real time. The present invention establishes the positioning accuracy optimization problem in the air-ground integrated collaborative positioning scenario and decomposes it, and constructs collaborative positioning schemes based on hierarchical game and multi-agent proximal strategy optimization algorithm respectively, and solves the power allocation of users through the semi-definite programming transformation method.
[0113] like Figure 1-Figure 3 As shown, the present invention includes the following steps:
[0114] Step 1: Within the drone-assisted air-ground integrated collaborative positioning network area, construct an air-ground integrated collaborative positioning network model consisting of three types of nodes: base stations, users, and drones. Users are ground-based mobile devices to be located in the network; base stations are deployed on the ground for positioning; and drones are positioning-assisting devices hovering in the air with known positions and altitudes, serving as drone nodes in the model. Users and base stations are responsible for positioning, while drones are responsible for assisting positioning. Users, base stations, and drones are all connected via wireless links, enabling mutual transmission of positioning signals, relative position measurement, and information exchange.
[0115] In the air-ground integrated collaborative positioning model based on hierarchical game and multi-agent proximal strategy optimization of the present invention, let is the set of all base stations, the i-th base station is a i , the position of the i-th base station is p i =(x i ,y i ,0),x i 、y i are the horizontal and vertical coordinates of the i-th base station respectively. is a user set, the kth user is u k , the position of the kth user is q k =(x k ,y k ,0),x k 、y k are the horizontal and vertical coordinates of the kth user respectively. Represents a set of drones, the lth drone is v l , the position of the lth UAV is s l =(x l ,y l ,z l ), where x l 、y l They represent the horizontal and vertical coordinates of the projection of the lth UAV on the ground, z l Indicates the flight altitude of the lth UAV.
[0116] The operation process of the collaborative positioning network model of the present invention is as follows: First, the user performs relative position measurement with the base station, obtains its own preliminary position estimate, and calculates its own positioning error, and the base station positioning process is completed. Secondly, the user submits a positioning assistance request to the drone. At this time, the drone selects a suitable user for auxiliary positioning, and allocates reasonable bandwidth resources for auxiliary positioning among the users who have established a connection. The user receives the positioning signal from the drone and solves its own position estimate and error, and the drone positioning process is completed. Finally, the user selects other suitable users to establish a collaborative positioning connection based on the position estimate and error solved in the first two steps, as well as the remaining energy of the mobile power supply carried by the user, and decides the power resources allocated to the established connection to minimize its own positioning error. At this point, the user collaborative positioning process is completed. The user's positioning error is greatly reduced by the relative position measurement signals received by the user from the base station, drone and other users.
[0117] Step 2: Based on the parameters in the scenario, a positioning signal model and a user positioning error model are established, as well as a multi-dimensional positioning resource optimization problem with the goal of minimizing collaborative positioning error under resource-constrained conditions. The optimization problem is then decomposed into three sub-problems: user collaboration, user-UAV collaboration, and power allocation, based on different types of node decisions. Resource constraints include user power, UAV bandwidth, and the number of connections between the two. The specific process is as follows:
[0118] Step 2.1: Build the positioning signal model and user positioning error model. The specific process is as follows:
[0119] Step 2.1.1, define the spatial geometric relationship of nodes in the collaborative positioning network model. Let the kth user u k With the i-th base station a i The distance is The azimuth is The kth user u k With the jth user u j The distance is The azimuth is The kth user u k With the first drone v l The distance is The azimuth is It is worth noting that due to the spatial deployment characteristics of drones, the range-direction matrix (RDM) from the drone contains the pitch angle information of the drone toward the user:
[0120]
[0121] in, For the kth user u kWith the first drone v l The unit direction vector of ; For the kth user u k For the lth drone v l The pitch angle; T is the transpose sign;
[0122] Step 2.1.2: Define the communication and positioning signal resources carried by the nodes in the cooperative positioning network model. Let the transmission power of the i-th base station be P a The kth user u k To the jth user u j The transmission power is P k,j (j≠k). The maximum user transmission power is P u,max ,Due to the mobility of users, each user carries a mobile battery for power supply to ensure positioning and other tasks, the kth user u k The total energy budget for the localization task is is the duration of the positioning task. The collaboration bandwidth between users is a fixed value β u , and the collaboration bandwidth between each user is equal. Define a Boolean variable c k,j , represents the kth user u k Whether a positioning relationship is established with the jth user. k To the lth drone v l The transmission power is The first drone v l To the kth user u k The allocated bandwidth is β l,k , the maximum transmission bandwidth of the UAV is β v,max . Define the drone positioning selection variable The variable representing the positioning selection of the lth UAV for the jth user is a Boolean value. When the value is 1, it means that the lth UAV establishes a positioning signal connection relationship with the jth user. When the value is 0, it means that the lth UAV does not establish a positioning connection with the jth user.
[0123] Step 2.1.3: Establish the collaborative positioning network model positioning signal model. The positioning signal model includes the receiving signal model, the geometric relationship model between the signal propagation delay and the target position, and the Fisher information matrix, as shown in formulas (2), (3), and (11). The wireless positioning network includes multiple reference nodes and multiple user nodes. Users achieve positioning by receiving broadband signals transmitted by the base station. The receiving signal model is:
[0124]
[0125] Among them, r j (t) is the waveform of the signal received by the jth user at time t; L jis the total number of paths; s(·) is the waveform of the transmitted signal; and Represents the jth user The complex gain and signal propagation delay of each path; j (t) is the additive white Gaussian noise of the jth user at time t; the geometric relationship model between the signal propagation delay and the target position p is:
[0126]
[0127] Where c is a constant related to the channel environment; p j is the position of the jth user; For the jth user The delay deviation of the path, where the delay error of the j-th user line-of-sight path is The delay error of the jth user's non-line-of-sight path The joint estimation parameters of the user position include the target position (to be estimated) p and the multipath parameters of the jth user The total parameter vector θ is:
[0128]
[0129] Among them, N b is the number of multipath parameters;
[0130] After receiving positioning signals from base stations, drones, and other users, users calculate the equivalent Fisher information matrix (EFIM) as a positioning parameter estimation method to quantify the estimation error of the positioning parameter data on their own position. Assume that the target position is The probability density function (PDF) of the observed positioning signal z is f(z|p), then the Fisher information matrix FIM is defined as the negative expectation of the Hessian matrix of the log-likelihood function:
[0131]
[0132] Where J(·) is the Fisher information matrix; For expectations;
[0133] The FIM can be divided into position parameter blocks J according to the total parameter vector θ. pp , multipath parameter block J κκ and the first cross term J pκ , the second cross term J κp :
[0134]
[0135] The equivalent Fisher information matrix J is obtained by eliminating multipath parameter interference through Schur complement eff :
[0136]
[0137] EFIM integrates the effects of multipath propagation, and the trace of its inverse matrix is the squared position error bound (SPEB) of the user position estimate:
[0138]
[0139] in, is the squared position error bound of the user position estimate; tr{·} is the trace of the matrix;
[0140] In the time of arrival (TOA) positioning model, the transmission power P at time t is t Influencing measurement error variance through signal-to-noise ratio
[0141]
[0142] Among them, SNR is the signal-to-noise ratio of the positioning signal received by the user;
[0143] At this time, FIM can be expressed as:
[0144]
[0145] Among them, P t,j is the transmission power of the jth user at time t; u j is the unit direction vector defining the positioning error direction for the jth user. Clearly, increasing transmit power linearly increases FIM, thereby reducing CRLB. FIM and its equivalent, EFIM, are used in wireless positioning networks to quantify target position estimation errors. This is achieved by jointly estimating position and network topology and optimizing the signal-to-noise ratio using transmit power to reduce estimation errors.
[0146] Step 2.1.4: Construct the user positioning error model. The user positioning error model is mainly represented by EFIM, which specifically includes the Fisher information matrix received by the user from the base station, the equivalent Fisher information matrix received by the user from other users, and the equivalent Fisher information matrix received by the user from the drone, as shown in formulas (14), (17), and (20), respectively.
[0147] In the air-ground collaborative positioning scenario of the present invention, in one time slot, a user needs to go through the following three steps to complete the positioning task: first, receive the positioning signal from the base station and establish an initial position estimate for itself; second, try to establish a connection with the drone and receive auxiliary positioning services from the drone; for some users in dense urban buildings or high mountain valleys, it is impossible to obtain a positioning signal of sufficient strength from the base station or establish a reliable connection with the drone. Relative position measurement will be used to obtain relative position information to further correct its own position estimate. The EFIM received by the user from the base station is expressed as:
[0148]
[0149] Where, J A (·) is the equivalent Fisher information matrix received by the user from the base station; N a is the total number of users; For the kth user u k Positioning information strength between the i-th base station; For the kth user u k With the i-th base station a i The plane direction angle between represents the i-th base station a i With the kth user u k The angle information between determines the direction of the change in user positioning error. The formula is:
[0150]
[0151] in, is the i-th base station a i With the kth user u k This vector makes the positioning signal directionally selective and further indicates the directional relationship between the user and the base station.
[0152] In order to describe the amount of ranging information carried by the signal from this direction, the positioning information strength is introduced, and the formula is:
[0153]
[0154] Describes the amount of information contained in the base station positioning signal and determines the gain value of the user positioning accuracy in this direction. k,i represents the communication channel gain between the kth user and the ith user, the positioning information strength provided by each base station and the transmission power P a Positively correlated with path loss Negative correlation.
[0155] The EFIM (i.e., the equivalent Fisher information matrix received by a user from other users) contained in the cooperative positioning signal between users is expressed as:
[0156]
[0157] Among them, J U (·) is the equivalent Fisher information matrix received by the user from other users, which determines the positioning accuracy gain obtained by the user in the ground cooperation part; K is the total number of users; For the kth user u k With the jth user u j The strength of the positioning information between them determines the position information gain from a certain direction obtained in user collaborative positioning; For the kth user u k For the jth user u j Pitch angle; represents the kth user u k With the jth user u j The plane angle information of each received cooperative signal determines the direction in which the user receives the positioning signal in the cooperative positioning step. The formula is:
[0158]
[0159] in, For the kth user u k With the jth user u j The unit direction vector of the collaborative signal between users is expressed as:
[0160]
[0161] Among them, ξ k,j represents the communication channel gain between the kth user and the jth user; P k,j is the collaborative power between the kth user and the jth user; c k,j is the connection variable between the kth user and the jth user; the collaboration power P k,j and concatenate variable c k,j Jointly determine the intensity of collaboration, c k,j It is a binary variable used to determine whether two users have established relative position measurement. If it is established successfully, it is 1, otherwise it is 0. Affects information gain.
[0162] The EFIM received by the user from the drone is expressed as:
[0163]
[0164] Among them, J V(·) is the equivalent Fisher information matrix received by the user from the UAV; N v is the total number of drones; For the kth user u k With the first drone v l The positioning information strength between For the kth user u k With the first drone v l Plane angle information of each received cooperative signal;
[0165] In a complete positioning cycle, the total EFIM of the positioning signals received by the user from the base station, drone and other users is expressed as the sum of the position information of each positioning signal:
[0166] J(u k )=J A (u k )+J U (u k )+J V (u k ) (twenty two);
[0167] Where J(·) is the total equivalent Fisher information matrix;
[0168] For this collaborative positioning network model, the squared position error bound is used as a measurement standard for positioning accuracy. The squared position error bound of the kth user is defined as:
[0169]
[0170] in, is the estimated location of the kth user; p k is the real location of the kth user; is the total equivalent Fisher information matrix of the kth user;
[0171] The more location information a node obtains, the larger the eigenvalue of the user's total equivalent Fisher information matrix, and the smaller the node's squared location error. In this case, the user's squared location error bound is expressed as the trace of the inverse matrix of the user's own positioning information matrix:
[0172]
[0173] in, is the squared position error bound of the user; this value varies with the changes in node selection and power and bandwidth allocation between users and between drones.
[0174] Step 2.2: Construct the overall optimization problem of the collaborative positioning network model; the details are as follows:
[0175] In the UAV-assisted air-ground integrated collaborative positioning scenario of the present invention, the optimization goal is to minimize the overall positioning error under the constraints of user transmission power and UAV transmission bandwidth. Due to the fixed deployment position and transmission power of the base station, the EFIMJ contained in the positioning signal from the base station is A (u k ) is a known quantity and is not affected by P k,j 、c k,j 、 and β l,k Impact. The collaborative transmission power P between users k,j , connect variable c k,j , the transmission power of users and drones and bandwidth allocation β l,k , influence J U (u k ) and J V (u k ). The optimization goal is to adjust P k,j 、c k,j 、 and β l,k , minimize the user's average positioning error and the positioning error. The overall optimization problem is as follows:
[0176]
[0177] in, It is an overall optimization problem; the goal of this optimization problem is to optimize the transmission power distribution between users and drones under the premise of meeting a series of power, energy, connectivity and bandwidth constraints to minimize the average positioning error of users and thus reduce the positioning error. C1, C2, C3, C4, C5, C6, C7, C8, and C9 are different constraints. Power constraints C1 and C3 ensure that the transmission power of each user and the entire system does not exceed the set upper limit in any time period to avoid system overload or user energy exhaustion. Energy constraint C2 considers the total energy consumption of the user to ensure that the user's energy usage is within the allowable range during the entire mission, which is suitable for battery-powered mobile devices. Bandwidth constraint C6 reasonably allocates spectrum resources to avoid excessive use or conflict of bandwidth resources, thereby ensuring the quality of positioning signals. Taking into account the limitations of users' ability to participate in collaborative positioning, C7, C8, and C9 control the number of nodes each user can connect to at the same time. P total_max N is the maximum positioning signal transmission power that can be tolerated in the positioning network; c The maximum number of other users a user can connect to at the same time; The maximum number of users that the drone can connect to at the same time.
[0178] Step 2.3: Decouple the overall optimization problem according to different optimization variables and decouple the overall optimization problem into three sub-problems: user collaboration, user-UAV collaboration, and user power allocation.
[0179] The air-ground hybrid collaborative positioning optimization problem of the present invention is essentially a mixed integer nonlinear programming problem, and there is an obvious coupling and joint optimization requirement between the power and bandwidth variables. Specifically, the connection selection and bandwidth allocation decision of the UAV and the connection selection and power allocation decision of the user affect each other, forming a complex optimization relationship. In order to improve the solution efficiency and conform to the system workflow, the original problem is decoupled into three interrelated sub-problems. By fixing the remaining variables at each stage of the system operation and dynamically optimizing the current key variables, a progressive exploration of the entire solution space is achieved. The specific decomposition scheme is as follows:
[0180] Step 2.3.1, construct the user-drone collaboration sub-problem;
[0181] When the user tries to connect to the drone, the drone’s decision is modeled as the air-ground link connection selection and bandwidth allocation sub-problem. Under the condition of fixing other variables, the user-drone collaboration sub-problem is modeled as
[0182]
[0183] in, is the kth user u at the current moment n k The equivalent Fisher information matrix received from other users, that is, the user based on the current ground connection status and power The amount of position information obtained in ground collaborative positioning is fixed in the current sub-problem. l,k and There is a nonlinear relationship between the two variables This problem is a mixed integer programming problem, and there are and Due to the nonlinear coupling between them, this problem belongs to the mixed integer nonlinear programming problem.
[0184] Step 2.3.2: Construct the user collaboration sub-problem.
[0185] In the stage of user collaborative connection selection, the user's connection decision is modeled as a user collaborative sub-problem. After fixing the air-ground link parameters, the user collaborative sub-problem is modeled as
[0186]
[0187] in, is the kth user u at time n+1 k The equivalent Fisher information matrix received from the drone is the updated drone-related EFIM. Since it only contains Boolean variables selected by the user connection, this problem belongs to integer programming. For K users, each user can choose at most N c The search space increases exponentially with the number of users, which is an NP-Hard problem.
[0188] Step 2.3.3: Construct the user power allocation sub-problem.
[0189] Finally, the collaborative positioning weight allocation process among users is modeled as a user power allocation sub-problem. After the fixed connection state and UAV bandwidth allocation, the user power allocation sub-problem is modeled as
[0190]
[0191] It can be proved that this problem is a convex optimization problem. is the kth user u at time n+1 k The equivalent Fisher information matrix received from other users, that is, the EFIM received by the updated user from other users.
[0192] Step 3: Use hierarchical game and multi-agent proximal strategy optimization to obtain the optimal collaborative positioning solution between different types of agents. That is, construct a collaborative positioning algorithm based on hierarchical game and multi-agent proximal strategy optimization to achieve alternating solutions to the three sub-problems and obtain the optimal collaborative positioning decision. The specific process is as follows:
[0193] Step 3.1: Transform and solve the user power allocation sub-problem. The specific process is as follows:
[0194] question Although it is a convex problem, its form contains a lot of matrix inversion operations, which increases the difficulty and complexity of its solution and makes the solution process unstable. In order to solve this problem more efficiently, Convert it to a semidefinite programming problem.
[0195] exist In the transformation, by introducing auxiliary variables General Rewritten as:
[0196]
[0197] At the same time, requirements Using Schur's complement lemma, this condition can be equivalently expressed as a block matrix constraint:
[0198]
[0199] According to Schur's complement lemma, the feasible solutions of the two types of problems have a one-to-one correspondence with the objective values.
[0200] Step 3.2: Construct a hierarchical game model to solve the user-drone collaboration subproblem and the user collaboration subproblem. The specific process is as follows:
[0201] Step 3.2.1: Construct energy-sensitive UAV connection decision as the upper-level game to select the appropriate air-ground positioning link and the allocated bandwidth resources on the established link. Specifically:
[0202] A drone collection consisting of L drones As a set of auctioneers, the maximum positioning bandwidth limit that each drone can provide (i.e., the maximum transmission bandwidth of each drone) is β v,max According to the UAV bandwidth constraint, the continuous bandwidth resources are divided into M orthogonal sub-channels to form the bandwidth block set of the lth UAV Among them, b l,M is the Mth bandwidth block of the lth UAV; the capacity of each bandwidth block is Δβ=β v,max / M. User collection As a bidder, the user submits a positioning assistance request to a drone within its communication range and participates in the auction process by submitting its own bid and information such as positioning error and remaining energy to the drone. The drone collects this information and decides the allocation of positioning services to the user based on a certain mechanism.
[0203] Each user's bidding strategy can be formally represented as a triple:
[0204]
[0205] Among them, k is the bidding strategy of the kth user; Indicates the subset of drones requesting a connection; is the bandwidth demand vector of the kth user, w k,l represents the bandwidth request of the kth user to the lth UAV; The bid vector for the unit service of the kth user, π k,l is the price paid by the kth user to the lth drone. This model satisfies the Boolean connection variable and bandwidth allocation variable β l,k The corresponding relationship, that is If and only if And β l,k =w k,l . Boolean connection variables express.
[0206] The objective function of the drone as the auctioneer must consider both the improvement of positioning accuracy and the fairness of positioning service distribution, so as to minimize the user's positioning error and its own cost consumption:
[0207]
[0208] Among them, ω is the trade-off coefficient between positioning error and auction revenue, J(u k ) is the kth user u k The total equivalent Fisher information matrix of ; Select variables for drone positioning, specifically the kth user u k Select variables for the positioning of the lth UAV.
[0209] The design of user bidding strategy reflects the impact of user residual energy on the competitiveness of drone positioning services. To this end, an energy-sensitive nonlinear bidding upper limit model is proposed:
[0210]
[0211] in, The maximum bid from the kth user to the lth user; is the γ-th power of the remaining current positioning energy of the k-th user; is the γth power of the remaining current positioning energy of the lth user; π max is the user's maximum bid; γ is the energy sensitivity coefficient. When the energy sensitivity coefficient γ is greater than 1, the model grants high-energy users an exponentially increasing bidding advantage while preserving the basic positioning capabilities of low-energy residual users through normalization, achieving an effective balance between positioning accuracy and fairness. By introducing a positive correlation between the bid cap and the user's residual energy, drones can provide positioning services to users fairly. High-energy residual users have an advantage in the competition for drone positioning services, and this mechanism encourages them to participate in drone positioning assistance while protecting low-energy users. The drone and user positioning process is shown in Table 1.
[0212] Table 1 Pseudo code table of energy-sensitive air-ground assisted positioning mechanism
[0213]
[0214] Step 3.2.2: Construct a user network formation game model as the lower-level game to achieve efficient reference node selection for ground collaborative positioning. Specifically:
[0215] Step 3.2.2.1: Establishment of user network formation game.
[0216] Collect users As a player in the game, each user can act as a target node to send a connection request, or as a positioning reference node to receive and decide whether to accept the connection request.
[0217] The goal of each user is to reduce the system positioning error as much as possible through collaboration during the game. Under this goal, the user independently decides whether to accept the positioning connection request from the target user. k , whose positioning strategy can be expressed as a binary vector s k :
[0218] s k =(c k,1 ,c k,2 ,…,c k,k-1 ,c k,k+1 ,…,c k,K ) (34);
[0219] Among them, the Boolean variable c k,j ∈{0,1},j=1,2,...K represents the kth user u k Whether a positioning relationship is established with the jth user, that is, u k Whether to accept the positioning request from the jth user. Subject to the original problem constraints:
[0220]
[0221] That is u k The number of positioning services provided at the same time cannot exceed N c , therefore, u k The strategy space can be written as S k :
[0222]
[0223] Each j-th user u receiving the positioning request j As a reference node, it evaluates whether it agrees with the kth user u based on its characteristic function k (Target user) Establish positioning collaboration.
[0224] If the kth user agrees to establish a positioning collaboration with the jth user, which can improve its own feature function, then the positioning collaboration is established, that is, c k,j =1; if the characteristic function cannot be improved, the positioning connection request from the jth user is rejected, i.e. c k,j =0.
[0225] Step 3.2.2.2. To motivate collaboration on the ground, design a characteristic function for each user in the game:
[0226]
[0227] in, is the characteristic function; u r is the rth user, serving as the reference node; represents the rth user u r The set of target users with established connections, ΔJ -1 (u k ) represents u k The positioning error decreases. k For the rth user u r for u k The power consumed in providing the service. E r is the rth user u r of the remaining energy.
[0228] This feature function encourages user collaboration through the following mechanisms:
[0229] 1) Revenue and contribution are positively correlated: the rth user u r The benefits of the positioning service are proportional to the total error reduction of the collaborative user group, which motivates it to actively provide positioning services to other users with relatively large errors.
[0230] 2) Energy consumption penalty term: the power consumption term ∑P in the denominator k This requires users to weigh benefits and energy consumption when collaborating to avoid excessive energy consumption.
[0231] 3) Residual energy adjustment: Residual energy E r The higher the value, the higher the user's ability to provide positioning for other users and the higher the benefits. This makes users with sufficient energy more willing to participate in the ground collaborative positioning process, thereby reducing the positioning error of surrounding users.
[0232] This characteristic function constructs a mechanism to promote user collaboration in a competitive manner by dynamically balancing positioning accuracy, energy consumption cost and energy status, and encourages users to gradually reduce the system's positioning error.
[0233] Step 3.2.2.3: Execute the user network formation game process.
[0234] Through multiple rounds of game processes, users eventually establish a collaborative network. In this network, users obtain stable reference node connections and efficient positioning weight distribution, so that the average positioning accuracy of the system in this network reaches the optimal level.
[0235] In the iterative process of forming the user network game of the present invention, the connection status of all users is first initialized randomly or based on experience. And calculate the initial network G (0) The characteristic function value of , where is the initial Boolean variable. In each iteration, each user acts as the target user and sends a positioning connection request to all other users within its communication range. Each other user who receives the connection request acts as a reference node and evaluates whether to agree to establish a collaborative relationship with the target user based on its characteristic function. If this link decision increases its own, it agrees to the current connection request. According to the decision of the reference node, the connection relationship {c k,j}. The characteristic function values of all reference nodes are recalculated. If no user changes their connection decision in the current iteration, the game is considered to have converged and the iteration is terminated; otherwise, the next iteration is continued. The game process is shown in Table 2.
[0236] Table 2 Pseudo code of user collaborative positioning mechanism based on user network formation game model
[0237]
[0238] Step 3.3: Apply the multi-agent proximal strategy optimization algorithm to optimize the positioning request decision in the user network formation game model, and form an efficient collaborative positioning network topology between users on the ground. The specific process is as follows:
[0239] Step 3.3.1. Obtain the user's current location estimate, positioning error, and remaining energy value.
[0240] Step 3.3.2: Apply the multi-agent proximal strategy optimization algorithm to obtain the user's optimal request decision; the specific process is as follows:
[0241] First, the state space Defined as: User's remaining energy E k 、Current number of connections c k,j , positioning square position error bound Neighbor distance matrix, the neighbor distance matrix consists of the distances between all users, each element in the matrix is the kth user u k With the jth user u j distance
[0242] Then, define each user's action as:
[0243]
[0244] in, is the action of the kth user at time t, indicating whether the kth user accepts the connection request from other users; c' k,K is the connection indicator variable between the kth user and the Kth user; the actions of all agents (users) in the same iteration step constitute the joint action space
[0245] Finally, the reward is defined as the characteristic function of the network formation game, which encourages users to autonomously choose other users who can bring higher benefits to themselves;
[0246]
[0247] in, is the reward for the kth user at time t;
[0248] The pseudo code of the collaborative positioning method based on hierarchical game and multi-agent proximal strategy optimization is shown in Table 3:
[0249] Table 3 Pseudo code of collaborative positioning method based on hierarchical game and multi-agent proximal strategy optimization
[0250]
[0251]
[0252] Among them, s is the state of the environment in the current reinforcement learning interaction round; a is the action; r is the reward; s' is the next state the environment enters after receiving the action.
[0253] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by technicians in this technical field within the essential scope of the present invention should also fall within the scope of protection of the present invention.
Claims
1. A collaborative localization method based on hierarchical game and multi-agent proximal strategy optimization, characterized in that: The steps include: Step 1: In the UAV-assisted air-ground integrated collaborative positioning network area, build an air-ground integrated collaborative positioning network model consisting of three types of nodes: base stations, users, and UAVs; Step 2: Establish the positioning signal model, user positioning error model, and the overall optimization problem of the collaborative positioning network model, and decouple the overall optimization problem into three sub-problems: user collaboration, user-UAV collaboration, and power allocation. Step 3: Construct a collaborative positioning algorithm based on hierarchical game and multi-agent proximal strategy optimization to achieve alternating solutions to the three sub-problems and obtain the optimal collaborative positioning decision.
2. The collaborative positioning method based on hierarchical game and multi-agent proximal strategy optimization according to claim 1 is characterized in that: In step 1, the user is a mobile device to be located in the network and is located on the ground; the base station is deployed on the ground for positioning; the drone is a positioning auxiliary device that hovers in the air with a known position and altitude, and serves as a drone node in the model; the user and the base station undertake the positioning task, and the drone undertakes the positioning auxiliary task; users, base stations and users, and drones and users are all connected through wireless links.
3. The collaborative positioning method based on hierarchical game and multi-agent proximal strategy optimization according to claim 1 is characterized in that: The specific process of step 2 is: Step 2.1, construct positioning signal model and user positioning error model; Step 2.2: Construct the overall optimization problem of the collaborative positioning network model; Step 2.3: Decouple the overall optimization problem according to different optimization variables and decouple the overall optimization problem into three sub-problems: user collaboration, user-UAV collaboration, and user power allocation.
4. The collaborative positioning method based on hierarchical game and multi-agent proximal strategy optimization according to claim 3 is characterized in that: The specific process of step 2.1 is as follows: Step 2.1.1, define the spatial geometric relationship of nodes in the collaborative positioning network model; let the kth user u k With the i-th base station a i The distance is The azimuth is The kth user u k With the jth user u j The distance is The azimuth is The kth user u k With the first drone v l The distance is The azimuth is The measured direction matrix from the drone contains the pitch angle information of the drone towards the user: in, For the kth user u k With the first drone v l The unit direction vector of ; For the kth user u k For the lth drone v l The pitch angle; T is the transpose sign; Step 2.1.2: Define the communication and positioning signal resources carried by the nodes in the cooperative positioning network model; let the transmission power of the i-th base station be P a ; kth user u k To the jth user u j The transmission power is P k,j ; The maximum transmission power of the user is P u,max ; kth user u k The total energy budget for the localization task is is the duration of the positioning task; the collaboration bandwidth between users is a fixed value β u , and the collaboration bandwidth between each user is equal; define the Boolean variable c k,j , represents the kth user u k Whether a positioning relationship is established with the jth user; the kth user u k To the lth drone v l The transmission power is The first drone v l To the kth user u k The allocated bandwidth is β l,k , the maximum transmission bandwidth of the UAV is β v,max ; Define drone positioning selection variables represents the positioning selection variable of the l-th UAV to the j-th user; Step 2.1.3: Establish a collaborative positioning network model and a positioning signal model. The positioning signal model includes a receiving signal model, a geometric relationship model between signal propagation delay and target position, and a Fisher information matrix. The receiving signal model is: Among them, r j (t) is the waveform of the signal received by the jth user at time t; L j is the total number of paths; s(·) is the waveform of the transmitted signal; and Represents the jth user The complex gain and signal propagation delay of each path; j (t) is the additive Gaussian white noise of the jth user at time t; The geometric relationship model between signal propagation delay and target position p is: Where c is a constant related to the channel environment; p j is the position of the jth user; For the jth user The delay deviation of each path; The Fisher information matrix is: Where J(·) is the Fisher information matrix; N b is the number of multipath parameters; P t,j is the transmission power of the jth user at time t; u j is the unit direction vector defining the positioning error direction of the jth user; Step 2.1.4: Construct a user positioning error model, including the Fisher information matrix received by the user from the base station, the equivalent Fisher information matrix received by the user from other users, and the equivalent Fisher information matrix received by the user from the drone. The equivalent Fisher information matrix received by the user from the base station is: Where, J A (·) is the equivalent Fisher information matrix received by the user from the base station; N a is the total number of users; For the kth user u k Positioning information strength between the i-th base station; For the kth user u k With the i-th base station a i The plane direction angle between represents the i-th base station a i With the kth user u k Angle information between them; The equivalent Fisher information matrix that a user receives from other users is: Among them, J U (·) is the equivalent Fisher information matrix received by a user from other users; K is the total number of users; For the kth user u k With the jth user u j The positioning information strength between For the kth user u k For the jth user u j Pitch angle; represents the kth user u k With the jth user u j Plane angle information of each received cooperative signal; The equivalent Fisher information matrix received by the user from the drone is: Among them, J V (·) is the equivalent Fisher information matrix received by the user from the UAV; N v is the total number of drones; For the kth user u k With the first drone v l The positioning information strength between For the kth user u k With the first drone v l Plane angle information of each received cooperative signal; The total equivalent Fisher information matrix is: J(u k )=J A (you k )+J U (you k )+J V (you k ); Where J(·) is the total equivalent Fisher information matrix; For the cooperative positioning network model, the squared position error bound is used as the measurement standard for positioning accuracy; the squared position error bound of the kth user is defined as: in, is the estimated location of the kth user; p k is the real location of the kth user; is the squared position error bound of the user position estimate; tr{·} is the trace of the matrix; is the total equivalent Fisher information matrix of the kth user; The more location information a node obtains, the larger the eigenvalue of the user's total equivalent Fisher information matrix, and the smaller the node's squared location error. At this time, the user's squared location error bound is expressed as the trace of the inverse matrix of the user's own positioning information matrix: in, is the squared position error bound of the user.
5. The collaborative positioning method based on hierarchical game and multi-agent proximal strategy optimization according to claim 4 is characterized in that: In step 2.2, the overall optimization problem is: in, is the overall optimization problem; C1, C2, C3, C4, C5, C6, C7, C8, and C9 are different constraints; P total_max N is the maximum positioning signal transmission power that can be tolerated in the positioning network; c The maximum number of other users a user can connect to at the same time; The maximum number of users that can connect to the drone at the same time; The specific process of step 2.3 is as follows: Step 2.3.1: The user-drone collaboration sub-problem is modeled as in, is the kth user u at the current moment n k The equivalent Fisher information matrix received from other users; Step 2.3.2: User collaboration sub-problem is modeled as in, is the kth user u at time n+1 k The equivalent Fisher information matrix received from the UAV; Step 2.3.3: The user power allocation sub-problem is modeled as in, is the kth user u at time n+1 k The equivalent Fisher information matrix received from other users.
6. The collaborative positioning method based on hierarchical game and multi-agent proximal strategy optimization according to claim 5, characterized in that: The specific process of step 3 is: Step 3.1: Transform and solve the user power allocation sub-problem. The specific process is: Converted to a semidefinite programming problem, In the transformation, by introducing the auxiliary variable X k >0, the formula Rewritten as: Also requires X k ≥[J(u k )] -1 ; Using Schur's complement lemma, this condition can be expressed equivalently as a block matrix constraint: According to Schur's complement lemma, the feasible solutions of the two types of problems have a one-to-one correspondence with the objective values; Step 3.2: Construct a hierarchical game model to solve the user-drone collaboration subproblem and the user collaboration subproblem. Step 3.3: Apply the multi-agent proximal strategy optimization algorithm to optimize the positioning request decision in the user network formation game model, and form an efficient collaborative positioning network topology structure among users on the ground.
7. The collaborative positioning method based on hierarchical game and multi-agent proximal strategy optimization according to claim 6, characterized in that: The specific process of step 3.2 is: Step 3.2.
1. Construct the energy-sensitive UAV connection decision as the upper-level game to select the appropriate air-ground positioning link and the allocated bandwidth resources on the established link; Step 3.2.2: Construct a user network formation game model as the lower-level game to achieve efficient reference node selection for ground collaborative positioning.
8. The collaborative positioning method based on hierarchical game and multi-agent proximal strategy optimization according to claim 7, characterized in that: The specific process of step 3.2.1 is as follows: A drone collection consisting of L drones As a set of auctioneers, the maximum transmission bandwidth of each drone is β v,max According to the bandwidth constraints of the UAV, the continuous bandwidth resources are divided into M orthogonal sub-channels to form the bandwidth block set of the lth UAV Among them, b l,M is the Mth bandwidth block of the lth UAV; the capacity of each bandwidth block is Δβ=β v,max / M; user collection As a bidder, the user submits a positioning assistance request to a drone within its communication range. The user participates in the auction by submitting its bid, positioning error, and remaining energy information to the drone. The drone collects this information and decides on the allocation of positioning services to the user based on the mechanism. Each user's bidding strategy is formally represented as a triple: Among them, k is the bidding strategy of the kth user; Indicates the subset of drones requesting connection; w k is the bandwidth demand vector of the kth user; π k The bid vector for the unit service of the kth user; The objective function of the drone as the auctioneer is: Among them, ω is the trade-off coefficient between positioning error and auction revenue, J(u k ) is the kth user u k The total equivalent Fisher information matrix of ; For the kth user u k Variables for positioning selection of the lth UAV; π k,l is the value of the k-th user's payment to the l-th drone; An energy-sensitive nonlinear bidding cap model is proposed in the user bidding strategy: in, The maximum bid from the kth user to the lth user; is the γ-th power of the remaining current positioning energy of the k-th user; is the γth power of the remaining current positioning energy of the lth user; π max is the user's maximum bid; γ is the energy sensitivity coefficient.
9. The collaborative positioning method based on hierarchical game and multi-agent proximal strategy optimization according to claim 8, characterized in that: The specific process of step 3.2.2 is as follows: Step 3.2.2.1, the establishment of user network formation game; the user collection As a player in the game, each user can act as a target node to send a connection request, and can also act as a positioning reference node to receive and decide whether to accept the connection request. The goal of each user is to reduce the positioning error of the system as much as possible through collaboration during the game. Under this goal, the user independently decides whether to accept the positioning connection request from the target user; for the kth user u k , whose positioning strategy is represented as a binary vector s k : s k =(c k,1 ,c k,2 ,…,c k,k-1 ,c k,k+1 ,…,c k,K ); The restrictions are: Then u k Strategy space writing S k : Each j-th user u receiving the positioning request j As a reference node, it evaluates whether it agrees with the kth user u based on the feature function. k Establishing location collaboration; Step 3.2.2.
2. Design a feature function for each user in the game: in, is the characteristic function; u r is the rth user, serving as the reference node; represents the rth user u r The set of target users with established connections; ΔJ -1 (u k ) represents u k The decrease of positioning error; P k For the rth user u r for u k The power consumed in providing services; E r is the rth user u r The remaining energy; Step 3.2.2.3: Execute the user network formation game process; specifically: In the iterative process of forming the user network game, the connection status of all users is first initialized randomly or based on experience. And calculate the initial network G (0) The characteristic function value of is the initial Boolean variable; in each iteration, each user acts as the target user and sends a positioning connection request to all other users within its communication range; each other user who receives the connection request acts as a reference node and evaluates whether to agree to establish a collaborative relationship with the target user based on its characteristic function. If this link decision increases its own, it agrees to the current connection request; according to the decision of the reference node, the connection relationship {c k,j }; Recalculate the characteristic function values of all reference nodes; If in the current iteration, all users have not changed their connection decisions, the game is considered to have converged and the iteration is terminated; otherwise, continue to the next round of iteration.
10. The collaborative positioning method based on hierarchical game and multi-agent proximal strategy optimization according to claim 9, characterized in that: The specific process of step 3.3 is as follows: Step 3.3.1, obtain the user's current location estimate, positioning error and remaining energy value; Step 3.3.2: Apply the multi-agent proximal strategy optimization algorithm to obtain the user's optimal request decision; the specific process is as follows: First, the state space Defined as: User's remaining energy E k 、Current number of connections c k,j , squared position error bound Neighbor distance matrix, the neighbor distance matrix consists of the distances between all users, each element in the matrix is the kth user u k With the jth user u j distance Then, define each user's action as: in, is the action of the kth user at time t, indicating whether the kth user accepts the connection request from other users; c' k,K is the connection indicator variable between the kth user and the Kth user; the actions of all users in the same iteration step constitute the joint action space Finally, the reward is defined as a characteristic function of the network formation game, which encourages users to autonomously choose other users who can bring them higher benefits. in, is the reward for the kth user at time t.
Citation Information
Patent Citations
Power adaptive allocation method based on hierarchical game model
CN112543498A
Unmanned aerial vehicle cluster collaborative target selection method for heterogeneous tasks
CN114205778A
Unmanned aerial vehicle interference node power distribution method based on reconfigurable intelligent surface
CN117054978A
Processing method and apparatus, device, and storage medium
WO2023246504A1