Collaborative positioning method based on hierarchical game and multi-agent proximal policy optimization

A collaborative localization method based on hierarchical game theory and multi-agent near-end strategy optimization solves the problem of insufficient wireless positioning accuracy in complex environments. By rationally allocating resources, it reduces system errors and optimizes energy consumption and bandwidth usage.

CN120703685BActive Publication Date: 2025-12-05SHANDONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510970787.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-12-05
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

In complex environments, traditional wireless positioning technology suffers from insufficient positioning accuracy, especially in indoor areas, urban canyons, and areas with signal obstruction such as overpasses. Furthermore, resource-constrained users tend to be selfish, necessitating a balance between collaborative resource investment and positioning benefits.

Method used

A cooperative localization method based on hierarchical game theory and multi-agent near-end strategy optimization is adopted. By constructing an integrated air-ground cooperative localization network, the power and bandwidth resources of base stations, users and UAVs are rationally allocated, and the cooperative localization algorithm is optimized to reduce system errors.

Benefits of technology

This approach reduces the average positioning error of the system by enabling collaborative positioning for different types of users, thus meeting the positioning accuracy requirements of different users while minimizing system energy consumption and bandwidth usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120703685B_ABST
    Figure CN120703685B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on hierarchical game and multi-agent near end strategy optimization's cooperative positioning method, belong to wireless positioning technical field, including the following steps: construct an air-ground integrated cooperative positioning network model by base station, user, unmanned aerial vehicle three types of nodes composition;Establish positioning signal model and user positioning error model, and the overall optimization problem of cooperative positioning network model, and the overall optimization problem is decoupled into user cooperation, user and unmanned aerial vehicle cooperation and power allocation three sub-problems;Construct cooperative positioning algorithm based on hierarchical game and multi-agent near end strategy optimization, realize the alternate solution of three sub-problems and obtain optimal cooperative positioning decision.The application adopts hierarchical game and multi-agent near end strategy optimization joint solution user's cooperative positioning and user and unmanned aerial vehicle cooperative positioning problem, to meet the positioning needs of different types of users, improve positioning accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless positioning technology, specifically relating to a cooperative positioning method based on hierarchical game theory and multi-agent proximal strategy optimization. Background Technology

[0002] In the era of 6G mobile communication, the explosive growth of smart IoT devices has created an urgent need for high-precision, low-power, real-time positioning services. However, traditional wireless positioning technologies face severe challenges in complex scenarios: satellite positioning suffers from significant accuracy deficiencies in indoor environments, urban canyons, and areas with signal obstruction such as overpasses. To overcome this bottleneck, collaborative positioning technology has emerged, achieving a leap in positioning performance by constructing a multi-device collaborative network. This technology divides positioning network nodes into base stations (known locations) and users (locations to be determined), innovatively utilizing complementary information between nodes for collaborative positioning. By exchanging multi-dimensional measurement data such as distance, angle, and signal strength in real time, and leveraging advanced data fusion algorithms, the system can significantly reduce positioning errors. The types of nodes participating in the collaboration are diverse, including environmental sensors, vehicle terminals, mobile devices, drones, and satellites, forming an integrated terrestrial and space-based three-dimensional positioning network.

[0003] However, two key challenges exist in complex environments: First, some users experience significant positioning errors due to insufficient base station information, and indiscriminate collaboration may lead to error propagation. Second, resource-constrained users exhibit selfish tendencies, necessitating a balance between collaborative resource investment and positioning benefits. In particular, while drones, as highly mobile platforms, can enhance user positioning through air-to-ground collaboration, multi-dimensional decision-making algorithms are needed to coordinate ground and air-to-ground cooperation and optimize intelligent allocation strategies for resources such as power and bandwidth. Solving these problems will directly impact the accuracy and fairness of collaborative positioning systems.

[0004] In terms of algorithms, complex systems often require multiple agents to cooperate or compete to achieve common or individual goals. Hierarchical game theory and multi-agent deep reinforcement learning are two important methods for solving such problems. Hierarchical game theory decomposes the decision-making process into multiple levels, each corresponding to a different policy granularity. For example, in a cooperative localization system, higher-level game theory might determine the node selection decision for a drone, while lower-level game theory optimizes the specific localization algorithm. This hierarchical structure reduces computational complexity and adapts to the heterogeneity of different agents, ensuring the system reaches a stable state in competition or cooperation. Multi-agent deep reinforcement learning combines deep neural networks with reinforcement learning, enabling multiple agents to autonomously learn optimal strategies in dynamic environments. The combination of hierarchical game theory and multi-agent deep reinforcement learning can further enhance the adaptability of agent systems. For example, in a drone-assisted cooperative localization network, higher-level game theory can optimize drone decision allocation, while reinforcement learning is used to dynamically adjust the localization strategy, thereby improving overall accuracy and efficiency. Therefore, collaborative positioning in an integrated air-ground network needs to consider user node selection and resource allocation. It is necessary to use hierarchical game theory and multi-agent near-end strategy optimization to jointly solve the collaborative positioning problems of users and users and UAVs, so as to meet the positioning needs of different types of users and improve the positioning accuracy of the system. Summary of the Invention

[0005] To address the aforementioned issues, this invention proposes a cooperative localization method based on hierarchical game theory and multi-agent proximal strategy optimization. This method employs hierarchical game theory to coordinate ground and air-to-ground localization collaboration, and optimizes power allocation through a semi-positive definite transformation method. Ultimately, it minimizes the system's localization error.

[0006] The technical solution of the present invention is as follows:

[0007] A cooperative localization method based on hierarchical game theory and multi-agent proximal policy optimization includes the following steps:

[0008] Step 1: In the UAV-assisted air-ground integrated cooperative positioning network area, construct an air-ground integrated cooperative positioning network model consisting of three types of nodes: base stations, users, and UAVs.

[0009] Step 2: Establish the overall optimization problem of the positioning signal model, user positioning error model, and cooperative positioning network model, and decouple the overall optimization problem into three sub-problems: user cooperation, user-UAV cooperation, and power allocation.

[0010] Step 3: Construct a cooperative localization algorithm based on hierarchical game theory and multi-agent proximal policy optimization, realize the alternating solution of the three sub-problems and obtain the optimal cooperative localization decision.

[0011] Furthermore, in step 1, the user is a mobile device to be located in the network, located on the ground; the base station is deployed on the ground for positioning; the drone is a positioning assistance device that hovers in the air with a known position and altitude, serving as a drone node in the model; the user and the base station undertake the positioning task, while the drone undertakes the positioning assistance task; users are connected to each other, base stations to users, and drones to users via wireless links.

[0012] Furthermore, the specific process of step 2 is as follows:

[0013] Step 2.1: Construct the positioning signal model and the user positioning error model;

[0014] Step 2.2: Constructing the overall optimization problem of the collaborative localization network model;

[0015] Step 2.3: Decouple the overall optimization problem for different optimization variables, and decouple the overall optimization problem into three sub-problems: user collaboration, user-drone collaboration, and user power allocation.

[0016] Furthermore, the specific process of step 2.1 is as follows:

[0017] Step 2.1.1: Define the spatial geometric relationships of nodes in the cooperative localization network model; let the k-th user u k With the i-th base station a i The distance is Azimuth angle is The kth user u k With the j-th user u j The distance is Azimuth angle is The kth user u k With the lth drone v l The distance is Azimuth angle is The measurement orientation matrix from the drone contains information about the drone's tilt angle towards the user:

[0018]

[0019] in, For the k-th user u k With the lth drone v l The unit direction vector; For the k-th user u k For the l-th drone v l The pitch angle; T is the transpose symbol;

[0020] Step 2.1.2: Define the communication and positioning signal resources carried by nodes in the cooperative positioning network model; let the transmit power of the i-th base station be P. aThe kth user u k To the j-th user u j The transmission power is P k,j The user's maximum transmit power is P. u,max The kth user u k The total energy budget for the positioning mission is The duration of the location task; the collaborative bandwidth between users is a fixed value β. u Furthermore, the collaborative bandwidth between each user is equal; define a Boolean variable c. k,j , representing the k-th user u k Whether a location relationship is established with the j-th user; the k-th user u k To the l-th drone v l The transmission power is The l-th drone v l To the k-th user u k The allocated bandwidth is β l,k The maximum transmit bandwidth of the drone is β v,max Define the UAV positioning selection variable. This represents the location selection variable for the l-th drone relative to the j-th user;

[0021] Step 2.1.3: Establish a cooperative positioning network model and a positioning signal model. The positioning signal model includes a received signal model, a geometric relationship model between signal propagation delay and target position, and a Fisher information matrix. The received signal model is as follows:

[0022]

[0023] Where, r j (t) represents the waveform of the signal received by the j-th user at time t; L j s represents the total number of paths; s(·) represents the waveform of the transmitted signal; and They represent the j-th user and the j-th... Complex gain and signal propagation delay of each path; z j (t) represents the additive white Gaussian noise of the j-th user at time t;

[0024] The geometric relationship between signal propagation delay and target position p is as follows:

[0025]

[0026] In the formula, c is a constant related to the channel environment; p j Let j be the position of the j-th user; For the j-th user The time delay deviation of the path;

[0027] Fisher's information matrix is ​​as follows:

[0028]

[0029] Where J(·) is the Fisher information matrix; N b P represents the number of multipath parameters. t,j Let u be the transmit power of the j-th user at time t; j Let the unit direction vector be defined to indicate the positioning error direction of the j-th user.

[0030] Step 2.1.4: Construct a user positioning error model, including the Fisher information matrix received by the user from the base station, the equivalent Fisher information matrix received by the user from other users, and the equivalent Fisher information matrix received by the user from the UAV.

[0031] The equivalent Fisher information matrix received by the user from the base station is:

[0032]

[0033] In the formula, J A (·) represents the equivalent Fisher information matrix received by the user from the base station; N a Total number of users; For the k-th user u k The strength of the positioning information between the i-th base station and the i-th base station; For the k-th user u k With the i-th base station a i Planar orientation angle between them; Indicates the i-th base station a i With the kth user u k Angle information between them;

[0034] The equivalent Fisher information matrix received by the user from other users is:

[0035]

[0036] Among them, J U (·) represents the equivalent Fisher information matrix received by the user from other users; K is the total number of users; For the k-th user u k With the j-th user u j The strength of the positioning information between them; For the k-th user u k For the j-th user u j The pitch angle; Represents the k-th user u k With the j-th user u j Planar angle information for each received cooperative signal;

[0037] The equivalent Fisher information matrix received by the user from the drone is:

[0038]

[0039] Among them, J V (·) represents the equivalent Fisher information matrix received by the user from the UAV; N v The total number of drones; For the k-th user u k With the lth drone v l The strength of the positioning information between them; For the k-th user u k With the lth drone v l Planar angle information for each received cooperative signal;

[0040] The total equivalent Fisher information matrix is:

[0041] J(u k ) = J A (u k )+J U (u k )+J V (u k );

[0042] Where J(·) is the total equivalent Fisher information matrix;

[0043] For the cooperative positioning network model, the squared position error bound is used as the measurement standard for positioning accuracy; the squared position error bound for the k-th user is defined as:

[0044]

[0045] in, p represents the estimated location of the k-th user. k This represents the actual location of the k-th user. tr{·} represents the squared location error bound for the user's location estimation; tr{·} is the trace of the matrix; The total equivalent Fisher information matrix for the k-th user;

[0046] The more location information a node acquires, the larger the eigenvalues ​​of the user's total equivalent Fischer information matrix, and the smaller the node's squared location error. In this case, the user's squared location error bound is represented as the trace of the inverse of the user's own location information matrix.

[0047]

[0048] in, This is the user's squared position error bound.

[0049] Furthermore, in step 2.2, the overall optimization problem is:

[0050]

[0051] in, This is a global optimization problem; C1, C2, C3, C4, C5, C6, C7, C8, and C9 are different constraints; P total_max N represents the maximum positioning signal transmission power that the positioning network can withstand. c This is the maximum number of other users a user can connect to at the same time. The maximum number of users that a drone can connect to at the same time;

[0052] The specific process of step 2.3 is as follows:

[0053] Step 2.3.1, Modeling the sub-problem of user-drone collaboration as follows:

[0054]

[0055] in, For the k-th user u at the current time n k The received equivalent Fischer information matrix from other users; Step 2.3.2, the user collaboration sub-problem is modeled as follows:

[0056]

[0057] in, For the k-th user u at time n+1 k The received equivalent Fischer information matrix from the UAV;

[0058] Step 2.3.3, Modeling the User Power Allocation Subproblem as follows:

[0059]

[0060] in, For the k-th user u at time n+1 k The received equivalent Fischer information matrix from other users.

[0061] Furthermore, the specific process of step 3 is as follows:

[0062] Step 3.1: Transform and solve the user power allocation subproblem; the specific process is as follows: Transform it into a semidefinite programming problem, in During the transformation, auxiliary variables are introduced. The formula Rewritten as:

[0063]

[0064] At the same time, it is required Using Schul's complement lemma, this condition is equivalently expressed as a block matrix constraint:

[0065]

[0066] By Schuler's lemma, there is a one-to-one correspondence between feasible solutions and objective values ​​for both types of problems;

[0067] Step 3.2: Construct a hierarchical game model to solve the user-drone cooperation subproblem and the user cooperation subproblem;

[0068] Step 3.3: Apply the multi-agent near-end policy optimization algorithm to optimize the location request decision in the user network forming game model, and form an efficient collaborative location network topology among users on the ground.

[0069] Furthermore, the specific process of step 3.2 is as follows:

[0070] Step 3.2.1: Construct an energy-sensitive UAV connectivity decision as an upper-level game to select appropriate air-to-ground positioning links and bandwidth resources allocated on existing links;

[0071] Step 3.2.2: Construct a user network to form a game model as the lower-level game, so as to achieve efficient selection of reference nodes for ground collaborative positioning.

[0072] Furthermore, the specific process of step 3.2.1 is as follows:

[0073] A drone ensemble consisting of L drones As a group of auction participants, the maximum launch bandwidth of each drone is β. v,max Based on the bandwidth constraints of the UAV, the continuous bandwidth resources are divided into M orthogonal sub-channels, forming the bandwidth block set of the l-th UAV. Among them, b l,M Let M be the bandwidth block of the l-th UAV; the capacity of each bandwidth block is Δβ = β. v,max / M; User set As a bidder, a user submits a location assistance request to a drone within their communication range. By submitting their bid, location error, and remaining energy information to the drone, the user participates in the auction process. The drone collects this information and determines the allocation of location services to the user based on the mechanism.

[0074] Each user's bidding strategy is formally represented as a triple:

[0075]

[0076] Among them, Ψ k The bidding strategy for the k-th user; Indicates a subset of drones requesting a connection; w k Let π be the bandwidth demand vector for the k-th user. k The bid vector serving the unit of the k-th user;

[0077] The objective function for the drone as the auctioneer is:

[0078]

[0079] Where ω is the trade-off coefficient between positioning error and auction revenue, J(u k ) represents the k-th user u k The total equivalent Fisher information matrix; For the k-th user u k Variables for the positioning of the l-th UAV; π k,l The value delivered by the k-th user to the l-th drone;

[0080] An energy-sensitive nonlinear bid cap model is proposed for user bidding strategies:

[0081]

[0082] in, The maximum bid from the k-th user to the l-th user; The remaining energy for the current location of the k-th user is the power of γ. γ is the remaining energy at the current location of the l-th user; π max γ represents the user's maximum bid; γ is the energy sensitivity coefficient.

[0083] Furthermore, the specific process of step 3.2.2 is as follows:

[0084] Step 3.2.2.1: Establishing a user network and game theory; setting up the user set. As players in the game, each user can both send connection requests as a target node and receive and decide whether to accept connection requests as a positioning reference node.

[0085] Each user's goal is to minimize the system's positioning error through cooperation during the game. Under this goal, each user autonomously decides whether to accept a positioning connection request from the target user. For the k-th user u... k Its positioning strategy is represented by a binary vector s k :

[0086] sk =(c k,1 ,c k,2 ,…,c k,k-1 ,c k,k+1 ,…,c k,K );

[0087] The restrictions are:

[0088]

[0089] Then u k Strategy Space Writing S k :

[0090]

[0091] For each user u that receives a location request j As a reference node, it evaluates whether to agree with the k-th user u based on the feature function. k Establish location-based collaboration;

[0092] Step 3.2.2.2: Design a feature function for each user in the game:

[0093]

[0094] in, The characteristic function; u r For the r-th user, serve as the reference node; Represents the r-th user u r The set of target users with established connections; ΔJ -1 (u k ) represents u k The decrease in positioning error; P k For the r-th user u r For u k The power consumed to provide the service; E r It is the r-th user u r The remaining energy;

[0095] Step 3.2.2.3: Execute the user network to form a game process; specifically:

[0096] In the iterative process of forming a user network game, the connection states of all users are first initialized randomly or based on experience. And calculate the initial network G (0) The characteristic function values, where, `c` is an initial Boolean variable; in each iteration, each user, as the target user, sends a location connection request to all other users within its communication range; each other user receiving the connection request acts as a reference node, evaluating whether to agree to establish a cooperative relationship with the target user based on its characteristic function. If this link decision results in its own increment, then it agrees to the current connection request; based on the decision of the reference nodes, the connection relationship `c` is updated. k,j}; Recalculate the characteristic function values ​​of all reference nodes; If none of the users have changed their connection decisions in the current iteration, the game is considered to have converged, and the iteration is terminated; otherwise, continue to the next iteration.

[0097] Furthermore, the specific process of step 3.3 is as follows:

[0098] Step 3.3.1: Obtain the user's current location estimate, positioning error, and remaining energy value;

[0099] Step 3.3.2: Apply the multi-agent proximal policy optimization algorithm to obtain the best user request decision; the specific process is as follows:

[0100] First, the state space Defined as: User's remaining energy E k Current number of connections c k,j Square position error boundary The neighbor distance matrix consists of the distances between all users, with each element representing the k-th user u. k With the j-th user u j distance

[0101] Then, define each user's action as:

[0102]

[0103] in, Let c' be the action of the k-th user at time t, indicating whether the k-th user accepts connection requests from other users; k,K The connection indicator variable is the connection between the k-th user and the k-th user; the actions of all users in the same iteration step constitute the joint action space.

[0104] Finally, the reward is defined as a characteristic function of the network forming game, thereby incentivizing users to autonomously choose other users who bring them higher returns.

[0105]

[0106] in, Let be the reward for the k-th user at time t.

[0107] The beneficial technical effects of this invention are as follows: This invention considers an integrated air-to-ground cooperative positioning network scenario consisting of base stations, users, and drones. In this scenario, a cooperative positioning method based on hierarchical game theory and multi-agent near-end policy optimization is constructed, rationally allocating the power and bandwidth resources of base stations and drones to users. From the perspective of reducing the system's average positioning error, this invention decouples the complex network scenario into simpler sub-problems and introduces an energy-sensitive air-to-ground positioning mechanism, a network formation game model, and a multi-agent near-end policy optimization algorithm to solve the coordination of air-to-ground decisions in the integrated air-to-ground cooperative positioning network. This allows for meeting the positioning accuracy requirements of different users while minimizing system energy consumption and bandwidth usage. Attached Figure Description

[0108] Figure 1 This is a flowchart of the cooperative localization method based on hierarchical game theory and multi-agent proximal strategy optimization according to the present invention.

[0109] Figure 2 This is a schematic diagram of the air-ground integrated cooperative positioning network scenario in this invention.

[0110] Figure 3 This is a flowchart of the cooperative localization algorithm based on hierarchical game theory and multi-agent proximal strategy optimization in this invention. Detailed Implementation

[0111] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:

[0112] In air-ground integrated cooperative positioning networks, decisions regarding positioning power allocation, bandwidth allocation, and reference node selection between users and UAVs are coupled and mutually influential, and users exhibit an imbalance in positioning accuracy and power allocation. To more fully utilize the power and bandwidth resources of users and UAVs and incentivize user participation in positioning tasks to improve the overall positioning accuracy of the system, a cooperative positioning algorithm framework based on hierarchical game theory is constructed. A cooperative positioning algorithm based on hierarchical game theory and multi-agent near-end policy optimization is proposed to obtain the optimal power and bandwidth allocation and node selection strategies for different nodes in the network in real time. This invention establishes and decomposes the positioning accuracy optimization problem in the air-ground integrated cooperative positioning scenario, constructs cooperative positioning schemes based on hierarchical game theory and multi-agent near-end policy optimization algorithms, and solves the user's power allocation using a semidefinite programming transformation method.

[0113] like Figures 1-3 As shown, the present invention includes the following steps:

[0114] Step 1: In the UAV-assisted air-ground integrated cooperative positioning network area, construct an air-ground integrated cooperative positioning network model consisting of three types of nodes: base stations, users, and UAVs. Users are mobile devices to be located in the network, located on the ground; base stations are deployed on the ground for positioning; UAVs are positioning assistance devices hovering in the air with known positions and altitudes, serving as UAV nodes in the model. Users and base stations undertake the positioning task, while UAVs undertake the positioning assistance task. Users are connected to each other, base stations to users, and UAVs to users via wireless links, enabling the transmission of positioning signals and the measurement of relative positions and information exchange.

[0115] In the air-ground integrated cooperative positioning model based on hierarchical game theory and multi-agent proximal strategy optimization of this invention, let Let a be the set of all base stations, and let a be the i-th base station. i The location of the i-th base station is p. i =(x i ,y i ,0), x i y i These are the horizontal and vertical coordinates of the i-th base station, respectively. Let u be the set of users. k The position of the kth user is q. k =(x k ,y k ,0), x k y k These are the x and y coordinates of the k-th user, respectively. Let v represent a set of drones, where the l-th drone is v. l The position of the l-th drone is s l =(x l ,y l ,z l ), where x l y l Let z represent the x and y coordinates of the l-th UAV's projection on the ground, respectively. l This represents the flight altitude of the l-th drone.

[0116] The cooperative positioning network model of this invention operates as follows: First, the user measures its relative position with the base station to obtain a preliminary position estimate and calculates its positioning error; the base station positioning process is complete. Second, the user submits a positioning assistance request to the drone. The drone selects a suitable user for assisted positioning and allocates reasonable bandwidth resources to users with established connections for this purpose. The user receives the positioning signal from the drone and calculates its own position estimate and error; the drone positioning process is complete. Finally, based on the position estimate and error calculated in the first two steps, and the remaining energy of its portable power supply, the user selects other suitable users to establish a cooperative positioning connection and determines the power resources allocated to the established connections to minimize its own positioning error; thus, the user cooperative positioning process is complete. By utilizing the relative position measurement signals received by the user from the base station, drone, and other users, the user's positioning error is significantly reduced.

[0117] Step 2: Based on the parameters in the scenario, establish a positioning signal model and a user positioning error model, as well as a multi-dimensional positioning resource optimization problem with the objective of minimizing cooperative positioning error under resource-constrained conditions. The optimization problem is then decomposed for different types of node decisions, decoupling it into three sub-problems: user cooperation, user-UAV cooperation, and power allocation. Resource constraints include user power, UAV bandwidth, and the number of connections between the two. The specific process is as follows:

[0118] Step 2.1: Construct the positioning signal model and the user positioning error model. The specific process is as follows:

[0119] Step 2.1.1: Define the spatial geometric relationships of nodes in the collaborative localization network model. Let the k-th user u... k With the i-th base station a i The distance is Azimuth angle is The kth user u k With the j-th user u j The distance is Azimuth angle is The kth user u k With the lth drone v l The distance is Azimuth angle is It is worth noting that, due to the spatial deployment characteristics of drones, the range-direction matrix (RDM) from the drone contains information about the drone's tilt angle towards the user:

[0120]

[0121] in, For the k-th user u kWith the lth drone v l The unit direction vector; For the k-th user u k For the l-th drone v l The pitch angle; T is the transpose symbol;

[0122] Step 2.1.2: Define the communication and positioning signal resources carried by nodes in the cooperative positioning network model. Let the transmit power of the i-th base station be P. a The kth user u k To the j-th user u j The transmission power is P k,j (j≠k). The user's maximum transmit power is P. u,max Due to user mobility, each user carries a portable battery for power supply, positioning, and other tasks. The k-th user u... k The total energy budget for the positioning mission is The duration of the location task. The collaboration bandwidth between users is a fixed value β. u Furthermore, the collaborative bandwidth is equal for each user. Define a Boolean variable c. k,j , representing the k-th user u k Whether a location relationship has been established with the j-th user. The k-th user u k To the l-th drone v l The transmission power is The l-th drone v l To the k-th user u k The allocated bandwidth is β l,k The maximum transmit bandwidth of the drone is β v,max Define the drone positioning selection variable. The variable representing the positioning selection of the l-th drone for the j-th user is a Boolean value. A value of 1 indicates that the l-th drone establishes a positioning signal connection with the j-th user, while a value of 0 indicates that the l-th drone does not establish a positioning connection with the j-th user.

[0123] Step 2.1.3: Establish a cooperative positioning network model and a positioning signal model. The positioning signal model includes a received signal model, a geometric relationship model between signal propagation delay and target position, and a Fisher information matrix, as shown in formulas (2), (3), and (11) respectively. The wireless positioning network contains multiple reference nodes and multiple user nodes. Users achieve positioning by receiving broadband signals transmitted by base stations. The received signal model is as follows:

[0124]

[0125] Where, r j (t) represents the waveform of the signal received by the j-th user at time t; L js represents the total number of paths; s(·) represents the waveform of the transmitted signal; and They represent the j-th user and the j-th... Complex gain and signal propagation delay of each path; z j (t) represents the additive white Gaussian noise of the j-th user at time t; the geometric relationship model between signal propagation delay and target position p is as follows:

[0126]

[0127] In the formula, c is a constant related to the channel environment; p j Let j be the position of the j-th user; For the j-th user The time delay deviation of each path, where the time delay error of the j-th user line-of-sight path is... Delay error of the non-line-of-sight path for the j-th user The joint estimation parameters for user locations include the target location (location to be estimated) p and the multipath parameters of the j-th user. The total parameter vector θ is:

[0128]

[0129] Where, N b This represents the number of multipath parameters.

[0130] After receiving location signals from base stations, drones, and other users, users quantify the estimation error of their own location based on the location signal data by calculating the Equivalent Fisher Information Matrix (EFIM) as a method of estimating location parameters. Let the target location be... If the probability density function (PDF) of the observed location signal z is f(z|p), then the Fisher information matrix (FIM) is defined as the negative expectation of the log-likelihood function Hessian matrix:

[0131]

[0132] Where J(·) is the Fisher information matrix; For expectations;

[0133] FIM can be divided into location parameter blocks J based on the total parameter vector θ. pp Multipath parameter block J κκ and the first cross term J pκ The second cross term J κp :

[0134]

[0135] By eliminating multipath parameter interference through Schur complement, the equivalent Fisher information matrix J is obtained. eff :

[0136]

[0137] EFIM incorporates the effects of multipath propagation, and the trace of its inverse matrix is ​​the squared position error bound (SPEB) for user position estimation.

[0138]

[0139] in, tr{·} represents the squared location error bound for the user's location estimation; tr{·} is the trace of the matrix;

[0140] In the Time of Arrival (TOA) positioning model, the transmit power P at time t is... t The influence of signal-to-noise ratio on measurement error variance

[0141]

[0142] Wherein, SNR is the signal-to-noise ratio of the positioning signal received by the user;

[0143] At this point, FIM can be represented as:

[0144]

[0145] Among them, P t,j Let u be the transmit power of the j-th user at time t; j Let FIM be the unit direction vector defining the positioning error direction for the j-th user. Clearly, increasing transmit power linearly increases FIM, thereby reducing CRLB. FIM and its equivalent, EFIM, are used in wireless positioning networks to quantify target location estimation errors by jointly estimating the location and network topology, and by optimizing the signal-to-noise ratio using transmit power to reduce estimation errors.

[0146] Step 2.1.4: Construct the user positioning error model. The user positioning error model is mainly represented by EFIM, which includes the Fisher information matrix received by the user from the base station, the equivalent Fisher information matrix received by the user from other users, and the equivalent Fisher information matrix received by the user from the UAV, as shown in formulas (14), (17) and (20) respectively.

[0147] In the air-to-ground cooperative positioning scenario of this invention, within a time slot, a user needs to go through the following three steps to complete the positioning task: First, receiving a positioning signal from the base station to establish an initial position estimate; second, attempting to establish a connection with the drone and receiving assisted positioning services from the drone; for some users located in dense urban buildings or high mountain valleys, where a sufficiently strong positioning signal cannot be obtained from the base station and a reliable connection with the drone cannot be established, relative position measurement will be used to obtain relative position information, thereby further correcting their position estimate. The EFIM received by the user from the base station is represented as:

[0148]

[0149] In the formula, J A (·) represents the equivalent Fisher information matrix received by the user from the base station; N a Total number of users; For the k-th user u k The strength of the positioning information between the i-th base station and the i-th base station; For the k-th user u k With the i-th base station a i Planar orientation angle between them; Indicates the i-th base station a i With the kth user u k The angular information between them determines the directionality of the user's positioning error change, as shown in the formula:

[0150]

[0151] in, For the i-th base station a i With the kth user u k The unit direction vector between the user and the base station gives the positioning signal directional selectivity and further represents the directional relationship between the user and the base station.

[0152] To describe the amount of ranging information carried by the signal from this direction, we introduce the positioning information strength, as shown in the formula:

[0153]

[0154] The amount of information contained in the base station positioning signal determines the gain value of the user's positioning accuracy in this direction. Where, ξ k,i This represents the communication channel gain between the k-th user and the i-th user, and the location information strength and transmission power P provided by each base station. a Positive correlation with path loss Negative correlation.

[0155] The EFIM (equivalent Fisher information matrix received by a user from other users) contained in the inter-user cooperative positioning signal is represented as follows:

[0156]

[0157] Among them, J U (·) represents the equivalent Fisher information matrix received by the user from other users, which determines the gain in positioning accuracy obtained by the user in the ground cooperation part; K is the total number of users; For the k-th user u k With the j-th user u j The strength of the positioning information between users determines the gain of the positioning information from a certain direction obtained in collaborative positioning. For the k-th user u k For the j-th user u j The pitch angle; Represents the k-th user u k With the j-th user u j The planar angle information of each received cooperative signal determines the direction of the positioning signal received by the user during the cooperative positioning step, as shown in the formula:

[0158]

[0159] in, For the k-th user u k With the j-th user u j The unit direction vector; the location information strength of the collaborative signal between users is represented as:

[0160]

[0161] Where, ξ k,j P represents the communication channel gain between the k-th user and the j-th user; k,j c represents the collaborative power between the k-th user and the j-th user. k,j The connection variable between the k-th user and the j-th user; the collaborative power P k,j and the connection variable c k,j The intensity of cooperation is jointly determined, c k,j This is a binary variable used to determine whether two users have established a relative position measurement; it returns 1 if successful and 0 otherwise. It affects information gain.

[0162] The EFIM received by the user from the drone is represented as follows:

[0163]

[0164] Among them, J V(·) represents the equivalent Fisher information matrix received by the user from the UAV; N v The total number of drones; For the k-th user u k With the lth drone v l The strength of the positioning information between them; For the k-th user u k With the lth drone v l Planar angle information for each received cooperative signal;

[0165] Within a complete positioning cycle, the total EFIM (Effective Frame Information Model) of the positioning signals received by a user from base stations, drones, and other user positioning signals is represented as the sum of the location information of each positioning signal:

[0166] J(u k ) = J A (u k )+J U (u k )+J V (u k ) (twenty two);

[0167] Where J(·) is the total equivalent Fisher information matrix;

[0168] For this cooperative positioning network model, the squared position error bound is used as the measurement standard for positioning accuracy. The squared position error bound for the k-th user is defined as:

[0169]

[0170] in, p represents the estimated location of the k-th user. k This represents the actual location of the k-th user. The total equivalent Fisher information matrix for the k-th user;

[0171] The more location information a node acquires, the larger the eigenvalues ​​of the user's total equivalent Fischer information matrix, and the smaller the node's squared location error. In this case, the user's squared location error bound can be represented as the trace of the inverse of the user's own location information matrix:

[0172]

[0173] in, This is the squared position error bound for the user; this value varies with changes in node selection and power and bandwidth allocation between users, and between drones and users.

[0174] Step 2.2: Constructing the overall optimization problem of the cooperative localization network model; details are as follows:

[0175] In the UAV-assisted air-to-ground integrated cooperative positioning scenario of this invention, the optimization objective is to minimize the overall positioning error under the constraints of user transmit power and UAV transmit bandwidth. Due to the fixed deployment location and transmit power of the base station, the positioning signal from the base station contains EFIMJ. A (u k ) is a known quantity, and is not affected by P. k,j c k,j , and β l,k Impact. Inter-user cooperative transmit power P k,j , Connect variable c k,j Transmission power of users and drones and bandwidth allocation β l,k Impact J U (u k ) and J V (u k The optimization objective is achieved by adjusting P. k,j c k,j , and β l,k The goal is to minimize the user's average positioning error boundary positioning error. The resulting overall optimization problem is as follows:

[0176]

[0177] in, This is a holistic optimization problem. The goal of this optimization problem is to optimize the transmission power allocation between users and drones while satisfying a series of power, energy, connectivity, and bandwidth constraints, in order to minimize the average positioning error of users and thus reduce positioning error. C1, C2, C3, C4, C5, C6, C7, C8, and C9 are different constraints. Power constraints C1 and C3 ensure that the transmission power of each user and the entire system does not exceed the set upper limit at any time period, avoiding system overload or user energy depletion. Energy constraint C2 considers the total energy consumption of users, ensuring that the energy use of users is within the allowable range throughout the mission, applicable to battery-powered mobile devices. Bandwidth constraint C6 reasonably allocates spectrum resources to avoid excessive use or conflict of bandwidth resources, thereby ensuring the quality of positioning signals. Considering the limitations of users' ability to participate in cooperative positioning, C7, C8, and C9 control the number of nodes that each user can connect to simultaneously. total_max N represents the maximum positioning signal transmission power that the positioning network can withstand. c This is the maximum number of other users a user can connect to at the same time. This represents the maximum number of users that a drone can connect to at any given time.

[0178] Step 2.3: Decouple the overall optimization problem for different optimization variables, and decouple the overall optimization problem into three sub-problems: user collaboration, user-drone collaboration, and user power allocation.

[0179] The air-to-ground hybrid cooperative positioning optimization problem of this invention is essentially a mixed-integer nonlinear programming problem, with significant coupling and joint optimization requirements between power and bandwidth variables. Specifically, the connection selection and bandwidth allocation decisions of the UAV interact with the connection selection and power allocation decisions of the user, forming a complex optimization relationship. To improve solution efficiency and align with the system workflow, the original problem is decoupled into three interrelated sub-problems. By fixing the remaining variables and dynamically optimizing the current key variables at each stage of system operation, a gradual exploration of the entire solution space is achieved. The specific decomposition scheme is as follows:

[0180] Step 2.3.1: Construct the sub-problem of user-drone collaboration;

[0181] During the user's attempt to connect to the drone, the drone's decision-making is modeled as a sub-problem of air-to-ground link connection selection and bandwidth allocation. With other variables fixed, the user-drone collaboration sub-problem is modeled as follows:

[0182]

[0183] in, For the k-th user u at the current time n k The received equivalent Fischer information matrix from other users, i.e., the user's information based on the current ground connection state. and power The amount of location information obtained in ground-based cooperative localization remains constant in the current subproblem. This is because the bandwidth variable β in the objective function... l,k and The nonlinear relationship exists, and there are two discrete variables. This problem is a mixed integer programming problem, and there are differences between the objective function variables. and The nonlinear coupling between them means that this problem belongs to the category of mixed-integer nonlinear programming problems.

[0184] Step 2.3.2: Construct the user collaboration sub-problem.

[0185] During the user collaboration connection selection phase, the user's connection decision is modeled as a user collaboration sub-problem. After fixing the air-to-ground link parameters, the user collaboration sub-problem is modeled as follows:

[0186]

[0187] in, For the k-th user u at time n+1 k The received equivalent Fischer information matrix from the drone is the updated drone-related EFIM. Since it only contains Boolean variables representing user connection choices, this problem is an integer programming problem. For K users, each user can choose at most N... c The search space increases exponentially with the number of users, making it an NP-hard problem.

[0188] Step 2.3.3: Construct the user power allocation subproblem.

[0189] Finally, the collaborative positioning weight allocation process among users is modeled as a user power allocation subproblem. After fixed connectivity and UAV bandwidth allocation, the user power allocation subproblem is modeled as follows:

[0190]

[0191] It can be proven that this problem is a convex optimization problem. Among them, For the k-th user u at time n+1 k The received equivalent Fischer information matrix from other users, i.e., the updated user's received EFIM from other users.

[0192] Step 3: Employ hierarchical game theory and multi-agent proximal strategy optimization to obtain the optimal cooperative localization scheme among different types of agents. This involves constructing a cooperative localization algorithm based on hierarchical game theory and multi-agent proximal strategy optimization, alternately solving the three sub-problems to obtain the optimal cooperative localization decision. The specific process is as follows:

[0193] Step 3.1: Transform and solve the user power allocation subproblem. The specific process is as follows:

[0194] question Although it is a convex problem, its form involves numerous matrix inversion operations, increasing its difficulty and complexity, and making the solution process unstable. To solve this problem more efficiently, [the following is proposed:] This is transformed into a semidefinite programming problem.

[0195] exist During the transformation, auxiliary variables are introduced. The formula Rewritten as:

[0196]

[0197] At the same time, it is required Using Schur complement lemma, this condition can be equivalently expressed as a block matrix constraint:

[0198]

[0199] By Schuler's lemma, there is a one-to-one correspondence between feasible solutions and objective values ​​for both types of problems.

[0200] Step 3.2: Construct a hierarchical game model to solve the user-drone cooperation subproblem and the user cooperation subproblem. The specific process is as follows:

[0201] Step 3.2.1: Construct an energy-sensitive UAV connectivity decision mechanism as a higher-level game to select appropriate air-to-ground positioning links and bandwidth resources allocated on existing links. Specifically:

[0202] A drone ensemble consisting of L drones As a pool of auction participants, each drone can provide a maximum positioning bandwidth limit (i.e., the maximum transmission bandwidth of each drone) of β. v,max Based on the bandwidth constraints of the UAV, the continuous bandwidth resources are divided into M orthogonal sub-channels, forming the bandwidth block set of the l-th UAV. Among them, b l,M Let M be the bandwidth block of the l-th UAV; the capacity of each bandwidth block is Δβ = β. v,max / M. User set As a bidder, a user submits a location assistance request to a drone within their communication range. By submitting their bid, location error, and remaining energy information to the drone, the user participates in the auction process. The drone collects this information and, based on a certain mechanism, decides to allocate location services to the user.

[0203] Each user's bidding strategy can be formally represented as a triple:

[0204]

[0205] Among them, Ψ k The bidding strategy for the k-th user; This represents a subset of drones requesting a connection. Let w be the bandwidth demand vector for the k-th user. k,l This represents the bandwidth request amount from the k-th user to the l-th drone; The bid vector for serving the unit of the k-th user, π k,l Let the output value be the value from the k-th user to the l-th drone. This model satisfies Boolean connection variables. With bandwidth allocation variable β l,k The correspondence, that is If and only if And β l,k =w k,l Boolean concatenation variables express.

[0206] The objective function of the auctioneer for drones must simultaneously consider improving positioning accuracy and ensuring the fairness of positioning service allocation, so as to minimize the user's positioning error and the auctioneer's own cost.

[0207]

[0208] Where ω is the trade-off coefficient between positioning error and auction revenue, J(u k ) represents the k-th user u k The total equivalent Fisher information matrix; Select variables for drone positioning, specifically the k-th user u. k Select variables for the positioning of the l-th drone.

[0209] The user bidding strategy design reflects the impact of users' remaining energy on their competitiveness in drone positioning services. Therefore, an energy-sensitive nonlinear bid ceiling model is proposed:

[0210]

[0211] in, The maximum bid from the k-th user to the l-th user; The remaining energy for the current location of the k-th user is the power of γ. γ is the remaining energy at the current location of the l-th user; π max The maximum bid is denoted by γ, and the energy sensitivity coefficient is γ. When the energy sensitivity coefficient γ > 1, this model allows high-energy users to gain an exponentially increasing bid advantage, while preserving the basic positioning capabilities of low-energy users through normalization, achieving an effective balance between positioning accuracy and fairness. By introducing a positive correlation mechanism between the bid cap and the user's remaining energy, the model ensures that drones provide positioning services to users fairly. High-energy users have an advantage in the competition for drone positioning services; this mechanism encourages them to participate in receiving drone positioning assistance while protecting low-energy users. The drone and user positioning process is shown in Table 1.

[0212] Table 1. Pseudocode table of energy-sensitive air-to-ground assisted positioning mechanism

[0213]

[0214] Step 3.2.2: Construct a user network to form a game model as the lower-level game, in order to achieve efficient reference node selection for ground-based collaborative positioning. Specifically:

[0215] Step 3.2.2.1: Establishment of user network game formation.

[0216] User set As players in a game, each user can either initiate a connection request as a target node or receive and decide whether to accept a connection request as a reference node.

[0217] Each user's goal is to minimize the system's positioning error through cooperation during the game. Under this goal, each user autonomously decides whether to accept a positioning connection request from a target user. For the k-th user u... k Its positioning strategy can be represented as a binary vector s k :

[0218] s k =(c k,1 ,c k,2 ,…,c k,k-1 ,c k,k+1 ,…,c k,K (34);

[0219] Wherein, the Boolean variable c k,j ∈{0,1},j=1,2,...K represents the k-th user u k Whether a location relationship is established with the j-th user, i.e., u k Should we accept a location request from the j-th user? (Constrained by the original problem)

[0220]

[0221] That is u k The number of location services provided at any given time cannot exceed N. c Therefore, u k The strategy space can be written as S k :

[0222]

[0223] For each user u that receives a location request j As a reference node, its characteristic function is used to evaluate whether it agrees with the k-th user u. k Establish location-based collaboration with (target users).

[0224] If the k-th user agrees to establish location collaboration with the j-th user, which can improve its own feature function, then location collaboration is established, i.e., c k,j =1; if it cannot improve its own characteristic function, then reject the location connection request from the j-th user, i.e., c k,j =0.

[0225] Step 3.2.2.2: To incentivize cooperation on the ground, design a feature function for each user in the game:

[0226]

[0227] in, The characteristic function; u r For the r-th user, serve as the reference node; Represents the r-th user u r The set of target users with established connections, ΔJ -1 (u k ) represents u k The decrease in positioning error. P k For the r-th user u r For u k The power consumed to provide the service. E r It is the r-th user u r The remaining energy.

[0228] This feature function incentivizes user collaboration through the following mechanism:

[0229] 1) Revenue and contribution are positively correlated: for the r-th user u r The benefits are proportional to the total reduction in error for the collaborative user group, incentivizing them to actively provide location services to other users with relatively large errors.

[0230] 2) Energy consumption penalty term: Power consumption term ∑P in the denominator k This forces users to weigh the benefits against energy consumption when collaborating, in order to avoid excessive energy consumption.

[0231] 3) Residual energy regulation: Residual energy E r The higher the energy level, the greater the user's ability to provide location services to other users and the greater the benefits they receive. This makes users with sufficient energy more willing to participate in the ground-based collaborative positioning process, thereby reducing positioning errors for surrounding users.

[0232] This characteristic function constructs a mechanism that promotes user collaboration through competition by dynamically balancing positioning accuracy, energy consumption cost, and energy state, thereby incentivizing users to gradually reduce the system's positioning error.

[0233] Step 3.2.2.3: Execute the user network to form a game process.

[0234] Through multiple rounds of game theory, users eventually establish a collaborative network. Under this network, users obtain stable reference node connections and efficient positioning weight allocation, enabling the system's average positioning accuracy to reach its optimal level.

[0235] In the user network formation game iteration process of the present invention, the connection states of all users are first initialized randomly or based on experience. And calculate the initial network G (0) The characteristic function values, where, `c` is the initial Boolean variable. In each iteration, each user, as the target user, sends a location connection request to all other users within its communication range. Each other user receiving the connection request acts as a reference node, evaluating whether to agree to establish a cooperative relationship with the target user based on its characteristic function. If this link decision results in its own increment, then it agrees to the current connection request. Based on the reference node's decision, the connection relationship `c` is updated. k,j Recalculate the characteristic function values ​​of all reference nodes. If none of the users change their connection decisions in the current iteration, the game is considered converged, and the iteration terminates; otherwise, continue to the next iteration. The game process is shown in Table 2.

[0236] Table 2. Pseudocode for User Collaborative Localization Mechanism Based on User Network Formation Game Theory Model

[0237]

[0238] Step 3.3: Apply the multi-agent near-end policy optimization algorithm to optimize the location request decision in the user network formation game model, forming an efficient collaborative location network topology among users on the ground. The specific process is as follows:

[0239] Step 3.3.1: Obtain the user's current location estimate, positioning error, and remaining energy value.

[0240] Step 3.3.2: Apply the multi-agent proximal policy optimization algorithm to obtain the best user request decision; the specific process is as follows:

[0241] First, the state space Defined as: User's remaining energy E k Current number of connections c k,j Positioning Square Position Error Boundary The neighbor distance matrix consists of the distances between all users, with each element representing the k-th user u. k With the j-th user u j distance

[0242] Then, define each user's action as:

[0243]

[0244] in, Let c' be the action of the k-th user at time t, indicating whether the k-th user accepts connection requests from other users; k,K The connection indicator variable is the connection between the k-th user and the k-th user; the actions of all agents (users) in the same iteration step constitute the joint action space.

[0245] Finally, the reward is defined as a characteristic function of the network forming game, thereby incentivizing users to autonomously choose other users who can bring them higher returns.

[0246]

[0247] in, Let this be the reward for the k-th user at time t;

[0248] The pseudocode for the cooperative localization method based on hierarchical game theory and multi-agent proximal strategy optimization in this invention is shown in Table 3:

[0249] Table 3. Pseudocode for the cooperative localization method based on hierarchical game theory and multi-agent proximal policy optimization.

[0250]

[0251]

[0252] Where s is the state of the environment in the current interaction round of reinforcement learning; a is the action; r is the reward; and s' is the next state the environment enters after receiving the action.

[0253] Of course, the above description is not intended to limit the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the present invention should also fall within the protection scope of the present invention.

Claims

1. A cooperative positioning method based on hierarchical game and multi-agent proximal policy optimization, characterized in that, The method comprises the following steps: Step 1, in the UAV-assisted air-ground integrated cooperative positioning network area, an air-ground integrated cooperative positioning network model composed of three types of nodes of base stations, users and unmanned aerial vehicles is constructed; Step 2, a positioning signal model and a user positioning error model are established, and the overall optimization problem of the cooperative positioning network model is established, and the overall optimization problem is decoupled into three sub-problems of user cooperation, user and unmanned aerial vehicle cooperation and power allocation; Step 3, a cooperative positioning algorithm based on hierarchical game and multi-agent proximal policy optimization is constructed, and the three sub-problems are alternately solved to obtain the optimal cooperative positioning decision.

2. The cooperative positioning method based on hierarchical game and multi-agent proximal policy optimization according to claim 1, characterized in that, In step 1, the user is a mobile device to be positioned in the network and is located on the ground; the base station is deployed on the ground and is used for positioning; the unmanned aerial vehicle is a positioning auxiliary device hovering in the air and the position and height of which are known, serving as an unmanned aerial vehicle node in the model; the user and the base station undertake the positioning task, and the unmanned aerial vehicle undertakes the positioning auxiliary task; the user and the user, the base station and the user, and the unmanned aerial vehicle and the user are connected through a wireless link.

3. The cooperative positioning method based on hierarchical game and multi-agent proximal policy optimization according to claim 1, characterized in that, The specific process of step 2 is as follows: Step 2.1, a positioning signal model and a user positioning error model are constructed; Step 2.2, the overall optimization problem of the cooperative positioning network model is constructed; Step 2.3, the overall optimization problem is decoupled into three sub-problems of user cooperation, user and unmanned aerial vehicle cooperation and user power allocation according to different optimization variables.

4. The cooperative positioning method based on hierarchical game and multi-agent proximal policy optimization according to claim 3, characterized in that, The specific process of step 2.1 is as follows: Step 2.1.1, define the spatial geometric relationship of nodes in the cooperative positioning network model; let the distance between the kth user u k and the ith base station a i be and the azimuth be The distance between the kth user u k and the jth user u j is and the azimuth is The distance between the kth user u k and the lth unmanned aerial vehicle v l is and the azimuth is The measurement direction matrix from the unmanned aerial vehicle contains the depression angle information of the unmanned aerial vehicle to the user: wherein is a unit direction vector of the kth user u k to the lth UAV v l ; is a pitch angle of the kth user u k to the lth UAV v l ; T is a transpose symbol; Step 2.1.2, define the communication and positioning signal resources carried by the nodes in the cooperative positioning network model; let the transmit power of the ith base station be P a ; the transmit power of the kth user u k to the jth user u j be P k,j ; the maximum transmit power of the user be P u,max ; the total energy budget of the kth user u k for the positioning task be the time duration for the positioning task; the cooperative bandwidth between users is a fixed value β u , and the cooperative bandwidth between each user is equal; define a Boolean variable c k,j , which represents whether the kth user u k establishes a positioning relationship with the jth user; the transmit power of the kth user u k to the lth UAV v l is the bandwidth allocated by the lth UAV v l to the kth user u k is β l,k , and the maximum transmit bandwidth of the UAV is β v,max ; define the UAV positioning selection variable to represent the positioning selection variable of the lth UAV to the jth user; Step 2.1.3, a positioning signal model of the cooperative positioning network model is established, and the positioning signal model comprises a received signal model, a geometric relationship model of signal propagation delay and a target position, and a Fisher information matrix; the received signal model is as follows: Where, r j (t) represents the waveform of the signal received by the j-th user at time t; L j s represents the total number of paths; s(·) represents the waveform of the transmitted signal; and They represent the j-th user and the j-th... Complex gain and signal propagation delay of each path; z j (t) represents the additive white Gaussian noise of the j-th user at time t; The geometric relationship model of signal propagation delay and the target position p is as follows: where c is a constant related to the channel environment; p j is the position of the jth user; is the position of the jth user is the delay deviation of the jth user The Fisher information matrix is as follows: where J(·) is the Fisher information matrix; N b is the number of multipath parameters; P t,j is the transmit power of the jth user at time t; u j is the unit direction vector defining the direction of the positioning error of the jth user; Step 2.1.4, a user positioning error model is constructed, comprising a Fisher information matrix received by the user from the base station, an equivalent Fisher information matrix received by the user from other users, and an equivalent Fisher information matrix received by the user from the unmanned aerial vehicle; The equivalent Fisher information matrix received by the user from the base station is as follows: where J A (·) is the equivalent Fisher information matrix accepted by the user from the base station; N a is the total number of users; is the positioning information strength between the kth user u k and the ith base station; is the angle information between the kth user u k and the ith base station a i ; represents the angle information between the ith base station a i and the kth user u k ; The equivalent Fisher information matrix received by the user from other users is as follows: where J U is the equivalent Fisher information matrix received by the user from other users; K is the total number of users; is the positioning information intensity between the kth user u k and the jth user u j ; is the pitch angle of the kth user u k to the jth user u j ; represents the plane angle information of each cooperative signal received by the kth user u k and the jth user u j ; The equivalent Fisher information matrix received by the user from the unmanned aerial vehicle is as follows: where J V (·) is the equivalent Fisher information matrix received by the user from the UAVs; N v is the total number of UAVs; is the positioning information strength between the kth user u k and the lth UAV v l ; is the plane angle information of each cooperative signal received by the kth user u k and the lth UAV v l ; The total equivalent Fisher information matrix is as follows: J(u k ) = J A (u k )+ J U (u k )+ J V (u k ) Wherein, J(·) is the total equivalent Fisher information matrix; For the cooperative positioning network model, the square position error bound is used as the measurement standard of positioning accuracy; the square position error bound of the kth user is defined as follows: wherein is the estimated position of the kth user; p k is the true position of the kth user; is the squared position error bound for the user position estimate; tr{•} is the trace of a matrix; is the total equivalent Fisher information matrix for the kth user; The more position information a node obtains, the greater the eigenvalue of the total equivalent Fisher information matrix of the user, and the smaller the node square position error; at this time, the square position error bound of the user is represented as the trace of the inverse matrix of the positioning information matrix possessed by the user itself: wherein is the user's square position error bound.

5. The cooperative positioning method based on hierarchical game and multi-agent proximal policy optimization according to claim 4, characterized in that, In step 2.2, the overall optimization problem is as follows: wherein, is an overall optimization problem; C1, C2, C3, C4, C5, C6, C7, C8, C9 are different constraint conditions; P total_max is the maximum positioning signal transmission power that can be tolerated in the positioning network; N c is the maximum number of other users that a user can connect at the same time; is the maximum number of users that a UAV can connect at the same time; The specific process of step 2.3 is as follows: Step 2.3.1, User Cooperation with UAV Subproblem Modeling as wherein, is the kth user u at the current time n k received equivalent Fisher information matrix from other users; Step 2.3.2, user collaboration sub-problem modeling is modeled as wherein, the kth user u at time n+1 k received equivalent Fisher information matrix from the drone; Step 2.3.3, the user power allocation sub-problem is modeled as wherein, the kth user u at time n + 1 k received equivalent Fisher information matrix from other users.

6. The cooperative positioning method based on hierarchical game and multi-agent proximal policy optimization according to claim 5, characterized in that, The specific process of step 3 is as follows: Step 3.1, transform and solve the user power allocation sub-problem; the detailed process is as follows: transform the user power allocation sub-problem into a semi-definite programming problem, and solve the semi-definite programming problem by using the interior point method. The semi-definite programming problem is transformed into a linear programming problem by using the Cholesky factorization method. In the transformation, auxiliary variables X k are introduced, and the formula is rewritten as: X k ≥[J(u k )] -1 ; using the Schur complement lemma, this condition is equivalent to the block matrix constraint: According to the Schur complement lemma, the feasible solutions of the two types of problems have a one-to-one correspondence with the target values; Step 3.2, a hierarchical game model is constructed to solve the user and unmanned aerial vehicle cooperation sub-problem and the user cooperation sub-problem; Step 3.3, applying multi-agent proximal policy optimization algorithm to optimize positioning request decision in user network formation game model, forming efficient cooperative positioning network topology among users on the ground.

7. The cooperative positioning method based on hierarchical game and multi-agent proximal policy optimization according to claim 6, characterized in that, The specific process of step 3.2 is as follows: Step 3.2.1, constructing energy-sensitive UAV connection decision as an upper game to select appropriate air-ground positioning links and bandwidth resources allocated on established links; Step 3.2.2, constructing a user network formation game model as a lower game to achieve efficient reference node selection for ground cooperative positioning. 8.The method of claim 7, wherein, The specific process of step 3.2.1 is as follows: A drone set composed of L drones As an auctioneer set, the maximum transmission bandwidth of each drone is β v,max ; according to the bandwidth constraint of the drone, the continuous bandwidth resource is divided into M orthogonal sub-channels to form a bandwidth block set of the lth drone Where b l,M is the Mth bandwidth block of the lth drone; the capacity of each bandwidth block is Δβ=β v,max / M; a user set As a bidder, the user submits a positioning assistance request to the drone within its own communication range, participates in the auction process by submitting its own bid and positioning error, and residual energy information to the drone, and the drone collects this information and allocates the positioning service to the user according to the mechanism. The bidding strategy of each user is formally represented as a triple: wherein, Ψ k is the bidding strategy of the kth user; denotes the subset of drones requesting a connection; w k is the bandwidth demand vector of the kth user; π k is the bid vector per unit of service of the kth user; The objective function of the UAV as the auctioneer is: where ω is the weighting coefficient of the positioning error and the auction revenue, J(u k ) is the total equivalent Fisher information matrix of the kth user u k ; is the positioning selection variable of the kth user u k to the lth UAV; π k,l is the bid value of the kth user to the lth UAV; An energy-sensitive nonlinear upper limit model is proposed in the user bidding strategy: wherein, is the maximum bid of the kth user to the lth user; is the kth user's current positioning energy remaining raised to the power of γ; is the lth user's current positioning energy remaining raised to the power of γ; π max is the maximum bid of the user; and γ is an energy sensitivity coefficient. 9.The method of claim 8, wherein, The specific process of step 3.2.2 is as follows: Step 3.2.2.1, establishment of user network formation game; set of users as a player of the game; each user can both send a connection request as a target node and receive and decide whether to accept a connection request as a positioning reference node; The target of each user is to reduce the positioning error of the system as much as possible through cooperation in the game process, under this target, the user decides whether to accept the positioning connection request from the target user autonomously; for the kth user u k The positioning strategy of the user is represented as a binary vector s k : s k = (c k,1 ,c k,2 ,…,c k,k-1 ,c k,k+1 ,…,c k,K ) ; The constraint condition is: Then u k The policy space of S k : The jth user u j As a reference node, based on the characteristic function, whether to agree with the kth user u k Establish a positioning collaboration; Step 3.2.2.2, designing a characteristic function for each user in the game: wherein, is a characteristic function; u r is the rth user as a reference node; denotes the rth user u r is a set of target users having established connections; ΔJ -1 (u k ) denotes u k a positioning error reduction amount; P k is the rth user u r is u k a power providing service consumption; E r is the residual energy of the rth user u r ; Step 3.2.2.3, executing the user network formation game process; specifically: In the iterative process of user network formation game, first, initialize the connection state of all users randomly or based on experience and calculate the characteristic function value of the initial network G (0) , wherein, is the initial Boolean variable; in each iteration, each user as the target user, sends a positioning connection request to all other users within its communication range; each other user receiving the connection request as a reference node, based on its characteristic function evaluation whether to agree to establish a cooperative relationship with the target user, if this link decision makes itself increase, agree to the current connection request; according to the decision of the reference node, update the connection relationship{c k,j}; Recalculate the characteristic function value of all reference nodes; if in the current iteration, all users do not change their connection decision, it is considered that the game has converged, and the iteration is terminated; otherwise, continue the next iteration. 10.The method of claim 9, wherein, The specific process of step 3.3 is as follows: Step 3.3.1, obtaining the current position estimate, positioning error, and remaining energy value of the user; Step 3.3.2, applying multi-agent proximal policy optimization algorithm to obtain the best request decision of the user; the specific process is as follows: First, the state space is defined as: the user's remaining energy E k , the current number of connections c k,j , the square position error bound The neighbor distance matrix is composed of the distances between all users, and each element in the matrix is the distance between the kth user u k and the jth user u j ​ Then, the action of each user is defined as: wherein, is the action of the kth user at time t, indicates whether the kth user accepts a connection request from another user; c k,K is the connection indication variable of the kth user to the kth user; the actions of all users in the same iteration step form the joint action space Finally, the reward is defined as the characteristic function form of the network formation game, thereby encouraging users to independently select other users who bring higher benefits to themselves; wherein, is the reward of the kth user at time t.

Citation Information

Patent Citations

  • Power adaptive allocation method based on hierarchical game model

    CN112543498A

  • Unmanned aerial vehicle interference node power distribution method based on reconfigurable intelligent surface

    CN117054978A