Differentiated task-oriented sensing integrated multi-unmanned aerial vehicle sensing method and system
By building a differentiated perception task performance evaluation system and multi-agent near-end strategy optimization algorithm, dynamically adjusting the allocation of drone resources, solving the problem of resource waste in the collaborative perception network of multiple drones, achieving efficient and energy-saving perception tasks, suitable for smart cities and intelligent transportation.
Patent Information
- Application Number
- CN202510812613.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The existing technology has failed to effectively allocate flexible resources for perception tasks of different types and accuracy requirements in multi-UAV collaborative perception networks, resulting in resource waste and reduced task completion rates, and is unable to meet the efficient, flexible and energy-saving needs of applications such as smart cities and intelligent transportation.
Build a differentiated perception task performance evaluation system, adopt Markov decision-making process and multi-agent proximity strategy optimization algorithm, dynamically adjust the drone trajectory, task association and synesthesia power control to ensure that high-precision tasks receive more resource support, while low-precision tasks only allocate necessary resources, and optimized decisions are achieved through the MAPPO-ABAM algorithm.
It improves the overall completion rate of tasks, reduces redundant perception and resource waste, reduces system energy consumption, and provides efficient, flexible and energy-saving perception solutions suitable for environmental perception applications such as smart cities and smart transportation.
Smart Images

Figure CN120353255A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communication and sensing, and particularly to a communication-sensing integrated multi-UAV sensing method and system for differentiated tasks. Background Technique
[0002] The statements in this part only provide background technical information related to the present disclosure and do not necessarily constitute prior art.
[0003] In the upcoming 5G-Advanced and 6G networks, environmental sensing services will play an unprecedentedly important role in supporting key applications such as smart cities and intelligent transportation. Due to the advantages of high mobility, flexible deployment, and line-of-sight links, unmanned aerial vehicles (UAVs) are integrated into traditional cellular networks to provide flexible and efficient sensing services. However, due to the limited computing and storage capabilities of UAVs themselves, it is necessary to transmit sensing data to the ground central control unit (CU) via a communication link to assist in real-time and high-quality decision-making. The integrated sensing and communication (ISAC) technology can simultaneously implement communication and radar sensing functions on a single software and hardware platform. Integrating the ISAC technology into UAVs can not only reduce the load of traditional separate communication and sensing systems and extend the service time of UAVs, but also achieve higher-quality communication and sensing services through the line-of-sight link between UAVs and the ground. A single UAV with ISAC function has limitations in coverage. A collaborative communication-sensing integrated network constructed by multiple ISAC UAVs can not only sense a large area and multiple targets in a short time, but also have multi-source information to improve sensing accuracy, while achieving higher-rate data transmission, significantly improving the overall system performance, so that the ground CU receives more accurate and timely data to assist in decision-making.
[0004] Although existing research has explored multi-ISAC UAV collaborative sensing networks, there are generally deficiencies: it is usually defaulted that all sensing tasks within the task area have the same type (such as target detection or positioning), accuracy requirements, and data update frequency, and the flexible allocation of ISAC UAVs and the allocation of their communication and sensing resources have not been carried out according to the differentiated characteristics of sensing tasks. However, in actual scenarios, multiple sensing tasks within a large area often show significant heterogeneity. For example, in the intelligent transportation scenario, there are different types of tasks such as zebra crossing pedestrian detection and moving vehicle tracking and positioning. And for high-speed moving or safety-sensitive targets, higher sensing accuracy and more frequent data updates are required. Therefore, multiple UAVs are needed to cooperate to complete the tasks. For static or non-critical targets, the requirements for accuracy and update frequency are relatively low, and a single UAV can meet the needs. Ignoring these differences between tasks may lead to unreasonable resource allocation: on the one hand, low-demand tasks may be allocated too many UAVs and their communication and sensing resources, resulting in redundant sensing and data transmission, wasting system energy and resources; on the other hand, high-demand tasks may not be fully guaranteed due to insufficient remaining resources, thus reducing the overall task completion rate and affecting the quality of sensing information and decision-making efficiency of the central unit (CU). Summary of the Invention
[0005] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a communication and sensing integrated multi-UAV sensing method and system for differentiated tasks, which can assist multiple ISAC UAVs to make autonomous decisions in a collaborative sensing environment, dynamically adjust the trajectory, task association, and communication and sensing power control schemes, ensure that high-precision tasks obtain more UAV resource support, while low-precision tasks are only allocated necessary resources, thereby effectively avoiding redundant sensing and resource waste, improving the overall task completion rate, reducing system energy consumption at the same time, and providing an efficient, flexible and energy-saving solution for future intelligent sensing applications.
[0006] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions: In the first aspect, the present invention provides a communication and sensing integrated multi-UAV sensing method for differentiated tasks, including: Construct a multi-communication and sensing integrated UAV collaborative sensing system architecture, and propose the communication and sensing working modes of the communication and sensing integrated UAVs under this system architecture; Based on the communication and sensing working modes, construct an air-ground communication model and an active sensing model of the communication and sensing integrated UAVs; Construct a differentiated sensing task performance evaluation system, which defines the differentiated task completion rate based on the detection probability of the detection task and the Cramer-Rao lower bound of the positioning task, and obtains the overall network coordinated sensing performance index based on the differentiated task completion rate and the average sensing cost; Under the condition of meeting multiple constraints, an optimization problem is obtained by maximizing the task completion rate and minimizing the average perception cost; the optimization problem is modeled as a Markov decision process, and the partial observation state, action, and reward are defined. The multi-agent proximal policy optimization algorithm combined with the action branch and the mask mechanism is used to solve the Markov decision process, and a cooperative perception strategy for communication-aware integrated drones is obtained.
[0007] A further technical solution, the differential task completion rate is expressed as:
[0008] Among them, represents the task completion rate of the time slot, represents the number of differential tasks, represents the task whether it needs to be executed in the time slot, represents the task the flag indicating whether it is completed.
[0009] A further technical solution, the average perception cost includes flight power consumption, signal transmission power consumption, and signal processing power consumption, and is expressed as:
[0010] Among them, represents the number of communication-aware integrated drones, represents the normalized flight power consumption, represents the th ISAC drone whether it needs to communicate in the time slot, represents the th ISAC drone whether it needs to sense in the time slot, represents the communication-aware signal transmission power of the drone, represents the processing power consumption of the sensed echo signal.
[0011] A further technical solution, the optimization problem is expressed as:
[0012] Among them, represents the th ISAC drone whether it selects to execute the task in the time slot , represents the communication-aware signal transmission power of the drone, represents the th drone in Flight distance of the time slot, Indicates the flight angle, 、 Indicates the weighting coefficient, Indicates the task completion rate, Indicates the average sensing cost, Indicates the th drone at the horizontal coordinate of the position in the time slot, Indicates the th drone at the vertical coordinate of the position in the time slot, Indicates the maximum horizontal abscissa value that the drone can reach, Indicates the maximum horizontal ordinate value that the drone can reach, Indicates the th ISAC drone whether to communicate in the time slot, Indicates the communication rate, Indicates the minimum communication rate threshold, Indicates the distance between two drones, Indicates the minimum safe distance for drones to fly between each other, Indicates the maximum flight distance of the drone within the time slot , Indicates the maximum transmission power of the drone in the time slot .
[0013] Further technical solution, the partial observation state of the th drone before the start of the time slot is expressed as:
[0014] Among them, 、 、 Indicates the normalized drone coordinates; Indicates the sensing task index executed by the drone in the previous time slot, Indicates the th drone whether to perform sensing in the time slot, Indicates the th drone whether to communicate in the time slot, Indicates the task in the time slot whether it needs to be executed, Marks the th sensing task in the time slot whether it needs to be executed; The drone determines its state before the time slot starts and obtains an action to be executed during this time slot. The action to be executed by the nth drone during the time slot is represented as:
[0015] wherein, and and respectively represent the discretized values of the flight angle, flight distance, and signal transmission power of the drone.
[0016] In a further technical solution, all drones simultaneously execute their respective actions synchronously during the time slot. After the actions are executed, corresponding rewards will be fed back according to their decisions , and the state will then transition to . The reward is represented as:
[0017] wherein, represents the reward shared by all drones for a part of the task completion rate and sensing cost, and and respectively represent the individual penalties of the drones.
[0018] In a further technical solution, in the multi-agent proximal policy optimization algorithm that combines the action branch and the mask mechanism, each communication and sensing integrated drone is regarded as an agent. The agent includes an actor network and a critic network, and an action branch architecture and an action mask mechanism are introduced into the actor network.
[0019] In a second aspect, the present invention provides a communication and sensing integrated multi-drone sensing system for differentiated tasks, including: A system architecture construction module, which is configured to: construct a collaborative sensing system architecture for multi-communication and sensing integrated drones, and propose the communication and sensing working modes of the communication and sensing integrated drones under this system architecture; A communication and sensing model construction module, which is configured to: construct an air-ground communication model and an active sensing model of the communication and sensing integrated drones based on the communication and sensing working modes; A performance evaluation system construction module, which is configured to: construct a performance evaluation system for differentiated sensing tasks. This evaluation system defines the differentiated task completion rate based on the detection probability of the detection task and the Cramer-Rao lower bound of the positioning task, and obtains the overall network coordinated sensing performance index based on the differentiated task completion rate and the average sensing cost; A decision-making model construction module, which is configured to: under the condition of meeting multiple constraint conditions, maximize the task completion rate and minimize the average perception cost to obtain an optimization problem; model the optimization problem as a Markov decision process, and define partial observation states, actions, and rewards. An optimization decision-making module, which is configured to: use a multi-agent proximal policy optimization algorithm combining an action branch and a masking mechanism to solve the Markov decision process, and obtain a cooperative perception strategy for integrated communication and perception unmanned aerial vehicles.
[0020] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in the integrated communication and perception multi-unmanned aerial vehicle perception method for differentiated tasks described in the first aspect are implemented.
[0021] In a fourth aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in the integrated communication and perception multi-unmanned aerial vehicle perception method for differentiated tasks described in the first aspect are implemented.
[0022] The above one or more technical solutions have the following beneficial effects: For multiple differentiated perception tasks, the cooperative perception framework proposed by the present invention can finely match the specific requirements of each perception task, ensure that high-precision tasks obtain more unmanned aerial vehicle resource support, while low-precision tasks are only allocated necessary resources, which can improve the overall task completion rate, reduce redundant perception and waste of resources and energy at the same time, and reduce the overall task execution cost. The present invention establishes an optimization problem of jointly optimizing unmanned aerial vehicle task association, trajectory planning, and communication and perception power control to maximize the task completion rate within the service time while minimizing the perception cost, and proposes an optimization algorithm based on multi-agent proximal policy optimization combined with an action branch and a masking mechanism (MAPPO-ABAM) to solve this problem. Through centralized training and distributed execution, the trained actor network is deployed on each ISAC unmanned aerial vehicle, enabling it to make autonomous decisions and dynamically adjust the trajectory, task association, and communication and perception power control scheme, which can reduce the computational pressure on the central processor and obtain superior performance with a high task completion rate and low perception cost.
[0023] The present invention constructs a unified index system for differential perception tasks, introduces a perception completion rate index based on the joint detection probability and the joint Cramer-Rao lower bound, breaks through the limitations of the traditional single-task evaluation system, realizes the unified quantification of heterogeneous accuracy requirements across task types and different numbers of associated drones, and defines the execution cost of perception tasks at the same time; designs a joint optimization model based on spatio-temporal coupling, deeply integrates the UAV trajectory planning, task association decision-making and communication-sensing power control, and through the three-dimensional constraint modeling of flight energy consumption - perception efficiency - communication quality, realizes the optimal spatio-temporal matching of multi-dimensional resources, ensures the maximization of the differential perception task completion rate while minimizing the perception cost; aiming at the problem of difficult rapid solution brought by the large-scale mixed action space and the high dynamicity of UAVs, based on the optimization algorithm of multi-agent proximal policy optimization combined with action branch and masking mechanism (MAPPO-ABAM), innovatively uses the action branch structure to decouple the mixed action space, combines the dynamic masking mechanism to avoid ineffective action exploration, and alleviates the high complexity of centralized computing through the distributed decision-making network; the present invention can assist multiple ISAC UAVs to make autonomous decisions in a collaborative perception environment, dynamically adjust the trajectory, task association and communication-sensing power control schemes, ensure that high-precision tasks obtain more UAV resource support, while low-precision tasks are only allocated necessary resources, thereby effectively avoiding redundant perception and resource waste, improving the overall task completion rate, and reducing the system energy consumption at the same time, providing an efficient, flexible and energy-saving solution for future intelligent perception applications.
[0024] Aiming at the application requirements of environmental perception such as smart cities and intelligent transportation, the present invention proposes a communication-sensing integrated multi-UAV collaborative perception framework for differential perception tasks, which can be used for communication-sensing resource scheduling, dynamic joint detection and positioning, intelligent transportation monitoring, etc. The UAVs empowered by ISAC technology actively sense multiple perception tasks with different types, accuracy requirements, and data update frequencies in a large-area target region and transmit the perception data to the ground CU to support precise and efficient decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0026] Figure 1 The flowchart of the communication-sensing integrated multi-UAV perception method for differential tasks in the embodiment of the present invention; Figure 2 is the architecture diagram of the multi-communication-sensing integrated UAV collaborative perception system in the embodiment of the present invention; Figure 3 is the communication and perception working mode diagram of the communication-sensing integrated UAV in the embodiment of the present invention; Figure 4It is the algorithm architecture diagram of the embodiment of the present invention based on Multi-Agent Proximal Policy Optimization (MAPPO) combined with Action Branch and Masking Mechanism (ABAM); Figure 5 It is the reward convergence graph of the algorithm of the embodiment of the present invention; Figure 6 It is the multi-ISAC UAV trajectory planning result graph obtained by the method of the embodiment of the present invention; Figure 7 It is the comparison graph of the perception task completion rate, objective function value, and perception cost value obtained by the algorithm of the embodiment of the present invention and the existing optimization method. Detailed implementation manners
[0027] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0028] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments of the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0029] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0030] Explanation of professional terms: ISAC: Integration of Sensing and Communication; CU: Central Control Unit; CRB: Cramér-Rao Lower Bound; PD: Detection Probability; MAPPO: Multi-Agent Proximal Policy Optimization; ABAM: Action Branch and Action Masking Mechanism.
[0031] Aiming at the problems of the existing technology, a communication and sensing integrated multi-UAV collaborative sensing framework for differentiated tasks is proposed. It can not only precisely match the specific requirements of each sensing task, improve the overall task completion rate, but also avoid the waste of resources and energy, reduce the overall task execution cost, and provide an efficient, flexible and energy-saving technical solution for future intelligent applications. The main challenges of this problem include: 1) Different types of sensing tasks use different performance metrics. The positioning task is mostly measured by the Cramer-Rao lower bound (CRB), while the detection task focuses on the detection probability (PD). At the same time, to avoid resource waste, high-precision tasks may require multiple UAVs to be associated, while low-precision tasks only need a single UAV to participate. Therefore, how to construct a unified index system that can reflect the overall collaborative sensing performance and be compatible with multiple task types and different UAV association numbers is a core challenge. 2) In the multi-ISAC UAV collaborative sensing network, the UAVs need to simultaneously meet the sensing requirements of multiple tasks and the communication requirements of transmitting sensing data to the central control unit, and also need to maintain a safe distance from other UAVs. The trajectory design of multiple UAVs has a high degree of coupling and complexity. 3) The number of optimization decision variables of multiple UAVs is large, resulting in a huge combined optimization variable space and high computational complexity. At the same time, the high-speed mobility of UAVs requires finding the optimal solution quickly within a limited time, posing challenges to the algorithm solving efficiency and real-time performance.
[0032] Embodiment 1 As Figure 1 shown, this embodiment discloses a communication and sensing integrated multi-UAV sensing method for differentiated tasks, which includes the following steps: S1: Construct a multi-communication and sensing integrated UAV collaborative sensing system architecture, and propose the communication and sensing working modes of the communication and sensing integrated UAVs under this system architecture; In this embodiment, as Figure 2 shown, the multi-communication and sensing integrated (ISAC) UAV collaborative sensing system architecture includes: a number of UAVs with ISAC units, and the set of all UAVs is , a ground control center (CU), a number of differentiated sensing tasks, and the set is , and a triple is used to describe the sensing task differences, and the set of triples is , where represents the type of the sensing task ( represents the detection task, represents the positioning task), represents the sensing interval, that is, after how many time slots the task needs to be sensed again, Represents the perception accuracy requirement, which represents the lowest detection probability threshold for the detection task and the highest Cramer-Rao lower bound threshold for the localization task. The service time of the ISAC drone is time slots, and the time slot set is denoted as , where the th drone's position at the th time slot is denoted as , where all drones fly at a fixed altitude within the time slots, and its movement model is given by the following formula:
[0033] where, is the abscissa of the th drone's position at the th time slot, is the ordinate of the th drone's position at the th time slot, is the flight distance of the th drone at the th time slot, is the flight angle.
[0034] In this embodiment, as Figure 3 shown, at each time slot, each integrated communication and sensing (ISAC) drone can choose to perform at most one sensing task or not perform any sensing task. Due to the different accuracy requirements of sensing tasks, each sensing task can be performed by one or more ISAC drones. In addition, taking the th ISAC drone as an example, if it performs a sensing task at the th time slot, then it must transmit the sensing data obtained from this task at the th time slot. Therefore, at the th time slot, each ISAC drone may have four working modes as follows:
[0035] where, represents whether the th ISAC drone will communicate at the th time slot, represents whether the th ISAC drone will sense at the th time slot.
[0036] S2: Construct an air-ground communication model and an active sensing model for the integrated communication and sensing drone based on the communication and sensing working modes; In this embodiment, the frequency bands are divided among multiple ISAC drones using orthogonal frequency division multiple access (OFDMA). The communication and sensing functions of a single drone share a frequency band, and its bandwidth is , so the interference between multiple drones is not considered when constructing the air-ground communication model and the active sensing model. The air-ground communication model mainly refers to the channel model when the ISAC drone transmits the acquired sensing data to the ground CU. After the drone obtains the sensing data of the task in the time slot, it will transmit the data to the ground CU by emitting communication or ISAC signals in the time slot. The communication between the drone and the ground uses a practical probability channel model. According to the line-of-sight (LOS) channel probability and the non-line-of-sight (NLOS) channel probability between the th ISAC drone and the CU, the channel gain can be obtained. According to the channel gain, the sensing data transmission rate between the th ISAC drone and the CU can be obtained:
[0037] where represents the communication rate, represents the transmission power of the drone's communication signal, and represents the noise power.
[0038] Meanwhile, to ensure the timeliness and quality of the sensing data transmission, the data transmission rate must be greater than the minimum communication rate threshold .
[0039] The air-ground communication model describes the communication process of the drone transmitting sensing data to the ground CU, and gives the calculation expression of the communication rate (i.e., the data transmission rate) and the minimum rate threshold. Optimization variables such as the position and transmission power of the drone will affect the communication rate.
[0040] In different working modes, the ISAC drone actively emits radar sensing signals or integrated communication and sensing signals, receives the sensing echoes reflected by the sensing targets, and processes the echoes to detect or locate potential targets in the task area. In the time slot, if the th ISAC drone selects to sense the task , then the power of the sensing echo it receives is , and the sensing signal-to-noise ratio can be obtained according to the power of the sensing echo:
[0041] where is the The ISAC drone in time slot for the task perceived signal-to-noise ratio, is the perceived signal processing gain, represents the perceived noise.
[0042] The active perception model describes the process by which the drone perceives the ground differential perception task and gives the power of the received echo and the perceived signal-to-noise ratio, which can be used to calculate the detection probability of the detection task and the Cramér-Rao lower bound of the positioning task, and further calculate the performance of the differential perception task.
[0043] S3: Construct a performance evaluation system for differential perception tasks. This evaluation system defines the differential task completion rate based on the detection probability of the detection task and the Cramér-Rao lower bound of the positioning task, and obtains the overall network coordinated perception performance index based on the differential task completion rate and the average perception cost; In this embodiment, a unified index system that can be compatible with multiple task types and different numbers of associated drones is defined respectively, and further the perception task completion rate index is obtained as an index to evaluate the overall collaborative perception performance of multiple ISAC drones. In addition, the cost of the collaborative perception task is defined and a weighted index is formed with the perception task completion rate to guide the system to achieve the maximum task completion rate with the minimum cost.
[0044] First, for the detection task, the detection probability is used as an index to measure the perception accuracy. The detection probability can represent the probability that the drone successfully detects the target. The th ISAC drone for the task The detection probability of perception can be expressed as:
[0045] Among them, represents whether the th ISAC drone selects to execute the task in the time slot , if it is 0, then the corresponding detection probability is also 0; represents the threshold for determining whether the target exists; is the right-tail function of the standard normal distribution.
[0046] According to the detection probability of a single drone for the task, the joint detection probability of the task is defined as:
[0047] Among them, represents the joint detection probability of the task .
[0048] Regardless of the task In this metric applies whether the time slot is associated with one or multiple UAVs. When only one UAV is performing a mission the values of the remaining UAVs are all 0, and the detection probability of a single UAV is obtained at this time.
[0049] For the positioning task, the Cramer-Rao lower bound (CRB) is used in the present invention to measure the accuracy of sensing and positioning. The Cramer-Rao lower bound refers to the lower bound of the mean square error of an unbiased estimator that can be obtained, and can be obtained by taking the inverse of the Fisher information matrix. Since it is difficult to directly obtain the observation of the position coordinates of the sensing target, its Fisher information matrix can be indirectly solved through the Fisher information matrix of the distance estimation of the ISAC UAV for the mission and the chain rule. The Fisher information matrix of the coordinate estimation is expressed as:
[0050] wherein represents the coordinates of the positioning task, refers to the set of distances between multiple ISAC UAVs and the mission and is the Jacobian matrix of represents the Fisher information matrix of the distance, and the , th elements in the matrix can be calculated by the following formula:
[0051] wherein represents the th element in the set , and represents the covariance matrix of represents the trace of the matrix, represents the inverse of the covariance matrix. Further, the Cramer-Rao lower bound that can be obtained is:
[0052] wherein and are the diagonal element values of , and is used as the metric for positioning accuracy. Since the coordinate values need to be estimated using the observed distances, when the mission needs to be sensed in a time slot, its position can be estimated only when at least two or more UAVs are associated. When the number of associated UAVs exceeds two, regardless of the specific number of associated UAVs, the above-mentioned All the indicators are applicable. However, if only one UAV is associated, the data is invalid, and at this time is set to a very large value, indicating that the task is not completed.
[0053] Based on the above indicators, the present invention defines the differential task completion rate as an indicator for evaluating the overall perception performance. First, define the identification of whether the task is completed , which is expressed as:
[0054] Among them, represents the precision threshold. When the task is a detection task, if the detection probability is greater than the required detection probability precision, then is 1, marking that the task is completed. And when the task is a positioning task, if is less than the given precision threshold , then it marks that the task is completed. The overall task completion rate of the time slot can be obtained by the following formula:
[0055] Among them, represents whether the task needs to be executed in the time slot. Note that here represents the effective task completion rate. Since the perception intervals of multiple tasks are different, if the task does not need to be executed in the time slot, then even if the perception accuracy requirement is met, it will not be included in the task completion rate.
[0056] is to measure the cost paid by the ISAC UAV when performing the perception task. The average perception cost of multiple UAVs in the time slot consists of three parts: flight power consumption, signal transmission power consumption, and signal processing power consumption, and is calculated by the following formula:
[0057] Since the magnitudes of the flight power consumption and the signal transmission power consumption are quite different, the flight power consumption is normalized and expressed as , represents the communication perception signal transmission power of the UAV, represents the processing power consumption of the perception echo. Once a UAV emits a perception signal or an ISAC signal, a perception echo processing power consumption will be generated.
[0058] Finally, The weighted sum of the task completion rate and the average sensing cost of each time slot is used as an indicator to evaluate the overall collaborative sensing performance of the network and the objective function of the following optimization problem, which is expressed as the following formula:
[0059] where, and are the weighting coefficients.
[0060] S4: Under the condition of satisfying multiple constraints, maximize the task completion rate and minimize the average sensing cost to obtain an optimization problem; model the optimization problem as a Markov decision process, and define the partially observable state, action, and reward; In this embodiment, based on the above-mentioned multi-ISAC UAV collaborative sensing system architecture, combined with the ISAC UAV air-ground communication model and the active sensing model, by solving the optimal UAV-task association scheme and the signal transmission power as well as the trajectory planning variables and , the optimal multi-ISAC UAV collaborative sensing strategy for different tasks is obtained. Under the condition of satisfying multiple constraints such as communication rate, multi-UAV safety collision avoidance constraint, flight constraint, and maximum transmission power, maximize the overall collaborative sensing performance index of the network, and achieve the maximum task completion rate of multiple different sensing tasks within service time slots, while minimizing the sensing cost, and avoid allocating too many resources to low-demand tasks, resulting in redundant sensing and waste of resources. That is to say, maximizing the task completion rate and minimizing the average sensing cost represent the best overall collaborative sensing performance of the network. The optimization problem is established as follows:
[0061] where, represents the maximum achievable horizontal abscissa value of the UAV, represents the maximum achievable horizontal ordinate value of the UAV, represents the distance between two UAVs, represents the minimum safe distance between the flights of UAVs, represents the maximum flight distance of the UAV within the time slot , represents the maximum transmission power of the UAV in the time slot .
[0062] Among them, the optimization variables and constraint conditions apply to each , where C1 restricts the flight range of the UAV to a fixed area , C2 is the data transmission rate constraint to ensure communication timeliness and quality, C3 is the safety distance constraint for multiple UAVs, where is the minimum interval, C4 is the flight distance constraint of the UAV within a single time slot, which cannot exceed , C5 is the flight angle constraint of the UAV, and C6 restricts that when the UAV sends communication signals, sensing signals, or integrated communication and sensing signals, the power cannot exceed .
[0063] Since each ISAC UAV can only observe part of the state and cannot obtain the global state of the cooperative sensing network, including the positions of other UAVs, etc., the present invention constructs a partially observable Markov decision process to simulate the task execution process of multiple ISAC UAVs within multiple time slots. The partially observable Markov process includes a state space, an action space, and a reward function.
[0064] Each UAV will have a partial observation of itself before the start of a time slot. The observation state of the th UAV before the start of the time slot is expressed as:
[0065] Among them, , , represent the normalized UAV coordinates; represents the index of the sensing task executed by the UAV in the previous time slot. If no sensing is performed, then ; marks whether the th sensing task needs to be executed in the time slot. The set of partial observations of all UAVs is the overall state of the system before the start of the time slot .
[0066] The UAV will obtain an action that needs to be executed in this time slot according to the state before the start of the time slot. The action executed by the th UAV in the time slot includes the flight angle, distance of the UAV, the index of the sensing task executed in this time slot (including the case of not performing sensing), and the signal transmission power, which can be expressed as:
[0067] Among them, , , respectively represent the discretized values of the flight angle, flight distance, and signal transmission power of the UAV. Each UAV is in the First, it executes flight according to the decision of its flight angle and distance within the time slot. After reaching the target position, it determines whether hovering communication and sensing are required for this time slot according to and , and then executes the corresponding operations. Since the optimization variables involve the mixed solution of discrete and continuous variables, the computational complexity is relatively high. Therefore, the flight angle, flight distance, and transmission power of the UAV are discretized, and are represented by , , respectively, where , , represent the discretization levels respectively. The action sets of all UAVs in the time slot are represented as .
[0068] All UAVs simultaneously execute their respective actions synchronously in the time slot. After the actions are executed, the system will feedback the corresponding reward , and then undergoes a state transition to . The reward is calculated by the following formula:
[0069] Among them, all UAVs share a part of the reward for the task completion rate and sensing cost , which is calculated by the differential sensing task performance evaluation system given above. In addition, , , represent the individual penalties of the UAVs respectively, where is used to constrain the communication power not to be lower than the threshold, penalizes the behavior of the UAV flying beyond the boundary, penalizes the situation where the UAV collides with other UAVs.
[0070] In the time slots, each UAV continuously executes the following loop: observe the current state, select and execute actions based on the state, obtain the corresponding reward, and then transfer to the state of the next time slot. This process continues until the service time ends.
[0071] S5: Solve the Markov decision process by using a multi-agent proximal policy optimization algorithm combined with an action branch and a mask mechanism to obtain a multi-communication and sensing integrated UAV collaborative sensing strategy; In this embodiment, as Figure 4As shown in the figure, the present invention uses the Multi-Agent Proximal Policy Optimization with Action Branch and Masking Mechanism (MAPPO-ABAM) algorithm to train the decision-making model to find the optimal action based on the partial observation variables of the UAV itself, so as to effectively guide multiple ISAC UAVs to perform collaborative perception and communication services for differential perception tasks.
[0072] In the MAPPO-ABAM algorithm, each ISAC UAV is regarded as an agent, and the agent contains an actor network and a critic network. The actor network is the core network that guides the actions of the UAV. Its input is the observation state of the UAV and its output is the action of the UAV while the critic network is used to evaluate the long-term return of the state-action pair. The present invention uses an actor network and a critic network with three hidden layers, and the structure is input layer - fully connected layer (MLP) - recurrent neural network layer (GRU) - MLP - output layer, where the MLP hidden layer contains 128 neurons and the GRU hidden layer contains 64 neurons. The algorithm has two stages: centralized training and distributed execution. Centralized training is mainly to collect global states and rewards, calculate the state value function to assist in training the optimal policy. After training is completed, the actor network can be directly deployed to each UAV, and each agent makes independent decisions only relying on the trained actor network and local observations, reducing the computational pressure on the central computing unit.
[0073] In the centralized training stage, the present invention uses three items for training: clipped loss, mean squared error loss, and entropy reward, which are used to limit the update amplitude of the actor network, optimize the state value function estimation, and improve the exploration ability of the network respectively. The clipped loss is calculated by the following formula:
[0074] where represents the approximate expectation, is the ratio of the old and new policies, is the Generalized Advantage Estimation (GAE) value, is the clipping threshold. For the multi-dimensional discrete action space, the present invention introduces an action branch structure in the actor network, so that actions in different dimensions have independent output branches, avoiding the curse of dimensionality problem caused by dimension combinations and reducing the training difficulty. The approximate expectation means using the sampled samples at time step to approximately calculate the expectation of the loss. For each time step draw data from the experience pool (batch), and use these data to estimate the expectation of the entire loss.
[0075] For each action dimension, such as the flight angle, the ratio of the new and old policies is calculated as follows:
[0076] Where, represents the current actor network, represents the old actor network.
[0077] The clipping losses corresponding to other action dimensions are calculated in the same form, and the losses of all action dimensions are averaged to obtain the overall clipping loss function. The overall loss function obtained after adding the mean squared error loss and the entropy reward is:
[0078] Where, is the entropy reward, is the mean squared error loss, 、 respectively represent the hyperparameters that control the weights of the entropy reward and the value function loss, represents the th dimension of the action.
[0079] According to the overall loss function, the present invention uses the gradient descent method to update and learn the parameters of the actor network and the critic network to obtain the optimal actor network. During the training process, the present invention uses multi-round iterative optimization. Specifically, in each training round, the agent executes the policy in the environment, collects trajectory data including states, actions, rewards, and state value estimates, and stores them in the trajectory storage pool. Subsequently, trajectory data is sampled from the storage pool, the advantage estimate and the target state value are calculated, and the actor network and the critic network are gradient-updated using the above loss function. This process continues to iterate until the policy converges. In addition, the present invention introduces an action masking mechanism in the actor network to prevent the agent from selecting invalid actions. For example, in the task selection action branch, when a certain task does not need to sense in the current time slot, the probability of the corresponding action is set to a very small value (such as ), and then the probabilities of the remaining actions are renormalized to ensure reasonable sampling of valid actions.
[0080] During the distributed execution phase, the trained actor network will be deployed to each drone. At the beginning of each time slot, all drones will observe and obtain the local state vector including drone position, task status, communication requirements, etc., and then input this vector into the actor network to obtain the action value. Each drone will make decisions autonomously based on its actor network. After receiving the action instruction, each drone will fly based on the flight model for 1 second. After the flight ends, the drone will hover at the target location and determine whether to communicate in the current time slot according to the sensing task execution situation in the previous time slot. At the same time, according to the action instruction in the current time slot, the drone will decide whether to execute the sensing task and which specific task to execute. Subsequently, the drone will perform hovering communication and sensing according to the corresponding working mode, and the hovering communication and sensing processes will refer to the above-mentioned ISAC drone air-ground communication model and active sensing model. For the drones that choose to execute the sensing task in the current time slot, they will receive the sensing echo data after sending the signal and send the data to the CU through the communication link in the next time slot. After the actions in this time slot are executed, the system state changes, and the drone enters the next time slot, repeating the operations of action acquisition - action execution, and this operation loops until the completion of service time slots.
[0081] The decision variables of the Markov decision process are the flight angle, flight distance, and communication and sensing signal transmission power of the drone in each time slot, as well as whether to perform sensing in each time slot. If sensing is performed, which sensing task to choose, and the index of the task will be given. Each drone is regarded as an agent in the MAPPO-ABAM algorithm and will carry a trained actor network. Using this actor network, each drone can output the above-mentioned several decision variables according to its partial observation state at the beginning of the time slot to achieve autonomous decision-making. The decision variables correspond to the dynamic adjustment of the trajectory, task association, and communication and sensing power control scheme.
[0082] Experimental description: The present invention implements the proposed communication and sensing integrated multi-drone collaborative sensing framework for differentiated tasks. The simulation platform used in the experiment is Pycharm 2022.2.2, the code language is Python3.9, and the deep learning framework used is pytorch2.4.1. The simulation environment consists of 4 drones and 4 differentiated tasks, including two detection tasks and two localization tasks. One of the two localization tasks is static and the other is dynamic, and so are the detection tasks. Moreover, the sensing accuracy requirements and sensing frequencies of each task are different. The proposed MAPPO-ABAM optimization algorithm in the present invention can be trained to convergence in the simulation environment, Figure 5 which is the convergence curve graph of the algorithm proposed in the present invention and two comparative benchmark algorithms, namely the PPO algorithm and the MADDPG algorithm. Figure 6The figure of the UAV obtained when deploying the trained actor network to multiple ISAC UAVs for autonomous decision-making. Figure 7 Compared the performance of two benchmark algorithms, namely Proximal Policy Optimization (PPO) and Multi-Agent Deep Deterministic Policy Gradient (MADDPG), and the proposed MAPPO-ABAM algorithm of the present invention in terms of the differential perception task completion rate, perception cost, and objective function value (i.e., the weighted function value of the perception task completion rate and perception cost) within the service time.
[0083] From Figure 5 It can be observed that the proposed MAPPO-ABAM algorithm of the present invention has the fastest convergence speed compared to the two benchmark algorithms. This is mainly due to the introduction of an action branch architecture in the actor network, which decouples the multi-dimensional action space, improves the learning flexibility of multi-dimensional actions, and avoids the curse of dimensionality caused by the large-scale combinatorial action space, enabling the algorithm to converge quickly and maintain excellent performance in complex scenarios. From Figure 6 It can be seen that under the preset differential tasks, multiple ISAC UAVs can autonomously form a task group according to the task location. UAV 3 hovers near Task 2 to perform low-precision detection; meanwhile, UAV 1, UAV 2, and UAV 4 continuously approach Task 1 and Task 3 to cooperate in perception for high-precision detection and positioning. In addition, UAV 1 and UAV 4 will also appropriately adjust their trajectories to stay close to Task 4 to provide low-precision positioning services when needed. From Figure 7 It can be found that the proposed MAPPO-ABAM algorithm of the present invention has the highest perception task completion rate and the lowest perception cost compared to the two benchmark algorithms. Both the actor network and the critic network of MAPPO-ABAM adopt GRU hidden layers to capture time dependencies and optimize long-term decisions. In a cooperative perception environment, the GRU layer can learn the perception interval patterns of differential tasks and plan UAV trajectories in advance, thereby reducing flight energy consumption. In addition, the proposed action masking mechanism effectively avoids the ISAC UAVs from being associated with perception tasks that do not need to be executed in the current time slot, fundamentally eliminating invalid task selection, preventing resource waste, and ensuring that critical tasks are given priority. The results show that within the given service time, the proposed scheme of the present invention achieves a task completion rate of 96.8085% at the lowest perception cost, while the task completion rates of the MADDPG and PPO algorithms are only 62.7660% and 77.6596% respectively. In the scenario of integrated communication and perception multi-UAV cooperative perception for differential tasks, the proposed scheme of the present invention demonstrates significant performance advantages.
[0084] Example Two This example discloses an integrated communication and perception multi-UAV perception system for differential tasks, including: System architecture construction module, which is configured to: construct a multi-communication and sensing integrated UAV collaborative sensing system architecture, and propose the communication and sensing working modes of the communication and sensing integrated UAV under this system architecture; Communication and sensing model construction module, which is configured to: construct an air-ground communication model and an active sensing model of the communication and sensing integrated UAV based on the communication and sensing working modes; Performance evaluation system construction module, which is configured to: construct a differential sensing task performance evaluation system. This evaluation system defines the differential task completion rate based on the detection probability of the detection task and the Cramer-Rao lower bound of the positioning task, and obtains the overall network coordinated sensing performance index based on the differential task completion rate and the average sensing cost; Decision model construction module, which is configured to: under the satisfaction of multiple constraint conditions, maximize the task completion rate and minimize the average sensing cost to obtain an optimization problem; model the optimization problem as a Markov decision process, and define the partial observation state, action, and reward; Optimization decision module, which is configured to: use the multi-agent proximal policy optimization algorithm combined with the action branch and mask mechanism to solve the Markov decision process, and obtain the collaborative sensing strategy of the multi-communication and sensing integrated UAV.
[0085] Embodiment III The purpose of this embodiment is to provide a computing device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method in Embodiment I are implemented.
[0086] Embodiment IV The purpose of this embodiment is to provide a computer-readable storage medium. A computer-readable storage medium stores a computer program, and when the program is executed by a processor, the steps of the method in Embodiment I are executed.
[0087] The steps involved in the devices in the above Embodiments III and IV correspond to those in Method Embodiment I, and the specific implementation manners can refer to the relevant description part of Embodiment I. The term "computer-readable storage medium" should be understood to include a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0088] Those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0089] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0090] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A synesthetic integrated multi-UAV perception method for differentiated tasks, characterized in that Including: Construct a multi-communication and sensing integrated UAV collaborative sensing system architecture, and propose the communication and sensing working modes of the communication and sensing integrated UAV under this system architecture; Construct an air-ground communication model and an active sensing model of the communication and sensing integrated UAV based on the communication and sensing working modes; Construct a differential sensing task performance evaluation system, which defines the differential task completion rate based on the detection probability of the detection task and the Cramer-Rao lower bound of the positioning task, and obtains the overall network coordinated sensing performance index based on the differential task completion rate and the average sensing cost; Under the condition of meeting multiple constraint conditions, maximize the task completion rate and minimize the average sensing cost to obtain an optimization problem; model the optimization problem as a Markov decision process, and define the partial observation state, action and reward; Use the multi-agent proximal policy optimization algorithm combined with the action branch and mask mechanism to solve the Markov decision process, and obtain the collaborative sensing strategy of the multi-communication and sensing integrated UAV.
2. The integrated multi-UAV sensing method for differentiated tasks according to claim 1, wherein The differential task completion rate is expressed as: Among them, represents the task completion rate of the time slot, represents the number of differentiated tasks, represents the task in whether the time slot needs to be executed, represents the task the flag indicating whether it is completed.
3. The synesthesia-integrated multi-UAV perception method for differentiated tasks according to claim 1, wherein, The average sensing cost includes flight power consumption, signal transmission power consumption, and signal processing power consumption, and is expressed as: Among them, represents the number of communication-sensing integrated UAVs, represents the normalized flight power consumption, represents the th ISAC UAV whether it needs to communicate in the time slot, represents the th ISAC UAV whether it needs to perform sensing in the time slot, represents the communication-sensing signal transmission power of the UAV, represents the processing power consumption of the sensed echo signal.
4. The integrated multi-UAV perception method for synesthesia-oriented differentiated tasks according to claim 1, characterized in that The optimization problem is expressed as: Among them, represents whether the th ISAC drone selects to execute a task in the time slot, , represents the transmission power of the communication sensing signal of the drone, represents the th drone's flight distance in the time slot, represents the flight angle, , represents the weighting coefficient, represents the task completion rate, represents the average sensing cost, represents the th drone's horizontal abscissa position in the time slot, represents the th drone's vertical ordinate position in the time slot, represents the maximum horizontal abscissa value that the drone can reach, represents the maximum horizontal ordinate value that the drone can reach, represents whether the th ISAC drone will communicate in the time slot, represents the communication rate, represents the minimum communication rate threshold, represents the distance between two drones, represents the minimum safe distance between drones during flight, represents the maximum flight distance of the drone within the time slot , represents the maximum transmission power of the drone in the time slot .
5. The integrated multi-UAV perception method for differentiated tasks according to claim 1, characterized in that The partial observation state of the first UAV before the start of the time slot is expressed as: Among them, , , represent the normalized UAV coordinates; represents the sensing task index executed by the UAV in the previous time slot, represents whether the th UAV needs to perform sensing in the time slot, represents whether the th UAV needs to communicate in the time slot, represents whether the task needs to be executed in the time slot, marks whether the sensing tasks need to be executed in the time slot; The drone determines its state before the start of the time slot to obtain an action to be executed during this time slot. The action to be executed by the nth drone during the time slot is denoted as: Among them, , , respectively represent the values after discretization of the flight angle, flight distance, and signal transmission power of the drone.
6. The synesthesia-integrated multi-UAV perception method for differentiated tasks according to claim 1, wherein All drones are simultaneously The time slots execute their respective actions synchronously. After the actions are completed, the corresponding rewards will be fed back according to their decisions. , and the state is transferred to , the reward is expressed as: Among them, indicates that all drones share a part of the reward for task completion rate and sensing cost, , , respectively represent the individual penalties of the drones.
7. The integrated multi-UAV perception method for differentiated tasks according to claim 1, characterized in that In the multi-agent proximal policy optimization algorithm combined with the action branch and mask mechanism, each communication and sensing integrated UAV is regarded as an agent, and the agent includes an actor network and a critic network. An action branch structure and an action mask mechanism are introduced into the actor network.
8. A synesthetic integrated multi-UAV perception system for differentiated tasks, characterized in that, Including: A system architecture construction module, which is configured to: construct a multi-communication and sensing integrated UAV collaborative sensing system architecture, and propose the communication and sensing working modes of the communication and sensing integrated UAV under this system architecture; A communication and sensing model construction module, which is configured to: construct an air-ground communication model and an active sensing model of the communication and sensing integrated UAV based on the communication and sensing working modes; A performance evaluation system construction module, which is configured to: construct a differential sensing task performance evaluation system, which defines the differential task completion rate based on the detection probability of the detection task and the Cramer-Rao lower bound of the positioning task, and obtains the overall network coordinated sensing performance index based on the differential task completion rate and the average sensing cost; A decision model construction module, which is configured to: under the condition of meeting multiple constraint conditions, maximize the network task completion rate and minimize the average sensing cost to obtain an optimization problem; model the optimization problem as a Markov decision process, and define the partial observation state, action and reward; An optimization decision module, which is configured to: use the multi-agent proximal policy optimization algorithm combined with the action branch and mask mechanism to solve the Markov decision process, and obtain the collaborative sensing strategy of the multi-communication and sensing integrated UAV.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in the communication and sensing integrated multi-UAV sensing method for differential tasks described in any one of claims 1-7.
10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the communication and sensing integrated multi-UAV sensing method for differential tasks described in any one of claims 1-7.
Citation Information
Patent Citations
Mobile crowd sensing coverage optimization method based on multi-modal trajectory
CN114429580A
Method for remotely searching vehicle and vehicle
CN115334148A
Intelligent Internet of Things multi-modal data sensing method
CN116628441A
RIS-assisted unmanned aerial vehicle network-based communication and sensing integrated system and method
CN119277319A
Communication sensing task allocation method and device based on multi-target asynchronous strategy
CN120029326A
Cited By
Multi-moving-target tracking method and system of communication and sensing integrated unmanned aerial vehicle
CN122218684A