Multi-unmanned aerial vehicle perception method and system for differentiated tasks
By building a synergistic integrated multi-UAV perception system for differentiated tasks and dynamically adjusting UAV resource allocation, the problem of unreasonable resource allocation in the multi-ISAC UAV collaborative perception network is solved, the task completion rate is improved and the system energy consumption is reduced.
Patent Information
- Application Number
- CN202510812613.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The existing multi-ISAC UAV collaborative perception network fails to flexibly allocate resources based on the differentiated characteristics of different perception tasks, resulting in irrational resource allocation, redundant perception and resource waste, or insufficient resources for high-demand tasks, affecting the overall task completion rate and system energy consumption.
Build a synergistic integrated multi-UAV perception system for differentiated tasks. By building a collaborative perception architecture, dynamically adjusting trajectories and synergistic power control, and adopting a multi-agent proximal strategy optimization algorithm, the system dynamically adjusts UAV resource allocation to meet high-precision mission requirements and reduce resource waste in low-precision missions.
It improves the overall task completion rate, reduces system energy consumption, achieves efficient and flexible resource scheduling, and adapts to various perception task requirements.
Smart Images

Figure CN120353255B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of wireless communications and perception, and in particular to a synaesthesia-integrated multi-UAV perception method and system for differentiated tasks. Background Art
[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.
[0003] In the upcoming 5G-Advanced and 6G networks, environmental perception services will play an unprecedented role, supporting key applications such as smart cities and intelligent transportation. Drones (UAVs), with their high maneuverability, flexible deployment, and line-of-sight links, are being integrated into traditional cellular networks to provide flexible and efficient perception services. However, due to the limited computing and storage capabilities of UAVs, communication links are required to transmit perception data to a ground-based central control unit (CU) to facilitate real-time, high-quality decision-making. Integrated Communication and Perception (ISAC) technology enables both communication and radar perception functions on a single hardware and software platform. Integrating ISAC technology into UAVs not only reduces the load on traditional separate communication and perception systems, extending UAV service life, but also enables higher-quality communication and perception services through line-of-sight links between the UAV and the ground. While a single UAV with ISAC functionality has limited coverage, a collaborative communication and perception network constructed with multiple ISAC UAVs can not only perceive large areas and multiple targets in a short period of time, but also utilize multi-source information to improve perception accuracy and achieve higher data rates, significantly improving overall system performance. This allows the ground-based CU to receive more accurate and timely data to aid decision-making.
[0004] While existing research has explored collaborative perception networks for multiple ISAC drones, they suffer from common shortcomings. They often assume that all perception tasks within a mission area have the same type (e.g., target detection or localization), accuracy requirements, and data update frequency. This fails to flexibly allocate ISAC drones and their communication and perception resources to address the diverse characteristics of these tasks. However, in real-world scenarios, multiple perception tasks across large areas often exhibit significant heterogeneity. For example, in smart transportation scenarios, different types of tasks exist, such as pedestrian detection at zebra crossings and vehicle tracking and localization. High-speed or safety-sensitive targets require higher perception accuracy and more frequent data updates, necessitating the coordinated completion of multiple drones. However, static or non-critical targets require lower accuracy and update frequency, which can be met by a single drone. Ignoring these differences between tasks can lead to inappropriate resource allocation: low-demand tasks may be over-allocated with an excessive number of drones and their corresponding communication and perception resources, resulting in redundant perception and data transmission, wasting system energy and resources. High-demand tasks may be under-resourced due to insufficient remaining resources, thus reducing overall mission completion and impacting the perception information quality and decision-making efficiency of the central unit (CU). Summary of the Invention
[0005] To overcome the shortcomings of the above-mentioned existing technologies, the present invention provides a synaesthesia-integrated multi-UAV perception method and system for differentiated tasks, which can assist multiple ISAC UAVs in making autonomous decisions in a collaborative perception environment, dynamically adjust trajectories, task associations and synaesthesia power control schemes, ensure that high-precision tasks obtain more UAV resource support, and low-precision tasks are only allocated necessary resources, thereby effectively avoiding redundant perception and resource waste, improving the overall task completion rate, and reducing system energy consumption, providing an efficient, flexible and energy-saving solution for future intelligent perception applications.
[0006] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:
[0007] In a first aspect, the present invention provides a synaesthesia-integrated multi-UAV perception method for differentiated tasks, comprising:
[0008] Construct a multi-communication and perception integrated UAV collaborative perception system architecture, and propose the communication and perception working mode of the communication and perception integrated UAV under this system architecture;
[0009] Based on the communication and perception working modes, an air-to-ground communication model and an active perception model of a communication-perception integrated UAV are constructed;
[0010] Construct a differentiated perception task performance evaluation system. This evaluation system defines the differentiated task completion rate based on the detection probability of the detection task and the Cramer-Rao lower bound of the localization task. Based on the differentiated task completion rate and the average perception cost, the overall coordinated perception performance index of the network is obtained.
[0011] Under the conditions of satisfying multiple constraints, the optimization problem is obtained by maximizing the task completion rate and minimizing the average perception cost; the optimization problem is modeled as a Markov decision process, and some observation states, actions and rewards are defined;
[0012] The Markov decision process is solved by a multi-agent proximal strategy optimization algorithm combining action branching and masking mechanism, and a multi-communication and perception integrated UAV collaborative perception strategy is obtained.
[0013] In a further technical solution, the differentiated task completion rate is expressed as:
[0014]
[0015] in, express The task completion rate of the time slot, represents the number of differentiation tasks, Indicates a task exist Whether the time slot needs to be executed, Indicates a task Indicates whether the task is completed.
[0016] In a further technical solution, the average perception cost includes flight power consumption, signal transmission power consumption, and signal processing power consumption, and is expressed as:
[0017]
[0018] in, represents the number of communication and perception integrated UAVs, represents the normalized flight power consumption, Indicates the Is the ISAC drone to be Time slots for communication, Indicates the Is the ISAC drone to be Time slot sensing, Indicates the UAV’s communication perception signal transmission power, Indicates the processing power consumption of sensing echo signals.
[0019] Further technical solution, the optimization problem is expressed as:
[0020]
[0021] in, Indicates the ISAC drones in Whether the time slot is selected to execute the task , Indicates the UAV’s communication perception signal transmission power, Indicates the A drone in Flight distance of the time slot, Indicates the flight angle, 、 represents the weighting coefficient, represents the task completion rate, represents the average perceived cost, Indicates the A drone in The horizontal coordinate of the time slot position, Indicates the A drone in The vertical coordinate of the time slot position, Indicates the maximum horizontal coordinate value that the drone can reach. Indicates the maximum horizontal vertical coordinate value that the drone can reach. Indicates the Is the ISAC drone to be Time slots for communication, Indicates the communication rate, Indicates the minimum communication rate threshold, represents the distance between the two drones, Indicates the minimum safe distance between drones. Indicates that the drone is in the time slot The maximum flight distance within Indicates that the drone is in the time slot The maximum transmit power.
[0022] Further technical solutions, A drone in The partial observation state before the start of the time slot is expressed as:
[0023]
[0024] in, 、 、 represents the normalized drone coordinates; Indicates the perception task index performed by the UAV in the previous time slot, Indicates the Is the drone going to Time slot sensing, Indicates the Is the drone going to Time slots for communication, Indicates a task exist Whether the time slot needs to be executed, Marks The perception task Whether the time slot needs to be executed;
[0025] The drone is based on its status before the time slot starts. Get an action that needs to be performed in this time slot, A drone in The actions performed by a time slot are expressed as:
[0026]
[0027] in, 、 、 They represent the discretized values of the drone’s flight angle, flight distance, and signal transmission power respectively.
[0028] A further technical solution is that all drones are simultaneously The time slots execute their respective actions synchronously. After the actions are completed, the corresponding rewards will be fed back according to their decisions. , and the state is transferred to , the reward is expressed as:
[0029]
[0030] in, Indicates that all drones share a portion of the task completion rate and perception cost rewards, 、 、 They represent the individual penalties of drones respectively.
[0031] A further technical solution is that in the multi-agent proximal strategy optimization algorithm that combines action branching and masking mechanism, each communication and perception integrated drone is regarded as an agent, and the agent includes an actor network and a critic network, and the action branching architecture and action masking mechanism are introduced into the actor network.
[0032] In a second aspect, the present invention provides a synaesthesia-integrated multi-UAV perception system for differentiated tasks, comprising:
[0033] The system architecture building module is configured to: build a multi-communication and perception integrated UAV collaborative perception system architecture, and propose the communication and perception working mode of the communication and perception integrated UAV under this system architecture;
[0034] A communication and perception model building module is configured to: build an air-to-ground communication model and an active perception model of the communication and perception integrated UAV based on the communication and perception working mode;
[0035] A performance evaluation system construction module is configured to: construct a differentiated perception task performance evaluation system, wherein the evaluation system defines a differentiated task completion rate based on the detection probability of the detection task and the Cramer-Rao lower bound of the positioning task, and obtains a network overall coordinated perception performance indicator based on the differentiated task completion rate and the average perception cost;
[0036] A decision model building module is configured to: obtain an optimization problem by maximizing the task completion rate and minimizing the average perceived cost under multiple constraints; model the optimization problem as a Markov decision process, defining some observation states, actions, and rewards;
[0037] The optimization decision module is configured to solve the Markov decision process using a multi-agent proximal strategy optimization algorithm that combines action branching and masking mechanisms to obtain a multi-communication perception integrated UAV collaborative perception strategy.
[0038] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the synaesthesia-integrated multi-UAV perception method for differentiated tasks as described in the first aspect.
[0039] In a fourth aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the synaesthesia-integrated multi-UAV perception method for differentiated tasks as described in the first aspect are implemented.
[0040] One or more of the above technical solutions have the following beneficial effects:
[0041] For multiple differentiated perception tasks, the collaborative perception framework proposed in this paper precisely matches the specific requirements of each perception task, ensuring that high-precision tasks receive more UAV resources while allocating only necessary resources to low-precision tasks. This improves overall task completion rates while reducing redundant perception, resource, and energy waste, and overall mission execution costs. This paper addresses the optimization problem of jointly optimizing UAV task association, trajectory planning, and communication perception power control to maximize task completion within the service time while minimizing perception costs. A multi-agent proximal policy optimization algorithm combined with action branching and masking mechanism (MAPPO-ABAM) is proposed to solve this problem. This algorithm deploys a trained actor network on each ISAC UAV through centralized training and distributed execution, enabling autonomous decision-making and dynamic adjustment of trajectory, task association, and communication perception power control schemes. This reduces the computational burden on the central processing unit (CPU), achieving superior performance with high task completion rates and low perception costs.
[0042] The present invention constructs a unified indicator system for differentiated perception tasks, introduces a perception completion rate indicator based on joint detection probability and joint Cramer-Rao lower bound, breaks through the limitations of the traditional single-task evaluation system, and realizes the unified quantification of heterogeneous accuracy requirements across task types and different numbers of associated drones, while defining the execution cost of the perception task; designs a joint optimization model based on spatiotemporal coupling, deeply integrates drone trajectory planning, task association decision-making and synaesthesia power control, and achieves optimal spatiotemporal matching of multi-dimensional resources through three-dimensional constraint modeling of flight energy consumption-perception efficiency-communication quality, ensuring maximization of the completion rate of differentiated perception tasks while minimizing perception costs; and addresses the difficulty of quickly solving the problems brought about by large-scale mixed action spaces and high dynamics of drones. To address the problem, an optimization algorithm based on multi-agent proximal policy optimization combined with action branching and masking mechanism (MAPPO-ABAM) was proposed. The action branching architecture was innovatively used to decouple the hybrid action space, and the dynamic masking mechanism was combined to avoid invalid action exploration. The high complexity of centralized computing was alleviated through a distributed decision-making network. The present invention can assist multiple ISAC drones in making autonomous decisions in a collaborative perception environment, dynamically adjust trajectories, task associations, and synaesthesia power control schemes, ensure that high-precision tasks obtain more drone resource support, and low-precision tasks are allocated only necessary resources, thereby effectively avoiding redundant perception and resource waste, improving the overall task completion rate, and reducing system energy consumption, providing an efficient, flexible, and energy-saving solution for future intelligent perception applications.
[0043] To address the needs of environmental perception applications such as smart cities and intelligent transportation, this paper proposes a multi-UAV collaborative perception framework that integrates communication and perception for differentiated perception tasks. This framework can be used for communication and perception resource scheduling, dynamic joint detection and positioning, and intelligent traffic monitoring. UAVs, enabled by ISAC technology, actively perceive multiple perception tasks within a large target area, varying in type, accuracy requirements, and data update frequency. These perception data are then transmitted to ground-based control units (CUs), enabling accurate and efficient decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0045] Figure 1 Flowchart of the synaesthesia-integrated multi-UAV perception method for differentiated tasks according to an embodiment of the present invention;
[0046] Figure 2 This is an architecture diagram of a multi-communication and perception integrated UAV collaborative perception system according to an embodiment of the present invention;
[0047] Figure 3 This is a diagram of the communication and perception working mode of the communication and perception integrated drone according to an embodiment of the present invention;
[0048] Figure 4 This is a diagram of the algorithm architecture of an embodiment of the present invention based on Multi-Agent Proximal Policy Optimization (MAPPO) combined with Action Branching and Masking Mechanism (ABAM);
[0049] Figure 5 is a reward convergence graph of the algorithm according to an embodiment of the present invention;
[0050] Figure 6 : This is a multi-ISAC UAV trajectory planning result diagram obtained by the method of the embodiment of the present invention;
[0051] Figure 7 This is a comparison chart of the perception task completion rate, objective function value, and perception cost value obtained by the algorithm of the embodiment of the present invention and the existing optimization method. DETAILED DESCRIPTION
[0052] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0053] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0054] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.
[0055] Explanation of professional terms:
[0056] ISAC: integrated communication and perception; CU: central control unit; CRB: Cramer-Rao lower bound; PD: detection probability; MAPPO: multi-agent proximal policy optimization; ABAM: action branching and action mask mechanism.
[0057] To address existing technical challenges, a multi-UAV collaborative perception framework with integrated communication and perception for differentiated tasks is proposed. This framework not only precisely matches the specific requirements of each perception task, improving overall task completion rates, but also avoids resource and energy waste and reduces overall task execution costs. This provides an efficient, flexible, and energy-efficient technical solution for future intelligent applications. Key challenges include: 1) Different perception tasks utilize different performance metrics. Localization tasks are often measured using the Cramer-Rao lower bound (CRB), while detection tasks focus on the probability of detection (PD). Furthermore, to avoid resource waste, high-precision tasks may require the association of multiple UAVs, while low-precision tasks only require a single UAV. Therefore, constructing a unified metric system that reflects overall collaborative perception performance while being compatible with multiple task types and varying numbers of associated UAVs is a key challenge. 2) In a multi-ISAC UAV collaborative perception network, UAVs must simultaneously meet the perception requirements of multiple tasks and the communication requirements for transmitting perception data to a central control unit, while maintaining a safe distance from other UAVs. This makes trajectory design for multiple UAVs highly coupled and complex. 3) The large number of optimization decision variables for multiple UAVs results in a vast combinatorial optimization variable space and high computational complexity. At the same time, the high-speed mobility of drones requires finding the optimal solution quickly within a limited time, which poses a challenge to the algorithm's solution efficiency and real-time performance.
[0058] Example 1
[0059] like Figure 1 As shown, this embodiment discloses a synaesthesia-integrated multi-UAV perception method for differentiated tasks, which includes the following steps:
[0060] S1: Construct a multi-communication and perception integrated UAV collaborative perception system architecture, and propose the communication and perception working mode of the communication and perception integrated UAV under this system architecture;
[0061] In this embodiment, if Figure 2 As shown in the figure, the ISAC UAV collaborative perception system architecture includes: UAVs with ISAC units, the set of all UAVs is , a ground control center (CU), differentiated perception tasks, the set is , triples are used to describe the differences in perception tasks, and the triple set is ,in Represents the type of perception task ( Represents the detection task, represents the positioning task), Represents the perception interval, that is, the task After how many time slots it needs to be sensed again, Represents the perception accuracy requirement, represents the lowest detection probability threshold for detection tasks, and represents the highest Cramer-Rao lower bound threshold for positioning tasks. The service time of the ISAC drone is time slots, and the time slot set is expressed as , among which A drone in The position of the time slot is represented by , where all drones are Flying at a fixed altitude within a time slot , whose mobility model is given by the following formula:
[0062]
[0063] in, For the A drone in The horizontal coordinate of the time slot position, For the A drone in The vertical coordinate of the time slot position, For the A drone in Flight distance of the time slot, is the flight angle.
[0064] In this embodiment, if Figure 3As shown in Figure 2, in each time slot, each ISAC UAV can choose to perform at most one perception task, or not perform any perception task. However, due to the different accuracy requirements of the perception tasks, each perception task can be performed by one or more ISAC UAVs. For example, if the ISAC drone A sensing task is performed in the time slot, then The perception data obtained by this task must be transmitted in the time slot. During the time slot, each ISAC drone may have four working modes as follows:
[0065]
[0066] in, Indicates the Is the ISAC drone to be Time slots for communication, Indicates the Is the ISAC drone to be Time slot sensing.
[0067] S2: Building an air-to-ground communication model and an active perception model for a communication-perception integrated UAV based on the communication and perception working modes;
[0068] In this embodiment, multiple ISAC drones use orthogonal frequency division multiple access to divide the frequency band. The communication and perception functions of a single drone share the frequency band, and its bandwidth is Therefore, the interference between multiple drones is not considered when constructing the air-ground communication model and active perception model. The air-ground communication model mainly refers to the channel model when the ISAC drone transmits the acquired perception data to the ground CU. After the time slot obtains the perception data of the task, it will The communication between the UAV and the ground adopts a practical probabilistic channel model. The probability of line-of-sight channel between the ISAC drone and the CU and non-line-of-sight channel probability The channel gain can be calculated , according to the channel gain, we can get Perception data transmission rate between ISAC drones and CU:
[0069]
[0070] in, Indicates the communication rate, Indicates the UAV’s communication signal transmission power, Represents the noise power.
[0071] At the same time, in order to ensure the timeliness and quality of sensor data transmission, the data transmission rate must be greater than the minimum communication rate threshold .
[0072] The air-ground communication model describes the communication process of the UAV transmitting perception data to the ground CU, and provides the calculation expression and minimum rate threshold of the communication rate (i.e., data transmission rate). Optimization variables such as the UAV's position and transmit power will affect the communication rate.
[0073] In different working modes, the ISAC UAV actively transmits radar sensing signals or interawareness integrated signals, receives sensing echoes reflected from sensing targets and processes the echoes to detect or locate potential targets in the mission area. If the time slot ISAC drones selected for mission For perception, the perception echo power it receives is , according to the perceived echo power, the perceived signal-to-noise ratio can be obtained as:
[0074]
[0075] in, For the ISAC drones in Time slot to task The perceptual signal-to-noise ratio, is the perceptual signal processing gain, represents the perceived noise.
[0076] The active perception model describes the process of the UAV's perception of the ground differentiated perception task, and gives the power and perception signal-to-noise ratio of the received echo. This power and perception signal-to-noise ratio can be used to calculate the detection probability of the detection task and the Cramer-Rao lower bound of the positioning task, thereby further calculating the performance of the differentiated perception task.
[0077] S3: Construct a differentiated perception task performance evaluation system. This evaluation system defines the differentiated task completion rate based on the detection probability of the detection task and the Cramer-Rao lower bound of the localization task. Based on the differentiated task completion rate and the average perception cost, the overall coordinated perception performance index of the network is obtained.
[0078] In this embodiment, a unified indicator system is defined that is compatible with multiple mission types and different numbers of associated drones. Furthermore, a perception task completion rate indicator is derived to evaluate the overall collaborative perception performance of multiple ISAC drones. Furthermore, the cost of the collaborative perception task is defined and weighted with the perception task completion rate to form a weighted indicator, guiding the system to achieve the maximum task completion rate at the minimum cost.
[0079] First, for the detection task, the detection probability is used as an indicator to measure the perception accuracy. The detection probability can represent the probability that the drone successfully detects the target. ISAC drones for missions The detection probability of perception can be expressed as:
[0080]
[0081] in, Indicates the ISAC drones in Whether the time slot is selected to execute the task , if it is 0, then the corresponding detection probability is also 0; Indicates the threshold for determining whether the target exists; is the right tail function of the standard normal distribution.
[0082] Define the task based on the detection probability of a single drone for the task The joint detection probability is:
[0083]
[0084] in, Indicates a task The joint detection probability of .
[0085] Regardless of the task exist This indicator applies whether the time slot is associated with one or multiple drones. When only one drone performs the mission When the rest of the drones are all 0, and what we get at this time is the detection probability of a single drone.
[0086] For positioning tasks, this paper uses the Cramer-Rao lower bound (CRB) to measure the accuracy of perceptual positioning. The Cramer-Rao lower bound is the lower bound of the mean square error of an available unbiased estimator and can be obtained by inverting the Fisher information matrix. Since the position coordinate observations of the perceived target are difficult to obtain directly, its Fisher information matrix can be indirectly solved using the Fisher information matrix of the distance estimation performed by the ISAC UAV for the task and the chain rule. The Fisher information matrix of the coordinate estimation is expressed as:
[0087]
[0088] in, represents the coordinates of the positioning task, Refers to multiple ISAC drones and missions The set of distances between means right The Jacobian matrix of The Fisher information matrix representing the distance, the first 、 The elements can be calculated by the following formula:
[0089]
[0090] in, Representing a collection The elements, and express The covariance matrix of represents the trace of the matrix, represents the inverse of the covariance matrix. The Cramer-Rao lower bound that can be obtained is:
[0091]
[0092] in, and yes The diagonal element values of As a measure of positioning accuracy. Since the coordinate value needs to be estimated using the observation distance, when the task In the time slot When it needs to be sensed, at least two or more drones must be associated to estimate its position. When the number of associated drones exceeds two, no matter how many there are, the above However, if only one drone is associated, the data is invalid. Set to a large value to indicate that the task is not completed.
[0093] Based on the above indicators, the present invention defines the differentiated task completion rate as an indicator for evaluating the overall perception performance. First, define the task Completed indicator , expressed as:
[0094]
[0095] in, Indicates the accuracy threshold. For detection tasks, if the detection probability is greater than the required detection probability accuracy, then When it is 1, it indicates that the task is completed. For positioning tasks, if Less than a given accuracy threshold , then it marks the completion of the task. The overall task completion rate of the time slot can be obtained by the following formula:
[0096]
[0097] in, Indicates a task exist Whether the time slot needs to be executed. Note here It represents the effective task completion rate. Since the perception intervals of multiple tasks are different, if the task exist If a time slot does not need to be executed, it will not be counted in the task completion rate even if the perception accuracy requirement is met.
[0098] To measure the cost of ISAC UAVs when performing sensing tasks, multiple UAVs were The average perception cost of a time slot is composed of three parts: flight power consumption, signal transmission power consumption, and signal processing power consumption, and is calculated using the following formula:
[0099]
[0100] Since the magnitude of flight power consumption and signal transmission power consumption is quite different, the flight power consumption is normalized and expressed as , Indicates the UAV’s communication perception signal transmission power, Represents the processing power consumption of the perception echo. Once a drone transmits a perception signal or ISAC signal, a perception echo processing power consumption will be generated.
[0101] Eventually The weighted sum of the completion rate of the time slot sensing task and the average sensing cost is used as an indicator to evaluate the overall collaborative sensing performance of the network and the objective function of the following optimization problem, which is expressed as the following formula:
[0102]
[0103] in, and is the weighting coefficient.
[0104] S4: Under multiple constraints, the optimization problem is to maximize the task completion rate and minimize the average perception cost. The optimization problem is modeled as a Markov decision process, defining some observation states, actions, and rewards.
[0105] In this embodiment, based on the above-mentioned multi-ISAC UAV cooperative perception system architecture, combined with the ISAC UAV air-ground communication model and active perception model, the optimal UAV and task association solution is solved. , signal transmission power and trajectory planning variables and , obtain the optimal multi-ISAC UAV collaborative perception strategy for differentiated tasks. Under the multiple constraints of communication rate, multi-UAV safety collision avoidance constraints, flight constraints and maximum transmission power, maximize the overall collaborative perception performance index of the network and achieve multiple differentiated perception tasks in The maximum task completion rate within each service time slot is achieved while minimizing the perception cost, avoiding allocating too many resources to low-demand tasks, which would result in redundant perception and resource waste. In other words, maximizing the task completion rate and minimizing the average perception cost represents the best overall collaborative perception performance of the network. The optimization problem is formulated as follows:
[0106]
[0107] in, Indicates the maximum horizontal coordinate value that the drone can reach. Indicates the maximum horizontal vertical coordinate value that the drone can reach. represents the distance between the two drones, Indicates the minimum safe distance between drones. Indicates that the drone is in the time slot The maximum flight distance within Indicates that the drone is in the time slot The maximum transmit power.
[0108] Among them, the optimization variables and constraints are All are applicable, among which C1 limits the flight range of the drone to a fixed area , C2 is the data transmission rate constraint to ensure the timeliness and quality of communication, C3 is the safety distance constraint of multiple drones, where Minimum interval, C4 is the flight distance constraint of the UAV in a single time slot, which cannot exceed , C5 is the UAV flight angle constraint, C6 constrains the UAV to send communication signals, perception signals or synaesthesia integration signals, the power cannot exceed .
[0109] Since each ISAC drone can only observe part of the state and cannot obtain the global state of the collaborative perception network, including the positions of other drones, the present invention constructs a partially observed Markov decision process to simulate the task execution process of multiple ISAC drones in multiple time slots. The partially observed Markov process includes a state space, an action space, and a reward function.
[0110] Each UAV will have a partial observation of itself before the time slot begins. A drone in The observed state before the start of the time slot is expressed as:
[0111]
[0112] in, 、 、 represents the normalized drone coordinates; Indicates the index of the perception task performed by the drone in the previous time slot. If no perception is performed, then ; Marks The perception task Whether the time slot needs to be executed. The partial observation set of all drones is the system in The overall state before the time slot starts .
[0113] The drone is based on its status before the time slot starts. An action will be obtained that needs to be performed in the time slot. A drone in The action performed in a time slot includes the drone's flight angle, distance, the perception task index performed in this time slot (including the case where perception is not performed), and the signal transmission power, which can be expressed as:
[0114]
[0115] in, 、 、 Respectively represent the discretized values of the drone’s flight angle, flight distance, and signal transmission power. First, the flight is executed according to the decision of its flight angle and distance in the time slot. After arriving at the target position, the and Determine whether hovering communication and perception are required in this time slot, and then perform the corresponding operation. Since the optimization variables involve the mixed solution of discrete and continuous variables, the computational complexity is high, so the flight angle, flight distance and transmission power of the UAV are discretized and used respectively. 、 、 represents the discretized value, where 、 、 Represent the discretization series respectively, all UAVs are The action set of a time slot is expressed as .
[0116] All drones are simultaneously The time slots execute their respective actions synchronously. After the actions are completed, the system will feedback the corresponding rewards based on their decisions. , and the state is transferred to The reward is calculated using the following formula:
[0117]
[0118] Among them, all drones share a part of the task completion rate and perception cost rewards , calculated by the differentiated perception task performance evaluation system given above. In addition, 、 、 Represents the individual penalties of the drone, where Used to constrain the communication power not to be lower than the threshold, Penalize drones that fly beyond the boundaries, Penalties are imposed for drones that collide with other drones.
[0119] exist During each time slot, each drone continuously executes the following loop: observe the current state, select and execute an action based on the state, obtain the corresponding reward, and then transfer to the state of the next time slot. This process continues until the service time ends.
[0120] S5: A multi-agent proximal strategy optimization algorithm combining action branching and masking mechanism is used to solve the Markov decision process, and a multi-communication and perception integrated UAV collaborative perception strategy is obtained;
[0121] In this embodiment, if Figure 4 As shown in the figure, the present invention adopts the multi-agent proximal policy optimization (MAPPO-ABAM) algorithm combined with action branching and mask mechanism to train the decision model to find the optimal action based on the drone's own partial observation variables, thereby effectively guiding the collaborative perception and communication services of multiple ISAC drones for differentiated perception tasks.
[0122] In the MAPPO-ABAM algorithm, each ISAC drone is considered as an intelligent agent, which contains an actor network. and critic network, the actor network is the core network that guides the drone’s actions, and its input is the observed state of the drone , the output is the action of the drone , while the critic network is used to evaluate the long-term rewards of state-action pairs. This invention uses a three-layer actor network and critic network with a structure of input layer-fully connected layer (MLP)-recurrent neural network layer (GRU)-MLP-output layer. The MLP hidden layer contains 128 neurons, while the GRU hidden layer contains 64 neurons. The algorithm has two stages: centralized training and distributed execution. The centralized training is mainly used to collect global states and rewards and calculate the state value function. , to assist in training the optimal strategy. After the training is completed, the actor network can be directly deployed to each drone. Each intelligent agent only relies on the trained actor network and local observations to make independent decisions, reducing the computational pressure of the central computing unit.
[0123] During the centralized training phase, the present invention uses clipping loss, mean square error loss, and entropy reward for training, which are used to limit the update amplitude of the actor network, optimize the state value function estimation, and improve the network's exploratory ability. Clipping loss is calculated by the following formula:
[0124]
[0125] in, represents the approximate expectation, is the ratio of the new and old strategies, is the generalized odds estimate (GAE), is the clipping threshold. The present invention introduces an action branch architecture into the actor network, so that actions of different dimensions have independent output branches, avoiding the dimensionality explosion problem caused by dimension combination and reducing the training difficulty. The sampling samples are used to approximate the loss expectation, for each time step Extract data from the experience pool (batch) and use this data to estimate the expectation of the entire loss.
[0126] For each action dimension, such as flight angle, the ratio of the new and old strategies is calculated as follows:
[0127]
[0128] in, represents the current actor network, Represents the old actor network.
[0129] The clipping losses corresponding to other action dimensions are calculated in the same way, and the losses of all action dimensions are averaged to obtain the overall clipping loss function. The overall loss function obtained after adding the mean square error loss and entropy reward is:
[0130]
[0131] in, is the entropy reward, is the mean square error loss, 、 Represent the hyperparameters that control the weights of entropy reward and value function loss, Indicates the action dimensions.
[0132] According to the overall loss function, the present invention uses the gradient descent method to update the parameters of the actor network and the critic network to obtain the optimal actor network. During the training process, the present invention uses multiple rounds of iterative optimization. Specifically, in each training round, the agent executes the strategy in the environment, collects trajectory data containing state, action, reward and state value estimation, and stores it in the trajectory storage pool. Subsequently, the trajectory data is sampled from the storage pool, the advantage estimate and the target state value are calculated, and the actor network and the critic network are gradient updated using the above-mentioned loss function. This process continues to iterate until the strategy converges. In addition, the present invention introduces an action mask mechanism in the actor network to avoid the agent from selecting invalid actions. For example, in the task selection action branch, when a task does not require perception in the current time slot, the probability of the corresponding action is set to a minimum value. (like ), and then renormalize the remaining action probabilities to ensure reasonable sampling of valid actions.
[0133] During the distributed execution phase, the trained actor network is deployed to each drone. At the beginning of each time slot, all drones observe and acquire a local state vector, including the drone's position, mission status, and communication requirements. This vector is then fed into the actor network to obtain an action value. Each drone then makes autonomous decisions based on its actor network. After receiving the action instruction, each drone will fly according to the flight model for one second. After the flight, the drone will hover at the target location and determine whether to communicate in the current time slot based on the performance of the perception task in the previous time slot. Simultaneously, based on the action instruction in the current time slot, the drone will decide whether to perform the perception task and which specific task to perform. The drone will then perform hovering communication and perception according to the corresponding operating mode. The hovering communication and perception process will refer to the ISAC drone air-to-ground communication model and active perception model described above. For drones selected to perform the perception task in the current time slot, they will receive perception echo data after sending a signal and transmit the data to the CU via the communication link in the next time slot. After the action in this time slot is executed, the system state changes, and the drone enters the next time slot, repeating the action acquisition-action execution cycle. This cycle continues until service time slots completed.
[0134] The decision variables in the Markov decision process are the drone's flight angle, flight distance, and synaesthesia signal transmission power in each time slot. Furthermore, the decision variables include whether to perform perception in each time slot, and if so, which perception task to select, along with the task index. In the MAPPO-ABAM algorithm, each drone is considered an intelligent agent and equipped with a trained actor network. This actor network enables each drone to output the aforementioned decision variables based on its partial observation state at the beginning of the time slot, enabling autonomous decision-making. These decision variables correspond to dynamic trajectory adjustment, task association, and communication perception power control schemes.
[0135] Experimental description:
[0136] The present invention implements the proposed communication perception integrated multi-UAV collaborative perception framework for differentiated tasks. The simulation platform used in the experiment is Pycharm 2022.2.2, the code language is Python3.9, and the deep learning framework used is pytorch2.4.1. The simulation environment consists of 4 UAVs and 4 differentiated tasks, including two detection tasks and two positioning tasks. The two positioning tasks are static and dynamic, and the detection tasks are also. In addition, the perception accuracy requirements and perception frequencies of each task are different. The MAPPO-ABAM optimization algorithm proposed in the present invention can be trained to convergence in a simulation environment. Figure 5 This is the convergence curve diagram of the algorithm proposed in this invention and two comparative benchmark algorithms, namely the PPO algorithm and the MADDPG algorithm. Figure 6This is the drone image obtained when the trained actor network is deployed on multiple ISAC drones for autonomous decision-making. Figure 7 The performance of two benchmark algorithms, namely the proximal policy optimization algorithm (PPO) and the multi-agent deep deterministic policy gradient algorithm (MADDPG), and the MAPPO-ABAM algorithm proposed in this invention, in terms of differentiated perception task completion rate, perception cost, and objective function value (i.e., the weighted function value of the perception task completion rate and perception cost) within the service time are compared.
[0137] from Figure 5 It can be observed that the MAPPO-ABAM algorithm proposed in this paper has the fastest convergence speed compared to the two benchmark algorithms. This is mainly due to the introduction of the action branch architecture in the actor network. This architecture improves the learning flexibility of multi-dimensional actions by decoupling the multi-dimensional action space, while avoiding the dimensionality explosion problem caused by large-scale combined action spaces, enabling the algorithm to converge quickly and maintain excellent performance in complex scenarios. Figure 6 It can be seen that under the preset differentiated tasks, multiple ISAC drones can autonomously form a task group based on the task location. Drone 3 hovers near Task 2 to perform low-precision detection; at the same time, Drone 1, Drone 2, and Drone 4 continue to approach Task 1 and Task 3, and collaboratively perceive to achieve high-precision detection and positioning. In addition, Drone 1 and Drone 4 will also adjust their trajectories appropriately to stay close to Task 4 to provide low-precision positioning services when needed. Figure 7 It can be seen that the proposed MAPPO-ABAM algorithm achieves the highest perception task completion rate and the lowest perception cost compared to the two baseline algorithms. MAPPO-ABAM's actor and critic networks both utilize GRU hidden layers to capture temporal dependencies and optimize long-term decision making. In a collaborative perception environment, the GRU layers are able to learn the perception interval patterns of differentiated tasks and plan drone trajectories in advance, thereby reducing flight energy consumption. Furthermore, the proposed action masking mechanism effectively prevents ISAC drones from being associated with perception tasks that are not required in the current time slot, fundamentally eliminating ineffective task selection, preventing resource waste, and ensuring that critical tasks are prioritized. Results show that within a given service time, the proposed solution achieves a 96.8085% task completion rate at the lowest perception cost, compared to only 62.7660% and 77.6596% for the MADDPG and PPO algorithms, respectively. In the collaborative perception scenario of multiple drones with integrated communication and perception for differentiated tasks, the proposed solution demonstrates significant performance advantages.
[0138] Example 2
[0139] This embodiment discloses a synaesthesia-integrated multi-UAV perception system for differentiated tasks, including:
[0140] The system architecture building module is configured to: build a multi-communication and perception integrated UAV collaborative perception system architecture, and propose the communication and perception working mode of the communication and perception integrated UAV under this system architecture;
[0141] A communication and perception model building module is configured to: build an air-to-ground communication model and an active perception model of the communication and perception integrated UAV based on the communication and perception working mode;
[0142] A performance evaluation system construction module is configured to: construct a differentiated perception task performance evaluation system, wherein the evaluation system defines a differentiated task completion rate based on the detection probability of the detection task and the Cramer-Rao lower bound of the positioning task, and obtains a network overall coordinated perception performance indicator based on the differentiated task completion rate and the average perception cost;
[0143] A decision model building module is configured to: obtain an optimization problem by maximizing the task completion rate and minimizing the average perceived cost under multiple constraints; model the optimization problem as a Markov decision process, defining some observation states, actions, and rewards;
[0144] The optimization decision module is configured to solve the Markov decision process using a multi-agent proximal strategy optimization algorithm that combines action branching and masking mechanisms to obtain a multi-communication perception integrated UAV collaborative perception strategy.
[0145] Example 3
[0146] The purpose of this embodiment is to provide a computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method of embodiment 1 when executing the program.
[0147] Example 4
[0148] The purpose of this embodiment is to provide a computer-readable storage medium, a computer-readable storage medium having a computer program stored thereon, which performs the steps of the method of embodiment 1 when executed by a processor.
[0149] The steps involved in the apparatuses of Examples 3 and 4 above correspond to those of Method Example 1. For detailed implementation, please refer to the relevant description of Example 1. The term "computer-readable storage medium" should be understood to mean a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and causing the processor to perform any of the methods of the present invention.
[0150] Those skilled in the art will appreciate that the modules or steps of the present invention described above can be implemented using a general-purpose computer device. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. Alternatively, they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0151] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
[0152] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.
Claims
1. A synaesthesia-integrated multi-UAV perception method for differentiated tasks, characterized by: include: Construct a multi-communication and perception integrated UAV collaborative perception system architecture, and propose the communication and perception working mode of the communication and perception integrated UAV under this system architecture; Based on the communication and perception working modes, an air-to-ground communication model and an active perception model of a communication-perception integrated UAV are constructed; Construct a differentiated perception task performance evaluation system. This evaluation system defines the differentiated task completion rate based on the detection probability of the detection task and the Cramer-Rao lower bound of the localization task. Based on the differentiated task completion rate and the average perception cost, the overall coordinated perception performance index of the network is obtained. Under the conditions of satisfying multiple constraints, the optimization problem is obtained by maximizing the task completion rate and minimizing the average perception cost. The optimization problem is modeled as a Markov decision process, defining some observation states, actions and rewards. The optimization problem is expressed as: in, Indicates the ISAC drones in Whether the time slot is selected to execute the task , Indicates the UAV’s communication perception signal transmission power, Indicates the A drone in Flight distance of the time slot, Indicates the flight angle, 、 represents the weighting coefficient, represents the task completion rate, represents the average perceived cost, Indicates the A drone in The horizontal coordinate of the time slot position, Indicates the A drone in The vertical coordinate of the time slot position, Indicates the maximum horizontal coordinate value that the drone can reach. Indicates the maximum horizontal vertical coordinate value that the drone can reach. Indicates the Is the ISAC drone to be Time slots for communication, Indicates the communication rate, Indicates the minimum communication rate threshold, Indicates the distance between the two drones, Indicates the minimum safe distance between drones. Indicates that the drone is in the time slot The maximum flight distance within Indicates that the drone is in the time slot Maximum transmit power; The Markov decision process is solved by a multi-agent proximal strategy optimization algorithm combining action branching and masking mechanism, and a multi-communication and perception integrated UAV collaborative perception strategy is obtained.
2. The synaesthesia-integrated multi-UAV perception method for differentiated tasks according to claim 1 is characterized in that: The differentiation task completion rate is expressed as: in, express The task completion rate of the time slot, represents the number of differentiation tasks, Indicates a task exist Whether the time slot needs to be executed, Indicates a task Indicates whether the task is completed.
3. The synaesthesia-integrated multi-UAV perception method for differentiated tasks according to claim 1, characterized in that: The average perception cost includes flight power consumption, signal transmission power consumption, and signal processing power consumption, and is expressed as: in, represents the number of communication and perception integrated UAVs, represents the normalized flight power consumption, Indicates the Is the ISAC drone to be Time slots for communication, Indicates the Is the ISAC drone to be Time slot sensing, Indicates the UAV’s communication perception signal transmission power, Indicates the processing power consumption of sensing echo signals.
4. The synaesthesia-integrated multi-UAV perception method for differentiated tasks according to claim 1, characterized in that: No. A drone in The partial observation state before the start of the time slot is expressed as: in, 、 、 represents the normalized drone coordinates; Indicates the perception task index performed by the UAV in the previous time slot, Indicates the Is the drone going to Time slot sensing, Indicates the Is the drone going to Time slots for communication, Indicates a task exist Whether the time slot needs to be executed, Marks The perception task Whether the time slot needs to be executed; The drone is based on its status before the time slot starts. Get an action that needs to be performed in this time slot, A drone in The actions performed by a time slot are expressed as: in, 、 、 They represent the discretized values of the drone’s flight angle, flight distance, and signal transmission power respectively.
5. The synaesthesia-integrated multi-UAV perception method for differentiated tasks according to claim 1, characterized in that: All drones are simultaneously The time slots execute their respective actions synchronously. After the actions are completed, the corresponding rewards will be fed back according to their decisions. , and the state is transferred to , the reward is expressed as: in, Indicates that all drones share a portion of the task completion rate and perception cost rewards, 、 、 They represent the individual penalties of drones respectively.
6. The synaesthesia-integrated multi-UAV perception method for differentiated tasks according to claim 1, characterized in that: In the multi-agent proximal strategy optimization algorithm combining action branching and masking mechanism, each communication and perception integrated UAV is regarded as an agent, and the agent includes an actor network and a critic network. The action branching architecture and action masking mechanism are introduced into the actor network.
7. The synaesthesia-integrated multi-UAV perception system for differentiated tasks is characterized by: include: The system architecture building module is configured to: build a multi-communication and perception integrated UAV collaborative perception system architecture, and propose the communication and perception working mode of the communication and perception integrated UAV under this system architecture; A communication and perception model building module is configured to: build an air-to-ground communication model and an active perception model of the communication and perception integrated UAV based on the communication and perception working mode; A performance evaluation system construction module is configured to: construct a differentiated perception task performance evaluation system, wherein the evaluation system defines a differentiated task completion rate based on the detection probability of the detection task and the Cramer-Rao lower bound of the positioning task, and obtains a network overall coordinated perception performance indicator based on the differentiated task completion rate and the average perception cost; The decision model building module is configured to: maximize the network task completion rate and minimize the average perception cost under multiple constraints to obtain an optimization problem; model the optimization problem as a Markov decision process, define some observation states, actions and rewards; the optimization problem is expressed as: in, Indicates the ISAC drones in Whether the time slot is selected to execute the task , Indicates the UAV’s communication perception signal transmission power, Indicates the A drone in Flight distance of the time slot, Indicates the flight angle, 、 represents the weighting coefficient, represents the task completion rate, represents the average perceived cost, Indicates the A drone in The horizontal coordinate of the time slot position, Indicates the A drone in The vertical coordinate of the time slot position, Indicates the maximum horizontal coordinate value that the drone can reach. Indicates the maximum horizontal vertical coordinate value that the drone can reach. Indicates the Is the ISAC drone to be Time slots for communication, Indicates the communication rate, Indicates the minimum communication rate threshold, Indicates the distance between the two drones, Indicates the minimum safe distance between drones. Indicates that the drone is in the time slot The maximum flight distance within Indicates that the drone is in the time slot Maximum transmit power; The optimization decision module is configured to solve the Markov decision process using a multi-agent proximal strategy optimization algorithm that combines action branching and masking mechanisms to obtain a multi-communication perception integrated UAV collaborative perception strategy.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the synaesthesia-integrated multi-UAV perception method for differentiated tasks as described in any one of claims 1 to 6 are implemented.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the synaesthesia-integrated multi-UAV perception method for differentiated tasks as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method for remotely searching vehicle and vehicle
CN115334148A
RIS-assisted unmanned aerial vehicle network-based communication and sensing integrated system and method
CN119277319A