Beam adjustment method and device for multiple unmanned aerial vehicles, and electronic equipment
By dynamically adjusting the base station beam scheme and using the policy network and reward function to optimize the drone service decision, the problems of discontinuous coverage and unstable signal quality in drone network communications are solved, achieving more efficient resource utilization and stable communication quality.
Patent Information
- Application Number
- CN202511212119.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-27
AI Technical Summary
In the scenario of multi-UAV collaborative operation, UAV network communication has problems such as discontinuous communication coverage, unstable signal quality and inefficient resource scheduling, which existing technologies have failed to effectively solve.
By obtaining the pre-assigned beam scheme and location information of the UAV, the target encoder in the policy network is used to process the state vector, generate the association probability matrix and beam parameter vector, and dynamically adjust the base station beam to optimize the UAV service decision. The reward function is combined to optimize the policy network and value network to achieve intelligent optimization of the beam parameters.
It improves communication quality and resource utilization efficiency, enhances beam tracking stability and response speed, and solves the problems of discontinuous communication coverage and unstable signal quality of static beam solutions under high-speed movement of multiple UAVs and complex scenarios.
Smart Images

Figure CN120730318A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of wireless communication technology, and more specifically, to a beam adjustment method, device, and electronic equipment for multiple UAVs. Background Art
[0002] With the rapid development of drone technology and the rise of the low-altitude economy, aerial drone networks have gradually become a vital communications infrastructure connecting the ground and airspace, supporting a wide range of aerial applications such as logistics distribution, intelligent inspections, and real-time live broadcasting. However, the highly dynamic nature of drones and the complex and ever-changing airspace environment place higher demands on the reliability, coverage, and data transmission speed of drone network communications.
[0003] On the one hand, the flight trajectory of drones is changeable. Unlike the linear coverage optimization of high-speed railways or highways, it is difficult for operators to predict the precise flight path of drones, resulting in the beam tracking and optimization of ground base stations being difficult to adapt to the high-speed movement requirements of drones.
[0004] On the other hand, drones usually have the predictability of preset routes and trajectories. However, current aerial networks fail to fully utilize this feature, especially in scenarios where multiple drones work together. They still use relatively rigid beam configuration methods, such as fixed base station beams or static cell strategies, which hinder the improvement of signal quality and resource scheduling efficiency.
[0005] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0006] The embodiments of the present application provide a beam adjustment method, device, and electronic device for multiple UAVs to at least solve the technical problems of discontinuous communication coverage, unstable signal quality, and low resource scheduling efficiency caused by the static beam solution in related technologies under high-speed movement and complex scenarios of multiple UAVs.
[0007] According to one aspect of an embodiment of the present application, a beam adjustment method for multiple UAVs is provided, including: obtaining a pre-allocated beam scheme for the UAVs and position information of the UAVs during flight, wherein the pre-allocated beam scheme includes an initial beam allocation scheme corresponding to multiple UAVs; determining a state vector of the UAV based on the position information, and processing the state vector through a target encoder in a policy network to obtain an association probability matrix and a beam parameter vector, wherein the policy network is used to adjust the service decision of the base station beam for the UAV, the target encoder semantically enhances the state vector through a multi-head attention mechanism, and outputs an association probability matrix and a beam parameter vector through different branches, the association probability matrix is used to reflect the service relationship between the UAV and the base station beam, and the beam parameter vector is used to reflect the parameter configuration of the base station beam; and adjusting the pre-allocated beam scheme based on the association probability matrix and the beam parameter vector.
[0008] Optionally, determining the state vector of the UAV based on the position information includes: determining a first vector, a second vector and a third vector respectively based on the position information, wherein the first vector is used to reflect the angular coordinate information and speed change information of the UAV relative to the base station, the second vector is used to reflect the historical beam parameter information of the UAV, and the third vector is used to reflect the angle difference between the UAV and the base station beam; fusing the first vector, the second vector and the third vector to obtain the state vector.
[0009] Optionally, the state vector is processed by the target encoder in the policy network to obtain an association probability matrix and a beam parameter vector, including: determining the input features corresponding to the state vector, and determining the first feature matrix input to the policy network based on the input features; processing the first feature matrix by the target encoder in the policy network to output a second feature matrix; normalizing the second feature matrix by the first branch in the policy network to obtain an association probability matrix; and performing a pooling operation on the second feature matrix by the second branch in the policy network to obtain a beam parameter vector.
[0010] Optionally, the policy network is trained in the following manner: obtaining a simulation state vector and a simulation action vector of the UAV, wherein the simulation action vector is a simulation association probability matrix and a simulation beam parameter vector generated by a pre-allocation algorithm; processing the simulation state vector through the initial policy network to obtain a predicted action vector, wherein the predicted action vector includes the predicted association probability matrix and the predicted beam parameter vector output by the initial policy network; determining a first loss function required for training the initial policy network, and determining a first loss between the simulation action vector and the predicted action vector through the first loss function; based on the first loss, using the gradient descent method to determine the optimal parameters of the initial policy network to obtain the policy network.
[0011] Optionally, the method also includes: determining the action vector of the UAV based on the association probability matrix and the beam parameter vector, wherein the action vector is used to reflect the beam service decision taken by the UAV; splicing the action vector and the state vector to obtain a spliced vector, and processing the spliced vector through a value network to obtain an evaluation value, wherein the value network is used to evaluate the value of the beam service decision taken by the UAV; optimizing the strategy network and the value network based on the evaluation value and the reward function, wherein the reward function is used to quantify the signal quality, effective coverage and beam stability of the UAV after adopting the target beam allocation scheme, and the target beam allocation scheme is a pre-allocated beam scheme adjusted according to the association probability matrix and the beam parameter vector.
[0012] Optionally, the reward function is determined in the following manner: determining the service probability of the target beam to the target UAV, and determining the signal receiving power after the target UAV accesses the target beam, and determining a first reward function based on the service probability and the signal receiving power, wherein the target UAV is any one of the multiple UAVs, the target beam is the optimal beam corresponding to the target UAV in the target beam scheme, and the first reward function is used to quantify the signal quality received by the target UAV from the target beam; determining a second reward function based on the signal receiving power and the preset probability, wherein the second reward function is used to balance the signal strength and effective coverage of the target beam to the target UAV; determining a third reward function based on the beam configuration and historical beam configuration in the beam parameter vector, wherein the third reward function is used to suppress changes in beam parameters; and weighting the first reward function, the second reward function, and the third reward function to obtain a reward function.
[0013] Optionally, optimizing the policy network and the value network based on the evaluation value and the reward function includes: determining a target evaluation value based on the reward function, wherein the target evaluation value is used to represent the expected evaluation value under the action vector; determining a second loss function required for optimizing the value network, and determining a second loss between the evaluation value and the target evaluation value through the second loss function; optimizing the value network based on the second loss and a preset soft update rate, wherein the preset soft update rate is used to smooth the update process of the value network and the policy network; and determining a third loss function required for optimizing the policy network, and determining a third loss under the action vector through the third loss function; optimizing the policy network based on the third loss and the preset soft update rate.
[0014] According to another aspect of an embodiment of the present application, a beam adjustment device for multiple UAVs is also provided, including: an acquisition module for acquiring a pre-allocated beam scheme of the UAV and the position information of the UAV during flight, wherein the pre-allocated beam scheme includes an initial beam allocation scheme corresponding to multiple UAVs; a processing module for determining the state vector of the UAV based on the position information, and processing the state vector through a target encoder in a strategy network to obtain an association probability matrix and a beam parameter vector, wherein the strategy network is used to adjust the service decision of the base station beam to the UAV, the target encoder semantically enhances the state vector through a multi-head attention mechanism, and outputs an association probability matrix and a beam parameter vector through different branches, the association probability matrix is used to reflect the service relationship between the UAV and the base station beam, and the beam parameter vector is used to reflect the parameter configuration of the base station beam; an adjustment module is used to adjust the pre-allocated beam scheme based on the association probability matrix and the beam parameter vector.
[0015] According to another aspect of an embodiment of the present application, an electronic device is also provided, including: a memory and a processor, wherein the memory is used to store program instructions; the processor is connected to the memory and is used to execute the above-mentioned beam adjustment method for multiple drones.
[0016] According to another aspect of an embodiment of the present application, a non-volatile storage medium is further provided, wherein the non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the above-mentioned beam adjustment method for multiple UAVs by running the computer program.
[0017] According to another aspect of the embodiments of the present application, a computer program product is also provided, including computer instructions, which, when executed by a processor, implement the above-mentioned beam adjustment method for multiple drones.
[0018] In an embodiment of the present application, a pre-allocated beam scheme of a UAV and the position information of the UAV during flight are obtained, wherein the pre-allocated beam scheme includes an initial beam allocation scheme corresponding to multiple UAVs; a state vector of the UAV is determined based on the position information, and the state vector is processed by a target encoder in a policy network to obtain an association probability matrix and a beam parameter vector, wherein the policy network is used to adjust the service decision of the base station beam for the UAV, the target encoder semantically enhances the state vector through a multi-head attention mechanism, and outputs an association probability matrix and a beam parameter vector through different branches, the association probability matrix is used to reflect the service relationship between the UAV and the base station beam, and the beam parameter vector is used to reflect the parameter configuration of the base station beam; the pre-allocated beam scheme is adjusted according to the association probability matrix and the beam parameter vector, thereby achieving the purpose of intelligently optimizing the base station beam service decision for multiple UAVs, thereby achieving the technical effect of improving communication quality and resource utilization efficiency, enhancing beam tracking stability and response speed, and thus solving the technical problems of discontinuous communication coverage, unstable signal quality and low resource scheduling efficiency in the static beam scheme in the related art under high-speed movement and complex scenarios of multiple UAVs. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0020] Figure 1 1 is a hardware structure diagram of a computer terminal for implementing a beam adjustment method for multiple UAVs according to an embodiment of the present application;
[0021] Figure 2 is a flow chart of a beam adjustment method for multiple UAVs according to an embodiment of the present application;
[0022] Figure 3 is a structural diagram of an Actor-Critic network according to an embodiment of the present application;
[0023] Figure 4 This is a structural diagram of a beam adjustment device for multiple UAVs according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0026] First, some nouns or terms that appear in the process of explaining the embodiments of this application are subject to the following explanations:
[0027] SSB (Synchronization Signal Block): A signal structure in 5G communication systems used by terminal devices for initial access and cell measurement. The SSB includes the Primary Synchronization Signal (PSS), the Secondary Synchronization Signal (SSS), and the Physical Broadcast Channel (PBCH). These signals and channels help terminals synchronize and receive basic cell broadcast information.
[0028] Markov Decision Process (MDP): A mathematical model used to describe how decisions are made based on the current state to achieve optimal outcomes in an uncertain environment. MDP is a fundamental concept in reinforcement learning. It defines the state space, action space, transition probabilities, and reward function, providing a framework for intelligent agents to learn optimal strategies in an environment.
[0029] RSRP (Reference Signal Received Power): A key metric for measuring network signal quality in 4G and 5G communication systems. In wireless networks, higher RSRP values indicate stronger reference signal power received by the terminal, generally indicating better communication quality.
[0030] DDPG (Deep Deterministic Policy Gradient): An algorithm used in reinforcement learning for processing continuous action spaces. Compared to traditional policy-based reinforcement learning methods, DDPG combines an actor-critic architecture and uses two deep neural networks to learn the policy and value function, respectively. It is particularly suitable for solving continuous control problems in complex environments.
[0031] MQTT (Message Queuing Telemetry Transport): A lightweight publish / subscribe-based communication protocol widely used in the Internet of Things (IoT) and remote data transmission scenarios.
[0032] In order to solve the problem of poor transmission efficiency of drone network communication in related technologies, the present invention provides a multi-drone beam adjustment method, which can be run on Figure 1 Among the computer terminals shown, the computer terminal will be described below.
[0033] The beam adjustment method embodiment of multiple drones provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal for implementing a beam adjustment method for multiple UAVs is shown. Figure 1 As shown, the computer terminal 10 may include one or more processors (illustrated as 102a, 102b, ..., 102n in the figure) (the processor may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions via a wired and / or wireless network connection. In addition, it may also include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. Those skilled in the art will understand that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1More or fewer components than shown, or with Figure 1 Different configurations shown.
[0034] It should be noted that the one or more processors and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be fully or partially integrated into any of the other components of the computer terminal 10. As discussed in the embodiments of the present application, the data processing circuitry functions as a processor control (e.g., the selection of a variable resistor terminal path connected to an interface).
[0035] The memory 104 can be used to store software programs and modules for application software, such as the program instructions / data storage device corresponding to the beam adjustment method for multiple drones in the embodiments of the present application. The processor executes the software programs and modules stored in the memory 104 to perform various functional applications and data processing, thereby implementing the beam adjustment method for multiple drones described above. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include memory remotely located relative to the processor, and these remote memories may be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0036] The transmission module 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission module 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission module 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0037] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .
[0038] It should be noted that, in some optional embodiments, the above Figure 1 The computer terminal shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the computer terminal described above.
[0039] In the above-mentioned operating environment, an embodiment of the present application provides an embodiment of a beam adjustment method for multiple UAVs. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0040] Figure 2 This is a flow chart of a beam adjustment method for multiple UAVs according to an embodiment of the present application. Figure 2 As shown, the method includes the following steps:
[0041] Step S202: obtaining a pre-allocated beam plan of the UAV and location information of the UAV during flight, wherein the pre-allocated beam plan includes initial beam allocation plans corresponding to multiple UAVs.
[0042] In step S204, the state vector of the UAV is determined based on the position information, and the state vector is processed by the target encoder in the policy network to obtain the association probability matrix and the beam parameter vector. The policy network is used to adjust the service decision of the base station beam to the UAV. The target encoder semantically enhances the state vector through the multi-head attention mechanism and outputs the association probability matrix and beam parameter vector through different branches. The association probability matrix is used to reflect the service relationship between the UAV and the base station beam, and the beam parameter vector is used to reflect the parameter configuration of the base station beam.
[0043] Step S206: Adjust the pre-allocated beam scheme according to the correlation probability matrix and the beam parameter vector.
[0044] Through steps S202 to S206, the goal of intelligently optimizing base station beams for multi-UAV service decisions is achieved, thereby achieving the technical effects of improving communication quality and resource utilization efficiency, and enhancing beam tracking stability and response speed. This further solves the technical problems of intermittent communication coverage, unstable signal quality, and low resource scheduling efficiency caused by static beam solutions in related technologies under high-speed movement of multiple UAVs and complex scenarios. Detailed explanations are given below.
[0045] In an embodiment of the present application, taking the example of a single base station that can customize 7 SSB beams and serve N drones at the same time, the intelligent agent (including the policy network (Actor) and the value network (Critic)) dynamically adjusts the azimuth angle, downtilt angle, beam width and other parameters of the beam according to the real-time position and service quality of the drone to achieve the optimal coverage solution for multiple drones.
[0046] In the above step S202, the following two pieces of key information are collected:
[0047] The first is a pre-assigned beam solution for drones. For example, based on the drones' initial flight path information, trajectory prediction is performed and an initial beam allocation solution corresponding to multiple drones is generated. This means that each drone is assigned the base station beam that is most likely to provide the best service, and key parameters such as beam azimuth, downtilt, and beamwidth are initially configured.
[0048] Second, the drone's location information during flight. For example, using the MQTT protocol, the drone's current location and service quality indicators, such as signal strength (RSRP), are uploaded to the ground control center or base station in real time. The MQTT protocol's lightweight nature and publish / subscribe model ensure efficient and reliable data transmission, keeping information up to date even under limited bandwidth and unstable network conditions.
[0049] In step S204, the state vector of each drone is constructed based on its real-time position information and fed into the policy network (Actor) for processing. The Actor network innovatively introduces a shared encoder module and uses the Transformer encoder (i.e., target encoder) as its core structure. This structure globally models each state in the input sequence through a multi-head self-attention mechanism. This effectively mines the spatial distribution patterns between drones and their high-order interactions with the beam, and outputs two key vectors through different branches: the association probability matrix and the beam parameter vector.
[0050] It should be noted that in the embodiment of the present application, the SSB beam tracking problem of the base station to the drone can be modeled as a Markov decision process. By defining a reasonable state space, action space and reward function, the beam tracking problem can be converted into a form suitable for solving by a reinforcement learning algorithm, providing a learning framework for the adaptive optimization of beam parameters.
[0051] 1. State space.
[0052] Optionally, determining the state vector of the drone based on the location information includes: determining a first vector, a second vector, and a third vector based on the location information, respectively, wherein the first vector is used to reflect the angular coordinate information and speed change information of the drone relative to the base station, the second vector is used to reflect the historical beam parameter information of the drone, and the third vector is used to reflect the angular difference between the drone and the base station beam; and fusing the first vector, the second vector, and the third vector to obtain the state vector. The specific expression is as follows:
[0053]
[0054] Where, represents the state vector at time t, represents the first vector, represents the second vector, Represents the third vector.
[0055] For the first vector:
[0056]
[0057] Where, Indicates drone exist The coordinate information and speed information at the moment, namely:
[0058]
[0059] Where, and Indicates the angular coordinate information of the drone relative to the base station. and Indicates the speed change of the drone relative to the base station.
[0060] Specifically, in order to enhance the accuracy of the base station beam direction selection, this application adopts a modeling method based on the polar coordinate domain to construct an angular coordinate system with the north direction and the horizontal plane as the reference base to characterize the position relationship of the UAV relative to the base station. Specifically, the azimuth angle of each UAV at any time t is first calculated. and pitch angle , and subtract the mechanical azimuth of the base station antenna and downtilt angle , thus obtaining the relative angle coordinates of the drone relative to the current antenna direction:
[0061]
[0062]
[0063] By adopting relative angle representation, the intelligent agent only needs to pay attention to the angle difference of the drone relative to the main direction of the antenna, avoiding the perception of antenna orientation information, thereby simplifying the state space of reinforcement learning and improving learning efficiency.
[0064] and Specifically, it is the angular velocity of the UAV relative to the base station in the polar coordinate domain, which corresponds to the rate of change of the azimuth and pitch angle directions, respectively. The calculation method is:
[0065]
[0066]
[0067] For the second vector:
[0068]
[0069] Where, Indicates the parameter configuration of the seven SSB beams at the previous moment, including but not limited to the electronic azimuth, electronic pitch angle, and beamwidth information of each beam.
[0070] For the third vector:
[0071]
[0072] Where, and Respectively represent the azimuth angle difference matrix and pitch angle difference matrix between the UAV and the 7 beams. For example, the Row, No. Column Elements It's a drone With beam The azimuth difference, The definition is similar. This represents the operation of expanding a matrix row by row into a column vector. This angle difference feature can enhance the agent's ability to perceive the angular deviation between the beam and the target, thereby helping to improve the accuracy and convergence efficiency of the beam pointing strategy.
[0073] 2. Action space.
[0074] In the embodiment of the present application, the association probability matrix is innovatively introduced , which is used to explicitly model the service relationship between the UAV and the SSB beam. The association probability matrix It can be regarded as an attention mechanism with structural perception capabilities, which not only enables the intelligent agent to perceive the service intensity of each beam to the drone, but also actively learns the matching pattern between the beam and the drone during training, thereby realizing the joint optimization of the service relationship and beam parameters, rather than just passively adjusting the beam parameters.
[0075] Specifically, the association probability matrix No. Row, No. Elements of a column Indicates drone By beam The probability of service, and satisfying:
[0076]
[0077] Based on the association probability matrix , the action vector at time t can be determined :
[0078]
[0079] Where, , represents the beam information at time t; , Represents beam The electronic azimuth, Represents beam Electronic pitch angle, Represents beam beamwidth.
[0080] 3. Reward function.
[0081] In the embodiment of the present application, in order to realize the efficient tracking and dynamic optimization control of the base station SSB beam to multi-target UAVs, a reward function based on multi-objective joint optimization is also designed. , including: determining the service probability of the target beam to the target UAV, and determining the signal receiving power after the target UAV accesses the target beam, and determining a first reward function based on the service probability and the signal receiving power, wherein the target UAV is any one of the multiple UAVs, the target beam is the optimal beam corresponding to the target UAV in the target beam scheme, and the first reward function is used to quantify the signal quality received by the target UAV from the target beam; determining a second reward function based on the signal receiving power and the preset probability, wherein the second reward function is used to balance the signal strength and effective coverage of the target beam to the target UAV; determining a third reward function based on the beam configuration and the historical beam configuration in the beam parameter vector, wherein the third reward function is used to suppress the change of the beam parameters; and weighting the first reward function, the second reward function, and the third reward function to obtain a reward function.
[0082] The reward function The specific expression is as follows:
[0083]
[0084] Where, is the weighted coefficient used to balance the first reward function , the second reward function And the third reward function The optimization goal between .
[0085] For the first reward function , which is used to quantify the signal quality received by the drone. It is constructed using a soft correlation mechanism and is calculated as follows:
[0086]
[0087] Where, Beam (i.e. the target beam mentioned above) to the UAV (i.e. the target drone mentioned above), For drones Use beam The first reward function can achieve the signal strength-oriented learning goal, prompting the policy network to actively adjust the beam parameters to enhance the overall communication quality.
[0088] For the second reward function , which is used to encourage the strategy to achieve complete target coverage while ensuring the strength of the primary service beam. It is calculated as follows:
[0089]
[0090] Where, is the indicator function, when the drone Maximum receiving power Exceeding the set threshold When , it is effective coverage and gets 1 point, otherwise it gets 0. This second reward function can significantly reduce the probability of non-coverage for edge users, thereby enhancing the overall service integrity of the system.
[0091] For the third reward function , which is used to penalize the sudden change between the current beam configuration and the previous beam configuration, suppress the drastic jump of beam parameters, and avoid the delay caused by frequent adjustments to the communication system. The calculation method is as follows:
[0092]
[0093] This third reward function effectively improves the continuity and stability of the beam control strategy by constraining the change range of the beam parameters between two consecutive moments.
[0094] Optionally, the state vector is processed by the target encoder in the policy network to obtain an association probability matrix and a beam parameter vector, including: determining the input features corresponding to the state vector, and determining the first feature matrix input to the policy network based on the input features; processing the first feature matrix by the target encoder in the policy network to output a second feature matrix; normalizing the second feature matrix by the first branch in the policy network to obtain an association probability matrix; and performing a pooling operation on the second feature matrix by the second branch in the policy network to obtain a beam parameter vector.
[0095] In this embodiment of the application, based on the DDPG reinforcement learning framework, a value network (Critic) and a policy network (Actor) are constructed to jointly solve the optimal beam configuration strategy. In order to enhance the Actor network's ability to model the coupling between beam allocation and parameter control, a shared encoder module is innovatively introduced into the Actor network. This shared encoder module is based on the Transformer encoder and processes the input state vector. After unified feature extraction, the correlation probability matrix is input through two parallel branches respectively With the beam parameter vector , realizing the association probability matrix With the beam parameter vector The structural level joint learning significantly improves the coordination and consistency between the two types of decisions. The overall Actor-Critic network framework is as follows Figure 3 As shown, the process analysis is as follows:
[0096] 1. Determine input features.
[0097] First, based on the input state vector , construct the input features of each drone :
[0098]
[0099] Where, , indicating drone Coordinate information and speed information; , indicating drone Azimuth and elevation angle differences with the seven SSB beams; , represents the beam state, .
[0100] Then, in the Actor network, the input features are concatenated , get the first feature matrix of the input Transformer encoder , where N represents the input feature The number of
[0101] 2. Semantic enhancement.
[0102] The first feature matrix is encoded by the Transformer in the Actor network Process and output the second feature matrix :
[0103]
[0104] Among them, the Transformer encoder is composed of basic modules such as multi-head self-attention mechanism, residual connection, layer normalization structure and feedforward neural network, and the output second feature matrix The OK For drones The context-enhanced representation of contains its interactive semantic information in the global sequence.
[0105] 3. Branch processing.
[0106] The encoded second feature matrix Send to two branches and output the associated probability matrix respectively With the beam parameter vector .
[0107] For the association probability matrix , first transform the second characteristic matrix After the input is fed into the hidden layer, it is fed into a set of fully connected layers and then normalized by the Sigmoid function to obtain the soft association probability between each drone and each beam. The specific structure is defined as follows:
[0108]
[0109] Where, is the network weight matrix, is the bias vector, Represents the element-wise Sigmoid function.
[0110] For the beam parameter vector , first transform the second characteristic matrix After inputting the hidden layer, it is sent to the pooling layer for pooling operation, and then input into the fully connected layer to obtain the control parameters of each beam. The specific structure is defined as follows:
[0111]
[0112]
[0113] Where, Represents the global context feature vector obtained by pooling operation. This application adopts the average pooling method, that is, , this can also be extended to other pooling mechanisms; 、 denote the weight and bias of the beam parameter prediction branch, Represents the feature pooling operation.
[0114] 4. Determine the assessed value of the network.
[0115] In the embodiment of the present application, the Critic network adopts a multi-layer perceptron structure to transform the original state vector With motion vector (That is, the fused association probability matrix With the beam parameter vector ) are concatenated as input and processed through a series of fully connected layers and ReLU activation functions in the hidden layer, and the output of the last fully connected layer is taken as the evaluation value of the corresponding action , which is used to guide the optimization process of the Actor network and the Critic network in combination with the reward function.
[0116] In step S206, the pre-assigned beam scheme for the UAV is dynamically adjusted based on the correlation probability matrix and beam parameter vector obtained in step S204. This adjustment includes both reallocation of beam service objects and real-time optimization of beam parameters (such as azimuth, downtilt, and beamwidth).
[0117] In an embodiment of the present application, step S208 is also included: determining the action vector of the UAV based on the association probability matrix and the beam parameter vector, wherein the action vector is used to reflect the beam service decision taken by the UAV; splicing the action vector and the state vector to obtain a spliced vector, and processing the spliced vector through the value network to obtain an evaluation value, wherein the value network is used to evaluate the value of the beam service decision taken by the UAV; optimizing the strategy network and the value network based on the evaluation value and the reward function, wherein the reward function is used to quantify the signal quality, effective coverage and beam stability of the UAV after adopting the target beam allocation scheme, and the target beam allocation scheme is a pre-allocated beam scheme adjusted according to the association probability matrix and the beam parameter vector.
[0118] To accelerate the convergence of the policy actor network, guide the agent to quickly grasp the optimization direction in the early stages of training, and effectively reduce resource overhead during online training, this application proposes an agent training and optimization method that integrates expert policy supervised learning and reinforcement learning adaptive optimization. This method fully utilizes expert knowledge to implement policy network pre-training and combines reinforcement learning to achieve performance fine-tuning. It can efficiently complete the construction of the optimal service relationship between the drone and the base station beam and the configuration of the beam parameters. It specifically includes the following two stages:
[0119] 1. Behavior cloning pre-training (i.e., Actor network pre-training).
[0120] Optionally, the policy network is trained in the following manner: obtaining the simulated state vector and simulated action vector of the drone, wherein the simulated action vector is a simulated association probability matrix and a simulated beam parameter vector generated by a pre-allocation algorithm; processing the simulated state vector through the initial policy network to obtain a predicted action vector, wherein the predicted action vector includes the predicted association probability matrix and the predicted beam parameter vector output by the initial policy network; determining the first loss function required for training the initial policy network, and determining the first loss between the simulated action vector and the predicted action vector through the first loss function; based on the first loss, using the gradient descent method to determine the optimal parameters of the initial policy network to obtain the policy network. The specific process analysis is as follows:
[0121] 1. Construct training data.
[0122] Constructing sample data ,in, represents the simulation state vector of the UAV, represents the simulated motion vector of the drone, , including the association probability matrix generated by the pre-assignment algorithm With the beam parameter vector .
[0123] 2. Determine the loss function required to train the Actor network .
[0124] The mean square error loss is used in the behavior cloning stage. The specific expression is as follows:
[0125]
[0126] Where, That is the first loss function mentioned above, Represents the loss weight coefficient, which is used to control the learning balance of the two branches; Represents the predicted action vector output by the Actor network, including the predicted association probability matrix and the predicted beam parameter vector .
[0127] 3. Update the parameters of the Actor network.
[0128] The first loss between the simulated action vector and the predicted action vector is determined by the first loss function, and the Actor network parameters are updated using the gradient descent method based on the first loss.
[0129] 2. Reinforcement learning fine-tuning (i.e. optimizing the Actor-Critic network).
[0130] Optionally, optimizing the policy network and the value network based on the evaluation value and the reward function includes: determining a target evaluation value based on the reward function, wherein the target evaluation value is used to represent the expected evaluation value under the action vector; determining a second loss function required for optimizing the value network, and determining a second loss between the evaluation value and the target evaluation value through the second loss function; optimizing the value network based on the second loss and a preset soft update rate, wherein the preset soft update rate is used to smooth the update process of the value network and the policy network; and determining a third loss function required for optimizing the policy network, and determining a third loss under the action vector through the third loss function; optimizing the policy network based on the third loss and the preset soft update rate.
[0131] In this embodiment, the DDPG reinforcement learning framework is used to fine-tune the Actor network's strategy and simultaneously train the Critic network, enabling the agent to further optimize its strategy during its interaction with the environment, thus overcoming the limitations of expert strategies. The specific process is analyzed as follows:
[0132] 1. Determine the loss function required to optimize the Critic network .
[0133] The training of the critic network is based on the mean square time difference error, which is expressed as follows:
[0134]
[0135] Where, That is the second loss function mentioned above, Represents the Q value (true evaluation value) output by the Critic network, represents the target Q value (the expected evaluation value under this service decision), that is, the approximate "true" Q value. The specific expression is as follows:
[0136]
[0137] Where, represents the reward function; is the conversion factor; Represents the state-action value function of the current Critic network, with parameters , used to estimate the action value under the current beam service strategy, Represents the optimized Critic network parameters; is the Actor network, and the parameters are , used to output the optimal beam strategy under a given state, Represents the state vector at time t+1.
[0138] The training goal of the Critic network is to minimize , thereby improving the accuracy of Q-value estimation.
[0139] 2. Determine the loss function needed to optimize the Actor network .
[0140] The optimization goal of the Actor network is to maximize the action value under the current strategy. The specific expression is as follows:
[0141]
[0142] Where, That is, the third loss function mentioned above is used to improve the expected evaluation value of the current beam service strategy in the critic network through back propagation.
[0143] 3. Optimize the Actor-Critic network.
[0144] According to the second loss obtained by the second loss function and the third loss obtained by the third loss function, the experience replay mechanism and the soft update target network are used. The specific expressions are as follows:
[0145]
[0146]
[0147] Where, Indicates the preset soft update rate, Represents the optimized Critic network parameters, Represents the optimized Actor network parameters.
[0148] This application conducted a simulation verification in a scenario where 20 drones were operating simultaneously within the coverage area of an airspace base station, and compared the traditional static beam scheme with the dynamic beam adjustment scheme proposed in this application.
[0149] Verification results show that for planned drones with trajectory prediction information, the received signal strength increased by an average of 4.5dB, significantly enhancing coverage quality; while for unauthorized drones not included in trajectory planning, the received signal strength decreased by an average of 5.8dB, effectively suppressing the communication capabilities of unauthorized targets.
[0150] In the embodiment of the present application, by introducing the modeling method of the association probability matrix and the Markov decision process, the joint optimization of the service relationship and beam parameters between the base station beam and multiple drones is achieved. Specifically, the present application utilizes the structure-aware association modeling mechanism, and through the soft association mechanism and the state representation based on the polar coordinate domain, it not only enhances the accuracy of the intelligent agent's beam allocation decision, but also improves the response speed and stability of the beam tracking. In addition, by designing a strategic network structure that integrates a shared encoder, the present application can simultaneously optimize the beam parameters and service relationship on the basis of unified feature extraction, significantly improving the consistency and coordination of the decision, thereby improving the efficiency of beam tracking.
[0151] According to an embodiment of the present application, a multi-UAV beam adjustment device is provided. It should be noted that the multi-UAV beam adjustment device of the embodiment of the present application can be used to perform the multi-UAV beam adjustment method provided in the embodiment of the present application. The following describes the multi-UAV beam adjustment device provided in the embodiment of the present application.
[0152] Figure 4 This is a structural diagram of a multi-UAV beam adjustment device provided according to an embodiment of the present application. Figure 4 As shown, the device includes:
[0153] An acquisition module 40 is configured to acquire a pre-allocated beam plan of a UAV and position information of the UAV during flight, wherein the pre-allocated beam plan includes initial beam allocation plans corresponding to multiple UAVs;
[0154] Processing module 42 is used to determine the state vector of the UAV based on the position information, and process the state vector through the target encoder in the policy network to obtain an association probability matrix and a beam parameter vector. The policy network is used to adjust the service decision of the base station beam for the UAV. The target encoder performs semantic enhancement on the state vector through a multi-head attention mechanism and outputs an association probability matrix and a beam parameter vector through different branches. The association probability matrix is used to reflect the service relationship between the UAV and the base station beam, and the beam parameter vector is used to reflect the parameter configuration of the base station beam.
[0155] The adjustment module 44 is configured to adjust the pre-allocated beam scheme according to the correlation probability matrix and the beam parameter vector.
[0156] Through the acquisition module, processing module and adjustment module in the above-mentioned multi-UAV beam adjustment device, the purpose of intelligently optimizing the base station beam for multi-UAV service decision-making is achieved, thereby achieving the technical effect of improving communication quality and resource utilization efficiency, enhancing beam tracking stability and response speed, and thus solving the technical problems of discontinuous communication coverage, unstable signal quality and low resource scheduling efficiency in the static beam solution in related technologies under high-speed movement and complex scenarios of multiple UAVs.
[0157] In the beam adjustment device for multiple UAVs provided in an embodiment of the present application, the processing module is also used to determine a first vector, a second vector, and a third vector respectively based on the position information, wherein the first vector is used to reflect the angular coordinate information and speed change information of the UAV relative to the base station, the second vector is used to reflect the historical beam parameter information of the UAV, and the third vector is used to reflect the angular difference between the UAV and the base station beam; the first vector, the second vector, and the third vector are fused to obtain a state vector.
[0158] In the beam adjustment device for multiple drones provided in an embodiment of the present application, the processing module is also used to determine the input features corresponding to the state vector, and determine the first feature matrix of the input strategy network based on the input features; the first feature matrix is processed by the target encoder in the strategy network to output the second feature matrix; the second feature matrix is normalized by the first branch in the strategy network to obtain the association probability matrix; and the second feature matrix is pooled by the second branch in the strategy network to obtain the beam parameter vector.
[0159] In the beam adjustment device for multiple UAVs provided in an embodiment of the present application, the processing module is also used to determine the service probability of the target beam for the target UAV, and to determine the signal receiving power of the target UAV after accessing the target beam, and to determine a first reward function based on the service probability and the signal receiving power, wherein the target UAV is any one of the multiple UAVs, the target beam is the optimal beam corresponding to the target UAV in the target beam scheme, and the first reward function is used to quantify the signal quality received by the target UAV from the target beam; determine a second reward function based on the signal receiving power and the preset probability, wherein the second reward function is used to balance the signal strength and effective coverage of the target beam for the target UAV; determine a third reward function based on the beam configuration and historical beam configuration in the beam parameter vector, wherein the third reward function is used to suppress changes in beam parameters; and perform weighted processing on the first reward function, the second reward function, and the third reward function to obtain a reward function.
[0160] The beam adjustment device for multiple UAVs provided in an embodiment of the present application also includes a training module 46 for obtaining a simulation state vector and a simulation action vector of the UAV, wherein the simulation action vector is a simulation association probability matrix and a simulation beam parameter vector generated by a pre-allocation algorithm; the simulation state vector is processed by an initial strategy network to obtain a predicted action vector, wherein the predicted action vector includes the predicted association probability matrix and the predicted beam parameter vector output by the initial strategy network; the first loss function required for training the initial strategy network is determined, and the first loss between the simulation action vector and the predicted action vector is determined by the first loss function; based on the first loss, the optimal parameters of the initial strategy network are determined by the gradient descent method to obtain the strategy network.
[0161] The beam adjustment device for multiple UAVs provided in an embodiment of the present application also includes an optimization module 48, which is used to determine the action vector of the UAV based on the association probability matrix and the beam parameter vector, wherein the action vector is used to reflect the beam service decision taken by the UAV; splicing the action vector and the state vector to obtain a spliced vector, and processing the spliced vector through the value network to obtain an evaluation value, wherein the value network is used to evaluate the value of the beam service decision taken by the UAV; optimizing the strategy network and the value network based on the evaluation value and the reward function, wherein the reward function is used to quantify the signal quality, effective coverage and beam stability of the UAV after adopting the target beam allocation scheme, and the target beam allocation scheme is a pre-allocated beam scheme adjusted according to the association probability matrix and the beam parameter vector.
[0162] In the beam adjustment device for multiple UAVs provided in an embodiment of the present application, the optimization module is also used to determine a target evaluation value based on a reward function, wherein the target evaluation value is used to represent the expected evaluation value under an action vector; determine a second loss function required for optimizing the value network, and determine a second loss between the evaluation value and the target evaluation value through the second loss function; optimize the value network based on the second loss and a preset soft update rate, wherein the preset soft update rate is used to smooth the update process of the value network and the policy network; and determine a third loss function required for optimizing the policy network, and determine the third loss under the action vector through the third loss function; optimize the policy network based on the third loss and the preset soft update rate.
[0163] An embodiment of the present application also provides an electronic device, including: a memory and a processor, wherein the memory is used to store program instructions; the processor is connected to the memory and is used to execute the above-mentioned beam adjustment method for multiple drones.
[0164] It should be noted that the above electronic equipment is used to perform Figure 2The beam adjustment method for multiple UAVs shown in the figure, therefore the relevant explanations in the above beam adjustment method for multiple UAVs are also applicable to the electronic device and will not be repeated here.
[0165] An embodiment of the present application also provides a non-volatile storage medium, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the above-mentioned beam adjustment method for multiple drones by running the computer program.
[0166] It should be noted that the above non-volatile storage medium is used to execute Figure 2 The beam adjustment method for multiple UAVs shown in the figure, therefore the relevant explanations in the above beam adjustment method for multiple UAVs are also applicable to the non-volatile storage medium and will not be repeated here.
[0167] An embodiment of the present application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above-mentioned beam adjustment method for multiple drones.
[0168] It should be noted that the above-mentioned computer program product is used to execute Figure 2 The beam adjustment method for multiple UAVs shown in the figure, therefore the relevant explanations and instructions in the above-mentioned beam adjustment method for multiple UAVs are also applicable to the computer program product and will not be repeated here.
[0169] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0170] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0171] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0172] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0173] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0174] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program code.
[0175] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A beam adjustment method for multiple UAVs, characterized in that: include: Obtaining a pre-allocated beam plan for a UAV and location information of the UAV during flight, wherein the pre-allocated beam plan includes initial beam allocation plans corresponding to multiple UAVs; Determine the state vector of the UAV based on the position information, and process the state vector through the target encoder in the policy network to obtain an association probability matrix and a beam parameter vector, wherein the policy network is used to adjust the service decision of the base station beam for the UAV, and the target encoder performs semantic enhancement on the state vector through a multi-head attention mechanism, and outputs the association probability matrix and the beam parameter vector through different branches. The association probability matrix is used to reflect the service relationship between the UAV and the base station beam, and the beam parameter vector is used to reflect the parameter configuration of the base station beam; The pre-assigned beam scheme is adjusted according to the association probability matrix and the beam parameter vector.
2. The method according to claim 1, characterized in that Determining a state vector of the drone based on the position information includes: Determine a first vector, a second vector, and a third vector based on the position information, respectively, wherein the first vector is used to reflect the angular coordinate information and speed change information of the UAV relative to the base station, the second vector is used to reflect the historical beam parameter information of the UAV, and the third vector is used to reflect the angular difference between the UAV and the base station beam; The first vector, the second vector, and the third vector are fused to obtain the state vector.
3. The method according to claim 1, characterized in that The state vector is processed by the target encoder in the policy network to obtain the association probability matrix and beam parameter vector, including: Determining an input feature corresponding to the state vector, and determining a first feature matrix input into the policy network based on the input feature; Processing the first feature matrix through a target encoder in the strategy network to output a second feature matrix; Normalizing the second feature matrix through the first branch in the policy network to obtain the association probability matrix; And performing a pooling operation on the second feature matrix through a second branch in the strategy network to obtain the beam parameter vector.
4. The method according to claim 1, wherein The policy network is trained in the following way: Acquire a simulation state vector and a simulation action vector of the UAV, wherein the simulation action vector is a simulation association probability matrix and a simulation beam parameter vector generated by a pre-allocation algorithm; Processing the simulation state vector through an initial policy network to obtain a predicted action vector, wherein the predicted action vector includes a predicted association probability matrix and a predicted beam parameter vector output by the initial policy network; Determine a first loss function required for training the initial policy network, and determine a first loss between the simulated action vector and the predicted action vector using the first loss function; Based on the first loss, a gradient descent method is used to determine the optimal parameters of the initial policy network to obtain the policy network.
5. The method according to claim 1, wherein The method further comprises: Determining an action vector of the UAV based on the association probability matrix and the beam parameter vector, wherein the action vector is used to reflect the beam service decision taken by the UAV; splicing the action vector and the state vector to obtain a spliced vector, and processing the spliced vector through a value network to obtain an evaluation value, wherein the value network is used to evaluate the value of the beam serving decision taken by the UAV; The policy network and the value network are optimized based on the evaluation value and the reward function, wherein the reward function is used to quantify the signal quality, effective coverage, and beam stability of the UAV after adopting a target beam allocation scheme, and the target beam allocation scheme is a pre-allocated beam scheme adjusted based on the association probability matrix and the beam parameter vector.
6. The method according to claim 5, characterized in that The reward function is determined as follows: Determining a service probability of a target beam for a target UAV, and determining a signal reception power of the target UAV after accessing the target beam, and determining a first reward function based on the service probability and the signal reception power, wherein the target UAV is any one of the multiple UAVs, the target beam is an optimal beam corresponding to the target UAV in the target beam solution, and the first reward function is used to quantify a signal quality received by the target UAV from the target beam; Determining a second reward function based on the signal reception power and a preset probability, wherein the second reward function is used to balance the signal strength and effective coverage of the target beam to the target drone; determining a third reward function based on the beam configuration in the beam parameter vector and the historical beam configuration, wherein the third reward function is used to suppress changes in the beam parameters; The first reward function, the second reward function, and the third reward function are weighted to obtain the reward function.
7. The method according to claim 5, characterized in that Optimizing the policy network and the value network according to the evaluation value and the reward function includes: Determining a target evaluation value according to the reward function, wherein the target evaluation value is used to represent an expected evaluation value under the action vector; Determining a second loss function required for optimizing the value network, and determining a second loss between the evaluation value and the target evaluation value using the second loss function; Optimizing the value network according to the second loss and a preset soft update rate, wherein the preset soft update rate is used to smooth the update process of the value network and the policy network; and Determine a third loss function required to optimize the policy network, and determine a third loss under the action vector by using the third loss function; The policy network is optimized according to the third loss and the preset soft update rate.
8. A beam adjustment device for multiple UAVs, characterized in that: include: An acquisition module is configured to acquire a pre-allocated beam scheme of a UAV and position information of the UAV during flight, wherein the pre-allocated beam scheme includes initial beam allocation schemes corresponding to multiple UAVs; a processing module, configured to determine a state vector of the UAV based on the position information, and process the state vector through a target encoder in a policy network to obtain an association probability matrix and a beam parameter vector, wherein the policy network is used to adjust the service decision of the base station beam for the UAV, and the target encoder performs semantic enhancement on the state vector through a multi-head attention mechanism, and outputs the association probability matrix and the beam parameter vector through different branches, wherein the association probability matrix is used to reflect the service relationship between the UAV and the base station beam, and the beam parameter vector is used to reflect the parameter configuration of the base station beam; An adjustment module is configured to adjust the pre-allocated beam scheme according to the association probability matrix and the beam parameter vector.
9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory is used to store program instructions; The processor is connected to the memory and is used to execute the beam adjustment method for multiple UAVs according to any one of claims 1 to 7.
10. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the beam adjustment method for multiple UAVs according to any one of claims 1 to 7 by running the computer program.
11. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by the processor, the beam adjustment method for multiple UAVs according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
System taking unmanned aerial vehicle as networking relay in smart city and working method thereof
CN114641051A
Hybrid beam forming method based on residual optimization in unmanned aerial vehicle communication scene
CN116073873A
Intelligent beam scanning method for cellular-free sensing integrated system
CN118554982A
RIS-assisted unmanned aerial vehicle network-based communication and sensing integrated system and method
CN119277319A
Beam adjustment method and device for unmanned aerial vehicle, electronic equipment and computer program product
CN119364383A
Cited By
Communication method and device, electronic equipment and nonvolatile storage medium
CN121124918A
Communication method and apparatus, electronic device, and nonvolatile storage medium
CN121124918B
Multi-unmanned aerial vehicle beam allocation and tracking method
CN121815411A