Method and device for establishing unmanned aerial vehicle controller model based on isovariant network and geometric symmetry

By using a UAV controller model based on equivariant networks and geometric symmetry, the problems of large data requirements and single strategy in existing technologies are solved, achieving efficient and unified UAV aerobatic flight control and improving the stability and responsiveness of flight control.

CN121978968AActive Publication Date: 2026-05-05ZHEJIANG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2026-04-07
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing deep reinforcement learning-based UAV control methods require large amounts of data, have long training times, and employ simplistic strategies in complex flight missions. They are unable to generate a unified flight strategy with universal understanding capabilities, resulting in high computational costs and difficulty in achieving highly maneuverable aerobatic flight.

Method used

A UAV controller model based on equivariant networks and geometric symmetry is adopted. By combining the real-time state of the UAV and aerobatic commands with a feature-based linear modulation layer, an irreducible representation transformation layer and an equivariant multilayer perceptron, an efficient unified flight strategy is generated, reducing redundant learning processes and improving data utilization efficiency and strategy generalization ability.

Benefits of technology

It achieves a unified strategy for acquiring multiple aerobatic flight techniques through a single training session, reducing training costs and complexity, improving the stability and responsiveness of flight control, and reducing computational resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121978968A_ABST
    Figure CN121978968A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle controller model establishment method and device based on an equivariant network and geometric symmetry. Determining an observation vector of the unmanned aerial vehicle based on the real-time state of the unmanned aerial vehicle and the current stunt instruction; inputting the observation vector into a characterization linear modulation layer, generating a modulation parameter corresponding to the real-time state according to the current stunt instruction, and modulating the real-time state; inputting the modulated observation vector into an irreducible representation conversion layer, and generating irreducible representation of an SO (2) group corresponding to the real-time state; and inputting the irreducible representation into an equivariant multi-layer perceptron, predicting and generating a control instruction corresponding to the current stunt instruction, and obtaining the output of the unmanned aerial vehicle controller model. Based on reinforcement learning of one network, an unmanned aerial vehicle controller model with high decision-making efficiency and strong generalization ability is obtained, and the unmanned aerial vehicle control cost is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of unmanned aerial vehicle (UAV) control, and in particular to a method and apparatus for establishing a UAV controller model based on equivariant networks and geometric symmetry. Background Technology

[0002] Currently, with the rapid development of drone technology, flight missions are gradually shifting from simple navigation to highly dynamic and maneuverable aerobatic maneuvers. Drone flight controllers typically employ deep reinforcement learning-based methods for complex drone control. However, existing deep reinforcement learning-based methods still face significant challenges in complex flight missions.

[0003] On the one hand, current reinforcement learning methods are typically data-intensive, requiring a large number of training samples to converge. They have high requirements for both the quantity and quality of samples, resulting in long training times and low training efficiency. On the other hand, existing flight strategies are often highly specialized, meaning that a separate agent needs to be trained for each acrobatic maneuver (such as flips, rolls, and spins). This "one-many-one-strategy" approach is not only computationally expensive but also fails to generate a unified flight strategy with universal understanding.

[0004] Some existing deep learning techniques utilize multi-task reinforcement learning (MTRL) to learn a general model by leveraging commonalities across tasks. However, their progress in the field of drone aerobatic flight has been limited. This is partly because applying them to drone aerobatic flight requires traditional aerobatic generation methods to rely on preset waypoints, which in turn limits the drone's ability to perform truly high-maneuverability extreme maneuvers. Therefore, there is an urgent need for an efficient and universal drone flight controller to efficiently and accurately control various uniform aerobatic scenarios. Summary of the Invention

[0005] In view of this, this application provides a method and apparatus for establishing a UAV controller model based on equivariant networks and geometric symmetry, in order to solve the problems of large data requirements and simple strategies in traditional UAV controllers.

[0006] Specifically, this application is implemented through the following technical solution:

[0007] The first aspect of this application provides a method for establishing a UAV controller model based on equivariant networks and geometric symmetry, the method comprising:

[0008] The observation vector of the UAV is determined based on the UAV's real-time status and current aerobatic commands;

[0009] The observation vector is input into the feature-based linear modulation layer, and modulation parameters corresponding to the real-time state are generated according to the current special effects command to modulate the real-time state.

[0010] The modulated observation vector is input into the irreducible representation transformation layer to generate the irreducible representation of the SO(2) group corresponding to the real-time state;

[0011] The irreducible representation is input into an equivariant multilayer perceptron to predict and generate the control command corresponding to the current stunt command, thereby obtaining the output of the UAV controller model.

[0012] A second aspect of this application provides an apparatus for establishing a UAV controller model based on equivariant networks and geometric symmetry, the apparatus comprising:

[0013] A construction module is used to determine the observation vector of the UAV based on the UAV's real-time status and current aerobatic commands;

[0014] A feature-based linear modulation layer is used to input the observation vector into the feature-based linear modulation layer, generate modulation parameters corresponding to the real-time state according to the current special effects command, and modulate the real-time state;

[0015] An irreducible representation transformation layer is used to input the modulated observation vector into the irreducible representation transformation layer to generate an irreducible representation of the SO(2) group corresponding to the real-time state;

[0016] An equivariant multilayer perceptron is used to input the irreducible representation into the equivariant multilayer perceptron, predict and generate the control command corresponding to the current stunt command, and obtain the output of the UAV controller model.

[0017] This application provides a method and apparatus for establishing a UAV controller model based on equivariant networks and geometric symmetry. Compared with the traditional UAV controller model processing method for realizing UAV aerobatic flight control, the UAV controller model of this invention combines characteristic linear modulation, geometrically symmetric SO(2) geometric symmetry embedding network and equivariant multilayer perceptron. Specifically, the characteristic linear modulation layer modulates the observation vector, SO(2) generates the irreducible representation of the modulated vector, and the equivariant multilayer perceptron predicts the corresponding aerobatic command through the irreducible representation. The UAV performs the corresponding aerobatic action according to the instruction of the aerobatic command. By directly encoding the SO(2) rotational symmetry into the UAV controller model composed of the linear modulation layer and the equivariant multilayer perceptron, the redundancy of state-action pairs in the entire model learning process is reduced, the data utilization efficiency and the generalization ability of the strategy are greatly improved, and a unified strategy UAV controller model capable of performing multiple aerobatic flights can be obtained through one training, which greatly reduces the training cost and complexity, and effectively reduces the consumption of computing resources and the requirement for training samples. Attached Figure Description

[0018] Figure 1A flowchart illustrating the method for establishing an unmanned aerial vehicle (UAV) controller model based on equivariant networks and geometric symmetry provided in this application;

[0019] Figure 2 A schematic diagram of the device for establishing a UAV controller model based on equivariant networks and geometric symmetry provided in this application. Detailed Implementation

[0020] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.

[0021] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used herein are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0022] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0023] The following specific embodiments are given to illustrate the technical solution of this application in detail.

[0024] Figure 1 This is a flowchart of an embodiment of the method for establishing a UAV controller model based on equivariant networks and geometric symmetry provided in this application. Please refer to... Figure 1 This embodiment provides a method for establishing a UAV controller model based on equivariant networks and geometric symmetry, which may include:

[0025] S101. Determine the observation vector of the UAV based on the UAV's real-time status and current aerobatic commands.

[0026] Drones are flying objects used for aerial aerobatic performances. They are typically micro-drones, and multiple drones combine to fly according to prescribed maneuvers, achieving various aerobatic effects. After takeoff, the drones' real-time status information is monitored. This real-time status information includes at least the drone's external status, such as position, speed, and attitude, and at least its internal status, such as motor output status. By combining the internal and external real-time statuses, the drone's real-time status is accurately assessed, leading to more precise action commands. The current aerobatic command is the target flight maneuver issued in real time, including at least the type of maneuver the drone is expected to perform, such as a flip, and corresponding action parameters, such as the flip angle and height. The commands can be issued using structured data or text input, with semantic understanding extracting structured variables to obtain the final action type and parameters. Each current aerobatic command contains actions corresponding to its current action type, with various action types including at least flip, roll, and rotate. The current aerobatic command is the operation that controls the drone to perform the corresponding aerobatic maneuver, as per control instructions. For example, in this embodiment, it includes flip, roll, and rotate. The observation vector is a concatenation of the UAV's real-time state, aerobatic commands, and historical commands. It serves as the input to the overall UAV controller model, which consists of a feature-modulated linear modulation layer, a geometrically symmetric SO(2) geometrically symmetric embedded network layer, and an equivariant multilayer perceptron. The UAV's real-time state can be acquired through sensors, including physical motion state and equipment operating state. Physical motion state includes position, velocity, attitude, and angular velocity, while equipment operating state includes motor parameters and battery voltage. The acquired raw data is preprocessed, for example, using Kalman filtering and normalization, and then concatenated with the processed vector to obtain a 25-dimensional normalized vector.

[0027] Specifically, determining the observation vector of the UAV based on its real-time state and current aerobatic command includes: generating random noise according to the type of aerobatic maneuver of the UAV; superimposing the real-time state and the random noise to obtain a state vector; obtaining the historical control command corresponding to the historical aerobatic maneuver of the UAV at the moment before the current aerobatic command; and concatenating the state vector, the historical control command, and the current aerobatic command to generate the observation vector.

[0028] Random noise is a Gaussian-distributed random disturbance signal. Based on the current aerobatic maneuver type, Gaussian random noise is generated for the real-time state of the drone. Specifically, this can be achieved by adjusting the mean and standard deviation of the Gaussian distribution function to generate random noise conforming to a Gaussian distribution. Since the dynamic characteristics of each aerobatic maneuver type differ, a correspondence is established between the maneuver type and the mean and standard deviation. Different parameters are configured for random noise of different aerobatic maneuver types. The generated noise can simulate the dynamic disturbances of real-world scenarios, enhancing the robustness and generalization ability of the drone flight control model and avoiding overfitting to a single state. The number of dimensions of the random noise is consistent with the real-time state. Elements corresponding to each dimension are superimposed in a one-to-one correspondence manner. The state vector is a multi-dimensional vector obtained by element-wise superposition, such as a 25-dimensional vector, used to provide robustness enhancement data. The 25-dimensional normalized vector is superimposed element-wise with the 25-dimensional random noise to obtain a 25-dimensional state vector. This ensures that the state information of each dimension matches the corresponding noise vector, avoiding cross-dimensional interference.

[0029] In practice, for example, the superposition is achieved using element-wise addition, with the specific formula as follows:

[0030] ;

[0031] in, Represents the state vector; Indicates real-time status; represents random noise; i represents the dimension index.

[0032] Historical control commands are the control commands issued by the drone at the time step preceding the current aerobatic command. If the execution of control commands is real-time, i.e., without delay, the historical control commands have already been executed, and thus the historical control commands are the control commands issued by the drone at the time step preceding the current moment. If the execution of control commands is not real-time, i.e., all generated control commands first enter the control command queue and are executed in the order they enter the queue, then the historical control command is the nearest control command at the front of the queue corresponding to the current moment. Historical control commands are used to reflect the historical movement trends of the drone and avoid sudden changes in movement. Furthermore, the extraction of historical control commands is based on the drone's sampling clock. For example, a timestamp is generated every 0.005 seconds, and all sensor data and command issuance are bound to this timestamp. When the timestamp of the current aerobatic command is t, the previous moment is t-0.005 seconds. By matching the timestamps, the historical control commands at that moment can be accurately extracted.

[0033] In general, the observation vector is generated by concatenating the state vector, the historical control commands, and the current acrobatic command. The state vector provides the current state of the UAV's movement, the historical control commands provide the movement trends of the UAV's historical states, and the current acrobatic command provides the UAV's current mission objective. The concatenation of these three provides a comprehensive basis for generating control commands that meet the requirements of the acrobatic maneuvers. Furthermore, for example, concatenating the above vectors in dimensional order yields a 33-dimensional observation vector (25+4+4=33).

[0034] Understandably, the UAV controller model that combines a feature-based linear modulation layer, an irreducible representation transformation layer, and an equivariant multilayer perceptron can learn the relationship between the UAV's current flight state and aerobatic commands. By using historical control commands, it can improve the stability and responsiveness of flight control, further optimize the UAV's flight performance, and achieve more efficient execution of aerobatic missions.

[0035] S102. Input the observation vector into the feature-based linear modulation layer, generate modulation parameters corresponding to the real-time state according to the current special effects command, and modulate the real-time state.

[0036] The UAV controller model provided by this invention combines a feature-modulated linear modulation layer, an irreducible representation transformation layer, and an equivariant multilayer perceptron. The feature-modulated linear modulation layer is used to modulate the input observation vector, and the modulated parameters are related to the current stunt command. The irreducible representation transformation layer is used to transform the modulated vector into an irreducible representation. The equivariant multilayer perceptron is used to predict and generate the corresponding control command based on the irreducible representation. Figure 2 Please refer to the schematic diagram of the UAV controller model provided in this application. Figure 2 From the overall structure of the UAV controller model, the UAV controller model generally includes at least an execution network and an evaluation network. The execution network sequentially includes the characteristic linear modulation layer, the irreducible representation transformation layer, and the equivariant multilayer perceptron. The evaluation network includes a standard characteristic linear modulation layer, a standard irreducible representation transformation layer, a standard equivariant multilayer perceptron, and a value head, which is used to influence the prediction strategy of the execution network during the training phase. The standard characteristic linear modulation layer receives the same input as the characteristic linear modulation layer and selects a target value head from the value heads based on the input special commands. The value head is used to evaluate the real-time state of the UAV.

[0037] The value head comprises multiple value heads and a gating module. Each value head receives the output of the standard variable multilayer perceptron, meaning the inputs to all value heads are the same, but each value head evaluates a different type of stunt action and has a different reward function. For example, value head A evaluates the state value of a flipping action, while value head B evaluates the state value of a circling action. The gating module directly receives the current stunt command from the input signal, determines the stunt action type based on the command, selects a target gating channel based on the stunt action type, and outputs the target state value corresponding to that channel. Other gating channels in the module are not selected, and the calculated state value cannot be output. For each stunt action type, there is only one target gating channel. As an optional embodiment, each value head is an independent linear layer specifically designed to evaluate different stunt tasks. The value head result for the corresponding task is dynamically selected and output using an index operation based on the task ID. In practice, for example, task ID 1 corresponds to the hovering task, task ID 2 corresponds to the flipping task, and task ID 3 corresponds to the obstacle avoidance task. The gating system selects the corresponding state value based on the input task ID. If task ID is 1, the gating system selects the state value corresponding to the hovering task. If task ID is 2, the gating system selects the state value corresponding to the flipping task.

[0038] The number of value heads can be multiple, determined by the total number of UAV aerobatic flight maneuvers. The method includes: determining the UAV aerobatic flight mission; based on the number of maneuver types in the aerobatic flight mission; designing the same number of value heads according to the number of maneuver types; for each value head, determining the representation method of the target parameters in each maneuver type according to the maneuver type corresponding to the value head, with each maneuver type having the same representation format but different representation content; calculating the product of multiple sub-reward functions based on the target parameters, which serves as the reward function for one value head. Each mission type requires an independent output head to evaluate its corresponding state value or generate corresponding control commands. In this embodiment, for example, since the mission types are static target tracking, flip maneuver, and roll maneuver, the corresponding number of multiple heads is 3. The target parameters for each maneuver type are usually the basis for control targets and system feedback. Different maneuver types involve different target parameters. In specific implementation, for example, for a static target tracking mission, the target parameter is configured as the target position p. target [0, 0, 0] T And expected speed v des [0, 0, 0] T For rollover maneuvers, the target parameters are configured using the target position p. target For [0, 0, r] T Expected speed v des[v, 0, 0] T and desired angular velocity ω des = [0, ω, 0] T For roll maneuvering missions, the physical quantities are configured using the target position as p. target [0, 0, 0] T and the desired angular velocity is ω des [ω, 0, 0] T For the same target parameters, their representation forms are identical, such as a one-dimensional matrix with three elements. Using multiple independent value heads, and configuring corresponding physical quantities according to the different requirements of stunt missions, avoids interference between different missions. Through the calculation and continuous optimization of state values, the execution network is guided to select the optimal strategy. It is understandable that by dynamically determining the number of heads based on the number of mission types, it can flexibly adapt to different numbers of stunt missions without pre-defining all missions, automatically expanding according to actual needs. This allows it to handle more complex multi-task problems while maintaining the overall flexibility of the UAV controller model.

[0039] The entire drone controller model includes a training phase and a usage phase. During the training phase, the evaluation network calculates the evaluation value corresponding to the state of the control command output by the execution network. Based on the evaluation value, the execution network's prediction strategy is rewarded, thereby ensuring that the state corresponding to the control command of the execution network is consistent with the expected state corresponding to the input current acrobatic command. During the usage phase, the evaluation network is no longer used; only the execution network from the training sequence is used to predict the control command corresponding to the current acrobatic command.

[0040] Specifically, before determining the observation vector of the UAV based on its real-time state and current aerobatic commands, the method further includes: acquiring sample observation vectors, inputting the execution network and the evaluation network as input values; the execution network and the evaluation network respectively transforming the sample observation vectors based on weight matrices satisfying equivariance constraints, modulating the sample states in the sample observation vectors according to sample task commands in the sample observation vectors, and generating sample control commands corresponding to the sample task commands based on the modulated variables; the evaluation network calculating the value reward value corresponding to the reward function for each value head, selecting a target value head among the value heads based on the sample task commands, and outputting the value reward value corresponding to the target value head; and training the execution network based on the value reward value and a near-end policy optimization algorithm. Here, the reward function is determined according to the task type and the physical quantity, and the reward function is obtained by multiplying multiple sub-reward functions; the reward value is calculated according to the reward function, and the reward value is used for training.

[0041] To enable drone controller models to be applicable to various types of aerobatic maneuvers, this invention proposes a reward function with a unified structure but variable parameters. The reward function for each value head is composed of the product of multiple sub-items, as shown in the following formula:

[0042] ;

[0043] in, Indicates location-based rewards; Indicates speed reward; Indicates angular velocity bonus; This indicates that the command parameters match the reward. Furthermore, the formulas for each reward value are as follows:

[0044] , ;

[0045] , ;

[0046] , ;

[0047] , ;

[0048] Where k represents the number of iterations and the linear velocity error norm. Angular velocity error norm ; For instruction completion error; v real v represents the real-time linear velocity of the drone. des ω is the expected linear velocity for drone aerobatic flight. real For the real-time angular velocity of the drone, ω des The desired angular velocity for drone aerobatic maneuvers.

[0049] As an optional embodiment, the Proximal Policy Optimization (PPO) algorithm is used when training the execution network. The optimization objective of the Proximal Policy Optimization algorithm includes at least a weighted sum of policy loss, value function loss, and entropy reward. The policy loss is an evaluation of the loss between the output values ​​of the new and old policies. The new and old policies are the action probability distributions generated by the execution network in the same state. The new policy is the current version of the execution network, while the old policy is the version at the previous training moment. This is mainly used to limit the magnitude of policy updates and prevent training oscillations. The value function loss is used to constrain the consistency between the evaluation network output and the actual reported output; specifically, it is expressed as the square of the difference between the evaluation network output and the actual reported output. The entropy reward is used to encourage policy exploration and avoid premature convergence; specifically, it can be calculated as follows: L entropy = -E t [H(π θ (a t |s t ))], where in state s t Take action a t The reward value for outperforming the average level is -E. t [H(π θ (a t |s t [)] is added, H is a function term, E t These are the calculated coefficients at time t.

[0050] The objective function is the optimization objective of the near-end policy optimization algorithm. This embodiment introduces a shearing operation and, based on hyperparameter constraints, limits the policy update tolerance, i.e., the magnitude of a single policy update. The policy loss is as follows:

[0051] ;

[0052] in, Indicates the currently executed network parameters; This represents the empirical expectation of the time step along the sampling trajectory; This represents the probability ratio between the old and new strategies; R represents the advantage function at time step t, used to measure the superiority of an action relative to the average policy. t ( ) indicates a unified definition of reward; This represents a hyperparameter used to control the tolerance for policy updates. The objective function is minimized through continuous iterative optimization at time steps, and the execution network parameters are updated based on this objective function. As an optional implementation, the descent gradient is calculated in each iteration, and the algorithm determines how to update the parameters of the UAV controller model; this is known as gradient descent. Alternatively, gradient ascent can be used, calculating the gradient through backpropagation and updating the execution network parameters. This update occurs once per time step, allowing the parameters to gradually approach their optimum after multiple training iterations.

[0053] During training, on the execution network side, each layer of the execution network is equipped with initial parameters, and sample observation vectors are obtained and input into the execution network. The sample observation vectors enter the characteristic linear modulation layer to generate modulation vectors. The modulation vectors then obtain the irreducible representation and input it into the equivariant multilayer perceptron, outputting the control command corresponding to the current aerobatic command. The execution network can flexibly generate precise flight control strategies according to different input states and aerobatic commands, achieving multi-tasking and efficient flight control. On the evaluation network side, the evaluation network shares the same SO(2)-Equivariant MLP backbone structure as the execution network. The parameters are synchronously optimized through the backpropagation process of PPO. The evaluation network also inputs the observation vectors and enters the standard characteristic linear modulation layer to generate modulation vectors. The modulation vectors then obtain the irreducible representation and input it into the standard equivariant multilayer perceptron. Each value head calculates the reward value through a reward function. Specifically, the evaluation network and the execution network share the same SO(2)-Equivariant MLP backbone structure. The input vector includes at least: the relative state at the current time t; the historical action executed in the previous step, used to reflect control inertia; and the task command vector, containing the task type (Hover, Flip, Roll, Rotate) and amplitude parameters. Therefore, the evaluation network input is not a control command, but a combination of state, historical action, and task command, representing the complete observation information of the UAV at time t. In the evaluation network, each task corresponds to an independent value head, which calculates the UAV's current state s under task i. t The expected future cumulative reward (i.e., state value) at time t is used to evaluate the state output by the network, which represents the dynamic state (position, velocity, attitude, angular velocity, etc.) of the UAV at the current time t. The state output by the evaluation network is used to calculate the advantage function: A t = r t +γ·V(s t+1 ) - V(s t ), r tγ and V() are adjustment coefficients for the advantage function, and V() is the evaluation function, which guides the PPO policy gradient update in the execution network. The reward value is output through the gating output channel, showing the state value corresponding to the task type. The reward function is used to quantify the performance of a specific task in the current state, and can include, for example, the difference between the state and the goal, time or action cost, and task completion. The reward value is a value calculated by the reward function, usually varying in the range of [-1, 1]. A positive value indicates that the current state is favorable to the task, and a negative value indicates poor task performance. The gating mechanism is a mechanism used to weight and selectively output the reward values ​​of different tasks. The state value is the expected reward or performance of completing a specific task in the current state. It can be understood that the collaborative design of the execution network and the evaluation network enables the execution network to be trained more accurately based on the guidance information of the evaluation network. The evaluation network influences the training strategy of the execution network, adjusting and optimizing the control strategy in real time, thereby improving the stability of aerobatic flight and the accuracy of aerobatic maneuvers. The execution network is designed as a G-isovariable function, while the evaluation network is designed as a G-invariant function. Specifically, in one possible implementation, the G-equivariant function is as follows:

[0054] ;

[0055] in, Indicates from generate Functions of a process; and A group representation of the input space; and Indicate that the SO(2) group is respectively with and The associated vector space; g represents the group element of the SO(2) group; x represents the input modulation parameter. Where g∈G, x∈ .

[0056] In another possible implementation, for example, if the output represents For the expression of ordinariness, that is (g)=I, and the G-invariant function is as follows:

[0057] ;

[0058] As an optional embodiment, the sample observation vector corresponds to a single action type. In this case, there is also a single target value head, and the state value is the reward value calculated by the reward function corresponding to the target value head. As another optional embodiment, the sample observation vector corresponds to a composite action type, i.e., an organic combination of multiple actions at the same time. In this case, there are multiple target value heads. The gating module calculates the proportion of each action type in the composite type based on the instruction parameters of each action type in the current special action instruction in the sample observation vector. The gating module selects the gating channel corresponding to each action type, and uses the proportion as the weight to calculate the weighted sum of each target gating channel as the final state value.

[0059] The step of inputting the observation vector into the feature-based linear modulation layer and generating modulation parameters corresponding to the real-time state according to the current stunt command to modulate the real-time state includes: the multilayer perceptron in the feature-based linear modulation layer generating a first modulation parameter and a second modulation parameter according to the input current stunt command; and calculating the modulation value corresponding to the real-time state based on the product of the first modulation parameter and the real-time state and the sum of the second modulation parameter.

[0060] Specifically, the feature-modulated linear modulation layer is a feature modulation network layer driven by instructions. Its function is as an instruction adjustment module, generating scaling and offset parameters based on the instructions to adjust the intermediate activation values ​​of the equivariant multilayer perceptron. The scaling parameter adjusts the amplitude; a value greater than 1 strengthens the corresponding state feature, while a value less than 1 weakens it. The offset parameter corrects the numerical value of the state feature to make it closer to the baseline. The observation vector is input into the feature-modulated linear modulation layer to generate modulation parameters, namely the first modulation parameter and the second modulation parameter. Preferably, the first modulation parameter is the scaling parameter, and the second modulation parameter is the offset parameter. Then, the observation vector is modulated. The feature-modulated linear modulation layer integrates a multilayer perceptron, which is a feedforward neural network that learns any nonlinear mapping from input to output through fully connected layers and nonlinear activation. It includes an input layer, hidden layers, and an output layer. The vector corresponding to the current special skill instruction is input into the multilayer perceptron, and through pre-trained weights, the original parameters are output. These original parameters are then split to obtain the scaling and offset parameters. The observed vector is input into the feature-modulated linear modulation layer. The current aerobatic command is encoded and normalized to obtain a command vector. The multilayer perceptron enhances the modulation parameters related to the corresponding aerobatic maneuver type based on the command vector. It is understood that the feature-modulated linear modulation layer can modulate the current aerobatic command and real-time state vector into a feature form more suitable for the processing of the equivariant multilayer perceptron, enabling the flight strategy to more accurately match mission requirements and reduce control errors. Weights are assigned to each dimension of the state feature by scaling parameters, and an affine transformation is completed by combining the offset parameters. The scaling parameters, offset parameters, and the vector dimensions of the state feature are consistent. In one possible implementation, the input to the feature-modulated linear modulation layer includes not only the current UAV state but also the action calculated in the previous step, i.e., the historical control state. The two are concatenated to form the real-time state vector: x=[srel,a(t-1)].

[0061] The formula for calculating the modulation value of the real-time state vector is as follows:

[0062] FiLM(x, c) = x·γ(c) + β(c);

[0063] Where x represents the real-time input state; c represents the current input special move command; Indicates the scaling parameter; This represents the offset parameter.

[0064] S103. Input the modulated observation vector into the irreducible representation transformation layer to generate the irreducible representation of the SO(2) group corresponding to the real-time state.

[0065] The SO(2) group is a representation of the rotational symmetry of a UAV dynamic system about the Z-axis. It identifies the rotational symmetry about the gravity axis in the UAV dynamic model and formalizes it as a two-dimensional rotational SO(2) symmetry group. The group elements... Indicates the rotation angle around the Z-axis Operations, such as rotating the drone's yaw angle from 30° to 60°, correspond to... The group elements. When the UAV rotates around the Z-axis, its dynamic equations remain unchanged. An irreducible representation is the basic unit in group representation theory, referring to a group representation that cannot be further decomposed into smaller subspaces closed by group actions. Since the equivariant multilayer perceptron requires the output to change synchronously with the group transformation of the input, it is necessary to ensure that the execution network has equivariance, that is, when the input changes according to the rules of the symmetric group, the output will also change synchronously according to the same rules. However, if the modulation parameters are not converted into SO(2) irreducible representations, they cannot satisfy the equivariance constraints of the equivariant multilayer perceptron, which will cause the equivariant multilayer perceptron to fail to maintain rotational symmetry, thereby causing the UAV to lose control of its aerobatic maneuvers.

[0066] The step of inputting the modulated real-time state into the irreducible representation conversion layer to generate the irreducible representation of the SO(2) group corresponding to the real-time state includes: calculating the integer frequency corresponding to the real-time state as the index of the irreducible representation; calculating the product of the real-time state and a complex number to calculate the complex representation corresponding to the real-time state; determining the conversion method according to the magnitude of the integer frequency; converting the complex representation into a real representation according to the conversion method; and decomposing the real representation into the sum of the values ​​of multiple real irreducible representations to generate the irreducible representation of the symmetric SO(2) group.

[0067] First, the rotation angle and frequency are obtained, and a complex expression for the irreducible representation of the SO(2) group is constructed based on these values. The rotation angle is acquired through UAV sensors, such as real-time acquisition via the gyroscope and inertial measurement unit on the UAV. The complex expression constructed based on the rotation angle and frequency is used to transform the abstract rotation operation into a linear transformation in the complex domain, facilitating subsequent real-number representation.

[0068] In one possible implementation, for example, the complex expression for the irreducible representation of the SO(2) group is as follows:

[0069] ;

[0070] in Indicate irreducible representation; θ represents the rotation angle around the origin; k represents an integer frequency; z represents an integer.

[0071] Then, the complex expression is converted into a real number representation, where k=0 corresponds to a trivial representation, k>0 corresponds to a nontrivial representation, and the irreducible representation of the SO(2) group is obtained when k>0. The physical quantities of the UAV are divided into irreducible and trivial representations according to rotational symmetry, providing symmetry-compliant inputs for the equivariant multilayer perceptron. The trivial representation is a scalar physical quantity whose value remains unchanged when rotating around the Z-axis, such as Z-axis position and Z-axis velocity. The irreducible representation is a physical quantity that needs to be transformed according to vector laws when rotating around the Z-axis, such as X / Y axis position, velocity, angular velocity, and components of the attitude matrix. In one possible implementation, in order to satisfy the equivariance constraint of the equivariant multilayer perceptron, and since it operates in the real number domain, the complex expression is converted into a real number representation. For the frequency k=0, which is a trivial representation, its expression is:

[0072] ;

[0073] Furthermore, for frequencies k>0, by combining conjugates This allows us to construct a two-dimensional real irreducible representation. When k=1, this corresponds to a rotation of a standard two-dimensional vector. This representation acts on a two-dimensional real vector. ∈ Above, the irreducible representation transformation matrix (θ) is:

[0074] ;

[0075] Where y represents the characteristic component.

[0076] Finally, after converting the real representation to a real number representation, the method further includes decomposing the real number representation into a feature space, wherein the finite-dimensional representation space V of any SO(2) group can be decomposed into a direct sum of the above real irreducible representations:

[0077] ;

[0078] in, Represents the irreducible representation space with frequency k; This indicates the number of times it appears.

[0079] In this embodiment, the SO(2) symmetry group and its representation in linear space are defined. The 25-dimensional state vector is mapped to the corresponding representation space according to its transformation properties under SO(2) rotations. For example, the (X,Y) components of the relative positions in the state vector are assigned vector representations (k=1), while the Z component is assigned scalar representations (k=0). It can be understood that converting the rotation angle and frequency into complex number expressions, and then converting them through real number representations, ensures efficient modeling of the SO(2) rotational symmetry, which can significantly improve the control accuracy and efficiency of UAVs when performing different aerobatic flight missions.

[0080] S104. Input the irreducible representation into the equivariant multilayer perceptron to predict and generate the control command corresponding to the current stunt command, and obtain the output of the UAV controller model.

[0081] The equivariant multilayer perceptron comprises at least a combination of multiple alternating layers of equivariant linear layers and equivariant nonlinear activation functions; that is, the equivariant multilayer perceptron is composed of a series of alternatingly stacked equivariant linear layers and equivariant nonlinear activation functions. Specifically, alternating stacking refers to the alternating use of equivariant linear layers and equivariant nonlinear activation functions in the neural network. In each layer, an equivariant linear transformation is performed first, followed by processing the output of the equivariant linear transformation using an equivariant nonlinear activation function. The output of each layer serves as the input to the next layer until a control command is output.

[0082] The step of inputting the irreducible representation into an equivariant multilayer perceptron to predict and generate the control command corresponding to the current stunt command includes: transforming the irreducible representation based on the weight matrix of the equivariant linear layer; the equivariant nonlinear activation function layer decomposes the transformed irreducible representation into Fourier series components at multiple frequencies, where the amplitude of each Fourier series component at each frequency is the amplitude component at the corresponding frequency, and the phase is the same as the irreducible representation before decomposition; and generating control commands based on the mapping of each decomposed Fourier series component. The UAV controller model input is the observation vector input in step S101, which includes the real-time state of the UAV and the current stunt command that the UAV is expected to perform. This vector passes sequentially through a feature-modulated linear modulation layer, an irreducible representation transformation layer, and an equivariant multilayer perceptron to obtain the control command issued to achieve the current stunt command action. The UAV executes the control command generated in step S104, ultimately completing the action corresponding to the current stunt command. Specifically, after the observation vector is input to the feature-based linear modulation layer, the modulation parameters corresponding to the real-time state are generated according to the record in step S102, and the real-time state is modulated to obtain the modulated observation vector; the modulated observation vector is input to the irreducible representation transformation layer to generate the irreducible representation of the SO(2) group corresponding to the real-time state; finally, the irreducible representation is input to the equivariant multilayer perceptron to predict the corresponding control command.

[0083] An equivariant multilayer perceptron is a multilayer perceptron generator with group equivariance constraints. When the input undergoes a group transformation, such as a drone rotating around the Z-axis, the model's output will synchronously transform according to the corresponding group representation rules. After the irreducible representation enters the equivariant multilayer perceptron, it first passes through an equivariant linear layer. The weight moments of the equivariant linear layer must satisfy the equivariance constraint. Specifically, this is achieved by solving for the basis that satisfies the matrix constraint using Schur's lemma, which is then used to represent the equivariant weight matrix. The output of the equivariant linear layer is further input to an equivariant nonlinear activation function, such as ReLU or GELU, to decompose the irreducible representation of the input into Fourier series components at different frequencies. Finally, through multilayer stacking, more complex features are gradually extracted, ultimately learning the mapping process. The weight matrix of the equivariant linear layer generates the neural network output by performing a linear transformation with the input data. The equivariance constraint ensures that the result of the input features being transformed by rotation and then linearly transformed by the weight matrix is ​​completely consistent with the result of first undergoing a linear transformation and then rotation. The weight matrix W of the equivariant linear layer must satisfy the following constraints:

[0084] , ;

[0085] Solving for the weight matrices that satisfy the constraints using Schul's lemma yields a set of basis vectors for the linear space. The weight matrices are then expressed as a linear combination of these basis vectors. Schul's lemma is used to solve for weight matrices that satisfy equivariance constraints, finding matrices that meet the requirements. Through this lemma, all linear matrices satisfying group constraints can be solved, forming a linear space. The corresponding basis vectors are determined using Schul's lemma; for example, triviality indicates that the corresponding basis is... Non-trivial representations correspond to a basis that is and To ensure that the basis vectors are linearly independent, verify that each basis can independently represent different directions by calculating the determinant of the basis or directly checking its linear independence.

[0086] In this embodiment, any equivariant weight matrix W can be represented as a linear combination of the basis:

[0087] ;

[0088] in, Indicates trainable parameters; This represents the basis matrix.

[0089] The equivariant nonlinear activation function layer decomposes the transformed irreducible representation into Fourier series components at multiple frequencies. Specifically, it decomposes the irreducible representation into Fourier series components corresponding to different frequencies k. For details on Fourier series decomposition, please refer to existing technologies; details will not be elaborated here. Subsequently, the nonlinear function independently acts on the amplitude of each Fourier series component while maintaining phase invariance, thus strictly preserving equivariance while introducing nonlinearity. It is understandable that by using equivariant linear layers and equivariant nonlinear activation functions, the geometric symmetry of the input features can be effectively captured, reducing the processing of redundant information. This allows the model to learn useful feature representations more efficiently during training, achieving rapid convergence and improving training efficiency.

[0090] The method provided in this invention employs a proximal policy optimization algorithm for efficient parallel training in the NVIDIA IsaacGym simulator. After training, the resulting unified policy is deployed on a physical quadcopter drone for real-world flight testing. Experimental results show that the policy can reliably respond to commands and successfully execute various high-speed aerobatic maneuvers. To better simulate the drone in the simulator, this invention also includes establishing a dynamic model of the drone. The dynamic equations characterizing the real-time physical motion state can be expressed as:

[0091] ;

[0092] ;

[0093] ;

[0094] ;

[0095] in, Indicates location; ∈ Indicates speed; ∈ Indicates a gesture; Indicates the quality of the drone; Represented as , unit vector; Indicates total thrust; The aerodynamic damping coefficient matrix is ​​represented by g; The gravitational acceleration vector; ω ∈ R 3 Indicates angular velocity; This represents the rotational inertia matrix of the UAV; Indicates the total torque; Antisymmetric matrix operator representing angular velocity.

[0096] The dynamic characteristics of a quadcopter drone are determined by its position p ∈ R 3 Linear velocity v ∈ R 3 The attitude rotation matrix R ∈ SO(3), and the angular velocity ω in the body coordinate system. B ∈ R 3 The method further includes, after establishing the dynamic model, simulating the evolution of the real-time state parameters of the UAV based on the dynamic model, wherein the evolution is described by a set of nonlinear differential equations; identifying the rotational symmetry of the evolution about the gravity axis, i.e., SO(2) symmetry, and adding the irreducible representation layer based on SO(2) symmetry between the characteristic linear layer and the equivariant multilayer perceptron based on the rotational symmetry.

[0097] Preferably, after the UAV controller model has been trained and tested, the method further includes: inputting real-time observation vectors into the UAV controller model, wherein the UAV controller model predicts and outputs control commands to control the UAV to perform the stunt maneuver corresponding to the current stunt command. During the usage phase, the UAV controller model only includes an execution network and no evaluation network.

[0098] The method provided in this embodiment, compared to the traditional UAV controller model which only uses reinforcement learning, proposes a combined UAV controller model architecture of an equivariant multilayer perceptron, a feature-based linear modulation layer, and SO(2). Utilizing the rotational symmetry of the dynamic model around the gravity axis, i.e., the geometric symmetry of the UAV, a data-efficient unified flight strategy capable of learning various aerobatic maneuvers is constructed. The trained execution network can respond to action commands through a single model structure, flexibly switching and executing various high-difficulty maneuvers such as flips, rolls, and rotations. It possesses high practicality and interactivity, significantly improving the accuracy, real-time performance, and stability of the UAV controller model in controlling aerobatic flight maneuvers. By embedding SO(2) geometric symmetry into the network structure, the redundancy of state-action pairs during the learning process is effectively reduced, thereby significantly improving data utilization efficiency and the generalization ability of the strategy. By optimizing the control command generation process, it is ensured that the UAV controller model can accurately generate commands corresponding to the control attitude and position when performing high-dynamic aerobatic flight control, avoiding loss of control due to equivariant disruption, while effectively reducing computational resource consumption and training sample requirements. By integrating historical control commands, the control over the historical dynamics of the UAV is further enhanced, motion oscillations are reduced, and the robustness of the UAV controller model is improved. This allows the UAV controller model to quickly adapt to complex environments and improve flight stability, achieving higher efficiency and flight accuracy.

[0099] Figure 2This application provides a schematic diagram of a device for building a UAV controller model based on equivariant networks and geometric symmetry. Please refer to... Figure 2 The apparatus provided in this embodiment includes:

[0100] Module 210 is used to determine the observation vector of the UAV based on the UAV's real-time status and current aerobatic commands;

[0101] The feature-based linear modulation layer 220 is used to input the observation vector into the feature-based linear modulation layer, generate modulation parameters corresponding to the real-time state according to the current special effects command, and modulate the real-time state;

[0102] The irreducible representation transformation layer 230 is used to input the modulated observation vector into the irreducible representation transformation layer to generate the irreducible representation of the SO(2) group corresponding to the real-time state;

[0103] An equivariant multilayer perceptron 240 is used to input the irreducible representation into the equivariant multilayer perceptron, predict and generate the control command corresponding to the current stunt command, and obtain the output of the UAV controller model.

[0104] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for establishing a UAV controller model based on equivariant networks and geometric symmetry, characterized in that, The method includes: The observation vector of the UAV is determined based on the UAV's real-time status and current aerobatic commands; The observation vector is input into the feature-based linear modulation layer, and modulation parameters corresponding to the real-time state are generated according to the current special effects command to modulate the real-time state. The modulated observation vector is input into the irreducible representation transformation layer to generate the irreducible representation of the SO(2) group corresponding to the real-time state; The irreducible representation is input into an equivariant multilayer perceptron to predict and generate the control command corresponding to the current stunt command, thereby obtaining the output of the UAV controller model.

2. The method according to claim 1, characterized in that, The determination of the observation vector of the UAV based on its real-time status and current aerobatic commands includes: Random noise is generated based on the type of aerobatic maneuvers performed by the drone; The state vector is obtained by superimposing the real-time state and the random noise. Obtain the historical control commands corresponding to the historical stunts performed by the drone at the moment preceding the current stunt command; The observation vector is generated by concatenating the state vector, the historical control command, and the current special skill command.

3. The method according to claim 1, characterized in that, The method further includes an execution network and an evaluation network. The execution network sequentially includes the characteristic linear modulation layer, the irreducible representation transformation layer, and the equivariant multilayer perceptron. The evaluation network includes a standard characteristic linear modulation layer, a standard irreducible representation transformation layer, a standard equivariant multilayer perceptron, and a value head, used to influence the prediction strategy of the execution network during the training phase. The standard characteristic linear modulation layer receives the same input as the characteristic linear modulation layer and selects a target value head from the value heads based on the input special commands. The value head is used to evaluate the real-time state of the UAV.

4. The method according to claim 3, characterized in that, Before determining the observation vector of the UAV based on its real-time status and current aerobatic commands, the method further includes: Obtain the sample observation vector, and input the values ​​to the execution network and the evaluation network; The execution network and the evaluation network respectively change the sample observation vector based on the weight matrix that satisfies the equivariance constraint, modulate the sample state in the sample observation vector according to the sample task instruction in the sample observation vector, and generate the sample control instruction corresponding to the sample task instruction based on the modulated variable. The evaluation network calculates the value reward value corresponding to the reward function for each value head, selects the target value head in the value heads based on the sample task instruction, and outputs the value reward value corresponding to the target value head. The execution network is trained based on the value reward and the near-end policy optimization algorithm.

5. The method according to claim 3, characterized in that, The number of value heads is multiple, and the method includes: Determine the aerobatic flight missions for the drones; Based on the number of maneuver types in the aerobatic flight mission; Design the same number of value heads based on the number of the described action types; For each value head, the representation method of the target parameter in each action type is determined according to the action type corresponding to the value head. The representation format is the same for each action type, but the representation content is different. The product of multiple sub-reward functions is calculated based on the target parameters to form the reward function of a value head.

6. The method according to claim 1, characterized in that, Before determining the observation vector of the UAV based on its real-time status and current aerobatic commands, the method further includes: The reward function is determined based on the task type and physical quantity, wherein the reward function is obtained by multiplying multiple sub-reward functions; The reward value is calculated based on the reward function, and the reward value is used for training.

7. The method according to claim 1, characterized in that, The step of inputting the observation vector into the feature-based linear modulation layer, generating modulation parameters corresponding to the real-time state according to the current special effects command, and modulating the real-time state includes: The multilayer perceptron in the characteristic linear modulation layer generates a first modulation parameter and a second modulation parameter based on the input current special effect command; The modulation value corresponding to the real-time state is calculated based on the product of the first modulation parameter and the real-time state and the sum of the second modulation parameter.

8. The method according to claim 1, characterized in that, The step of inputting the modulated observation vector into the irreducible representation transformation layer to generate the irreducible representation of the SO(2) group corresponding to the real-time state includes: Calculate the integer frequency corresponding to the real-time state, and use it as the index of the irreducible representation; Calculate the product of the real-time state and the complex number, and calculate the complex number representation corresponding to the real-time state; The conversion method is determined based on the magnitude of the integer frequency, and the complex number representation is converted into a real number representation based on the conversion method. The real number representation is decomposed into the sum of the values ​​of multiple real irreducible representations, generating an irreducible representation of the symmetric SO(2) group.

9. The method according to claim 1, characterized in that, The equivariant multilayer perceptron includes at least a combination of multiple alternating equivariant linear layers and equivariant nonlinear activation function layers. The step of inputting the irreducible representation into the equivariant multilayer perceptron to predict and generate the control command corresponding to the current stunt command includes: The irreducible representation is transformed based on the weight matrix of the equivariant linear layer; The equivariant nonlinear activation function layer decomposes the transformed irreducible representation into Fourier series components at multiple frequencies. The amplitude of each Fourier series component at a frequency is the amplitude component at the corresponding frequency, and the phase is the same as the irreducible representation before decomposition. Control commands are generated based on the mapping of the decomposed Fourier series components.

10. A device for establishing a UAV controller model based on equivariant networks and geometric symmetry, characterized in that, The device includes: A construction module is used to determine the observation vector of the UAV based on the UAV's real-time status and current aerobatic commands; A feature-based linear modulation layer is used to input the observation vector into the feature-based linear modulation layer, generate modulation parameters corresponding to the real-time state according to the current special effects command, and modulate the real-time state; An irreducible representation transformation layer is used to input the modulated observation vector into the irreducible representation transformation layer to generate an irreducible representation of the SO(2) group corresponding to the real-time state; An equivariant multilayer perceptron is used to input the irreducible representation into the equivariant multilayer perceptron, predict and generate the control command corresponding to the current stunt command, and obtain the output of the UAV controller model.