Satellite cooperative tracking and pointing moving target method and system based on multi-agent reinforcement learning
By loading a policy network on the satellite for local observation and reward function optimization, multi-agent reinforcement learning-based satellite collaborative tracking is achieved, which solves the problems of computational delay and communication overhead of traditional satellite formations in dynamic target scenarios, improves responsiveness and tracking accuracy, reduces energy consumption, and has good scalability and fault tolerance.
Patent Information
- Application Number
- CN202511159827.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-08-19
AI Technical Summary
Traditional centralized satellite formation observation suffers from large computational delays, high communication overhead, and lacks online adjustment capabilities in dynamic multi-target scenarios, making it difficult to achieve efficient, reliable, and scalable decentralized collaborative decision-making for large-scale agile satellite formations.
A satellite collaborative tracking and aiming method for moving targets based on multi-agent reinforcement learning is adopted. By loading a policy network on each satellite, using local observation and reward functions for real-time attitude control, constructing parallel convolutional channels and fully connected channels, and outputting action probability distribution, end-to-end collaborative tracking and aiming decision-making is achieved.
It significantly shortens response delay, improves tracking continuity and accuracy, controls energy consumption and maneuverability safety, and has good scalability and fault tolerance to meet high-efficiency observation needs.
Smart Images

Figure CN120793231A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of aerospace remote sensing and satellite formation cooperative observation, and particularly relates to a satellite cooperative tracking and sighting moving target method and system based on multi-agent reinforcement learning. BACKGROUND
[0002] In recent years, agile satellites can realize rapid revisit and high-precision imaging among cross-regional and multi-type targets by relying on high-torque reaction flywheels and optimal control algorithms of attitude and orbit coupling, and are widely used in disaster emergency, maritime supervision, battlefield intelligence and other task scenarios. However, when facing observation requirements of multiple targets, high dynamics and strong timeliness in the global scope, a single satellite is difficult to balance high revisit frequency and long tracking time due to the limitations of orbit and field of view. In order to make up for this deficiency, the industry usually adopts multi-satellite formation cooperative observation, and adopts a two-stage paradigm of "centralized task allocation-independent execution": the ground or mother satellite solves the NP-hard scheduling problem after gathering global information, and then issues instructions to each member satellite.
[0003] This centralized scheme exposes several bottlenecks in the dynamic target scenario: the calculation time delay caused by centralized solving weakens the response speed to sudden events; the frequent global state feedback and instruction issuance occupy limited link resources, resulting in high communication load; the plan is highly dependent on prior information, and once the target trajectory mutates or the on-board execution is disturbed, the whole calculation needs to be recalculated, lacking flexibility and robustness; at the same time, as the formation scale increases, the calculation complexity of centralized optimization increases exponentially, and the scalability is limited.
[0004] Deep reinforcement learning has shown advantages of end-to-end optimization and online adaptation in attitude control and resource scheduling fields, but existing researches still follow the divide-and-conquer approach of "allocation first, then execution", which separates target selection and attitude tracking and sighting, and fails to fully tap the potential of decentralized decision-making and multi-agent cooperative learning; it also lacks systematic evaluation and verification in complex scenarios coupled by factors such as energy constraints, orbit differences and communication topology. Therefore, how to realize efficient, reliable and scalable decentralized cooperative decision-making of large-scale agile satellite formation under the premise of ensuring attitude accuracy and energy consumption constraints is still a key technical problem to be solved. SUMMARY
[0005] The purpose of the present application is to overcome the defects of the traditional centralized "task allocation-attitude control" mode in the dynamic multi-target scenario, such as large calculation time delay, high communication overhead and lack of online adjustment capability.
[0006] In order to achieve the above purpose, the present application proposes a satellite cooperative tracking and sighting moving target method based on multi-agent reinforcement learning, comprising: loading a trained policy network on each satellite and initializing an interaction scenario; Constructing local observations of satellites, inputting the strategy network, and outputting satellite tracking actions at each time point; The strategy network comprises a parallel convolution channel and a full connection channel; the convolution channel and the full connection channel are spliced in a feature dimension, output an action probability distribution through a fusion layer, and obtain a satellite tracking action through probability sampling.
[0007] As an improvement of the above method, the training process of the strategy network comprises: Step 1: generating satellite and target trajectories and expanding the attention region, screening available satellites and targets; Step 2: setting local observations and reward functions, aggregating satellite capabilities and training the strategy network by using generated scene data, and obtaining optimal global parameters through iterative updates.
[0008] As an improvement of the above method, the step 1 comprises: Step 1-1: randomly generating weight values for all targets and normalizing them; Step 1-2: regarding the targets falling into the attention region as observation targets; Step 1-3: for each satellite, considering the maneuvering capability of the agile satellite, first expanding the attention region by a certain angle, then judging whether the satellite is located in the expanded region, and regarding the satellite located in the expanded region as an observation satellite.
[0009] As an improvement of the above method, the step 2 comprises: Step 2-1: randomly initializing global strategy network parameters; Step 2-2: randomly setting initial attitude quaternion and angular velocity for each satellite, loading global strategy network parameters, and in subsequent scene training, all satellites in the scene share one strategy network parameter, which is called aggregated strategy network; Step 2-3: calculating the local observation of the satellite; Step 2-4: inputting the local observation into the strategy network to obtain the satellite tracking action, and determining the control torque according to the action; calculating the angular acceleration of the satellite, updating the angular velocity, and then updating the satellite attitude quaternion according to the angular velocity; Step 2-5: for each satellite, repeating steps 2-3 and 2-4 to obtain the respective local observation; Step 2-6: calculating the attitude accuracy reward and tracking effect reward, as well as the energy consumption penalty and angular velocity penalty of each satellite, weighting the calculation results to obtain the overall reward, and using a centralized evaluation network to evaluate the satellite action according to the overall reward; Step 2-7: updating the positions of the satellites and the targets at the next time point; Step 2-8: Steps 2-3 to 2-7 are executed in a loop within the same scenario until all the target tracking tasks are completed. Step 2-9: The multi-agent proximal policy optimization algorithm is used to update the parameters of the aggregated policy network and the centralized evaluation network. The overall reward and the evaluation output by the centralized evaluation network are considered to calculate the gradient, and the global parameters are updated. Step 2-10: Steps 2-2 to 2-9 are repeated until the policy network supports completing the cooperative tracking task in multiple heterogeneous scenarios.
[0010] As an improvement of the above method, the step 2-3 calculates the local observation of the satellite includes: Generate multiple desired attitude quaternions according to the vectors pointing to each target and other satellite footprints; Multiply the conjugate of these desired attitude quaternions with the current attitude quaternion to obtain multiple attitude error quaternions; Extract N groups of attitude error quaternions with the largest real parts and M groups of other satellite footprint attitude error quaternions for each satellite, and fill the insufficient positions with zero quaternions; Concatenate the normalized target weight and the current angular velocity with the extracted attitude error quaternions to form the local observation.
[0011] As an improvement of the above method, the attitude accuracy reward is obtained by summing the real part of the main tracking star quaternion error of each target after logarithmic scaling, and the expression is: ; Wherein, represents the attitude accuracy reward; represents the real part of the quaternion; is the set of all satellites participating in the target corresponding attitude error quaternion; is the total number of targets; the main tracking star is the satellite corresponding to the attitude error quaternion with the largest real part when calculating the attitude error quaternion of all satellites for a certain target.
[0012] As an improvement of the above method, the expression of the tracking effect reward is: ; Wherein, represents the tracking effect reward; is an indication variable of whether the target is successfully tracked; is the target normalized weight; is the total number of targets.
[0013] As an improvement of the above method, the expression of the energy consumption penalty is: ; wherein, represents an energy consumption penalty; is an energy consumption penalty coefficient; is a three-axis control torque vector made by the satellite; represents the modulus of the orientation vector.
[0014] As an improvement of the above method, the calculation formula of the angular velocity penalty is: ; wherein, represents an angular velocity penalty; is the current angular velocity of the satellite; is the maximum angular velocity; represents the modulus of the orientation vector.
[0015] The application also provides a satellite cooperative tracking and pointing moving target system based on multi-agent reinforcement learning, which is realized based on the above method, and the system comprises: a satellite tracking and pointing module for loading a policy network, constructing and inputting local observations of the satellite, and outputting satellite tracking and pointing actions at each time point; the policy network comprises a parallel convolution channel and a fully connected channel; after the convolution channel and the fully connected channel are spliced in the feature dimension, the action probability distribution is output through a fusion layer, and the satellite tracking and pointing action is obtained through probability sampling.
[0016] Compared with the prior art, the application has the following advantages: 1. The response time delay is significantly shortened; each satellite can output attitude control instructions in real time depending on local observations, without waiting for centralized scheduling, and can meet the high-time-efficiency observation demand in response to a sudden target.
[0017] 2. The tracking continuity and accuracy are improved; the parallel one-dimensional convolution channel and the weighted reward mechanism jointly act to enable the policy to dynamically balance the tracking accuracy and energy consumption in a multi-satellite and multi-target scenario.
[0018] 3. The energy consumption and maneuvering safety are controlled; the energy consumption penalty and the angular velocity constraint are introduced into the reward function to guide the policy to actively suppress invalid or excessive maneuvering, thereby effectively reducing the reaction flywheel dynamic load and power consumption while ensuring the attitude safety, and prolonging the on-orbit life of the satellite.
[0019] 4. Good scalability and fault tolerance; through the policy sharing and capability aggregation training framework, a single global policy can be applied to different formation scales and diversified task scenarios; when individual satellites are out of communication or the target trajectory changes suddenly, the remaining satellites can still be re-assigned and maintain stable tracking. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 A policy network structure diagram is shown. Figure 2 A capability aggregation training framework diagram is shown. DETAILED DESCRIPTION
[0021] The technical solutions of the present application will be described in detail below with reference to the accompanying drawings.
[0022] The satellite cooperative tracking and pointing moving target method and system based on multi-agent reinforcement learning provided by the present application is a cooperative tracking and pointing technology for real-time tracking and continuous observation of moving targets by multiple agile satellites in orbit. By relying only on local observation of each satellite and supplemented by a small amount of inter-satellite information interaction, target priority evaluation and attitude control decision are simultaneously completed, thereby significantly improving the rapid response capability and tracking accuracy for multiple targets, and effectively suppressing invalid maneuver to reduce energy consumption.
[0023] Embodiment 1 The satellite cooperative tracking and pointing moving target method based on multi-agent reinforcement learning provided by the present application includes: first, generating different satellites and targets, and dividing regions to construct multiple scenes; then constructing a convolutional parallel policy network; subsequently setting a local observation and reward function, using different scenes to train the network to obtain an optimal model through a capability aggregation training framework; finally, loading the model in the network, and running in the scene to be executed to obtain the tracking and pointing action made by each satellite at each time. Specifically, the following steps are included: First step: generating a scene; Step A: generating satellite and target trajectory and expanding the attention region, screening available satellites and targets, and providing basic data for task allocation.
[0024] This part lays the foundation for subsequent task allocation by generating satellite and target trajectories, constructing attention regions, and then screening targets and satellites that meet the conditions. The specific process includes: Step A1: generating satellite and target trajectory files, randomly generating weight values for all targets, and normalizing the weight values of all targets.
[0025] Step A2: calibrate the attention region in the form of latitude and longitude coordinates. The specific operation can be completed by calling the Polygon class in the PythonShapely library in clockwise / counterclockwise order to initialize the region.
[0026] Step A3: for each target, determine whether it falls within the attention region according to its latitude and longitude coordinates; only the region-in target is activated and added to the observation queue. The specific operation can be calculated by calling the Polygon.contains method.
[0027] Step A4: For each satellite, considering the maneuverability of agile satellites, the region of interest is first expanded in angle, and then it is determined whether the satellite is located in the expanded region. Specifically, first, the polygon center is obtained by means of Polygon.centroid , and let the original boundary point be at a distance of from the center. Then, the new vertex is obtained by expanding radially by degrees, as shown in equation (1).
[0028] (1) and the new polygon composed of all is taken as the expanded region of interest, and the Polygon.contains method is called again; if the satellite is located in the expanded region, it is activated and added to the list of available satellites.
[0029] Second step: build a convolutional parallel strategy network; Step B: parallel convolution channels in the traditional fully connected strategy network, output action probability after fusing two features, guide satellite selection and sighting action.
[0030] Based on the traditional multilayer perceptron strategy network, a convolution feature channel is added, forming a convolution fully connected parallel structure. The structure can simultaneously excavate the local correlation of the satellite local observation sequence, and the overall network is shown in Figure 1 . After the convolution channel (CNN) and the fully connected channel (MLP) are spliced in the feature dimension, the action probability distribution is output through the fusion layer (MIX). The specific process includes: Step B1: input the satellite's local observation (State) into the strategy network, where the local observation is a one-dimensional observation state. The observation is input into the convolution channel and the fully connected channel. The local observation includes: the attitude error quaternion of the satellite to the nearby target, the attitude error quaternion of the satellite to the nearby satellite subsatellite point, the normalized weight of the satellite to the nearby target, and the current angular velocity of the satellite.
[0031] Step B2: the convolution channel is composed of 2 one-dimensional convolution layers and max pooling layers stacked alternately, and each convolution kernel size is 3 and the pooling window size is 2. Finally, the convolution feature vector is obtained through the flattening operation.
[0032] Step B3: the fully connected channel adopts a 3-layer fully connected network, each layer is connected with a ReLU activation function, and the fully connected feature vector is output.
[0033] Step B4: concatenate the convolution feature vector and the fully connected feature vector to obtain the joint representation, and input it into the fusion fully connected layer to map to the action probability distribution (Prob) in the action space. The processing process is shown in equation (2).
[0034] (2) Step B5: Obtain the satellite tracking action based on probability sampling and transmit it to the satellite actuator to complete the status update.
[0035] Step 3: Model training; Step C: Set the local observation and reward function. The reward function includes attitude accuracy rewards, tracking performance rewards, energy consumption penalties, and angular velocity penalties. In the multi-scenario capability aggregation training architecture, each scenario not only maintains its own aggregated satellite capabilities during training but also updates them with the global satellite capabilities to achieve a stable collaborative tracking strategy.
[0036] For the satellite collaborative tracking task, we set up local observation and reward functions, use the generated scene data to aggregate satellite capabilities and train the tracking strategy network, and obtain the optimal global parameters through iterative updates. The training framework diagram is shown in the figure below. Figure 2 The specific process is as follows: Step C1: Randomly initialize the global policy network parameters.
[0037] Step C2: Initialize the interactive scene. Each satellite randomly sets its initial attitude quaternion and angular velocity. Then, load the global policy network parameters. In subsequent scene training, all satellites in the scene share the same policy network parameters, hence the name "aggregated policy network." This completes the initialization of the environment and network.
[0038] Step C3: Construct satellite local observations based on the scene state. First, calculate the vectors of the satellite pointing to each target and other satellite subsatellite points and generate multiple expected attitude quaternions based on them (the "target-pointing" expected attitude quaternion is used to perceive the attitude required for the surrounding satellites pointing to the target, and the "other satellite subsatellite points" expected attitude quaternion is used to perceive the relative positions of the surrounding satellites relative to the satellite). Then, multiply the conjugate of these expected attitude quaternions with the current attitude quaternion to obtain multiple attitude error quaternions; then, for each satellite, extract the three sets of target error quaternions with the largest real part (the largest real part means that the current satellite can be corrected after the least attitude adjustment). The error quaternions of the sub-satellite points of two other satellites (i.e., each satellite only considers the relative positions of the two satellites closest to itself) are combined (the application can pre-specify that each satellite only considers the nearest N targets and M satellites. However, in some small scenes, there may not be N targets or M other satellites. In this case, in order to keep the local observation length of the input strategy network unchanged, the missing parts are filled with quaternions of all zeros). Finally, the normalized target weight and current angular velocity are spliced with these error quaternions to form a local observation.
[0039] Step C4: The local observations obtained in the previous step are fed into the policy network in step 2 to generate actions, and the control torque is determined based on the actions. The angular acceleration is then calculated using the satellite's rigid body dynamics, and the angular velocity is updated. The satellite's attitude quaternion is then updated based on the angular velocity.
[0040] Step C5: For each satellite in the scene, repeat steps C3-C4 to obtain their respective local observations.
[0041] Step C6: Calculate the attitude accuracy reward. The attitude accuracy reward is calculated by summing the real part of the logarithmically scaled quaternion error of the primary tracking satellite for each target. The calculation process is shown in formula (3).
[0042] (3) in, Indicates taking the real part of the quaternion, To participate in the goal The corresponding attitude error quaternion in the set of all satellites, is the total number of targets. The primary tracking satellite calculates the error quaternion for a target for all satellites and selects the satellite corresponding to the error quaternion with the largest real part as the primary tracking satellite; that is, the satellite whose current attitude is closest to successfully tracking the target among all satellites.
[0043] Step C7: Calculate the tracking effect reward. The tracking effect reward is weighted summed according to the target successful tracking indicator. The calculation process is shown in formula (4).
[0044] (4) in, Target An indicator variable indicating whether the tracking was successful. is the target normalized weight.
[0045] Step C8: Calculate the energy consumption penalty for each satellite individually. The energy consumption penalty is the control moment norm, and the calculation process is shown in formula (5).
[0046] (5) in, is the energy consumption penalty coefficient, is the total number of satellites, The three-axis control torque vector for the satellite; The modulus of the orientation quantity.
[0047] Step C9: Calculate the angular velocity penalty for each satellite individually. The angular velocity penalty gives a negative reward when the angular velocity exceeds the threshold. The calculation process is shown in formula (6).
[0048] (6) wherein, is the current angular velocity of the satellite, is the maximum angular velocity.
[0049] Step C10: The above rewards are weighted to obtain an overall reward function. And the centralized evaluation network is used to evaluate the satellite action according to the overall reward. The centralized evaluation network is the central evaluation network in the multi-agent proximal policy optimization algorithm MAPPO.
[0050] Step C11: Update the positions of the satellite and the target at the next time.
[0051] Step C12: C3 to C11 are executed in the same scene in a loop until all the tracking tasks are declared to be completed.
[0052] Step C13: The multi-agent proximal policy optimization algorithm is used to update the parameters of the aggregated policy network and the centralized evaluation network, and the evaluation output by the centralized evaluation network is comprehensively considered. The Adam optimizer is used, the learning rate is set to 1e-3, and the gradient is calculated .
[0053] Step C14: Update the global parameters wherein, is the network parameter before updating, is the network parameter after updating, is the learning rate.
[0054] Step C15: Repeat steps C2 to C14 until the model can stably complete the cooperative tracking task in multiple heterogeneous scenes.
[0055] Fourth step: task execution; Step D: Each satellite loads the trained policy network and outputs the satellite attitude and tracking action in real time in the actual scene to complete the target cooperative tracking task. Among them, other satellite information and target information are transmitted through inter-satellite / earth-space communication, and only a small amount of position coordinates are transmitted.
[0056] Through the trained model, the tracking task is executed on the target task scene. The specific process is as follows: Step D1: Initialize the interaction scene. Set the initial attitude quaternion of each satellite to the direction of the load pointing to the center of the earth, and the angular velocity is set to 0. Then read the trained model parameters to complete the initialization of the environment and the network.
[0057] Step D2: through the interaction of the model and the environment until the end of the scene. Output satellite tracking action, satellite attitude, target tracking result, etc. at each time.
[0058] Embodiment 2 The application also provides a satellite cooperative tracking moving target system based on multi-agent reinforcement learning, which is realized based on the above method, and the system comprises: The satellite tracking module is used for loading the policy network, constructing and inputting the local observation of the satellite, and outputting the satellite tracking action at each time. The policy network comprises a convolution channel and a fully connected channel in parallel; the convolution channel and the fully connected channel are spliced in the feature dimension, output the action probability distribution through a fusion layer, and obtain the satellite tracking action through probability sampling.
[0059] The application breaks the traditional "two-stage" operation framework of "target assignment-attitude control" in series, proposes an end-to-end policy network based on local observation information, so that each satellite in the formation can complete the target priority determination and tracking decision synchronously under the condition of decentralization, thereby significantly shortening the instruction loop, improving the rapid response capability to sudden targets, and maintaining high system fault tolerance and robustness without a central node. In view of the multiple requirements of attitude accuracy, continuous tracking effect, energy consumption control and attitude safety in the process of multi-satellite cooperative observation, the application provides a multi-agent reward mechanism considering cooperative game and energy saving constraint; by weighting the tracking quality and energy consumption in the reinforcement learning process, each agent dynamically adjusts the observation attitude to meet the multi-target continuous tracking demand, while actively suppressing invalid or excessive attitude maneuver to reduce energy consumption. In order to improve the generalization of the strategy, the application further constructs a heterogeneous multi-scene capability aggregation training framework: in the real orbit dynamics simulation environment, multiple scenes covering different satellite scales, target densities and spatial distributions are introduced, information exchange and cooperative control are realized through strategy sharing and capability aggregation mechanism, the adaptation ability of each agent to complex environment is gradually strengthened, and the task execution efficiency and operation stability of the whole satellite formation are significantly improved on this basis.
[0060] The application can also provide a computer device comprising at least one processor, memory, at least one network interface and user interface. The various components in the device are coupled together through a bus system. It can be understood that the bus system is used to realize the connection communication between the components. In addition to the data bus, the bus system also includes power bus, control bus and state signal bus.
[0061] The user interface can include a display, a keyboard or a clicking device. For example, a mouse, a trackball, a touchpad or a touch screen, etc.
[0062] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (Read-Only Memory, ROM), a programmable read-only memory (Programmable ROM, PROM), an erasable programmable read-only memory (Erasable PROM, EPROM), an electrically erasable programmable read-only memory (Electrically EPROM, EEPROM) or a flash memory. The volatile memory can be a random access memory (Random Access Memory, RAM) used as an external cache. By way of example, but not by way of limitation, many forms of RAM are available, such as static random access memory (Static RAM, SRAM), dynamic random access memory (Dynamic RAM, DRAM), synchronous dynamic random access memory (Synchronous DRAM, SDRAM), double data rate synchronous dynamic random access memory (Double Data Rate SDRAM, DDR SDRAM), enhanced synchronous dynamic random access memory (Enhanced SDRAM, ESDRAM), synchronous link dynamic random access memory (Synchlink DRAM, SLDRAM) and direct memory bus random access memory (Direct Rambus RAM, DRRAM). The memory described herein is intended to include, but not limited to, these and any other suitable types of memory.
[0063] In some embodiments, the memory stores elements, executable modules or data structures, or a subset thereof, or an extended set thereof: an operating system and an application program.
[0064] Among them, the operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application program includes various application programs, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The program for implementing the method of the embodiments of the present disclosure can be included in the application program.
[0065] In the above-described embodiments, the processor can also be used to: execute the steps of the above method.
[0066] The method can be applied to a processor or implemented by the processor. The processor can be an integrated circuit chip having a signal processing capability. In implementation, the steps of the method can be completed by an integrated logic circuit of hardware in the processor or by an instruction in the form of software. The processor can be a general-purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The methods disclosed above can be implemented or executed by the processor. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed above can be directly embodied as a hardware code executed by the processor or a combination of hardware and software modules in the processor. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, or other mature storage mediums in the art. The storage medium is located in the storage memory, and the processor reads information in the storage memory and combines the hardware to complete the steps of the method.
[0067] It can be understood that the embodiments described in the present application can be implemented in hardware, software, firmware, middleware, microcode or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field-Programmable Gate Arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for executing functions described in the present application or a combination thereof.
[0068] For software implementation, the present application can be implemented by executing function modules (such as processes, functions, etc.) described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0069] The application can also provide a non-volatile storage medium for storing a computer program. When the computer program is executed by a processor, each step in the above method embodiment can be implemented.
[0070] Finally, it should be pointed out that the above embodiments are only used to illustrate the technical solutions of the application but not to limit the application. Although the application is described in detail with reference to the embodiments, those skilled in the art should understand that the technical solutions of the application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the application, and all of them should be covered in the scope of the claims of the application.
Claims
1. A method for satellite collaborative tracking of moving targets based on multi-agent reinforcement learning, comprising: Load the trained policy network on each satellite and initialize the interaction scenario; Construct local observations of the satellite, input them into the strategy network, and output the satellite tracking and aiming actions at each moment; The strategy network includes parallel convolutional channels and fully connected channels. After the convolutional channels and fully connected channels are spliced in the feature dimension, the action probability distribution is output through the fusion layer, and the satellite tracking action is obtained according to the probability sampling.
2. The satellite collaborative tracking and aiming method for moving targets based on multi-agent reinforcement learning according to claim 1 is characterized in that: The training process of the policy network includes: Step 1: Generate satellite and target trajectories and expand the area of interest, screening available satellites and targets; Step 2: Set up local observation and reward functions, use the generated scene data to aggregate satellite capabilities and train the policy network, and obtain the optimal global parameters through iterative updates.
3. The satellite collaborative tracking and aiming method for moving targets based on multi-agent reinforcement learning according to claim 2 is characterized in that: The step 1 comprises: Step 1-1: Randomly generate weight values for all targets and normalize them; Step 1-2: The target falling into the area of interest is regarded as the target to be observed; Steps 1-3: For each satellite, considering the maneuverability of the agile satellite, first expand the area of interest by a set angle, then determine whether the satellite is located in the expanded area, and use the satellite in the expanded area as the observation satellite.
4. The satellite collaborative tracking and aiming method for moving targets based on multi-agent reinforcement learning according to claim 2 is characterized in that: The step 2 includes: Step 2-1: Randomly initialize the global policy network parameters; Step 2-2: Each satellite randomly sets the initial attitude quaternion and angular velocity, loads the global policy network parameters, and in subsequent scene training, all satellites in the scene share the same policy network parameters, called the aggregated policy network; Step 2-3: Calculate the local observation of the satellite; Step 2-4: Input the local observation into the strategy network to obtain the satellite tracking action and determine the control torque based on the action; calculate the satellite's angular acceleration and update the angular velocity, and then update the satellite attitude quaternion based on the angular velocity; Step 2-5: For each satellite, repeat steps 2-3 and 2-4 to obtain their respective local observations; Step 2-6: Calculate the attitude accuracy reward and tracking effect reward, as well as the energy consumption penalty and angular velocity penalty for each satellite. Weight the calculated results to obtain the overall reward. Use the centralized evaluation network to evaluate the satellite action based on the overall reward. Step 2-7: Update the positions of the satellite and the target at the next moment; Step 2-8: Repeat steps 2-3 to 2-7 in the same scene until all tracking tasks are completed; Step 2-9: Use the multi-agent proximal policy optimization algorithm to update the parameters of the aggregated policy network and the centralized evaluation network. Take into account the overall reward and the evaluation calculation gradient output by the centralized evaluation network to update the global parameters. Step 2-10: Repeat steps 2-2 to 2-9 until the policy network supports collaborative tracking tasks in multiple heterogeneous scenarios.
5. The satellite collaborative tracking and aiming method for moving targets based on multi-agent reinforcement learning according to claim 4 is characterized in that: The steps 2-3 of calculating the local observation of the satellite include: Generate multiple desired attitude quaternions based on vectors pointing to each target and other subsatellite points; Multiply the conjugates of these desired attitude quaternions with the current attitude quaternion to obtain multiple attitude error quaternions; For each satellite, extract N groups of attitude error quaternions with the largest real part and M groups of attitude error quaternions of other satellites’ sub-satellite points, and fill the missing positions with zero quaternions; The normalized target weight and current angular velocity are concatenated with the extracted attitude error quaternion to form a local observation.
6. The satellite collaborative tracking and aiming method for moving targets based on multi-agent reinforcement learning according to claim 4 is characterized in that: The attitude accuracy reward is obtained by summing the real part of the logarithmic scaling of the quaternion error of the main tracking satellite of each target, and the expression is: ; in, represents the posture accuracy reward; Indicates taking the real part of the quaternion; To participate in the goal The corresponding attitude error quaternion in the set of all satellites; is the total number of targets; the main tracking star is to calculate the attitude error quaternion of all satellites with respect to a certain target, and select the satellite corresponding to the attitude error quaternion with the largest real number part as the main tracking star.
7. The satellite collaborative tracking and aiming method for moving targets based on multi-agent reinforcement learning according to claim 4 is characterized in that: The expression of the tracking effect reward is: ; in, Indicates tracking effect reward; Target An indicator variable indicating whether the tracking was successful; is the target normalized weight; is the overall target number.
8. The satellite collaborative tracking and aiming method for moving targets based on multi-agent reinforcement learning according to claim 4 is characterized in that: The expression of the energy consumption penalty is: ; in, represents energy consumption penalty; is the energy consumption penalty coefficient; The three-axis control torque vector for the satellite; The modulus of the orientation quantity.
9. The satellite collaborative tracking and aiming method for moving targets based on multi-agent reinforcement learning according to claim 4 is characterized in that: The calculation formula of the angular velocity penalty is: ; in, represents the angular velocity penalty; is the current angular velocity of the satellite; is the maximum angular velocity; The modulus of the orientation quantity.
10. A satellite collaborative tracking and aiming system for moving targets based on multi-agent reinforcement learning, implemented based on the method of any one of claims 1-9, characterized in that: The system comprises: Obtain a satellite tracking module to load the policy network, construct and input the local observation of the satellite, and output the satellite tracking action at each moment; and The policy network includes parallel convolutional channels and fully connected channels. After the convolutional channels and the fully connected channels are spliced in the feature dimension, the action probability distribution is output through the fusion layer, and the satellite tracking action is obtained according to the probability sampling.
Citation Information
Patent Citations
Reinforced learning-based spatial non-cooperative target parameter self-tuning tracking method
CN110850719A
Satellite space target collaborative observation distributed planning method based on multi-agent reinforcement learning
CN116187160A
Method and system for tracking and pointing moving target by satellite based on deep reinforcement learning
CN117699055A
Star group orbit pursuit decision-making method based on multi-near-end reinforcement learning
CN119962403A
Satellite scheduling method and device based on reinforcement learning, equipment and storage medium
CN120342462A
Cited By
Satellite intelligent orbital transfer decision-making system and method
CN121734694A