Cooperative tracking control method based on distributed reinforcement learning and related device

By employing a distributed reinforcement learning-based collaborative tracking control method in a multi-electro-optical theodolite system, autonomous decision-making and adaptive adjustment of each node were achieved, solving the problems of central node failure risk and adaptability to complex scenarios, and improving the system's robustness and tracking accuracy.

CN121995904APending Publication Date: 2026-05-08XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI
Filing Date
2026-03-18
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing multi-electro-optical theodolite collaborative tracking systems suffer from the risk of single-point failure at the central node, lack real-time adaptability, are unable to cope with complex dynamic scenarios, and lack independent decision-making capabilities at each node, resulting in insufficient system robustness and flexibility.

Method used

A cooperative tracking control method based on distributed reinforcement learning is adopted. By deploying an agent at each photoelectric theodolite node, the agent makes autonomous decisions using local observation information and communication information between neighboring nodes, generates multi-dimensional action vectors, and performs safety verification and smoothing through the instruction processing layer. Finally, the servo control execution layer drives the photoelectric theodolite to track.

Benefits of technology

It eliminates the risk of single-point failure at the central node, enhances the robustness and autonomy of the system, enables it to dynamically respond to complex scenarios, improves the accuracy and continuity of multi-target collaborative tracking, and ensures the stability and reliability of target tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121995904A_ABST
    Figure CN121995904A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of cooperative tracking, and discloses a cooperative tracking control method based on distributed reinforcement learning and a related device. The cooperative tracking control system based on distributed reinforcement learning is used for a distributed system formed by networking a plurality of photoelectric theodolites, the distributed system comprises a plurality of intelligent agents serving as independent tracking nodes, and each intelligent agent represents one photoelectric theodolite node. Each intelligent agent comprises a calculation processing unit connected with the servo control unit; the cooperative tracking control system comprises an intelligent agent reinforcement learning decision-making layer which is deployed in a calculation processing unit; the instruction processing layer is deployed in the calculation processing unit and is connected with the agent reinforcement learning decision-making layer; the servo control execution layer is deployed in a servo control unit and is connected with the instruction processing layer; according to the method, the adaptability of a multi-photoelectric theodolite system to a complex scene and the autonomy of cooperative tracking can be improved while the single-point fault risk of the center node is eliminated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of collaborative tracking technology, specifically to a collaborative tracking control method and related apparatus based on distributed reinforcement learning. Background Technology

[0002] Optical theodolites are key equipment in modern target range surveying, space target surveillance, and astronomical observation, using photoelectric sensors to measure angles and continuously track moving targets. As target speeds and maneuverability increase, a single theodolite, limited by its small field of view and blind spots, struggles to independently achieve continuous and stable tracking of high-speed maneuvering targets. To overcome the performance bottleneck of a single optical theodolite, research has shifted towards networking multiple theodolites to form a distributed system. By spatially distributing and interconnecting multiple theodolite nodes, they collaborate to form a sensing community, effectively achieving complementary fields of view, expanded observation range, and avoidance of blind spots. Using the angle intersection method, high-precision estimation of target spatial position can be achieved. In task allocation, multi-target tracking, dynamic role switching, and task reconfiguration are supported. While a single node failure can lead to system paralysis, a multi-node network system possesses strong fault tolerance and can adjust local strategies in scenarios with limited communication or sudden target changes. In applications such as intelligent traffic monitoring and aerospace target identification, multi-sensor network systems demonstrate superior overall performance compared to single-point observation systems.

[0003] However, most current mainstream multi-electro-optical theodolite collaborative tracking systems adopt a centralized control architecture. This architecture typically establishes a central node (such as a central server or master control station) responsible for aggregating data such as target miss distance and angle information collected by all theodolite nodes, running a fusion algorithm to generate a global situational awareness, uniformly calculating the expected tracking angle of each node, and then issuing control commands to each theodolite for execution. While this "centralized calculation, peripheral execution" model facilitates management and global optimal solution calculation, its inherent defects are also quite prominent. First, the central node is the "heart" of the entire system. Once it fails, is attacked, or the communication link is interrupted, the entire tracking network will be paralyzed, posing a serious single point of failure risk, and its survivability in high-confrontation battlefield environments is questionable. Second, this architecture lacks real-time adaptability. For complex dynamic scenarios such as the target's violent maneuvering or the temporary loss of signal due to cloud cover, the central node's command generation cycle is long and its strategies are fixed, making it difficult to respond to local changes in a timely manner, often leading to target loss or collaborative breakdown. Furthermore, each node passively executes instructions, lacks independent decision-making capabilities, and cannot make rapid adjustments based on local observation information, which restricts the overall robustness and flexibility of the system.

[0004] In summary, how to eliminate the risk of single-point failure at the central node while improving the adaptability of multi-electro-optical theodolite systems to complex scenarios and the autonomy of collaborative tracking has become an urgent technical challenge. Summary of the Invention

[0005] The purpose of this invention is to provide a collaborative tracking control method and related device based on distributed reinforcement learning, so as to overcome the problems existing in the prior art. It can eliminate the risk of single-point failure of the central node, while improving the adaptability of the multi-electro-optical theodolite system to complex scenes and the autonomy of collaborative tracking.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a cooperative tracking control system based on distributed reinforcement learning for a distributed system formed by networking several photoelectric theodolites. The distributed system includes several agents acting as independent tracking nodes, each agent representing a photoelectric theodolite node, and each agent including a computing processing unit connected to a servo control unit. The cooperative tracking control system includes: Several intelligent agents, each of which includes a computing processing unit connected to a servo control unit and a communication module, are deployed in different spatial locations as independent tracking nodes, and are interconnected with each other through the communication module. The agent reinforcement learning decision layer is deployed in the computing processing unit to generate multi-dimensional action vectors based on the agent's local observation information, communication information received from neighboring agents, and historical action sequences. The multi-dimensional action vectors include the agent's role settings, communication actions, and azimuth and pitch angle settings. The instruction processing layer, deployed in the computing processing unit, is connected to the agent's reinforcement learning decision layer. It is used to receive multi-dimensional action vectors, perform security checks and smoothing processes on the azimuth and pitch angle settings in sequence, and output the target angle value. The servo control execution layer, deployed within the servo control unit and connected to the instruction processing layer, is used to receive target angle values ​​and drive the servo control unit to track the target based on these target angle values.

[0007] Secondly, the present invention provides a distributed system formed by a network of photoelectric theodolites, including several photoelectric theodolites, including the aforementioned collaborative tracking control system, wherein the collaborative tracking control system is communicatively connected to the several photoelectric theodolites and is used to perform collaborative tracking control on the several photoelectric theodolites.

[0008] Thirdly, the present invention provides a cooperative tracking control method based on distributed reinforcement learning, comprising the following steps: Step 1: Each agent acquires its own local observation information and historical action sequence, receives communication information shared by neighboring agents, and constructs a joint observation state based on the local observation information, historical action sequence, and communication information. Step 2: Based on the joint observation state, the agent reinforcement learning decision layer generates a multi-dimensional action vector. Step 3: Receive the multi-dimensional motion vector through the instruction processing layer, perform safety verification and smoothing processing on the azimuth and pitch angle settings in sequence, and output the target angle value; Step 4: The servo control execution layer drives the agent to track the target based on the target angle value, and steps 1-4 are executed again until the tracking task ends.

[0009] In some embodiments, the local observation information includes: the agent's azimuth angle, the current value of the pitch angle, the target's miss distance, whether the target is within the field of view, and the relative position between the target and the agent.

[0010] In some embodiments, the generation of multi-dimensional action vectors by the agent reinforcement learning decision layer based on the joint observation state specifically includes: The joint observation state is input into the local Actor network through the agent reinforcement learning decision layer. The local Actor network maps the current policy and outputs a normalized multi-dimensional action vector.

[0011] In some embodiments, the training steps of the agent reinforcement learning decision layer specifically include: In the simulation environment, each of the agents performs steps 1-4 to generate empirical data; Using a centralized Critic network and team rewards, the Actor network of each agent is centrally trained based on empirical data to obtain the trained policy network parameters. The trained policy network parameters are deployed to each agent to complete the training of the agent's reinforcement learning decision layer.

[0012] In some embodiments, the smoothing process specifically includes: The azimuth and elevation angle settings after safety verification are filtered.

[0013] In some embodiments, the step of driving the agent to track the target via the servo control execution layer based on the target angle value specifically includes: The servo control execution layer generates a position deviation based on the target angle value, generates a speed command based on the position deviation, and drives the servo control unit to track the target by combining the speed command with the actual angular velocity.

[0014] Fourthly, the present invention provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.

[0015] Fifthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.

[0016] The above technical solution has the following advantages or beneficial effects: Firstly, this invention provides a cooperative tracking control system based on distributed reinforcement learning. Firstly, the invention employs a distributed architecture, eliminating the single-point-of-failure risk inherent in traditional centralized control. Each agent possesses independent decision-making capabilities. When some nodes fail or communication is disrupted, the remaining nodes can still continue the tracking task based on local observations and historical information, significantly improving the system's robustness and battlefield survivability. Secondly, this invention deeply integrates the adaptability and autonomous learning capabilities of reinforcement learning with the stability and reliability of servo control. Through the agent reinforcement learning decision layer, the system can dynamically learn the optimal cooperative strategy and adaptively adjust role allocation and communication behavior. Meanwhile, the safety verification and smoothing mechanisms of the instruction processing layer ensure the stability of servo instructions, effectively preventing severe jitter or actuator shocks caused by initial exploration in reinforcement learning or sudden environmental changes. Furthermore, through cooperative communication between agents, the system achieves enhanced global perception capabilities, enabling it to more effectively cope with complex scenarios such as target maneuvering or occlusion, significantly improving the accuracy and continuity of multi-target or multi-node cooperative tracking.

[0017] Secondly, this invention provides a distributed system formed by a network of photoelectric theodolites. By adopting the aforementioned collaborative tracking control system, the risk of single-point failure of the central node is eliminated, significantly improving the robustness and reliability of the system. At the same time, the system can effectively enhance the adaptability of multiple photoelectric theodolites in the face of complex scenarios, greatly improve the autonomy and intelligence level of collaborative tracking, and ensure the continuity and stability of target tracking. In addition, this invention uses the photoelectric theodolite as a distributed intelligent agent node, making full use of its high-precision measurement characteristics and deeply integrating it with reinforcement learning collaborative control.

[0018] Thirdly, this invention provides a cooperative tracking control method based on distributed reinforcement learning. First, this method achieves decentralized cooperative tracking through a distributed architecture. Each agent makes independent decisions based on local observations and neighbor information, eliminating the risk of single-point failure of the central node and significantly improving the system's robustness and survivability in complex environments. Second, this method achieves deep collaboration between reinforcement learning and servo control. The reinforcement learning decision layer endows the system with adaptive learning capabilities, enabling dynamic optimization of cooperative strategies; the safety verification and smoothing mechanisms of the instruction processing layer ensure the stability of servo control, effectively preventing sudden instruction changes from impacting the actuators. Furthermore, through information sharing and joint observation state construction among multiple agents, the system achieves global situational awareness, significantly improving the cooperative tracking accuracy and anti-occlusion capability of maneuvering targets, providing a reliable guarantee for continuous and stable tracking in highly dynamic scenarios.

[0019] In some embodiments, this invention introduces multi-dimensional local observation information such as azimuth angle, pitch angle, miss distance, field of view state, and relative position, enabling the agent to comprehensively perceive its own state and the relationship with the target. This significantly enhances the input richness of reinforcement learning decision-making, strengthens the agent's perception of target maneuvers and the environment, provides key data support for achieving accurate collaborative tracking and dynamic role allocation, and effectively improves tracking stability.

[0020] In some embodiments, this invention directly maps joint observation states to normalized action vectors through a local Actor network, achieving end-to-end collaborative decision-making. Normalization ensures uniformity in the scale of action outputs, improving the stability and convergence efficiency of reinforcement learning training. This invention enables agents to rapidly generate multi-dimensional instructions encompassing roles, communication, and angle control based on real-time perception, significantly improving decision response speed and the accuracy of collaborative tracking.

[0021] In some embodiments, this invention employs a centralized training and distributed execution framework, utilizing global information through a Critic network to evaluate joint policies, effectively addressing the non-stationarity problem. A team reward mechanism guides agents to learn cooperative optimal strategies, avoiding selfish behavior. Simulation-generated empirical data supports efficient offline training, and the trained lightweight Actor network can be directly deployed to each agent, balancing cooperative performance with online execution efficiency.

[0022] In some embodiments, the present invention smooths the setpoint using a filtering algorithm, effectively filtering out high-frequency noise and abrupt changes in the output of the reinforcement learning decision layer. This process significantly reduces instantaneous impact and mechanical wear on the servo system, avoiding tracking instability caused by severe command jitter. The smoothed angle input enables devices such as photoelectric theodolites to achieve smoother continuous tracking, improving the reliability and tracking accuracy of the system.

[0023] In some embodiments, the present invention converts the target angle value into a position deviation through a servo control execution layer, thereby generating a speed command. This command, combined with the actual angular velocity, forms a closed-loop control, achieving cascade control of the position and speed loops. This effectively improves the system's dynamic response performance and steady-state accuracy. The precise combination of speed feedforward and deviation adjustment ensures that equipment such as photoelectric theodolites can smoothly and quickly track moving targets.

[0024] Fourthly, the present invention provides a computer device that, through a processor executing a specific computer program, can efficiently implement the steps of the method of the present invention. When performing data processing tasks, the computer device can accurately perform numerical calculations and logical judgments, avoiding errors caused by human factors. At the same time, since the computer program has high stability and reliability, it can ensure the accuracy and consistency of the data processing results.

[0025] Fifthly, the present invention provides a computer-readable storage medium in which the steps of the method of the present invention are programmed into a computer program and stored on a computer-readable storage medium. Users can easily load these programs onto any compatible computer device and execute them without rewriting or converting the code, which greatly improves the convenience and flexibility of program execution. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of a cooperative tracking control system architecture based on distributed reinforcement learning, as shown in some embodiments of this specification. Figure 2 This is a schematic diagram of a cooperative tracking control method based on distributed reinforcement learning, as shown in some embodiments of this specification. Figure 3 This is a schematic diagram of the structure of a computer device according to some embodiments of this specification. Detailed Implementation

[0027] The present invention will be further described in detail below with reference to specific embodiments. These descriptions are for explanation purposes only and are not intended to limit the scope of the invention.

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] Most current mainstream multi-electro-optical theodolite collaborative tracking systems adopt a centralized control architecture. This architecture typically establishes a central node (such as a central server or master control station) responsible for aggregating data such as target miss distance and angle information collected by all theodolite nodes, running a fusion algorithm to generate a global situational awareness, uniformly calculating the expected tracking angle of each node, and then issuing control commands to each theodolite for execution. While this "centralized calculation, peripheral execution" model facilitates management and global optimal solution calculation, its inherent defects are also quite prominent. First, the central node is the "heart" of the entire system. Once it fails, is attacked, or the communication link is interrupted, the entire tracking network will be paralyzed, posing a serious single point of failure risk, and its survivability in high-confrontation battlefield environments is questionable. Second, this architecture lacks real-time adaptability. For complex dynamic scenarios such as the target's violent maneuvering or the temporary loss of signal due to cloud cover, the central node's command generation cycle is long and its strategies are fixed, making it difficult to respond to local changes in a timely manner, often leading to target loss or coordination breakdown. In addition, each node passively executes commands, lacking independent decision-making capabilities and unable to make rapid adjustments based on local observation information, which restricts the overall robustness and flexibility of the system.

[0031] Networking multiple photoelectric theodolites still faces the following challenges: Target maneuvering models and atmospheric disturbances are difficult to model accurately, requiring algorithms with strong robustness and adaptability; tracking tasks need to be allocated in real-time and dynamically based on the target's motion, station geometry, and the theodolite's own state; when multiple theodolites are tracking collaboratively, the timing and accuracy of the handover process affect the continuity of tracking; limited communication bandwidth prevents the sharing of global information, necessitating efficient communication strategies. Traditional collaborative tracking methods are mostly based on preset rules and optimal estimation theory. However, photoelectric theodolite systems involve multiple repetitive physical constraints such as zenith blind zones, dynamic accuracy, and servo limits. Traditional methods rely on accurate system modeling and prior environmental knowledge, which can complete tasks in simple scenarios but lack learning and adaptive capabilities, making it difficult to cope with complex, dynamic, and unknown environments. Furthermore, in highly dynamic, multi-target scenarios, computational complexity is high, making it difficult to meet real-time requirements.

[0032] This invention provides a cooperative tracking control system based on distributed reinforcement learning for a distributed system formed by networking several photoelectric theodolites. The distributed system includes several agents acting as independent tracking nodes, each agent representing a photoelectric theodolite node, and each agent includes a computing processing unit connected to a servo control unit. The cooperative tracking control system includes: The agent reinforcement learning decision layer is deployed in the computing processing unit to generate multi-dimensional action vectors based on the agent's local observation information, communication information received from neighboring agents, and historical action sequences. The multi-dimensional action vectors include the agent's role settings, communication actions, and azimuth and pitch angle settings. The instruction processing layer, deployed in the computing processing unit, is connected to the agent's reinforcement learning decision layer. It is used to receive multi-dimensional action vectors, perform security checks and smoothing processes on the azimuth and pitch angle settings in sequence, and output the target angle value. The servo control execution layer, deployed within the servo control unit and connected to the instruction processing layer, is used to receive target angle values ​​and drive the servo control unit to track the target based on these target angle values.

[0033] This invention provides a cooperative tracking control system based on distributed reinforcement learning. First, the invention employs a distributed architecture, eliminating the single-point-of-failure risk inherent in traditional centralized control. Each agent possesses independent decision-making capabilities. When some nodes fail or communication is disrupted, the remaining nodes can still continue the tracking task based on local observations and historical information, significantly improving the system's robustness and battlefield survivability. Second, this invention deeply integrates the adaptability and autonomous learning capabilities of reinforcement learning with the stability and reliability of servo control. Through the agent-based reinforcement learning decision layer, the system can dynamically learn the optimal cooperative strategy and adaptively adjust role allocation and communication behavior. Meanwhile, the safety verification and smoothing mechanisms of the instruction processing layer ensure the stability of servo commands, effectively preventing severe jitter or actuator shocks caused by initial exploration in reinforcement learning or sudden environmental changes. Furthermore, through cooperative communication between agents, the system achieves enhanced global perception capabilities, enabling it to more effectively handle complex scenarios such as target maneuvering or occlusion, significantly improving the accuracy and continuity of multi-target or multi-node cooperative tracking.

[0034] Example: This embodiment provides a cooperative tracking control system based on distributed reinforcement learning. See [link to relevant documentation]. Figure 1A distributed system for networking several photoelectric theodolites, the distributed system comprising several agents as independent tracking nodes, each agent representing a photoelectric theodolite node, and each agent including a computing processing unit connected to a servo control unit; the cooperative tracking control system includes: The agent reinforcement learning decision layer is deployed in the computing processing unit to generate multi-dimensional action vectors based on the agent's local observation information, communication information received from neighboring agents, and historical action sequences. The multi-dimensional action vectors include the agent's role settings, communication actions, and azimuth and pitch angle settings. The instruction processing layer, deployed in the computing processing unit, is connected to the agent's reinforcement learning decision layer. It is used to receive multi-dimensional action vectors, perform security checks and smoothing processes on the azimuth and pitch angle settings in sequence, and output the target angle value. The servo control execution layer, deployed within the servo control unit and connected to the instruction processing layer, is used to receive target angle values ​​and drive the servo control unit to track the target based on these target angle values.

[0035] In some embodiments, the cooperative tracking control system based on distributed reinforcement learning consists of N photoelectric theodolite nodes deployed at different locations and a communication network connecting them. After the nodes are deployed, a scanning observation mode is initiated. When a target is identified, the system switches to the cooperative tracking control framework proposed in this invention, entering a cooperative tracking control phase involving several intelligent agents. Each agent engages in limited data interaction through the communication network, relying on the control framework to achieve hierarchical management and control from the top layer to the middle layer to the bottom layer, forming a distributed cooperative tracking control system with both efficient perception and precise control capabilities.

[0036] In some embodiments, each photoelectric theodolite is an independent intelligent agent, which is equipped with a sensor unit, a servo control unit, a computing processing unit, and a communication module. The computing processing unit is connected to the sensor unit, the servo control unit, and the communication module, respectively.

[0037] In some embodiments, the plurality of said intelligent agents are interconnected through a communication network (communication module); the sensor unit includes its own local sensors.

[0038] In some embodiments, the agent reinforcement learning decision layer is the top layer, the instruction processing layer is the middle layer, and the servo control execution layer is the bottom layer.

[0039] In some embodiments, the agent reinforcement learning decision layer adopts a centralized training, distributed execution (CTDE) architecture. Each agent outputs a multi-dimensional action vector based on its own local sensor observations, communication information received from neighboring agents, and its own action history. The action vector includes the current node's role setting at the current moment, communication actions, and desired azimuth and pitch angle settings. Simultaneously, the system manages the handover process between agents according to preset interaction triggering conditions and protocols, ensuring smooth operation. Based on event triggering rules and the current strategy, the system dynamically schedules the communication behavior of the current node, determining the timing, target, and content of communication. Agents share team rewards during the training phase to learn collaborative strategies; during the execution phase, they make independent decisions.

[0040] In some embodiments, the specific algorithm for the agent's reinforcement learning decision layer is as follows: State space design: intelligent agent i At any moment t Input Includes local observation information (self-local observation). From neighboring agents j Received communication information (communication information observation) With historical action sequence .

[0041] Self-local observation Raw or preprocessed information from its own local sensors.

[0042] ; In the formula, Represents intelligent agents i The azimuth angle; Represents intelligent agents i The pitch angle; Indicates the miss distance in azimuth angle; The distance of the miss from the target, representing the pitch angle; Indicates whether it is within the line of sight of the photoelectric theodolite; Indicates the relative position of the target and the photoelectric theodolite on the horizontal axis; This indicates the relative position of the target and the photoelectric theodolite in terms of their vertical coordinates; This indicates the relative position of the vertical coordinates between the target and the photoelectric theodolite.

[0043] From neighboring agents j Received communication information The content should include at least the role settings of neighboring agents, the predicted location of the target, and tracking quality indicators.

[0044] Historical action sequence Including the agent itself recently The action vector output by the step.

[0045]

[0046] In the formula, Indicates an action.

[0047] Ultimately, the intelligent agent i The complete input is:

[0048] In the formula, Indicates the concatenation function; Indicates aggregation.

[0049] Motion space design: Each agent outputs a four-dimensional continuous action:

[0050] In the formula, This indicates the role settings, specifically the collaborative roles the intelligent agent intends to assume: primary tracker, auxiliary tracker, area searcher, or standby agent. ,in, For standby mode, remain stationary or in a low-power state. For regional searchers, execute pre-defined or learned search patterns. To assist trackers, it points to the predicted target area, provides redundant data, or prepares for takeover. The primary tracker is responsible for the main tracking task, keeping the target stable in the center of the field of view; This indicates a communication action, driven by individual agents. It represents the strength or priority of the node's willingness to communicate in the current period. This value, together with the event triggering mechanism, will ultimately determine whether a message is sent and the urgency of the message. ; Indicates the control action for the azimuth angle; Indicates the control action for pitch angle; and The range is limited to the mechanical angle limit of the photoelectric theodolite. Inside.

[0051] Reward function design: Design a multi-objective weighted reward function to simultaneously optimize tracking accuracy, handover smoothness, role rationality, and communication efficiency:

[0052] In the formula, Indicates the total reward; Indicates the weighting coefficient of the tracking reward; Indicates a tracking reward; Indicates the weighting coefficient of the handover reward; Indicates handover reward; Indicates the weighting coefficient of the role's reward; Indicates a character reward; The weighting coefficient representing the reward for communication efficiency; Indicates a reward for improved communication efficiency; This represents the weighting coefficient of the overall penalty; This indicates a comprehensive penalty. The aforementioned... Used to balance the relative importance of various indicators.

[0053] Tracking rewards Encourage the reduction of tracking error:

[0054] In the formula, Scale factor; A negative constant indicates a severe penalty for missing targets.

[0055] Handover Rewards Encourage a smooth handover: If the target is not lost during the handover process, and the new agent keeps the error within the threshold within ΔT after the handover, the handover is considered successful and a positive reward is obtained. In the formula, This indicates a positive reward; if the handover fails, In the formula, punish.

[0056] Character Rewards Encourage reasonable role allocation and avoid frequent switching. Grant additional rewards if the current agent successfully tracks for multiple consecutive periods. Impose minor penalties on unnecessary role switching.

[0057] Communication efficiency reward Encourage the effective use of communication resources and guide intelligent agents to communicate when information is valuable:

[0058] In the formula, Information value is indicated by factors such as reduced tracking error after communication, which suggests improved decision-making quality and higher information value. The communication bandwidth overhead is proportional to the frequency of message transmission. Indicates the first weight; This indicates the second weight.

[0059] Comprehensive penalty items This includes penalties for mechanical constraints and system failures. Negative rewards are applied to detected out-of-limit behaviors in angle, angular velocity, and angular acceleration. Furthermore, if all agents fail to track the target, a negative reward is applied at each step until any node captures the target.

[0060] Smooth handover mechanism: To ensure the continuity of target tracking during node handover, a smooth handover mechanism combining predictive and contingency measures is designed. The handover mechanism is triggered by two types of events: predictive handover triggering and contingency handover triggering.

[0061] Predictive cross-traffic triggering refers to initiating the action in advance when the target is about to leave the current primary tracker's field of view. The triggering condition is based on the prediction of the target's future position.

[0062] In the formula, This represents the main tracker (agent). i Based on current observations, the predicted target is... Position after time; Represents intelligent agents i The field of view; Represents intelligent agents j The field of view; This represents a predefined forecast time window.

[0063] The prediction condition indicates that the main tracker predicts the target will leave its field of view and enter another node (agent). j The field of view.

[0064] Emergency-type cross-traffic triggering refers to a situation where a problem occurs during current tracking, triggering an event that requires the primary tracker to be involved. i The continuous loss of targets exceeds the time threshold The following are possible causes: Detection of a main tracker server system failure, complete communication interruption, or other anomalies; the main tracker's tracking error consistently and significantly exceeding an acceptable threshold, while other nodes exhibit better tracking performance.

[0065] Once the handover trigger conditions are met, the system executes a standardized handover protocol to ensure an orderly and reliable process.

[0066] First, the current main tracker (agent). i Alternatively, a neighboring node that detects a fault broadcasts a handover request message, including the agent ID, trigger type, and the predicted state vector of the current target. Nodes receiving the handover information assess their own takeover capabilities, and the optimal candidate confirms the information with the broadcaster. Tracking control continues until the broadcaster receives the information, striving to keep the target within its field of view as much as possible. jThe agent begins to move, positioning its optical axis in advance to point towards the predicted boundary point, preparing for relay tracking, at the handover moment. Candidates (intelligent agents) j Officially put his character's actions Set as the primary tracker and begin closed-loop tracking based on its own sensors, the agent... i Change its role to either Assistant Tracker or Standby. After the handover is complete, the new primary tracker (agent)... j The system broadcasts a handover notification, updates the system status of all nodes, monitors the tracking performance over a period of time during the handover, and assesses whether the handover was successful.

[0067] Adaptive communication mechanism: The design employs an event-triggered adaptive communication mechanism. It abandons fixed-period broadcasting and adopts a hierarchical triggering mechanism based on event importance. Each event is assigned a priority, determining whether it is triggered and the urgency of transmission. Important priority events such as target loss, new target discovery, role change, and system anomalies require immediate broadcasting. Medium-priority events, such as significant changes in target state or tracking quality, are allowed to be sent with a small delay or after aggregation with other information. Low-priority events can be sent periodically at a fixed frequency, such as periodically synchronizing their own position, attitude, and role.

[0068] Simultaneously, based on the different roles of nodes and events, simple differentiated information should be designed to reduce unnecessary data transmission. For information from the primary tracker, the content should focus on the core information of the target; auxiliary trackers can transmit their own observations to supplement or verify the primary tracker's information; searchers report search status and findings.

[0069] Strengthen the learning operation mechanism: Reinforcement learning is a machine learning method that maximizes long-term cumulative rewards through trial and error by interacting with the system. In this system, each photoelectric theodolite acts as an agent, completing a closed-loop process of observation, decision-making, execution, and learning.

[0070] First, a multi-agent cooperative tracking training simulation system is constructed, simulating N photoelectric theodolite nodes, a moving target, and a communication network. During the training phase, a "centralized training" architecture is adopted. Each agent outputs actions based on local observations and communication within the simulation system. After each action is executed, the system calculates the team reward and system state for the next moment, using data such as observation information from each node. The generated empirical data is aggregated in a central trainer. The Critic network is responsible for evaluating the value of the system's joint actions and generating gradient signals; each agent's Actor network adjusts its own policy parameters based on these signals, thereby learning how to generate cooperative behaviors that yield higher rewards.

[0071] After training, the execution phase transitions to a "distributed execution" architecture. Each agent operates based on local observation information. and from neighboring agents j Received communication information It runs its pre-trained Actor network to generate actions in real time. This enables decentralized, autonomous, and collaborative tracking.

[0072] In some embodiments, the instruction processing layer is an instruction processing layer that receives the angle setting values ​​(azimuth angle and pitch angle setting values) output by the agent reinforcement learning decision layer, performs smoothing filtering and kinematic constraint verification on the azimuth angle and pitch angle setting values, and generates continuous and feasible target angle values.

[0073] In some embodiments, the input to the instruction processing layer is the azimuth and pitch angle settings output by the agent reinforcement learning decision layer. ], azimuth angle, and pitch angle control actions[ And mechanical constraint parameters (character settings).

[0074] Then, a safety check (kinematic constraint check) is performed, including angle, angular velocity, and angular acceleration checks, to avoid exceeding mechanical limits. Simultaneously, velocity smoothing is performed; if there are abrupt changes in control commands, they are smoothed and filtered to generate continuous, trackable target angle values. ],in, Indicates the target azimuth. This represents the target pitch angle, and the target angle value is sent to the underlying servo controller at a fixed frequency for execution.

[0075] In some embodiments, the servo control execution layer is a local servo control execution layer, with each node employing a servo control algorithm. It receives target angle values ​​from the instruction processing layer, calculates tracking errors in real time, and outputs torque commands to drive the azimuth and pitch axis servo motors of the intelligent agent, achieving rapid, stable, and high-precision pointing of the photoelectric theodolite (intelligent agent) line of sight to the target. After the photoelectric theodolite performs the pointing action, its relative geometric relationship with the target changes, thereby generating new local observation information and forming a closed loop.

[0076] In some embodiments, the servo control execution layer, based on mature servo control technology, is responsible for accurately executing the target angle value (angle command) generated by the command processing layer, thereby achieving high-precision motion control of the photoelectric theodolite. The command processing layer outputs [...]. The servo controller serves as the input to the servo control execution layer. It reads the actual angle feedback in real time and calculates the deviation between the actual angle and the set value. It drives the azimuth and pitch axis servo motors to achieve precise position tracking.

[0077] In some embodiments, the centralized training and distributed execution framework used in this embodiment can be replaced with other multi-agent reinforcement learning architectures, such as value decomposition networks.

[0078] In some embodiments, the multi-scale weighted reward function of this embodiment can be replaced by hierarchical reward, adaptive weighted reward, etc.

[0079] In some embodiments, the smooth handover mechanism of this embodiment may increase or decrease the triggering conditions and adjust the handover mechanism appropriately.

[0080] In some embodiments, this embodiment can be extended to more scenarios that require collaborative target tracking, such as radar networking systems, drone swarm collaborative tracking, and intelligent camera networks.

[0081] In one embodiment of the present invention, a distributed system formed by a network of photoelectric theodolites is provided, comprising a plurality of photoelectric theodolites, characterized in that it includes the aforementioned collaborative tracking control system, which is communicatively connected to the plurality of photoelectric theodolites and is used to perform collaborative tracking control on the plurality of photoelectric theodolites.

[0082] In one embodiment of the present invention, a cooperative tracking control method based on distributed reinforcement learning is provided, comprising the following steps: Step 1: According to the task requirements, deploy the intelligent agent (photoelectric theodolite node) in the designated area. After scanning the target, lock the target and prepare to start the collaborative observation task according to the collaborative tracking and control system.

[0083] Step 2, System Initialization: This includes agent initialization, simulation system initialization, Actor and Critic network initialization, experience replay buffer initialization, etc.

[0084] Step 3, Constructing a Joint Observation State: Each agent acquires its own local observation information through sensor units and receives communication information shared by neighboring agents through a communication link (communication network). Each agent fuses its local observation information, historical action sequences, and communication information to construct a joint observation state that includes global situational awareness. This serves as the input for top-level decision-making.

[0085] In some embodiments, the local observation information includes: the agent's azimuth angle, the current value of the pitch angle, the target's miss distance, whether the target is within the field of view, and the relative position between the target and the agent.

[0086] In some embodiments, step 3 specifically includes: intelligent agent i At any moment t Input Includes local observation information (self-local observation). From neighboring agents j Received communication information (communication information observation) With historical action sequence .

[0087] Self-local observation Raw or preprocessed information from its own local sensors.

[0088] ; In the formula, Represents intelligent agents i The azimuth angle; Represents intelligent agents i The pitch angle; Indicates the miss distance in azimuth angle; The distance of the miss from the target, representing the pitch angle; Indicates whether it is within the line of sight of the photoelectric theodolite; Indicates the relative position of the target and the photoelectric theodolite on the horizontal axis; This indicates the relative position of the target and the photoelectric theodolite in terms of their vertical coordinates; This indicates the relative position of the vertical coordinates between the target and the photoelectric theodolite.

[0089] From neighboring agents j Received communication information The content should include at least the role settings of neighboring agents, the predicted location of the target, and tracking quality indicators.

[0090] Historical action sequence Including the agent itself recently The action vector output by the step.

[0091]

[0092] In the formula, Indicates an action.

[0093] Ultimately, the intelligent agent i The complete input is:

[0094] In the formula, Indicates the concatenation function; Indicates aggregation.

[0095] Step 4, the agent's reinforcement learning decision layer generates action vectors: the agent's reinforcement learning decision layer generates joint observation states. The input is fed into the local Actor network, which maps the input according to the current policy and outputs a normalized multidimensional action vector (a high-dimensional continuous action vector).

[0096] In some embodiments, the multidimensional action vector includes the agent's role setting, communication actions, and azimuth and pitch angle settings.

[0097] In some embodiments, each agent outputs a four-dimensional continuous action:

[0098] In the formula, This indicates the role settings, specifically the collaborative roles the intelligent agent intends to assume: primary tracker, auxiliary tracker, area searcher, or standby agent. ,in, For standby mode, remain stationary or in a low-power state. For regional searchers, execute pre-defined or learned search patterns. To assist trackers, it points to the predicted target area, provides redundant data, or prepares for takeover. The primary tracker is responsible for the main tracking task, keeping the target stable in the center of the field of view; This indicates a communication action, driven by individual agents. It represents the strength or priority of the node's willingness to communicate in the current period. This value, together with the event triggering mechanism, will ultimately determine whether a message is sent and the urgency of the message. ; Indicates the control action for the azimuth angle; Indicates the control action for pitch angle; and The range is limited to the mechanical angle limit of the photoelectric theodolite. Inside.

[0099] Step 5: The instruction processing layer receives the multi-dimensional action vector output by the agent's reinforcement learning decision layer, and maps the multi-dimensional action vector into specific target angle values. During this process, the instruction processing layer performs safety verification (dynamic constraint verification) and smoothing processing on the azimuth and pitch angle settings, eliminates sudden signals that may cause mechanism oscillation, and ensures that the instructions output to the servo control execution layer are within the mechanical safety range.

[0100] In some embodiments, the smoothing process specifically includes: The azimuth and elevation angle settings after safety verification are filtered.

[0101] In some embodiments, the input to the instruction processing layer is the azimuth and pitch angle settings output by the agent reinforcement learning decision layer. ], azimuth angle, and pitch angle control actions[ And mechanical constraint parameters (character settings).

[0102] Then, a safety check (kinematic constraint check) is performed, including angle, angular velocity, and angular acceleration checks, to avoid exceeding mechanical limits. Simultaneously, velocity smoothing is performed; if there are abrupt changes in control commands, they are smoothed and filtered to generate continuous, trackable target angle values. ],in, Indicates the target azimuth. This represents the target pitch angle, and the target angle value is sent to the underlying servo controller at a fixed frequency for execution.

[0103] Step 6, Servo control execution layer execution: The servo control execution layer generates a position deviation based on the target angle value, generates a speed command based on the position deviation, and drives the servo control unit to track the target by combining the speed command with the actual angular velocity. Steps 1-6 are executed again until the tracking task ends.

[0104] In some embodiments, step 6 specifically includes: The servo control execution layer receives the target angle value sent by the instruction processing layer as the position loop input of the servo control execution layer. The PID controller calculates the position deviation to generate a speed command. Combined with the actual angular velocity feedback from the encoder and the current loop feedback, the servo motor is driven to rotate with high precision, so that the theodolite's line of sight is corrected in real time and accurately points to the target. Steps 1-6 are executed again until the tracking task ends.

[0105] In some embodiments, the servo control execution layer, based on mature servo control technology, is responsible for accurately executing the target angle value (angle command) generated by the command processing layer, thereby achieving high-precision motion control of the photoelectric theodolite. The command processing layer outputs [...]. The servo controller serves as the input to the servo control execution layer. It reads the actual angle feedback in real time and calculates the deviation between the actual angle and the set value. It drives the azimuth and pitch axis servo motors to achieve precise position tracking.

[0106] In some embodiments, the training steps of the agent reinforcement learning decision layer are described below. Figure 2 Specifically, it includes: (1) In the simulation environment, each of the agents performs steps 1-6 to generate experience data.

[0107] (2) Using a centralized Critic network and team rewards, the Actor network of each agent is trained in a centralized manner based on empirical data to obtain the trained policy network parameters.

[0108] In some embodiments, the Actor network for each agent is centrally trained based on empirical data using a centralized Critic network and team rewards, specifically including: Environmental Interaction and State Update: The simulation system or actual environment evolves the target's motion state based on the actions executed by all theodolite nodes. Each node re-observes the target, acquires information, calculates the current cooperative tracking accuracy and coverage, and generates a reward value for the current time step to evaluate the contribution of the current action to the cooperative task.

[0109] Experience storage: Each agent encapsulates the quadruple information of the current time step into an experience sample and stores it in the experience replay buffer to accumulate data support for subsequent network parameter updates.

[0110] Multi-layer network collaborative update: Random sampling is performed from the experience replay buffer to train the model. The Critic network is trained centrally using global information containing observations from all nodes to evaluate the value of joint actions. The feedback gradient of the Critic network is used to perform distributed updates on the Actor network of each node to continuously optimize the collaborative strategy of the top-level decision layer.

[0111] In some embodiments, the reward function is designed as follows: Design a multi-objective weighted reward function to simultaneously optimize tracking accuracy, handover smoothness, role rationality, and communication efficiency:

[0112] In the formula, Indicates the total reward; Indicates the weighting coefficient of the tracking reward; Indicates a tracking reward; Indicates the weighting coefficient of the handover reward; Indicates handover reward; Indicates the weighting coefficient of the role's reward; Indicates a character reward; The weighting coefficient representing the reward for communication efficiency; Indicates a reward for improved communication efficiency; This represents the weighting coefficient of the overall penalty; This indicates a comprehensive penalty. The aforementioned... Used to balance the relative importance of various indicators.

[0113] Tracking rewards Encourage the reduction of tracking error:

[0114] , Scale factor; A negative constant indicates a severe penalty for missing targets.

[0115] Handover Rewards Encourage a smooth handover: If the target is not lost during the handover process, and the new agent keeps the error within the threshold within ΔT after the handover, the handover is considered successful and a positive reward is obtained. In the formula, This indicates a positive reward; if the handover fails, In the formula, punish.

[0116] Character Rewards Encourage reasonable role allocation and avoid frequent switching. Grant additional rewards if the current agent successfully tracks for multiple consecutive periods. Impose minor penalties on unnecessary role switching.

[0117] Communication efficiency reward Encourage the effective use of communication resources and guide intelligent agents to communicate when information is valuable:

[0118] In the formula, Information value is indicated by factors such as reduced tracking error after communication, which suggests improved decision-making quality and higher information value. The communication bandwidth overhead is proportional to the frequency of message transmission. Indicates the first weight; This indicates the second weight.

[0119] Comprehensive penalty items This includes penalties for mechanical constraints and system failures. Negative rewards are applied to detected out-of-limit behaviors in angle, angular velocity, and angular acceleration. Furthermore, if all agents fail to track the target, a negative reward is applied at each step until any node captures the target.

[0120] This process employs a distributed execution strategy, where each node makes independent decisions based on its own observations, while a centralized Critic network during the training phase ensures consistency in collaboration.

[0121] (3) Determine whether the task has ended: If so, end the task, deploy the trained policy network parameters to each agent, and complete the training of the agent's reinforcement learning decision layer. Otherwise, repeat the training steps and enter the observation-decision-control loop of the next moment to continue performing the collaborative tracking task.

[0122] This method employs a distributed architecture, eliminating the risk of single-point failure at the central node. The failure of a single node only locally degrades system performance without causing a complete system crash; other nodes can quickly adapt through re-coordination. New theodolite nodes only need to integrate the same agent module and connect to the network to be integrated into the existing collaborative system through online learning or pre-trained models, without requiring reconstruction of the central algorithm. Furthermore, the decision-making model based on local observation and limited communication exhibits better robustness to network interruptions and communication delays. When communication temporarily fails, the agent can rely on its historical information and model to make reasonable independent decisions.

[0123] See Figure 3 In one embodiment of the present invention, a computer device is provided, comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, and is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in the computer storage medium to realize a corresponding method flow or corresponding function. The processor described in this embodiment of the present invention can be used for the operation of a cooperative tracking control method based on distributed reinforcement learning.

[0124] In one embodiment of the present invention, a computer-readable storage medium is provided, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device used to store programs and data. It is understood that the computer-readable storage medium here may include the built-in storage medium in the computer device, or it may include extended storage media supported by the computer device. The computer-readable storage medium provides storage space that stores the operating system of the terminal; and the storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions may be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here may be a high-speed RAM memory or a non-volatile memory, such as at least one disk storage device. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the cooperative tracking control method based on distributed reinforcement learning in the embodiment.

[0125] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0126] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0127] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxesFigure 1 The function specified in one or more boxes.

[0128] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A cooperative tracking control system based on distributed reinforcement learning, used in a distributed system formed by networking several photoelectric theodolites, characterized in that, The distributed system includes several intelligent agents as independent tracking nodes, each intelligent agent representing an optoelectronic theodolite node, and each intelligent agent including a computing processing unit connected to a servo control unit; The collaborative tracking control system includes: The agent reinforcement learning decision layer is deployed in the computing processing unit to generate multi-dimensional action vectors based on the agent's local observation information, communication information received from neighboring agents, and historical action sequences. The multi-dimensional action vectors include the agent's role settings, communication actions, and azimuth and pitch angle settings. The instruction processing layer, deployed in the computing processing unit, is connected to the agent's reinforcement learning decision layer. It is used to receive multi-dimensional action vectors, perform security checks and smoothing processes on the azimuth and pitch angle settings in sequence, and output the target angle value. The servo control execution layer, deployed within the servo control unit and connected to the instruction processing layer, is used to receive target angle values ​​and drive the servo control unit to track the target based on these target angle values.

2. A distributed system formed by networking photoelectric theodolites, comprising several photoelectric theodolites, characterized in that, The system includes the collaborative tracking control system as described in claim 1, which is communicatively connected to several photoelectric theodolites and is used for collaborative tracking control of the several photoelectric theodolites.

3. A cooperative tracking control method based on distributed reinforcement learning, characterized in that, The cooperative tracking control system based on distributed reinforcement learning as described in claim 1 includes the following steps: S1, each of the intelligent agents obtains its own local observation information and historical action sequence, receives communication information shared by neighboring intelligent agents, and constructs a joint observation state based on the local observation information, historical action sequence and communication information; S2, the agent reinforcement learning decision layer generates a multi-dimensional action vector based on the joint observation state; S3, receive the multi-dimensional motion vector through the instruction processing layer, perform safety verification and smoothing processing on the azimuth and pitch angle set values ​​in sequence, and output the target angle value; S4, the servo control execution layer drives the intelligent agent to track the target according to the target angle value, and executes S1-S4 again until the tracking task ends.

4. The cooperative tracking control method based on distributed reinforcement learning according to claim 3, characterized in that, The local observation information includes: the agent's azimuth angle, current pitch angle, target miss distance, whether the target is within the field of view, and the relative position between the target and the agent.

5. The cooperative tracking control method based on distributed reinforcement learning according to claim 3, characterized in that, The process of generating multi-dimensional action vectors based on the joint observation state through the agent reinforcement learning decision layer specifically includes: The joint observation state is input into the local Actor network through the agent reinforcement learning decision layer. The local Actor network maps the current policy and outputs a normalized multi-dimensional action vector.

6. The cooperative tracking control method based on distributed reinforcement learning according to claim 3, characterized in that, The training steps for the agent reinforcement learning decision layer specifically include: In the simulation environment, each of the agents executes S1-S4 to generate empirical data; Using a centralized Critic network and team rewards, the Actor network of each agent is centrally trained based on empirical data to obtain the trained policy network parameters. The trained policy network parameters are deployed to each agent to complete the training of the agent's reinforcement learning decision layer.

7. The cooperative tracking control method based on distributed reinforcement learning according to claim 3, characterized in that, The smoothing process specifically includes: The azimuth and elevation angle settings after safety verification are filtered.

8. The cooperative tracking control method based on distributed reinforcement learning according to claim 3, characterized in that, The step of driving the intelligent agent to track the target based on the target angle value through the servo control execution layer specifically includes: The servo control execution layer generates a position deviation based on the target angle value, generates a speed command based on the position deviation, and drives the servo control unit to track the target by combining the speed command with the actual angular velocity.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 3-8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 3-8.