A method and system for mobile ground station cooperative measurement and control deployment and resource scheduling

By combining multi-agent reinforcement learning and asynchronous Actor-Critic networks, the problem of collaborative deployment and resource scheduling of mobile ground stations under large-scale constellations is solved, achieving efficient resource scheduling, resolving resource conflicts and window waste, and improving computational efficiency and response speed.

CN122437586APending Publication Date: 2026-07-21XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIDIAN UNIV
Filing Date
2026-04-13
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing satellite telemetry and control technologies are insufficient for the coordinated deployment and resource scheduling of multiple mobile ground stations in large-scale constellations, resulting in frequent resource conflicts, low utilization rates, inability to adapt to dynamic scene changes, and high computational complexity and untimely response of traditional methods.

Method used

By employing a multi-agent reinforcement learning network combined with a dynamic constraint Voronoi diagram and an asynchronous Actor-Critic network, the responsibility domain and path of the mobile ground station are generated. The task weights are corrected by a time coefficient, and a deployment-scheduling closed-loop linkage mechanism is established to achieve dynamic collaborative deployment and scheduling.

Benefits of technology

It significantly improves resource utilization and task completion rate, reduces resource conflicts and window waste, improves computing efficiency and response speed, and adapts to complex dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122437586A_ABST
    Figure CN122437586A_ABST
Patent Text Reader

Abstract

The application discloses a kind of mobile ground station cooperation measurement and control deployment and resource scheduling method and system, belong to satellite measurement and control technical field.Method includes: multiple with dynamic constraint Voronoi diagram is combined multiple agent deterministic strategy gradient algorithm, complete mobile ground station's cooperation site selection and path deployment;Deployment result is used as input source, using the asynchronous Actor-Critic network of introduction time coefficient, the urgency and priority of task are jointly modeled, and the matching and execution order of measurement and control task are output;Deployment-scheduling closed-loop linkage mechanism is built, and scheduling result is reversely acted on deployment reward, and local re-planning is triggered when scene changes.The application solves the problems of low resource utilization and slow conflict resolution caused by limited coverage of traditional fixed stations and static deployment and scheduling through integrated optimization of space deployment and time scheduling, and maximizes measurement and control task benefits and completion rate in dynamic environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of satellite telemetry, tracking, and command (TT&C) technology, and specifically relates to a method and system for collaborative TT&C deployment and resource scheduling of mobile ground stations. Background Technology

[0002] The satellite tracking, telemetry, and command (TT&C) ground system is the core foundational support for the continuous operation of satellite missions, undertaking key tasks such as issuing TT&C commands, receiving telemetry data, measuring orbit and attitude, and ensuring on-orbit operation. In recent years, the scale of low-Earth orbit constellations has grown explosively, and the density and complexity of Earth observation and communication missions have continued to increase. The number of visible time windows between satellites and ground stations has increased significantly, and the visible time windows of different satellites and different ground stations highly overlap on the time axis. This has led to unprecedented conflicts and competition in the scheduling of TT&C resources, making it difficult for traditional TT&C technology systems to adapt to the needs of industry development.

[0003] Current telemetry, tracking, and command (TT&C) mission planning is based on the visible time windows of ground stations, employing a centralized, pre-planned approach to generate mission sequences and station allocation schemes. Mainstream technical approaches include rule-based and human experience-based scheduling, optimization modeling based on mathematical programming / constraint programming, and intelligent optimization or heuristic algorithms such as genetic algorithms, simulated annealing, and tabu search. These methods can only generate executable plans in static or quasi-static scenarios. Under dynamic conditions of increased constellation size, frequent mission arrivals, and continuously changing resource states, they expose numerous intractable problems: overlapping window conflicts are difficult to resolve quickly, leading to a large number of TT&C missions failing to execute normally; algorithms are prone to getting trapped in local optima, resulting in low utilization of TT&C resources and serious window waste; global replanning modes are costly and unresponsive, unable to adapt to real-time decision-making needs under dynamic disturbances. Furthermore, satellite TT&C mission planning is essentially a complex NP-hard problem with multiple constraint dimensions and a large search space, making it difficult for traditional methods to consistently obtain high-quality decision solutions within the project timeframe.

[0004] To alleviate the bottleneck caused by the limited number and coverage of fixed ground stations, mobile ground stations have been introduced as a mobile telemetry and control resource, hoping to achieve better visibility windows and higher mission benefits through location adjustments. However, existing technologies for the application of mobile ground stations have significant shortcomings. They treat deployment and scheduling separately, or only perform offline static deployment, lacking online coordination mechanisms that adapt to changing scenarios, thus failing to leverage the mobility advantages of mobile ground stations. Furthermore, they lack real-time decision support under incomplete information and dynamic disturbances (such as temporary task insertion, window updates, and changes in resource usage), failing to effectively realize the collaborative benefits of multiple mobile ground stations. The deployment results are disconnected from actual scheduling needs, and the core problems of resource conflicts and low utilization remain unresolved.

[0005] Furthermore, the existing satellite tracking and control system mainly relies on fixed tracking and control stations, lacking flexible redeployment capabilities. Limited by geographical location and the number of stations, it faces significant communication and computational pressures and has limited coverage, making it difficult to support the tracking and control needs of future large-scale on-orbit satellites. Traditional scheduling strategies, such as greedy algorithms and evolutionary optimization algorithms, show significantly lower overall tracking and control mission benefits with the dramatic increase in satellite scale, generating sparse tracking and control cycles that cannot meet the control requirements of observation duration, and thus cannot achieve optimal allocation and efficient utilization of tracking and control resources for large-scale constellations.

[0006] In summary, under the premise of limited resources, how to achieve collaborative dynamic deployment of multiple mobile ground stations, efficiently schedule telemetry and control resources, rationally select deployment locations and execution windows, balance task benefits and completion rates, realize collaborative allocation and conflict resolution among multiple tasks, and quickly respond to dynamic scene changes to complete real-time replanning has become a key technical problem that urgently needs to be solved in the field of large-scale constellation telemetry and control. There is an urgent need for a collaborative telemetry and control technology solution that integrates deployment and scheduling to break through the existing technical bottlenecks. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a method and system for collaborative telemetry and control deployment and resource scheduling of mobile ground stations, which addresses the shortcomings of the prior art and solves the technical problems of resource conflicts and low utilization in existing satellite telemetry and control missions.

[0008] The present invention adopts the following technical solution: A method for collaborative telemetry, tracking, and command (TT&C) deployment and resource scheduling of mobile ground stations includes the following steps: S1. Based on the service demand distribution within the task area, the multi-agent reinforcement learning network generates the responsibility domain of each mobile ground station using a dynamic constraint Voronoi diagram. Combined with the multi-agent deterministic policy gradient algorithm, it outputs the stopping point sequence and transfer path of each mobile ground station, thus completing the collaborative site selection and deployment of mobile ground stations. S2. Using the location trajectory and stopping point sequence output by the collaborative site selection deployment as input sources, determine the set of visible time windows; define a time coefficient for the task that increases as the current time approaches the end of the window, and adjust the task weight according to the time coefficient; use an asynchronous Actor-Critic network with introduced time coefficients to jointly model the ground station time window, task acceptance status, task priority, and shortest measurement and control duration constraint, and output the matching relationship between the measurement and control task and the ground station, the time window selection, and the execution order; S3. Generate a measurement and control action sequence and task allocation scheme according to the execution order, and build a deployment-scheduling closed-loop linkage mechanism: the scheduling result is used to apply to deployment rewards and strategy updates, and when a change in scene state is detected, a replanning mechanism is triggered to update the deployment and scheduling scheme.

[0009] Preferably, in step S1, the responsibility domain of each mobile ground station is generated using a dynamic constraint Voronoi diagram, specifically as follows: Using satellite-to-station visibility time window, mission priority, mission area environmental information, and mobile ground station status information as inputs, and under terrain, energy, and safety constraints, the responsibility domain of each mobile ground station is divided based on a comprehensive cost function. The comprehensive cost function is a weighted sum of the mobile ground station's maneuver time to reach the target location, maneuver energy consumption, service intensity at the location, environmental risk cost, and current load level. The service intensity is determined by the visibility time window and its weight, and the location... By the Comprehensive cost function for mobile ground station services for:

[0010] in, Indicates the first The mobile ground station has reached its location. maneuver time, Indicates motor energy consumption. Indicates position Service intensity at the location Indicates the environmental risk cost, Indicates the current load level. and These are the weighting coefficients.

[0011] Preferably, in step S1, after generating the responsibility domains of each mobile ground station, a candidate stopping point screening step is also included: A set of candidate stopping points is generated within each responsibility domain. The candidate stopping points consist of road network nodes, accessible area centers, service hotspot centers, and locations that meet the staying safety constraints. Candidate points that do not meet the terrain accessibility constraints, road reachability constraints, staying safety constraints, or remaining energy constraints are eliminated to obtain a set of valid candidate stopping points within the responsibility domain. The multi-agent deterministic policy gradient algorithm outputs a sequence of stopping points based on the set of valid candidate stopping points.

[0012] Preferably, in step S1, the multi-agent deterministic policy gradient algorithm adopts a centralized training and distributed execution mechanism, constructing an Actor network for each agent corresponding to a mobile ground station, and simultaneously constructing a centralized Critic network; during the training phase, the joint state and joint action are input into the Critic network, and during the execution phase, each agent outputs an action independently through the Actor network based on its own local state.

[0013] Preferably, in step S2, determining the set of visible time windows specifically involves: The coordinates and normal vectors of the ground station are transformed to the geocentric equatorial inertial coordinate system; the mapping relationship between satellite time and the angle of approach is derived based on the Kepler equations; the critical angle of approach corresponding to the connection and disconnection of the satellite-to-ground link is solved by combining the physical constraints of the minimum elevation angle of satellite-to-ground communication; the critical angle of approach is inversely solved to the corresponding time through the mapping relationship, the available visible time window of the mobile ground station and the target satellite is determined, and the visible time window set is integrated.

[0014] Preferably, in step S2, before joint modeling using the asynchronous Actor-Critic network with introduced time coefficients, the method further includes constructing a telemetry and control scheduling state. The telemetry and control scheduling state includes the available time window state of the ground station, the task execution state, the task time window information, the task priority vector, the task's shortest telemetry and control duration constraint, the task's remaining available time, and the task's remaining relaxation time. The task's remaining available time is the difference between the end time of the task window and the current time, and the task's remaining relaxation time is the difference between the task's remaining available time and the task's shortest telemetry and control duration.

[0015] Preferably, in step S2, the asynchronous Actor-Critic network is an asynchronous advantage Actor-Critic network, which constructs multiple parallel working threads. Each thread independently interacts with the scheduling environment replica and accumulates state, action, and reward samples. Each thread asynchronously calculates the policy gradient and value gradient and updates the global network parameters. The task weight is corrected according to the time coefficient, specifically by weighting the basic weight of the task with the time coefficient to obtain the comprehensive weight of the task. The comprehensive weight is used as the weight basis for the joint modeling of the asynchronous Actor-Critic network. For the Each task at time Define time coefficient for:

[0016] in, and Tasks The start and end times of the window. The time sensitivity coefficient, This is the adjustment constant.

[0017] Preferably, in step S3, the scheduling result is applied in reverse to the deployment rewards and policy updates, specifically as follows: The scheduling and execution results of the measurement and control tasks are fed back to the multi-agent reinforcement learning network. The scheduling and execution results are used as the basis for calculating the global reward function in the deployment phase. Based on the global reward function, the Actor network and Critic network of the multi-agent deterministic policy gradient algorithm are driven to perform parameter iterative updates.

[0018] Preferably, in step S3, the scene state change includes at least one of task insertion, time window update, road state change, environmental risk change, site resource state change, and ground station available time resource change; triggering a replanning mechanism to update the deployment and scheduling scheme specifically involves: recalculating the comprehensive cost function for the area affected by the scene state change and updating the responsibility domain division result for that area, and only performing local deployment adjustments for mobile ground stations in the affected area; updating the telemetry and control scheduling status in real time, re-outputting the scheduling decision using an asynchronous Actor-Critic network with a time coefficient, and generating a new telemetry and control action sequence and task allocation scheme based on the new scheduling decision.

[0019] Secondly, embodiments of the present invention provide a mobile ground station collaborative telemetry and control deployment and resource scheduling system, including: The deployment module is used to generate the responsibility domain of each mobile ground station based on the service demand distribution within the task area using a dynamic constraint Voronoi diagram. It combines a multi-agent deterministic policy gradient algorithm to output the stopping point sequence and transfer path of each mobile ground station, completes the collaborative site selection and deployment of mobile ground stations, and outputs the location trajectory and stopping point sequence of the collaborative site selection deployment. The scheduling module, connected to the deployment module, is used to receive the location trajectory and stopping point sequence, determine the set of visible time windows, define a time coefficient for the task that increases as the current time approaches the end of the window, and adjust the task weight according to the time coefficient. It uses an asynchronous Actor-Critic network with introduced time coefficient to jointly model the ground station time window, task acceptance status, task priority, and shortest measurement and control duration constraint, and outputs the matching relationship between the measurement and control task and the ground station, the time window selection, and the execution order. The linkage module is connected to the deployment module and the scheduling module respectively. It is used to receive the execution order and generate the measurement and control action sequence and task allocation scheme, feed back the scheduling result to the deployment module, so that the scheduling result can be applied to the deployment reward and strategy update, detect changes in the scene state and trigger the replanning mechanism, drive the deployment module to update the deployment scheme and drive the scheduling module to update the scheduling scheme.

[0020] Thirdly, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described mobile ground station collaborative telemetry and control deployment and resource scheduling method.

[0021] Fourthly, embodiments of the present invention provide a computer-readable storage medium including a computer program, which, when executed by a processor, implements the steps of the above-described mobile ground station collaborative telemetry and control deployment and resource scheduling method.

[0022] Fifthly, a chip includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described mobile ground station collaborative telemetry and control deployment and resource scheduling method.

[0023] In a sixth aspect, embodiments of the present invention provide an electronic device, including a computer program, which, when executed by the electronic device, implements the steps of the above-described mobile ground station collaborative telemetry and control deployment and resource scheduling method.

[0024] Compared with the prior art, the present invention has at least the following beneficial effects: A collaborative deployment and resource scheduling method for mobile ground stations in telemetry, tracking, and command (TT&C) is proposed. This method utilizes a dynamic constraint Voronoi diagram combined with the MADDPG algorithm to solve the collaborative site selection and path planning problem for multiple mobile ground stations, transforming spatial deployment into responsibility domain partitioning and path optimization. Subsequently, the deployment results are used as input, and an asynchronous Actor-Critic network with an introduced time coefficient is employed to solve the dynamic scheduling problem of TT&C tasks. Finally, a closed-loop linkage mechanism is established, enabling bidirectional iteration where scheduling feedback guides deployment optimization, and deployment provides a window for scheduling. Through the MADDPG algorithm, the system can achieve autonomous collaboration among multiple stations under terrain and environmental constraints, avoiding the "curse of dimensionality" problem that occurs in traditional centralized planning when the number of stations increases, significantly improving computational efficiency in large-scale scenarios. The introduced "time coefficient" dynamically reflects the urgency of tasks, allowing the scheduling strategy to automatically tilt towards high-urgency tasks before the window closes, effectively resolving conflicts caused by sudden task insertions. The closed-loop mechanism design allows the deployment strategy to be adjusted in real time based on scheduling benefits, avoiding the shortcomings of offline static deployment in adapting to dynamic environments. This online iterative optimization capability enables the system to maintain a high resource utilization rate and task completion rate even when faced with incomplete information and dynamic disturbances.

[0025] Furthermore, by quantifying physical constraints and mission requirements as costs, the defined responsibility domains not only consider spatial distance but also maneuverability and environmental risks. This decomposes the globally complex joint decision-making problem into multiple locally independent sub-problems, significantly reducing computational complexity. Each mobile station only needs to perform refined searches within its own responsibility domain, avoiding blind exploration across the entire area. This significantly improves the convergence speed and rationality of site selection and deployment, ensuring the maximization of overall telemetry and control benefits under limited energy and time constraints.

[0026] Furthermore, at the algorithmic level, by pre-screening invalid candidate points, the solution space is directly restricted to the feasible region, avoiding the algorithm outputting theoretically optimal but physically unreachable deployment locations. This reduces invalid search time and improves the algorithm's convergence efficiency. At the physical level, it ensures that the mobile ground station can actually reach and safely remain, fully considering the complexity of the actual terrain and the safety of the equipment. This "constraints are rules" design approach ensures that the paths and stopping points generated by the method fully conform to the requirements of the actual road network and geographical environment, eliminating the need for complex obstacle avoidance post-processing.

[0027] Furthermore, during the training phase, the Critic network utilizes global information for value evaluation, enabling the Actor network updates to take into account the influence of other agents and ensuring policy coordination. During the execution phase, each mobile station only needs to make independent decisions based on its own local observations, without relying on real-time instructions from the central node or high-frequency communication with other stations, exhibiting strong robustness and real-time performance. This approach leverages global information to improve training quality while retaining the flexibility and resilience of a distributed system, making it suitable for practical applications where communication between ground stations is limited or has high latency.

[0028] Furthermore, calculations based on the geocentric equatorial inertial coordinate system and Keplerian orbital dynamics fully consider the physical constraints of Earth's rotation, satellite orbital eccentricity, and minimum elevation angle on the communication link. This precise calculation based on a physical model can eliminate false visibility windows caused by geometric approximations, providing a reliable input source for subsequent scheduling algorithms. Accurate window prediction is a prerequisite for efficient scheduling, determining whether the telemetry and control mission can be successfully executed and avoiding mission failures or resource conflicts caused by window calculation errors.

[0029] Furthermore, introducing the feature of remaining relaxation time allows for a direct reflection of the scheduling margin of a task while meeting the minimum monitoring and control duration, enabling the algorithm to distinguish between urgent and lenient tasks. Combining task priority and resource status, the scheduling model can perceive the system's congestion level and task urgency in real time. This high-dimensional state description allows the reinforcement learning model to more accurately understand the current environment, thereby making decisions that better align with global interests. It solves the problems of coarse state descriptions and ineffective handling of overlapping time windows in traditional methods, improving the adaptability of the scheduling strategy in complex, time-varying environments.

[0030] Furthermore, the asynchronous parallel mechanism utilizes multi-threaded exploration of the environment, breaking the reliance on sample correlation in traditional deep reinforcement learning's experience replay pool. This improves sample utilization and convergence speed without requiring a complex distributed computing architecture. Most importantly, the introduction of a time coefficient transforms static task weights into dynamic weights, allowing the model to automatically prioritize tasks with closing windows during decision-making. This mimics the intuition of a human scheduler putting out fires, effectively preventing high-value tasks from timeouts due to delayed scheduling, thus maximizing overall task rewards in a dynamic environment.

[0031] Furthermore, by feeding back scheduling benefits to the deployment network, the location selection behavior of mobile stations no longer blindly pursues geometric centers or service hotspots, but rather pursues the maximization of schedulable benefits. If a location has many visible satellites but low actual benefits due to scheduling conflicts, the deployment network will reduce its preference for that location through parameter updates, ensuring the consistency of strategies between upper and lower layers, and achieving a globally optimal solution from the spatial-temporal dimension, rather than simply a local optimum.

[0032] Furthermore, by distinguishing between affected and unaffected areas, existing deployments are preserved, preventing network-wide disruptions caused by local disturbances. This incremental update mechanism is well-suited for highly dynamic and real-time scenarios like large-scale constellation telemetry and control, ensuring that the system maintains high service quality and mission continuity even with frequent mission arrivals or sudden environmental changes.

[0033] It is understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.

[0034] In summary, this invention addresses the rover site selection challenge by dynamically constraining Voronoi and MADDPG, optimizes telemetry and control scheduling using time coefficient correction and asynchronous AC, and constructs a bidirectional closed-loop feedback mechanism. Compared to existing technologies, it achieves a leap from static fragmentation to dynamic collaboration, effectively solving the problems of frequent resource conflicts, delayed response, and low returns under large-scale constellations, and significantly improving the intelligence level, resource utilization, and task completion rate of the telemetry and control system.

[0035] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0036] Figure 1 This is a flowchart of the method of the present invention; Figure 2 The following are comparison charts of satellite revenue in the simulation experiment: (a) is a comparison chart of revenue from 28 satellites, and (b) is a comparison chart of revenue from 40 satellites. Figure 3 A schematic diagram of a computer device provided in an embodiment of the present invention; Figure 4 This is a block diagram of a chip provided according to an embodiment of the present invention.

[0037] Among them, 60. Computer equipment; 61. Processor; 62. Memory; 63. Computer program; 600. Electronic device; 610. Processing unit; 620. Storage unit; 6201. Random access memory unit; 6202. Cache memory unit; 6203. Read-only memory unit; 6204. Program / utility; 6205. Program module; 630. Bus; 640. Display unit; 650. Input / output interface; 660. Network adapter; 700. External device. Detailed Implementation

[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0039] In the description of this invention, it should be understood that the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0040] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0041] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this invention generally indicates that the preceding and following objects have an "or" relationship.

[0042] It should be understood that although terms such as first, second, third, etc., may be used in the embodiments of the present invention to describe the preset range, these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from one another. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.

[0043] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."

[0044] The accompanying drawings illustrate various structural schematic diagrams according to embodiments disclosed in this invention. These drawings are not to scale, and some details have been enlarged for clarity, and some details may have been omitted. The shapes of the various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are merely exemplary and may deviate from reality due to manufacturing tolerances or technical limitations. Furthermore, those skilled in the art can design regions / layers with different shapes, sizes, and relative positions as needed.

[0045] This invention provides a method for collaborative deployment and resource scheduling of mobile ground stations for telemetry, tracking, and command (TT&C). Addressing issues such as frequent time window overlaps and conflicts, low ground resource utilization, lack of mobile deployment capabilities for fixed stations, and difficulty in timely replanning when scenarios change in large-scale constellation TT&C missions, this invention offers an integrated optimization method for dynamic deployment and resource scheduling of mobile ground stations. Under given terrain accessibility constraints and minimum elevation angle conditions, this invention uses the satellite-station visibility time window as the basic input to construct a unified decision-making closed loop for mobile station deployment and TT&C scheduling: multi-agent reinforcement learning enables collaborative site selection and path decision-making for multiple mobile stations, ensuring that each mobile station continuously selects locations with higher observation benefits within the planning period; and an online scheduling strategy based on Actor-Critic is adopted at the deployment points to jointly model constraints such as ground station time windows, task acceptance status, task windows, priorities, and minimum TT&C duration, outputting executable TT&C action sequences and task allocation schemes. The above technical solutions enable rapid response to changes in telemetry and control scenarios under dynamic disturbances and incomplete information conditions, reduce resource conflicts and window waste, improve the overall benefits and completion rate of telemetry and control tasks, and enhance the applicability and scalability of the method under multi-station, multi-satellite, complex constraints and large-scale task conditions.

[0046] Current satellite tracking, telemetry, and command (TT&C) mission scheduling is mostly based on fixed ground stations with pre-determined locations. Research focuses primarily on time window allocation, resource contention mitigation, and mission execution sequencing, making it difficult to proactively improve TT&C capabilities from a spatial deployment perspective. Unlike fixed ground stations, mobile ground stations have the capability to select and deploy locations, and their placement directly impacts the distribution of visible time windows, mission coverage, and subsequent scheduling benefits.

[0047] To address the characteristics of mobile ground stations—namely, their selectability, maneuverability, and dynamic reconfigurability of coverage—this paper models the spatial location deployment problem in conjunction with the telemetry, tracking, and command (TT&C) task scheduling problem. This allows for the interaction and collaborative optimization of ground station parking locations, utilization of visible time windows, and task execution sequencing. A dynamic constraint Voronoi diagram is used to generate responsibility domains and constrain candidate parking areas. MADDPG is employed to complete the collaborative deployment decision for multiple mobile ground stations. An asynchronous Actor-Critic network with a time coefficient is introduced to achieve online task scheduling under dynamic time window conditions, thereby improving the efficiency of TT&C resource utilization and the ability to ensure critical missions in complex scenarios.

[0048] Please see Figure 1 The present invention discloses a method for collaborative telemetry, monitoring, and control deployment and resource scheduling of mobile ground stations, comprising the following steps: S1. The multi-agent reinforcement learning network generates the responsibility domain division results, stopping point sequence and transfer path of each mobile ground station based on the satellite visibility time window, mission area environmental information and mobile ground station status information, and outputs the deployment location of each mobile ground station. Calculate the visible time window First, define the ground station. The coordinates in the geocentric rotating coordinate system are set as follows: The corresponding surface normal vector is denoted as To ensure the effective transmission of telemetry and control signals, the minimum elevation angle for satellite-to-ground communication is set as follows: Let the coordinates of the ground station in the geocentric equatorial inertial coordinate system (ECI) be... ,set up The Greenwich Mean Time at that moment is If we consider the Earth as an ellipsoid rotating at a uniform angular velocity, its average equatorial radius is... The polar radius is By transforming the coordinate system, we can obtain: (1) in, Let be the Earth's rotational angular velocity. Similarly, the transformation equation of the ground station normal vector in the ECI coordinate system can be obtained: (2) Assume the target satellite is in The angle of approach at time is . For satellites in The approximate angles at time t are subject to the following relationship in Kepler's equations: (3) The time can be derived from equation (3). Angle with near point Relationship equations: (4) Let the satellite's vector relative to the ground station be equal to the ground station's normal vector. The included angle between them is Then the following geometric relationship exists: (5) Based on minimum elevation angle Due to physical constraints, the geometric conditions for establishing a satellite-to-ground communication link are as follows: when the relative angle between the satellite and the ground station satisfies the visibility condition (i.e., the satellite is within the effective coverage area of ​​the ground station), the satellite and the ground station can communicate normally; otherwise, the communication link is broken. Based on the above connectivity determination conditions, the critical proximity angles at which the satellite-to-ground link is just established and just broken are calculated. The solution. Finally, through the established near-point angle and time... The mapping relationship is used to solve the critical proximity angle into the corresponding specific time, thereby accurately obtaining the available visible time window between the mobile ground station and the target satellite.

[0049] MADDPG Site Selection and Deployment Method under Dynamically Constrained Voronoi Diagram Mobile ground station site selection and deployment is used to determine the responsibility area, stationing location, and transfer path of each mobile ground station within the planning period. Taking satellite-to-station visibility time window, mission priority, mission area environmental information, and mobile ground station status information as inputs, and under terrain, energy, and safety constraints, it generates site selection and deployment results for collaborative services of multiple mobile ground stations. The outputs include the responsibility domain division, stationing point sequence, and transfer path for each mobile ground station, providing a deployment basis for subsequent telemetry, tracking, and command (TT&C) mission scheduling.

[0050] (1) Generate dynamic constraints Voronoi responsibility domain partitioning and candidate stopping points To achieve coordinated coverage of the mission area by multiple mobile ground stations, a responsibility domain is first generated based on the distribution of service demands within the mission area. Let the planning time be... The deployable area is , No. The responsibility domain of each mobile ground station is denoted as The responsibility domain is defined as follows: (6) in, Indicates position By the The comprehensive cost function for providing services from a mobile ground station is defined as follows: (7) In the formula, Indicates the first The mobile ground station has reached its location. maneuver time, Indicates motor energy consumption. Indicates position Service intensity at the location Indicates the environmental risk cost, Indicates the current load level. and These are the weighting coefficients.

[0051] Service intensity Determined by the visible time window and its weight, it is expressed as: (8) in, Indicates position At any moment The set of available service time windows Indicates the first The weight of each time window, Indicates the position of the time window. Coverage contribution coefficient at the location.

[0052] In this way, the mission area is divided into multiple responsibility domains according to the principle of minimizing the overall service cost, so that different mobile ground stations are given priority to be responsible for the measurement and control support tasks in their respective responsibility domains, thereby decomposing the global-local learning problem into multiple local collaborative deployment sub-problems.

[0053] After the responsibility domains are divided, a set of candidate stopping points is generated within each responsibility domain. Let the first... A mobile ground station at time The set of candidate stopping points is: (9) Candidate stopping points can consist of road network nodes, accessible area centers, service hotspot centers, and locations that satisfy the stay safety constraints. Candidate points that do not meet the terrain accessibility constraints, road accessibility constraints, stay safety constraints, or remaining energy constraints are eliminated, resulting in a set of valid stopping points within the responsibility domain. This set is used for subsequent multi-agent collaborative deployment decisions.

[0054] (2) MADDPG collaborative site selection and deployment Each mobile ground station is modeled as an intelligent agent, let the mobile ground station be... The total number is . No. An intelligent agent at time The local state is defined as The joint state is represented as , Indicates the first Current location of the mobile ground station Indicates its remaining energy. Indicates the current domain of responsibility. This represents the set of remaining locations to be covered within the responsibility domain.

[0055] Mobile ground station time The deployment location is determined by the time. The direction and distance of movement are jointly determined. Define the first... An intelligent agent at time The action is ,in, Joint actions are represented as .

[0056] Based on the action output, the first The location of each mobile ground station has been updated to: (10) (3) Reward function and network parameter update Based on the calculation results of the visible time window, at time... The set of time windows corresponding to candidate deployment locations and their weights can be obtained. Assume there are a total of [number missing] time windows within a decision-making period. A time window is defined, and the basic revenue item is: (11) in, Indicates the first Accept / reject decision within a time window Indicates the first Each time window has a weight, which is determined by the time window length, task priority, or a preset evaluation index.

[0057] To ensure that deployment decisions take into account benefits, energy consumption, risks, and load balancing, a global reward function is constructed; (12) in, Indicates the first The energy consumption cost of mobile ground stations Indicates the risk and cost. This indicates a penalty for unbalanced load or coordination conflicts. and These are the weighting coefficients.

[0058] Multi-agent cooperation adopts global reward As a training signal, the deployment decisions of each mobile ground station are aligned with the overall benefit optimization.

[0059] For each intelligent agent Construct an Actor network Used to output deterministic actions based on local states. Simultaneously, a centralized Critic network was constructed. During the training phase, the joint state and joint action are input to evaluate the value of multi-agent cooperative behavior. Let the joint state be... joint action The Critic network is based on For input and output This value is used to guide policy updates for each Actor network.

[0060] This architecture employs a centralized training and distributed execution mechanism. During the training phase, the Critic network utilizes joint states and joint actions to complete value assessment and policy learning; during the execution phase, each agent can independently output actions through the Actor network based solely on its own local state, thereby meeting the real-time requirements of online collaborative deployment of multiple mobile ground stations.

[0061] Establish an experience replay pool Samples generated during environmental interaction Write to the replay pool. On each update, retrieve from the experience replay pool. Extraction Batch training is performed on a group of samples.

[0062] For the There are three agents, and the objective value is defined as: (13) in, As a discount factor, and These represent the parameters of the target Actor network and the target Critic network, respectively.

[0063] The Critic network updates its parameters by minimizing the loss function: (14) The Actor network updates its parameters based on the gradient of the Critic's evaluation of the joint action. This will enhance the overall benefits of the joint deployment strategy.

[0064] To improve training stability, target networks are set for both the Actor and Critic networks, and the parameters of the target networks are updated using a soft update method: (15) in, This is the soft update coefficient.

[0065] After training is completed, during the deployment phase, each mobile ground station will operate according to its local conditions. Actions are directly output from the Actor network. The system obtains the deployment location for the next moment; then it updates the set of remaining locations to be covered within the responsibility domain, the visible time window and its weight information, and proceeds to the next round of decision-making. The entire process cyclically executes "state update - action output - location update - time window update", forming a dynamic deployment closed loop.

[0066] When task insertion, time window updates, road condition changes, environmental risk changes, or site resource status changes are detected, the comprehensive cost function and responsibility domain division results for the affected area are recalculated, and the local state of the corresponding agents is updated. Only local deployment adjustments are performed on the affected areas. Unaffected areas retain their original deployment results, thereby reducing overall replanning overhead and improving online response capabilities.

[0067] S2. Using the set of visible time windows corresponding to the deployment location and their weights as input, and using an asynchronous Actor-Critic network with introduced time coefficients, jointly model the ground station time window, task acceptance status, task priority and shortest measurement and control duration constraint, and output the matching relationship between the measurement and control task and the ground station, the time window selection and execution order. Measurement and Control Task Scheduling Optimization Model After completing the site selection and deployment of mobile ground stations, to address the telemetry, tracking, and command (TT&C) conflicts caused by overlapping visible time windows of multiple satellites under large-scale constellation conditions, a TT&C task scheduling optimization model incorporating a time coefficient was established. Using task time window information, available time resources at ground stations, task priority information, and the minimum TT&C duration constraint as inputs, an online scheduling model was constructed using the asynchronous advantage A3C method incorporating a time coefficient, ensuring that scheduling decisions simultaneously reflect both task importance and time urgency.

[0068] (1) Scheduling state and action space At the discrete decision time Constructing the measurement, control, and scheduling status Defined as: (16) in, Indicates the time of each ground station Available time window status, Indicates the task execution status. This indicates task time window information. Represents the task priority vector. This indicates the minimum control and measurement time constraint for the task. Indicates the remaining available time for the task. This indicates the remaining slack time for the task.

[0069] For the For each task, the remaining available time and remaining relaxation time are defined as follows: (17) in, Indicates task At the end of the window, Indicates task The shortest monitoring and control duration. Through state representation, a unified description of task schedulability, resource availability, window timeliness, and time urgency is provided. At the decision-making moment... The scheduling action is defined as follows: action This indicates the selection of the ground station, execution time window, and start time for a given telemetry and control mission, or outputs a rejection / delay decision. Actions that do not meet resource exclusivity constraints, time window constraints, minimum telemetry and control duration constraints, and execution feasibility conditions are constrained using action masks, and a feasible action discrimination function is defined: (18) Actor networks only satisfy the following conditions: Output policy probabilities on the action set.

[0070] (2) Construction of time coefficient and return function To depict the dynamic urgency of the task before the window closes, the first... Each task at time Define time coefficient for: (19) in, and Tasks The start and end times of the window. The time sensitivity coefficient, This is the adjustment constant.

[0071] As the current moment approaches the end of the mission window, Gradually increase the weight of urgent tasks in scheduling decisions.

[0072] Set a task At any moment The basic weight is The comprehensive weighting considering time urgency is defined as follows: (20) Execute action Afterwards, the environment returns immediate feedback. Define the instant reward function as follows: (twenty one) in, This indicates the number of task windows participating in scheduling within the current decision-making cycle. Indicates task At any moment Whether it is accepted for execution Indicates resource conflict penalties. This indicates a penalty for delaying the task. Indicates a resource switching penalty. The weight of the penalty item.

[0073] (twenty two) This reward function is used to uniformly measure task benefits, time urgency, and scheduling costs.

[0074] (3) A3C scheduling network and asynchronous training mechanism A scheduling policy network is constructed using the asynchronous advantage Actor-Critic algorithm. The Actor network is represented as follows: Used in state The output action probability distribution is shown below. The Critic network is represented as... This is used to estimate the value function of the current state. The cumulative discounted return is defined as: (twenty three) in, As a discount factor, Estimate the step size for the return. The advantage function is defined as: (twenty four) The Actor network updates its parameters based on the dominance function. This increases the probability of selecting high-dominance actions; the Critic network updates parameters based on the state value estimation error. This is to improve the accuracy of the value function approximation.

[0075] Multiple parallel worker threads are constructed, each interacting independently with a corresponding copy of the scheduling environment, accumulating state, action, and reward samples on a local trajectory. For any thread, in the state... downsampling action After execution, a state transition sample is obtained. .

[0076] Each thread independently calculates the policy gradient and value gradient, and updates the parameters of the global Actor network and Critic network asynchronously. The Critic network minimizes... Update parameters; Actor network based on dominance function Update the strategy parameters.

[0077] By using asynchronous parallel interaction and parameter updates, the training efficiency of the scheduling strategy and its adaptability to dynamic scenarios are improved.

[0078] S3. Generate a measurement and control action sequence and task allocation scheme according to the execution order, and trigger a replanning mechanism to update the deployment and scheduling scheme when a change in scene state is detected.

[0079] Based on the matching relationship between the telemetry and control tasks and the ground station, the time window selection, and the execution order output in step S2, a corresponding telemetry and control action sequence and task allocation scheme are generated. The scheme includes the final matching relationship between the tasks and the ground station, the execution time window of each task, the execution order, and the corresponding ground station resource occupation arrangement, supporting the online collaborative scheduling of multi-satellite telemetry and control tasks within the planning period.

[0080] Real-time detection of changes in scenario status, such as task insertion, time window updates, road condition changes, environmental risk changes, station resource status changes, ground station available time resources changes, and task window changes.

[0081] During the online operation phase, the status is updated in real time when a task arrives, the task window changes, resource usage changes, or the available time resources of the ground station change. The Actor network outputs scheduling actions based on the current state. The task involves completing ground station matching, time window selection, and execution order determination. After the action is executed, the task execution status, resource occupancy status, remaining available time, and time coefficient are updated, and then the next round of decision-making begins.

[0082] During the replanning process, the updated deployment results are used as the input source for the time window of the scheduling phase. At the same time, the updated scheduling results are used inversely to affect the reward function and MADDPG network parameter updates of the deployment phase. This allows the deployment decision to be iteratively optimized based on the scheduling results. The scheduling decision obtains better visibility time window support based on the updated deployment location, forming an online iterative optimization with a closed-loop linkage between deployment and scheduling. This continuously improves the measurement and control benefits and task completion rate in a dynamic environment.

[0083] The output of the telemetry, tracking, and command (TT&C) mission scheduling optimization module includes: the matching relationship between missions and ground stations, the selected execution time window for the missions, the mission execution order, and the corresponding resource allocation. These outputs are used to support the online collaborative scheduling of multi-satellite TT&C missions within the planning period.

[0084] In another embodiment of the present invention, a mobile ground station collaborative telemetry and control deployment and resource scheduling system is provided. This system can be used to implement the above-mentioned mobile ground station collaborative telemetry and control deployment and resource scheduling method. Specifically, the mobile ground station collaborative telemetry and control deployment and resource scheduling system includes a deployment module, a scheduling module, and a linkage module.

[0085] The deployment module is used to generate the responsibility domain of each mobile ground station based on the service demand distribution within the task area using a dynamic constraint Voronoi diagram. It combines a multi-agent deterministic policy gradient algorithm to output the stopping point sequence and transfer path of each mobile ground station, completes the collaborative site selection and deployment of mobile ground stations, and outputs the location trajectory and stopping point sequence of the collaborative site selection deployment. The scheduling module, connected to the deployment module, is used to receive the location trajectory and stopping point sequence, determine the set of visible time windows, define a time coefficient for the task that increases as the current time approaches the end of the window, and adjust the task weight according to the time coefficient. It uses an asynchronous Actor-Critic network with introduced time coefficient to jointly model the ground station time window, task acceptance status, task priority, and shortest measurement and control duration constraint, and outputs the matching relationship between the measurement and control task and the ground station, the time window selection, and the execution order. The linkage module is connected to the deployment module and the scheduling module respectively. It is used to receive the execution order and generate the measurement and control action sequence and task allocation scheme, feed back the scheduling result to the deployment module, so that the scheduling result can be applied to the deployment reward and strategy update, detect changes in the scene state and trigger the replanning mechanism, drive the deployment module to update the deployment scheme and drive the scheduling module to update the scheduling scheme.

[0086] This invention provides a terminal device comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to achieve corresponding method flows or corresponding functions. The processor described in this embodiment can be used in the operation of a mobile ground station collaborative telemetry and control deployment and resource scheduling method, including: A multi-agent reinforcement learning network generates the responsibility domain of each mobile ground station based on the service demand distribution within the task area using a dynamically constrained Voronoi diagram. Combined with a multi-agent deterministic policy gradient algorithm, it outputs the dwell point sequence and transfer path of each mobile ground station, completing the collaborative site selection and deployment of the mobile ground stations. The location trajectory and dwell point sequence output by the collaborative site selection and deployment are used as input sources to determine the set of visible time windows. A time coefficient is defined for each task, increasing as the current time approaches the end of the window, and the task weight is adjusted based on this time coefficient. An asynchronous Actor-Critic network with an introduced time coefficient is used to jointly model the ground station time window, task acceptance state, task priority, and shortest measurement and control duration constraint, outputting the matching relationship between measurement and control tasks and ground stations, time window selection, and execution order. Based on the execution order, a measurement and control action sequence and task allocation scheme are generated, and a deployment-scheduling closed-loop linkage mechanism is constructed: the scheduling result is used to influence deployment rewards and policy updates, and a replanning mechanism is triggered to update the deployment and scheduling scheme when a change in scene state is detected.

[0087] Please see Figure 3The terminal device is a computer device. In this embodiment, the computer device 60 includes a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When executed by the processor 61, the computer program 63 implements the mobile ground station collaborative telemetry and control deployment and resource scheduling method described in this embodiment. To avoid repetition, these details are not elaborated here. Alternatively, when executed by the processor 61, the computer program 63 implements the functions of each model / unit in the mobile ground station collaborative telemetry and control deployment and resource scheduling system of this embodiment. To avoid repetition, these details are not elaborated here.

[0088] Computer device 60 can be a desktop computer, laptop, handheld computer, cloud server, or other computing device. Computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art will understand that... Figure 3 This is merely an example of computer device 60 and does not constitute a limitation on computer device 60. It may include more or fewer components than shown, or combine certain components, or different components. For example, computer device may also include input / output devices, network access devices, buses, etc.

[0089] The processor 61 may be a Central Processing Unit (CPU), or other general-purpose processors, graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0090] The memory 62 can be an internal storage unit of the computer device 60, such as a hard disk or memory of the computer device 60. The memory 62 can also be an external storage device of the computer device 60, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the computer device 60.

[0091] Furthermore, the memory 62 may include both internal storage units of the computer device 60 and external storage devices. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 can also be used to temporarily store data that has been output or will be output.

[0092] Please see Figure 4 The terminal device is an electronic device 600, which is manifested in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including storage unit 620 and processing unit 610), a display unit 640, etc.

[0093] The storage unit stores program code, which can be executed by the processing unit 610 to perform the steps described in the method section of this specification according to various exemplary embodiments of the present invention. For example, the processing unit 610 can perform actions such as... Figure 1 The steps are shown in the figure.

[0094] Storage unit 620 may include a readable medium in the form of a volatile storage unit, such as random access memory (RAM) 6201 and / or cache memory 6202, and may further include a read-only memory (ROM) 6203.

[0095] Storage unit 620 may also include a program / utility 6204 having a set (at least one) program module 6205, such program module 6205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0096] Bus 630 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the multiple bus structures.

[0097] Electronic device 600 can also communicate with one or more external devices 700 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 600, and / or with any device that enables electronic device 600 to communicate with one or more other computing devices (e.g., router, modem). This communication can be performed via input / output interface 650. Furthermore, electronic device 600 can also communicate with one or more networks (e.g., local area network, wide area network, and / or public network, such as the Internet) via network adapter 660. Network adapter 660 can communicate with other modules of electronic device 600 via bus 630. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.

[0098] Example 4 This invention also provides a storage medium, specifically a computer-readable storage medium, which is a memory device in a terminal device for storing programs and data. It is understood that the computer-readable storage medium here can include both built-in storage media in the terminal device and extended storage media supported by the terminal device; it can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor, which can be one or more computer programs (including program code). More specific examples of the computer-readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, random access memory, read-only memory, erasable programmable read-only memory, optical fiber, portable compact disk read-only memory, optical storage device, magnetic storage device, or any suitable combination thereof.

[0099] Computer-readable storage media also include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium can also be any readable medium other than a readable storage medium that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, radio frequency, etc., or any suitable combination thereof.

[0100] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0101] One or more instructions stored in a computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the mobile ground station collaborative telemetry and control deployment and resource scheduling method in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by the processor in the following steps: A multi-agent reinforcement learning network generates the responsibility domain of each mobile ground station based on the service demand distribution within the task area using a dynamically constrained Voronoi diagram. Combined with a multi-agent deterministic policy gradient algorithm, it outputs the dwell point sequence and transfer path of each mobile ground station, completing the collaborative site selection and deployment of the mobile ground stations. The location trajectory and dwell point sequence output by the collaborative site selection and deployment are used as input sources to determine the set of visible time windows. A time coefficient is defined for each task, increasing as the current time approaches the end of the window, and the task weight is adjusted based on this time coefficient. An asynchronous Actor-Critic network with an introduced time coefficient is used to jointly model the ground station time window, task acceptance state, task priority, and shortest measurement and control duration constraint, outputting the matching relationship between measurement and control tasks and ground stations, time window selection, and execution order. Based on the execution order, a measurement and control action sequence and task allocation scheme are generated, and a deployment-scheduling closed-loop linkage mechanism is constructed: the scheduling result is used to influence deployment rewards and policy updates, and a replanning mechanism is triggered to update the deployment and scheduling scheme when a change in scene state is detected.

[0102] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0103] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0104] Alternative Option 1 The original scheme used MADDPG (Multi-Agent Deep Deterministic Policy Gradient Algorithm) for cooperative location selection and deployment, which can be replaced with MAPPO (Multi-Agent Proximal Policy Optimization). The MAPPO algorithm, through proximal policy optimization, effectively solves the mode collapse problem that easily occurs in the MADDPG algorithm during training, improving training stability. At the same time, MAPPO has higher sample utilization, achieving faster convergence speed with the same training samples and reducing computational overhead during the training phase. The core steps of this alternative remain unchanged; only the MADDPG network is replaced with a MAPPO network, and the network's policy update method and loss function are redefined. This alternative is suitable for engineering scenarios with high requirements for algorithm training stability.

[0105] Alternative Solution 2 The original scheme used an A3C (Asynchronous Advantage Actor-Critic) network to schedule telemetry and control tasks. This can be replaced with a hybrid scheme combining DDPG (Deep Deterministic Policy Gradient Algorithm) with time series forecasting. First, an LSTM (Long Short-Term Memory) network is used to predict the changing trends of the satellite's visible time window. The prediction results are then used as the state input to the DDPG network, which outputs the scheduling decision. This alternative scheme can anticipate the dynamic changes in the time window, making scheduling decisions more forward-looking. It is suitable for telemetry and control scenarios with frequent satellite orbit changes and highly dynamic visible time windows, further reducing the risk of mission timeouts and improving the completion rate of telemetry and control tasks.

[0106] Alternative Solution 3 The original solution employed a local replanning mechanism, which can be replaced with a hierarchical replanning mechanism, dividing replanning into two levels: global coarse planning and local fine planning. When a significant change in the scenario state is detected (such as abnormal resource status of multiple ground stations or large-scale task insertion), global coarse planning is executed to quickly adjust the responsibility domains of each mobile ground station and the overall deployment framework. When a minor change in the scenario state is detected (such as a single task window update), local fine planning is executed to adjust specific stopping points and scheduling schemes within the global framework. This alternative solution balances the efficiency and accuracy of replanning and is suitable for large-scale constellation telemetry and control projects with complex telemetry and control scenarios and diverse types of state changes.

[0107] The above alternative solutions do not deviate from the core concept of this invention. They only involve reasonable replacement of specific algorithms and mechanisms, and can all achieve the purpose of this invention, thereby improving the utilization rate of measurement and control resources and the task completion rate.

[0108] Please see Figure 2To verify the effectiveness and superiority of the method of this invention, a simulation experiment was designed for performance testing. The experiment selected 28 and 40 low-Earth orbit satellites as the tracking and control targets, setting 200 iterations within the planning cycle. The overall benefit of the tracking and control mission was used as the evaluation index. The method of this invention was compared with the traditional greedy strategy. The experimental results are as follows: Figure 2 (a) Comparison of revenue from 28 satellites Figure 2 (b) The comparison chart of returns for 40 satellites is shown. In the test scenario with 28 satellites, the overall return of the telemetry and control task using the greedy strategy gradually increases with the number of iterations, reaching approximately 780 after 200 iterations, with a slow convergence speed. After 100 iterations, the return growth tends to level off. The method of this invention converges rapidly to a stable value after 50 iterations, and the overall return reaches 820 after 200 iterations, an improvement of approximately 5.1% compared to the greedy strategy. In the test scenario with 40 satellites, the overall return of the telemetry and control task using the greedy strategy is approximately 800 after 200 iterations, with a slow rate of return growth throughout. The method of this invention achieves stable convergence after 60 iterations, and the overall return reaches 940 after 200 iterations, an improvement of approximately 17.5% compared to the greedy strategy. Furthermore, under both satellite scales, the return curves of the method of this invention show no significant fluctuations, while the return curve of the greedy strategy exhibits multiple small oscillations, indicating that the method of this invention has superior robustness. As the number of satellites increases from 28 to 40, the overall benefit of the method of this invention increases from 820 to 940, maintaining a stable growth trend, which proves that the method of this invention has good scalability in large-scale constellation telemetry and control scenarios.

[0109] In summary, this invention presents a method and system for collaborative telemetry, tracking, and command (TT&C) deployment and resource scheduling of mobile ground stations. Addressing core issues such as resource conflicts, low utilization rates, and poor dynamic response in large-scale constellation TT&C, it proposes an integrated deployment-scheduling collaborative TT&C scheme, achieving significant technical results. Firstly, it achieves coupled optimization of spatial deployment of mobile ground stations and time scheduling of TT&C tasks, overcoming the shortcomings of existing technologies that treat these two aspects separately. Through dynamic constraint Voronoi diagrams and multi-agent reinforcement learning, it achieves scientific collaborative site selection and deployment. Combined with an A3C network incorporating time coefficients, it achieves refined scheduling, ensuring that deployment locations provide optimal spatial support for scheduling. The scheduling results inversely optimize deployment strategies, significantly improving the overall utilization efficiency of TT&C resources. Secondly, it effectively solves the resource conflict problem caused by overlapping time windows. By accurately calculating visible time windows, dynamically adjusting task weights, and prioritizing urgent tasks, while combining action mask constraints on infeasible scheduling actions, it makes the allocation of TT&C tasks more reasonable. In simulation experiments with 28 and 40 satellites, it achieves higher overall benefits compared to traditional greedy strategies, and the benefits continue to increase steadily with the increase in satellite scale. Third, it significantly improves adaptability to dynamic scenarios. The designed local replanning mechanism only adjusts the affected area, avoiding the high overhead of global replanning. It can quickly respond to dynamic disturbances such as task insertion and resource status changes, solving the problem of untimely response in traditional replanning technologies. At the same time, the closed-loop linkage mechanism allows the solution to be continuously iterated and optimized, ensuring long-term operational measurement and control performance. Fourth, it balances engineering practicality and algorithmic efficiency. The multi-constraint screening of candidate stopping points and the centralized training and distributed execution mechanism of multi-agent algorithms make the technical solution conform to engineering constraints such as terrain, energy, and safety, while reducing computational and communication overhead, meeting the requirements of online real-time decision-making. It provides a feasible technical solution for the measurement and control support of large-scale low-Earth orbit constellations, and has good scalability and robustness.

[0110] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0111] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0112] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0113] In the embodiments provided by this invention, it should be understood that the disclosed devices / terminals and methods can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0114] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0115] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0116] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random-access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0117] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0118] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0119] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0120] The above content is only for illustrating the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solution based on the technical concept proposed in this invention shall fall within the scope of protection of the claims of this invention.

Claims

1. A method for collaborative telemetry, tracking, and command (TT&C) deployment and resource scheduling of mobile ground stations, characterized in that, Includes the following steps: S1. Based on the service demand distribution within the task area, the multi-agent reinforcement learning network generates the responsibility domain of each mobile ground station using a dynamic constraint Voronoi diagram. Combined with the multi-agent deterministic policy gradient algorithm, it outputs the stopping point sequence and transfer path of each mobile ground station, thus completing the collaborative site selection and deployment of mobile ground stations. S2. Using the location trajectory and stopping point sequence output by the collaborative site selection deployment as input sources, determine the set of visible time windows; The task is defined by a time coefficient that increases as the current time approaches the end of the window, and the task weight is adjusted according to the time coefficient. An asynchronous Actor-Critic network with an introduced time coefficient is used to jointly model the ground station time window, task acceptance status, task priority and shortest measurement and control duration constraint, and output the matching relationship between the measurement and control task and the ground station, the time window selection and execution order. S3. Generate a measurement and control action sequence and task allocation scheme according to the execution order, and build a deployment-scheduling closed-loop linkage mechanism: the scheduling result is used to apply to deployment rewards and strategy updates, and when a change in scene state is detected, a replanning mechanism is triggered to update the deployment and scheduling scheme.

2. The method for collaborative telemetry, tracking, and command deployment and resource scheduling of mobile ground stations according to claim 1, characterized in that, In step S1, the responsibility domain of each mobile ground station is generated using a dynamic constraint Voronoi diagram, specifically as follows: Using satellite-to-station visibility time window, mission priority, mission area environmental information, and mobile ground station status information as inputs, and under terrain, energy, and safety constraints, the responsibility domain of each mobile ground station is divided based on a comprehensive cost function. The comprehensive cost function is a weighted sum of the mobile ground station's maneuver time to reach the target location, maneuver energy consumption, service intensity at the location, environmental risk cost, and current load level. The service intensity is determined by the visibility time window and its weight, and the location... By the Comprehensive cost function for mobile ground station services for: in, Indicates the first The mobile ground station has reached its location. maneuver time, Indicates motor energy consumption. Indicates position Service intensity at the location Indicates the environmental risk cost, Indicates the current load level. and These are the weighting coefficients.

3. The method for collaborative telemetry, tracking, and command deployment and resource scheduling of mobile ground stations according to claim 1, characterized in that, In step S1, after generating the responsibility domain for each mobile ground station, the process also includes a candidate stopping point selection step: A set of candidate stopping points is generated within each responsibility domain. The candidate stopping points consist of road network nodes, accessible area centers, service hotspot centers, and locations that meet the staying safety constraints. Candidate points that do not meet the terrain accessibility constraints, road reachability constraints, staying safety constraints, or remaining energy constraints are eliminated to obtain a set of valid candidate stopping points within the responsibility domain. The multi-agent deterministic policy gradient algorithm outputs a sequence of stopping points based on the set of valid candidate stopping points.

4. The method for collaborative telemetry, tracking, and command deployment and resource scheduling of mobile ground stations according to claim 1, characterized in that, In step S1, the multi-agent deterministic policy gradient algorithm adopts a centralized training and distributed execution mechanism. An Actor network is constructed for each agent corresponding to a mobile ground station, and a centralized Critic network is constructed at the same time. During the training phase, the joint state and joint action are input into the Critic network. During the execution phase, each agent outputs an action independently through the Actor network based on its own local state.

5. The method for collaborative telemetry, tracking, and command deployment and resource scheduling of mobile ground stations according to claim 1, characterized in that, In step S2, the set of visible time windows is determined, specifically as follows: The coordinates and normal vectors of the ground station are transformed to the geocentric equatorial inertial coordinate system; the mapping relationship between satellite time and the angle of approach is derived based on the Kepler equations; the critical angle of approach corresponding to the connection and disconnection of the satellite-to-ground link is solved by combining the physical constraints of the minimum elevation angle of satellite-to-ground communication; the critical angle of approach is inversely solved to the corresponding time through the mapping relationship, the available visible time window of the mobile ground station and the target satellite is determined, and the visible time window set is integrated.

6. The method for collaborative telemetry, tracking, and command deployment and resource scheduling of mobile ground stations according to claim 1, characterized in that, In step S2, before joint modeling using the asynchronous Actor-Critic network with introduced time coefficients, the step further includes constructing the telemetry and control scheduling state. The telemetry and control scheduling state includes the available time window state of the ground station, the task execution state, the task time window information, the task priority vector, the task's shortest telemetry and control duration constraint, the task's remaining available time, and the task's remaining relaxation time. The task's remaining available time is the difference between the end time of the task window and the current time, and the task's remaining relaxation time is the difference between the task's remaining available time and the task's shortest telemetry and control duration.

7. The method for collaborative telemetry, tracking, and command deployment and resource scheduling of mobile ground stations according to claim 1, characterized in that, In step S2, the asynchronous Actor-Critic network is an asynchronous advantage Actor-Critic network, which constructs multiple parallel working threads. Each thread independently interacts with the scheduling environment replica and accumulates state, action, and reward samples. Each thread asynchronously calculates the policy gradient and value gradient and updates the global network parameters. The task weights are corrected according to the time coefficient. Specifically, the basic weights of the tasks are combined with the time coefficients to obtain the comprehensive weights of the tasks. The comprehensive weights are used as the weights for joint modeling of the asynchronous Actor-Critic network. For the The task at time Define time coefficient for: in, and Tasks The start and end times of the window. The time sensitivity coefficient, This is the adjustment constant.

8. The method for collaborative telemetry, tracking, and command deployment and resource scheduling of mobile ground stations according to claim 1, characterized in that, In step S3, the scheduling results are applied in reverse to the deployment rewards and policy updates, specifically as follows: The scheduling and execution results of the measurement and control tasks are fed back to the multi-agent reinforcement learning network. The scheduling and execution results are used as the basis for calculating the global reward function in the deployment phase. Based on the global reward function, the Actor network and Critic network of the multi-agent deterministic policy gradient algorithm are driven to perform parameter iterative updates.

9. The method for collaborative telemetry, tracking, and command deployment and resource scheduling of mobile ground stations according to claim 1, characterized in that, In step S3, the scene state change includes at least one of task insertion, time window update, road state change, environmental risk change, site resource state change, and ground station available time resource change; triggering the replanning mechanism to update the deployment and scheduling scheme, specifically: recalculating the comprehensive cost function for the area affected by the scene state change, updating the responsibility domain division result of the area, and only performing local deployment adjustment of mobile ground stations in the affected area; updating the telemetry and control scheduling status in real time, re-outputting the scheduling decision using an asynchronous Actor-Critic network with a time coefficient, and generating a new telemetry and control action sequence and task allocation scheme based on the new scheduling decision.

10. A mobile ground station collaborative telemetry, monitoring, and control deployment and resource scheduling system, characterized in that, include: The deployment module is used to generate the responsibility domain of each mobile ground station based on the service demand distribution within the task area using a dynamic constraint Voronoi diagram. It combines a multi-agent deterministic policy gradient algorithm to output the stopping point sequence and transfer path of each mobile ground station, completes the collaborative site selection and deployment of mobile ground stations, and outputs the location trajectory and stopping point sequence of the collaborative site selection deployment. The scheduling module, connected to the deployment module, is used to receive the location trajectory and stopping point sequence, determine the set of visible time windows, define a time coefficient for the task that increases as the current time approaches the end of the window, and adjust the task weight according to the time coefficient. It uses an asynchronous Actor-Critic network with introduced time coefficient to jointly model the ground station time window, task acceptance status, task priority, and shortest measurement and control duration constraint, and outputs the matching relationship between the measurement and control task and the ground station, the time window selection, and the execution order. The linkage module is connected to the deployment module and the scheduling module respectively. It is used to receive the execution order and generate the measurement and control action sequence and task allocation scheme, feed back the scheduling result to the deployment module, so that the scheduling result can be applied to the deployment reward and strategy update, detect changes in the scene state and trigger the replanning mechanism, drive the deployment module to update the deployment scheme and drive the scheduling module to update the scheduling scheme.