Message persistence state management system for intermittent satellite-ground links
Patent Information
- Application Number
- CN202610894221.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-06-22
AI Technical Summary
[0009]为解决上述间歇性星地链路环境下存在的数据冗余传输、星载资源受限、重复执行及传输状态不连续等技术问题,本发明提供了一种面向间歇性星地链路的消息持久化状态管理系统
[0020]本发明的有益效果是:与现有采用日志累积式同步、星载侧复杂排序协议及无持久化状态控制机制的星地通信系统相比,本发明通过构建面向间歇性链路的消息状态合并管理、地面顺序控制、星上轻量级去重以及全流程持久化恢复机制,在数据传输效率、星载资源占用、业务一致性及系统可靠性方面取得了显著提升,主要体现在以下方面:
Smart Images

Figure CN122419587B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of space-air information transmission and space-ground collaborative communication technology, and in particular to a message persistence state management system for intermittent space-ground links. Background Technology
[0002] With the rapid development of low-Earth orbit satellite constellations and integrated space-air-ground networks, satellites are gradually evolving from simple data acquisition terminals into distributed nodes with on-orbit computing and service processing capabilities. Continuous bidirectional interaction of mission control information, operational status data, telemetry logs, and service payload data is required between the ground system and the onboard system, thus forming a typical space-ground distributed collaborative processing architecture.
[0003] However, unlike traditional data centers or terrestrial internet environments, satellite-to-ground communication links are characterized by significant discontinuity and resource constraints: satellites only establish short-term connections with ground stations during transit; a single communication window typically lasts only a few minutes to a dozen minutes; uplink and downlink bandwidth are limited by antenna gain and power consumption, resulting in limited capacity; link round-trip latency is high, and bit error rate and packet loss rate fluctuate significantly; satellite-borne computing and storage resources are far lower than those of terrestrial servers.
[0004] The aforementioned characteristics make it difficult to directly apply traditional distributed consistency and reliable transmission mechanisms designed for stable links in a satellite-to-ground environment, thus giving rise to a series of new technical problems, which are as follows: 1. Offline backlog and outdated data conflict issues: Ground systems typically use log accumulation or full synchronization to store state change records. After a satellite has been offline for an extended period, all historical changes need to be sent sequentially upon entering the next transmission window. However, most control or configuration services only focus on the final state, and intermediate processes lack practical semantics. This results in a large amount of invalid historical data consuming valuable bandwidth resources, crowding out critical control and effective service transmission slots, and reducing link utilization efficiency. Existing technologies lack folding or overlay mechanisms oriented towards the final state.
[0005] 2. Limited onboard computing power leads to the unavailability of complex protocols: Traditional reliable transmission mechanisms rely on strategies such as sliding windows, out-of-order reordering, and cache sorting, requiring significant memory space and sustained computing power. Limited by the low-frequency processors, small memory capacity, and strict power consumption constraints of onboard platforms, such complex protocols struggle to operate stably for extended periods, easily leading to excessive resource consumption or degraded real-time performance. Existing solutions struggle to provide lightweight ordering and state control methods adapted to the low-computing-power environment of onboard systems.
[0006] 3. Non-idempotent replay risk under high packet loss environment: Due to the high link acknowledgment loss rate, the ground system frequently triggers retransmissions. If the application layer lacks a unified deduplication or idempotent control mechanism, duplicate packets may lead to repeated task execution, repeated configuration overwriting, or multiple triggering of control commands, thereby causing on-board service status disorder and unpredictable system behavior. Existing technologies cannot guarantee the uniqueness and consistency of service execution at the logical layer.
[0007] 4. Loss of transmission state due to connection interruption: In scenarios with intermittent links and process restarts, if the message sending state is only stored in memory, sent but unacknowledged data will be unrecognizable after a connection interruption, often requiring a full retransmission, resulting in additional bandwidth consumption and uncontrollable task latency. Existing systems lack a unified persistent state management and compensation recovery mechanism, making it difficult to guarantee the continuity of cross-window transmission.
[0008] Given that existing satellite-to-ground communication systems still employ data synchronization and transmission management mechanisms designed for continuously online networks, they fail to adapt to the operational characteristics of satellite-to-ground links, such as prolonged offline periods, short-window access, narrow bandwidth, high packet loss, and limited onboard resources. This results in low link resource utilization, heavy onboard processing burden, and difficulty in ensuring service consistency. Therefore, corresponding data management mechanisms oriented towards the final state are needed, along with lightweight sequence control and reliable transmission methods with lower computational and storage overhead, message management methods with deduplication and acknowledgment capabilities, and persistent transmission state management and compensation recovery mechanisms. These measures aim to reduce the transmission of redundant historical data and increase the proportion of effective data, ensure the uniqueness and consistency of service execution, and guarantee the continuity and recoverability of cross-window transmission. Summary of the Invention
[0009] To address the technical problems of redundant data transmission, limited onboard resources, repetitive execution, and discontinuous transmission status in intermittent satellite-to-ground link environments, this invention provides a message persistence state management system for intermittent satellite-to-ground links.
[0010] To achieve the above objectives, the present invention provides a message persistence state management system for intermittent satellite-to-ground links, the system comprising a ground gateway side and a satellite gateway side; The ground gateway includes a state merging and overlay module, a persistence manager, a pre-ordering locking engine, and a ground distributor. The state merging and overlay module is used to merge and overlay multiple state updates of the same service object based on a composite identifier during satellite offline periods, folding multiple intermediate states into a single final state. The persistence manager is used to implement external persistent storage for the entire message lifecycle, define and manage various atomic states, and write state changes to a local database in real time. The pre-ordering locking engine is used to assign sequence keys to messages with sequential dependencies and establish a mutex lock mechanism. The ground distributor is used to integrate state, sequence, and link quality information, driving the entire process of message readiness to confirmation. The satellite gateway side includes an ID mapping deduplication module, a stateless executor, and a state recovery manager. The ID mapping deduplication module maintains a set of executed message IDs on the satellite gateway side, identifies duplicate messages through hash matching, and returns only an acknowledgment receipt when a match is found; otherwise, it executes business logic and writes it into the set. The stateless executor directly calls the business processor to perform operations in the order of receipt. The state recovery manager enables the new master node to read the locally persisted set of executed message IDs to rebuild the idempotent index and restore the state of the business processor when a master node switch occurs on the satellite gateway.
[0011] Furthermore, when performing the merge and overwrite operation, the state merging and overwrite module searches the database based on a composite unique identifier composed of message type and resource key; if no corresponding record exists, a new READY record is added; if there is a READY or PENDING record that has not been sent, its payload and timestamp information are directly overwritten, and no new message copy is generated.
[0012] Furthermore, the atomic states managed by the persistence manager include: The READY state indicates that the message has been persisted to the database and is in the initial waiting-for-scheduling phase; The PENDING status indicates that the message has entered the sending queue and is waiting to depart on the ground. The IN_FLIGHT status during transmission indicates that the message has been sent via the optimized QUIC protocol and is being transmitted over the satellite-to-ground link or is awaiting satellite processing. The ACKED status indicates that an acknowledgment signal has been received from the satellite. The FAILED status indicates a transmission failure caused by exceeding the retry limit or a link malfunction.
[0013] Furthermore, the pre-ordering locking engine specifically executes the following process: the ground distributor retrieves the ready state task with the ordering_key of K1, and immediately locks K1 after obtaining the first message; the message is sent through the satellite-ground link, and the message status changes to the in-transmission state; during the locking period, subsequent messages of K1 are prohibited from leaving the database; the lock is released and the next message is allowed to enter the sending process only after receiving an acknowledgment from the satellite gateway.
[0014] Furthermore, the ground distributor adopts a coroutine model to achieve dynamic lifecycle management. When a satellite link is detected to be in an available window, a coroutine instance is created, and the coroutine is destroyed when the link leaves the window or an anomaly occurs.
[0015] Furthermore, the system adopts the QUIC protocol optimized for satellite-to-ground links as the transport layer; the optimized QUIC protocol extends the connection liveness timer to support long-term session suspension across window interruption periods and 0-RTT fast reconnection, taking into account the characteristics of long latency, high packet loss and intermittent access in satellite-to-ground communication; a customized congestion control algorithm is introduced to differentiate between channel noise and network congestion in packet loss; and multiplexing stream strategy orchestration is performed in conjunction with the application layer sequence key.
[0016] Furthermore, the ID mapping deduplication module uses a circular buffer or hash table to implement a fixed-capacity set of executed message IDs. When a hash match is found, it is determined to be a duplicate message, and the satellite gateway immediately generates an acknowledgment and returns it to the ground, while discarding the received payload. When a hash match is not found, it is determined that the message is arriving for the first time, and the satellite gateway distributes the payload to the corresponding service processor to perform control operations. After the operations are completed, the current message ID is written into the set.
[0017] Furthermore, the recovery process executed by the state recovery manager includes: when a satellite gateway node is promoted from a slave node to a master node or demoted from a master node to a slave node, the newly elected master node starts initialization, reads the set of executed message IDs and the status of incomplete transactions in the local persistent storage, rebuilds the memory index of the idempotent deduplication set and the state machine of the service processor, and receives data frames sent from the ground normally after the recovery is completed.
[0018] Furthermore, the system implements a state persistence and compensation recovery mechanism through the collaborative persistence manager on the ground gateway side and the state recovery manager on the satellite gateway side. This includes: when a link is abnormally interrupted, the distribution process is restarted, or the on-board master node is switched, the ground gateway scans the incomplete state according to the persistent records, rolls back the timed-out unacknowledged records to the ready state, and re-enters the transmission process or redirects to the new master node; the message ID on the satellite gateway side is immediately written to the locally persisted set of executed message IDs after the service is executed.
[0019] Furthermore, when the ground distributor is waiting for confirmation receipt, it makes a timeout judgment based on a dynamic threshold. The dynamic threshold is generated by a weighted adaptive algorithm based on real-time link round-trip time sampling and channel quality evaluation factors. The channel quality evaluation factors are jointly determined by time delay jitter, satellite overpass elevation angle and signal-to-noise ratio.
[0020] The beneficial effects of this invention are as follows: Compared with existing satellite-to-ground communication systems that employ log-accumulated synchronization, complex onboard sequencing protocols, and lack persistent state control mechanisms, this invention significantly improves data transmission efficiency, onboard resource consumption, service consistency, and system reliability by constructing message state merging management for intermittent links, ground sequence control, lightweight onboard deduplication, and a full-process persistent recovery mechanism. These improvements are mainly reflected in the following aspects: 1. Significantly reduces the amount of historical redundant data during offline periods and increases the proportion of effective data during the window period; by merging and covering multiple state changes of similar resources on the ground side, a large number of intermediate states generated during offline periods are folded into the final effective state, avoiding the bandwidth waste caused by synchronizing all historical logs one by one in traditional solutions; therefore, within a limited transit window, link resources can be prioritized for transmitting critical control information and real effective business data, significantly improving the effective payload ratio and window utilization, and enhancing overall transmission efficiency; 2. Significantly reduces the computational and storage resource consumption of the satellite-borne system; by moving the sequence control and dependency management logic to the ground side, the satellite-borne system no longer needs to maintain sliding windows, out-of-order buffers, or complex reordering algorithms, and only needs to process messages directly according to the order of reception; compared with the traditional scheme that relies on a complete protocol stack to achieve reliable transmission, this invention significantly reduces on-board memory usage and CPU processing load, reduces software complexity and power consumption pressure, and improves the operational stability and real-time performance of the satellite-borne system under resource-constrained conditions; 3. Enhance the uniqueness and consistency of business execution in high packet loss environments; by establishing a lightweight message identification and deduplication mechanism on the satellite, duplicate data is quickly identified and filtered, avoiding the problem of the same business being executed multiple times in retransmission scenarios; this mechanism effectively ensures the idempotency and uniqueness of business behaviors such as control commands, task scheduling and configuration operations, prevents repeated operations from causing system state disorder, and improves the determinism and reliability of satellite business logic. 4. Ensure the recovery of transmission status and the continuity of tasks in the event of disconnection; By persistently storing the message sending progress and confirmation status, and automatically rebuilding the transmission context after the link is restored or the node is restarted, incomplete data can continue to be processed from the interruption point, avoiding duplicate sending or task rollback; This mechanism significantly enhances the system's fault tolerance and continuous operation capability under intermittent link conditions, and improves the success rate of cross-window transmission and the overall reliability of task completion. 5. Reduce the overhead of invalid retransmissions and duplicate processing, and improve bandwidth utilization efficiency; through the synergistic effect of state folding, deduplication confirmation and fine-grained state control, the invalid data traffic caused by repeated sending, repeated confirmation or repeated execution is reduced, so that limited bandwidth resources are used more for actual business transmission, thereby improving the overall link utilization and system operating efficiency. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort: Figure 1 This is an overall framework diagram of the system of the present invention; Figure 2 This is a flowchart of the offline state merging and overwriting logic in this invention; Figure 3 This is a timing diagram of the pre-sorting mechanism of "ground-locked - satellite-stateless" in this invention; Figure 4 This is the execution flowchart of the ID mapping deduplication module in this invention; Figure 5 This is a flowchart of the state persistence and compensation recovery mechanism for distributed scenarios in this invention. Detailed Implementation
[0022] The present invention will now be described in detail with reference to the accompanying drawings. Unless otherwise specified, the features of the following embodiments and implementations can be combined with each other.
[0023] This invention is mainly applied to satellite-to-ground communication scenarios characterized by long latency, frequent interruptions, and short access windows between satellites and ground stations. By persistently recording and managing the transmission progress, confirmation status, and execution results during message transmission, it achieves continuous transmission status assurance and reliable delivery control in intermittent link environments.
[0024] See Figure 1 The present invention provides a message persistence state management system for intermittent satellite-to-ground links, the overall architecture of which is composed of a ground gateway side and a satellite gateway side.
[0025] The ground gateway side includes a state merging overlay module, a persistence manager, a pre-ordering locking engine, and a ground distributor, among which: The state merging and overlay module is responsible for handling state update storms during satellite offline periods. When the same service object generates multiple state updates, the system, based on a composite identifier retrieval database (a data storage and retrieval model designed in this invention, specifically a mechanism that uses multiple fields as a unique index to locate a record), directly overwrites the existing records to be transmitted with the payload instead of adding new ones. This folds multiple intermediate states into a single final state, significantly reducing redundant data during the window period and improving the effective transmission ratio.
[0026] The persistence manager is used to implement external persistent storage for the entire message lifecycle. It defines five atomic states: READY (ready state; the message has been persisted into the database and is initially waiting for scheduling), PENDING (pending state; the message has entered the sending queue and is waiting to depart on the ground), IN_FLIGHT (in-transit state; the message has been sent via the optimized QUIC protocol and is being transmitted through the satellite-to-ground link or waiting for satellite processing), ACKED (acknowledged state; the message has received an acknowledgment signal from the satellite), and FAILED (failed state; the message has exceeded the retry count or is experiencing a link failure). All changes are written to the local database in real time. This design makes the ground distributor process completely stateless, and any crash and restart can accurately restore the transmission progress. It is the core infrastructure for achieving fault self-healing and state recoverability.
[0027] The pre-ordering locking engine completely transfers the responsibility of sequence control to the ground side. It assigns an `ordering_key` to messages with sequence dependencies and establishes a mutex lock mechanism. The lock is acquired immediately after the first message is sent and unlocked upon the arrival of an ACK (acknowledgment receipt), ensuring that the same service sequence is naturally ordered at the link layer. The spaceborne end does not need to maintain a sliding window or reordering buffer, significantly reducing computational and memory overhead.
[0028] The ground distributor, acting as the central scheduling unit of the ground gateway, integrates multi-dimensional information such as status, sequence, and link quality to drive the entire process of messages from readiness to confirmation. It employs a coroutine model to implement dynamic lifecycle management: when a satellite link enters an available window, a coroutine instance is created; when the link leaves the window or an anomaly occurs, the coroutine is destroyed, avoiding long-term resource occupation and improving system scalability.
[0029] The transport layer employs the QUIC protocol optimized for satellite-to-ground links. While maintaining compatibility with the standard QUIC protocol framework, it has made adaptive improvements in the following dimensions to address the characteristics of long latency, high packet loss, and intermittent access in satellite-to-ground communication: The connection liveness timer has been extended to support long-term session suspension across window interruptions and 0-RTT fast reconnection; a customized congestion control algorithm has been introduced, incorporating a packet loss differentiation mechanism (distinguishing between channel noise and network congestion) to adapt to drastic bandwidth fluctuations; and multiplexing stream strategy orchestration has been combined with the application layer's ordering_key to offload the pressure of out-of-order reassembly at the satellite end.
[0030] The satellite gateway side includes an ID mapping deduplication module, a stateless actuator, and a state recovery manager, among which: The ID mapping deduplication module is the core of the on-board idempotent control, maintaining a fixed-capacity set of executed MessageIDs. Upon receiving a message, it performs a hash match; if a match is found, it's considered a duplicate, only an ACK is returned, and the payload is discarded; otherwise, execution is allowed. This design intercepts duplicates at the logic layer, avoiding business side effects, while simultaneously providing a rapid response to drive the ground state machine's progress.
[0031] Stateless executors are used to directly invoke business processors in the order they are received, performing operations such as container management and configuration updates. They do not maintain long-term contexts, do not cache historical messages, and release resources immediately upon completion. This stateless design enables a single node to support tens of thousands of TPS (Transactions Per Second), and recovery during Leader switching only requires reading local persistence, achieving extremely simple resource consumption. TPS represents the number of transactions the system can process per second, a core metric for measuring high-concurrency performance; Leader refers to the primary node.
[0032] The State Recovery Manager supports rapid recovery of the satellite gateway in failure scenarios such as Leader switching. After the new Leader starts, it reads the local set of executed MessageIDs, rebuilds the idempotent index, restores the business processor state, and seamlessly takes over subsequent message processing. Through local persistence and rapid reconstruction mechanisms, failover is transparent to the ground, ensuring business continuity.
[0033] This invention establishes a centralized state folding and sequential control mechanism on the ground side, constructs a lightweight idempotent execution and state confirmation mechanism on the spaceborne side, and achieves full-process recoverable state management through a persistent database, thereby forming a hierarchical state control system of "strong ground management and lightweight onboard execution." The solution includes: 1. The state merging and overriding logic for offline periods is executed through the state merging and overriding module.
[0034] A unified message state database is built on the ground gateway side. A persistence manager manages all messages to be sent persistently. A state merging and overriding module creates a composite unique identifier for each message, consisting of (message_type, resource_key), to uniquely represent similar operations on the same business object. Here, message_type is the message type, and resource_key is the resource key. When the ground system generates a new state update request while the satellite is offline, the system first searches the database based on the composite unique identifier: if no corresponding record exists, a new READY state record is added; if an incomplete READY or PENDING record already exists, its payload and timestamp information are directly overwritten without generating a new message copy.
[0035] By employing the above method, multiple intermediate state updates are folded into a single final state, achieving merged overlay management of states. For most control or configuration-related services, only the final state has semantic meaning; intermediate changes do not affect the final result, thus redundancy can be eliminated through overlay. This mechanism significantly reduces the amount of historical redundant data during offline periods, increases the proportion of effective data within the window period, and reduces bandwidth consumption.
[0036] See Figure 2 The offline state merging and overwriting logic in this invention includes: New message request received: The system receives a status update request from the application layer (business system). This request includes the message type, target resource identifier, and business payload. This step serves as the trigger point for the entire status management process, signifying that a new business status change has entered the processing pipeline. The system must immediately initiate persistence and merging decision logic to ensure reliable status persistence.
[0037] Extract the composite identifier: Extract a composite unique identifier composed of message_type and resource_key from the request. The design of this identifier follows the principle of business semantic uniqueness. message_type defines the operation category (such as configuration distribution, state synchronization), and resource_key locks the operation object (such as container instance, configuration file). The combination of the two uniquely represents "a certain type of operation on a certain object", providing a basis for judgment for merging subsequent operations of the same type.
[0038] Database Existence Retrieval: Based on the composite identifier, an index retrieval is performed in the persistent database to determine the existence of historical records. This retrieval covers all states (READY, PENDING, IN_FLIGHT, ACKED, FAILED), but subsequent merging branches are only operable for READY and PENDING state messages. This is because: IN_FLIGHT records have already been processed in the pipeline or on-board; ACKED records have been confirmed and do not need to be merged; and FAILED records have clearly failed and require a separate retry path. Retrieval performance relies on the database index of the composite identifier to ensure millisecond-level response times in high-frequency update scenarios.
[0039] If no branch exists, a new READY record is created. If a search fails, it's determined to be the first time this type of message has been accessed on this object. The system adds a new record with a READY status, fully writing metadata such as Payload, timestamp, and MessageID. This branch handles scenarios of "first report" or "previous cycle confirmed," ensuring the traceability of status history and laying the foundation for subsequent window period sending.
[0040] If a branch exists, a status check is performed. Upon a successful retrieval, the original record's status is further checked to determine if it is READY or PENDING. This check is a critical decision point, distinguishing between two scenarios: "pending transmission without confirmation" and "transmitted and being processed." READY / PENDING indicates the record is still waiting on the ground and has not yet received confirmation from satellite, thus meeting the conditions for merging. IN_FLIGHT / ACKED / FAILED indicates the record has left the database or is in its final state, requiring independent processing to avoid the uncertainty caused by overwriting the IN_FLIGHT state. When the original record is READY / PENDING, a status merging and overwriting process is performed. The system retains the original MessageID to ensure sorting continuity, directly overwriting the old content with the new payload, updating the timestamp to the latest version, and maintaining the original state (READY / PENDING). This mechanism folds multiple intermediate updates during offline periods into a single final state, eliminating redundant transmission based on the business assumption that "only the final state has semantic meaning," significantly improving the proportion of effective data during the window period. When the original record is not in a READY / PENDING state, it is considered that the previous cycle has ended and the current request belongs to a new update cycle; the system adds an independent READY record to avoid confusion with the historical final state; this design ensures the archiving integrity of ACKED / FAILED records, while supporting the continuous state evolution of the same object.
[0041] Waiting for transit window distribution: Regardless of whether it's a new record or an overlay, all records eventually enter the READY / PENDING state, waiting for the satellite transit window to arrive before triggering distribution. During this stage, the system continues to listen for new requests with the same composite identifier and continuously performs merging and overlay until the link becomes available, forming an optimized mode of "multiple rounds of merging during offline periods and single transmission during the transit window".
[0042] Second, the pre-sorting mechanism of "ground locking - satellite statelessness" is executed through the pre-sorting locking engine.
[0043] To avoid introducing complex out-of-order caching and reassembly logic at the satellite end, this invention centralizes the order control responsibility to the ground side. Specifically, a pre-ordering locking engine assigns an ordering_key to each type of message with sequential dependencies in the ground gateway, and establishes a mutex lock control mechanism based on this key: when the first message corresponding to a certain ordering_key is sent, the key is immediately locked; during the locking period, subsequent messages with the same key are prohibited from leaving the database; only when an ACK is received from the satellite is the lock released and the next message allowed to enter the sending process.
[0044] Through the aforementioned serialization control, the same service sequence is naturally transmitted in an ordered manner at the link layer. The sorting complexity is transferred from the satellite-based end to the ground-based end. Serialization on the transmitting side ensures natural ordering on the receiving side, thus avoiding the buffering and computational overhead required for out-of-order reordering. Therefore, the satellite-based end does not need to maintain a sliding window or reordering buffer; it only needs to execute directly according to the receiving order, significantly reducing onboard computation and memory usage.
[0045] See Figure 3 The pre-sorting mechanism of "ground-locked - satellite-stateless" includes: 1. The ground distributor initiates a query to the task database to retrieve the READY task with ordering_key K1. This query sorts by priority and timestamp, ensuring that the earliest ready message under the same key is processed first. The task database returns the first record Msg-Seq-1 that meets the conditions, containing complete metadata such as MessageID, Payload, and sorting number.
[0046] 2. The task database returns Msg-Seq-1 to the distributor. At this point, the message is still in the READY state and has not yet entered the sending process. The task database and the ground distributor's memory view are consistent, laying the foundation for subsequent state transitions and locking control.
[0047] 3. The ground distributor immediately locks ordering_key K1, marking the key as occupied. Lock information includes the acquisition timestamp and message identifier, written to the memory lock table and synchronized to the database transaction, ensuring lock visibility and mutual exclusion in multi-instance scenarios. The locking operation is executed atomically to prevent concurrent sending caused by multi-threaded contention.
[0048] 4. The ground distributor sends Msg-Seq-1 to the satellite gateway via the satellite-to-ground link. The message status changes from READY to IN_FLIGHT, and the ground distributor starts an ACK waiting timer. At this time, K1 is under lock protection, and subsequent messages with the same key are blocked, ensuring strict seriality of the link layer transmission.
[0049] 5. During the lockout period, the ground distributor implements a blocking strategy for subsequent READY messages of K1, prohibiting them from leaving the database and entering the transmission process. This design moves the ordering control forward to the ground side, ensuring natural ordering on the receiving side through serialization on the transmitting side, and avoiding the introduction of out-of-order buffering and reordering logic at the spaceborne end.
[0050] 6. The satellite gateway executes Msg-Seq-1 directly in the order of reception, eliminating the need to maintain a sliding window, reordering buffer, or out-of-order detection mechanism. This "stateless reception" design significantly reduces onboard computation and memory usage, and generates an acknowledgment receipt immediately upon completion.
[0051] 7. The satellite gateway returns MSG_ACK (Seq-1) to the ground distributor, carrying the acknowledgment MessageID and processing result. The ACK is transmitted via the satellite-to-ground link. Upon receiving it, the ground distributor determines that the transmission was successful and triggers the unlocking process. If the ACK is lost or times out, the ground distributor retains the lock and triggers a retry to ensure that the strict order of messages with the same key is not disrupted.
[0052] 8. After receiving the ACK, the ground distributor performs an unlock operation (Unlock(K1)) on K1, deleting the occupied flag from the memory lock table and synchronously updating the database to a released state. Unlocking is a key signal for pipeline advancement, indicating that the next message with the same key can enter the sending process.
[0053] 9. After unlocking, the ground distributor queries the database again to retrieve the next READY task (Seq-2) for K1 and returns Msg-Seq-2. This query is protected by a locking mechanism to ensure that no other instances concurrently acquire the task, maintaining global sequential consistency.
[0054] 10. The ground distributor sends Msg-Seq-2 to the satellite gateway, repeating the "lock-send-wait-unlock" cycle. This pipeline mechanism allows message sequences with the same ordering_key to be transmitted continuously within the window period, while different keys can be transmitted in parallel without interference, achieving a balance between throughput and order.
[0055] Third, to address the issue of duplicate retransmissions caused by high packet loss rates in satellite-to-ground links, an ID mapping deduplication module is introduced.
[0056] Specifically, the satellite maintains a fixed-capacity set of executed MessageIDs, which can be implemented using a circular buffer or a hash table structure to record the identifiers of the most recently successfully processed messages. When a new message is received: the ID hash is matched using the ID mapping deduplication module; if a match is found, it is determined to be a duplicate message, only an acknowledgment is returned, and the payload is discarded; if no match is found, business logic is executed and the message is written to the set.
[0057] Idempotent execution is achieved through deduplication at the logical layer, avoiding reliance on the reliability assumptions of the underlying ACK. This mechanism effectively prevents duplicate task triggering or abnormal state changes, ensuring the uniqueness and consistency of on-board service execution results.
[0058] See Figure 4 The on-board lightweight deduplication process based on ID mapping includes: 1. The satellite gateway receives data frames from the ground gateway via an intermittent satellite-to-ground link. This process involves RF signal demodulation, forward error correction decoding, and frame boundary synchronization, ultimately extracting the complete protocol data unit. Due to the high packet loss characteristics of the satellite-to-ground link, the frames received in this step may be the first transmission from the ground or retransmission frames triggered by ACK loss.
[0059] 2. The satellite gateway performs protocol parsing on the received frames and extracts the globally unique MessageID field from the frame header. This identifier is generated by the ground gateway and uses a UUID (Universally Unique Identifier) or snowflake algorithm to ensure system-wide uniqueness, serving as the core basis for subsequent idempotency determination. The extracted MessageID will then enter the hash calculation process to prepare for set queries.
[0060] 3. The satellite gateway performs a hash operation on the extracted MessageID and performs a matching query in a fixed-capacity executed set. This set is implemented using a hash table or a circular buffer, supporting fast lookup with O(1) time complexity. If the slot corresponding to the calculated hash index already contains the same MessageID, it is considered a hit; otherwise, it is considered a miss. This step is the key decision point of the entire idempotent control, directly determining the subsequent processing path of the message.
[0061] When a hash match is found, the message is determined to be a duplicate retransmission. The satellite gateway immediately generates an acknowledgment (ACK) and sends it back to the ground, while discarding the received payload without any service processing. This design is based on the assumption of "execution equals safety," avoiding service side effects caused by repeated execution. Simultaneously, the fast ACK response drives the ground state machine to unlock subsequent messages, maintaining the continuity of the transmission pipeline.
[0062] When a hash match fails, the message is considered to be arriving for the first time. The satellite gateway distributes the payload to the corresponding service processor to execute actual control operations, such as container start / stop, configuration updates, or telemetry data acquisition. Service execution may involve onboard operating system calls, file system writes, or network configuration changes, representing a major stage of onboard resource consumption. After execution, the process enters the ID registration phase to establish idempotent protection.
[0063] After successful execution of the business logic, the current MessageID is written to the executed set, marking the message as processed. The write operation is strongly bound to the business execution, ensuring atomicity (attackality means that all operations in a transaction either all succeed or none are executed). For hash table implementations, open addressing or chained collision resolution is used; for circular buffers, the oldest record is overwritten using a FIFO (First-In, First-Out) strategy. After writing, the MessageID will be recognized as processed for a period of time, preventing potential retransmissions and duplications.
[0064] After the service execution is completed and the MessageID is successfully registered, a FRAGMENT_ACK (segmented acknowledgment frame) is constructed. The ACK content includes the acknowledged MessageID, processing result status code, processing timestamp, and other metadata. This acknowledgment frame is transmitted back to the ground via the satellite-to-ground link, serving as the sole credential for the ground gateway to update its persistent state, release locked resources, and advance the transmission window, thus completing the entire transmission loop.
[0065] 4. The ACK and FRAGMENT_ACK segmented acknowledgment frames are returned to the ground gateway via the satellite-to-ground radio link. This process also faces the risks of link intermittency and high packet loss, but the consequences of ACK loss are mitigated by the ground timeout retransmission mechanism. The satellite gateway does not maintain the ACK transmission status, maintaining the simplicity of a stateless design. After receiving the ACK, the ground changes the corresponding message status from IN_FLIGHT to ACKED, triggering the scheduling and transmission of subsequent messages.
[0066] IV. State persistence and compensation recovery mechanism for distributed scenarios.
[0067] To ensure that transmission status is not lost in the event of link interruption or node switching, this invention establishes a unified persistent state model on both the ground and satellite sides. All messages have clearly defined status identifiers throughout their lifecycle, including READY, PENDING, IN_FLIGHT, ACKED, and FAILED, and are written to the local database in real time. When abnormal situations occur, such as abnormal link interruption, distribution process restart, or satellite leader switching, the ground gateway automatically scans for incomplete states based on persistent records: records that have timed out and not been acknowledged are rolled back to READY; the transmission process is re-entered; or the transmission is redirected to the new leader.
[0068] By externalizing and persistently storing the transmission state, the communication process becomes recoverable and traceable, ensuring eventual consistency. This mechanism avoids duplicate full transmissions or task interruptions due to state loss, improving business continuity across windows and in distributed environments.
[0069] See Figure 5 The state persistence and compensation recovery mechanism for distributed scenarios includes the ground gateway stage and the satellite gateway stage.
[0070] Ground gateway phase: 1. The application layer submits a status update request to the ground gateway. The system assigns a globally unique MessageID and a composite identifier composed of message_type and resource_key to the message to uniquely identify the service operation. This step marks the message officially entering its transmission lifecycle, and the system begins to track and manage it.
[0071] 2. The ground gateway initializes the message state to READY and writes it to the local persistent database (SQLite or RocksDB) in real time. The WAL (Write-Ahead Log) mode ensures the atomicity and crash safety of transactions. Even if the system fails at this moment, the message data will not be lost, fulfilling the reliability promise of "commit and it's done," laying the foundation for subsequent anomaly recovery.
[0072] 3. When the ground gateway detects that the satellite link has entered the available window, the ground distributor changes the message status from READY atomically through PENDING to IN_FLIGHT, and sends a data frame to the satellite gateway through the satellite-ground link. The IN_FLIGHT status clearly indicates that the message has left the database but has not yet received an acknowledgment. It is the key status to distinguish between "pending to be sent" and "sent but not acknowledged", and it is also the core basis for timeout judgment and compensation recovery.
[0073] 4. Determine if an ACK has been received: The ground distributor starts an ACK waiting timer, dynamically calculates the timeout threshold (2 times RTT for high-quality links, 3 times RTT for fluctuating links, and 5 times RTT for low-quality links), and waits for the satellite gateway to return an acknowledgment. This ACK determination node distinguishes between normal completion paths and abnormal recovery paths, determining the final course of the message lifecycle.
[0074] Upon receiving a valid ACK confirmation, the ground distributor updates the message status from IN_FLIGHT to the ACKED final state, releases the associated locked resources and memory usage, and archives the record to a history table for auditing purposes. This step marks the complete closure of the message transmission lifecycle, and the message is removed from the active processing set.
[0075] When no valid ACK confirmation is received, an anomaly detection is initiated. The ground distributor continuously monitors the operating environment, including link signal strength (detecting interruption), process heartbeat (detecting crash restart), and cluster leader tenure (detecting node switchover). Upon triggering any anomaly signal, a compensation and recovery process is initiated, suspending normal transmission and prioritizing state consistency.
[0076] 5. Upon triggering an anomaly, the ground distributor immediately scans the persistent database, retrieving incomplete records with a status of IN_FLIGHT or PENDING, and extracting key metadata such as MessageID, sending timestamp, retry count, and target node. This scan provides a complete data foundation for subsequent timeout assessment and recovery decisions.
[0077] For each incomplete record, calculate its actual waiting time (i.e., the difference between the current system time and the sending timestamp) and compare it with a dynamic threshold; The dynamic threshold is generated by a weighted adaptive algorithm based on real-time link round-trip time (RTT) sampling and channel quality evaluation factors: the system continuously monitors the smooth round-trip time (SRTT) fed back from the transport layer to characterize the basic physical delay characteristics, and introduces a channel correction factor jointly determined by delay jitter (RTTVAR), satellite overpass elevation angle, and signal-to-noise ratio; the threshold expands or contracts in real time with the dynamic changes in the satellite-to-ground geometric distance during satellite overpass, thereby constructing a precise packet loss judgment boundary at the logic layer to distinguish between "long link delay fluctuations" and "real packet loss".
[0078] Based on the comparison results, the timeout judgment is divided into two scenarios: "can wait for recovery" if it has not timed out, and "needs immediate compensation" if it has timed out, thereby driving differentiated recovery strategies.
[0079] For unacknowledged records that time out, the ground distributor atomically rolls back its state from IN_FLIGHT / PENDING to READY, releasing the mutex lock on the composite identifier, but retaining the latest payload and timestamp. This rollback operation ensures that message semantics are not lost, supporting precise continuation from the acknowledgment position during subsequent link windows, rather than a full retransmission.
[0080] For records that have not timed out, their IN_FLIGHT state and locked resources remain unchanged, and transmission is paused without any state change. This design is suitable for scenarios of momentary link interruption or brief process restart. After recovery, the system can directly continue waiting for ACK, avoiding unnecessary bandwidth waste caused by rollback retries.
[0081] 6. In a single-machine deployment scenario, the rolled-back READY record is re-injected into the sending queue, awaiting the next link entry window. In a cluster deployment scenario, incomplete records are redirected to the newly elected Leader, which seamlessly takes over after loading the shared database state. Both paths ultimately revert to the standard sending process, achieving self-healing from faults.
[0082] Satellite gateway phase: 7. The satellite gateway receives data frames transmitted via the satellite-to-ground link and completes signal demodulation, forward error correction decoding, frame boundary synchronization, and integrity verification. This step is the physical layer endpoint of satellite-to-ground transmission, and the extracted valid frames enter the subsequent processing pipeline.
[0083] 8. The satellite gateway parses data frames and directly calls the service processor to perform actual control operations (such as container start / stop, configuration update, and telemetry acquisition) in the order of reception. This design follows the principle of "stateless reception," eliminating the need to maintain sliding windows, reordering buffers, or out-of-order detection mechanisms, thus minimizing onboard resource consumption.
[0084] 9. Upon completion of the business execution, the MessageID is immediately written to a locally persistent set of executed MessageIDs (implemented by a hash table or circular buffer). This record is strongly bound to the business execution, ensuring that "a record is generated upon completion of execution," providing data support for idempotent deduplication determination and Leader switchover recovery.
[0085] 10. Construct a FRAGMENT_ACK acknowledgment frame, including the acknowledgment MessageID, processing result status code, and timestamp, and transmit it back to the ground via the satellite-to-ground link. Processing resources are released immediately after ACK generation. The satellite does not maintain the ACK transmission status; if lost, it relies on the ground timeout retransmission mechanism as a fallback, maintaining the simplicity of the design.
[0086] 11. The satellite gateway continuously monitors the Leader role status of this node in the cluster. When a heartbeat timeout triggers an election, and this node is promoted from a Follower to Leader, or demoted from Leader to Follower, it enters a state recovery process to ensure message processing continuity during role changes.
[0087] New Leader Read Status: When a newly elected Leader starts up and initializes, it first reads the set of executed MessageIDs and the status of incomplete transactions from the local persistent storage. This read operation reconstructs the global view on the satellite, ensuring that the new Leader can correctly identify processed messages (preventing duplicates) and take over subsequent processing (without missing new messages).
[0088] Context recovery: Based on the read persistent data, the new Leader reconstructs the memory index of the idempotent deduplicated set, the state machine of the business processor, and the wait queue for unacknowledged transactions. After recovery, the satellite gateway has full reception, processing, and acknowledgment response capabilities, presenting a service-available status to the outside world.
[0089] 12. After recovery, the satellite gateway normally receives data frames transmitted from the ground, and performs the standard procedures of deduplication, service processing, status recording, and ACK return. This step marks the final completion of the Leader handover recovery, and satellite-to-ground transmission enters a new steady-state cycle.
[0090] The above embodiments are only used to illustrate the design concept and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made based on the principles and design ideas disclosed in the present invention are within the protection scope of the present invention.
Claims
1. A message persistence state management system for intermittent satellite-to-ground links, characterized in that, The system includes a ground gateway side and a satellite gateway side; The ground gateway includes a state merging and overlay module, a persistence manager, a pre-sorting locking engine, and a ground distributor. The state merging and overlay module, during satellite offline periods, retrieves state update records for the same service object from the database based on a composite unique identifier composed of message type and resource key. If no corresponding record exists, a new READY record is added. If an incomplete READY record or a PENDING record exists, its payload and timestamp are directly overwritten without generating a new message copy, thus folding multiple intermediate states of the same service object into a single final state to be sent. The persistence manager manages the atomic states of messages from READY, PENDING, IN_FLIGHT, ACKED to FAILED. READY indicates the message has been persistently stored and is awaiting scheduling; PENDING indicates the message has entered the sending queue and is awaiting departure; IN_FLIGHT indicates the message has been sent and is in the satellite-to-ground link transmission or waiting for satellite processing; ACKED indicates a confirmation receipt has been received from the satellite gateway; and FAILED indicates that the number of retries has exceeded or transmission failed due to link abnormalities. Each state change is written to the local database in real time. The pre-ordering locking engine is used to assign a sequence key (ordering_key) to messages with sequential dependencies. After the ground distributor retrieves the ready state task with sequence key K1 and obtains the first message, it immediately establishes a mutex lock on K1. After sending the message, its state is changed to IN_FLIGHT. During the period when K1 is locked, subsequent messages with the same sequence key are prohibited from leaving the database. The mutex lock on K1 is released only after receiving an acknowledgment from the satellite gateway, and the next message is allowed to enter the sending process. The ground distributor is used to integrate message status, sequence key and link quality information. When waiting for confirmation, it generates a dynamic timeout threshold based on real-time link round-trip time sampling and channel quality evaluation factor through a weighted adaptive algorithm. The channel quality evaluation factor is jointly determined by time delay jitter, satellite overpass elevation angle and signal-to-noise ratio. When the timeout is not confirmed, the corresponding record is rolled back to READY state and then re-enters the transmission process or is redirected to a new master node. The satellite gateway side includes an ID mapping deduplication module, a stateless executor, and a state recovery manager. The ID mapping deduplication module maintains a fixed-capacity set of executed message IDs using a circular buffer or hash table. It identifies duplicate messages through hash matching. When a match is found, it is determined to be a duplicate message, an acknowledgment is immediately generated and returned to the ground, and the received payload is discarded. When a match is not found, it is determined to be the first arrival, and the payload is handed over to the stateless executor to directly call the corresponding service processor to perform control operations according to the receiving order. After the service execution is completed, the current message ID is written to the set of executed message IDs. The state recovery manager is used to enable the newly elected master node to read the set of executed message IDs and the state of incomplete transactions in the local persistent storage when the satellite gateway node is promoted from a slave node to a master node or demoted from a master node to a slave node. It reconstructs the memory index of the idempotent deduplication set and the state machine of the service processor. After the recovery is completed, it receives data frames sent by the ground. When a link abnormal interruption occurs, the ground distribution process restarts, or the on-board master node switches, the persistent manager on the ground gateway side and the state recovery manager on the satellite gateway side work together to perform state scanning and compensation recovery.
2. The message persistence state management system for intermittent satellite-to-ground links according to claim 1, characterized in that, The ground distributor uses a coroutine model to implement dynamic lifecycle management. When a satellite link is detected to be available, a coroutine instance is created, and the coroutine is destroyed when the link is out of the window or an anomaly occurs.
3. The message persistence state management system for intermittent satellite-to-ground links according to claim 1, characterized in that, The system employs the QUIC protocol optimized for satellite-to-ground links as the transport layer. This optimized QUIC protocol, designed for the long latency, high packet loss, and intermittent access characteristics of satellite-to-ground communication, extends the connection liveness timer to support long-term session suspension across window interruptions and 0-RTT fast reconnection. A customized congestion control algorithm is implemented, introducing a packet loss differentiation mechanism that distinguishes between channel noise and network congestion. Multiplexing stream strategy orchestration is performed in conjunction with the application layer's sequence key.
Citation Information
Patent Citations
Constellation satellite multi-channel management and control information interaction transmission system architecture and transmission method
CN115412148A
Reliable communication system based on publishing and subscribing mode in satellite-ground weak network environment
CN121077546A
Method and system for reducing message instances
CN1849831A