A satellite cloud-oriented distributed file system
By designing a metadata management and topology-aware replica placement strategy for the TiKV architecture in a satellite cloud environment, the problems of high communication overhead and high power consumption in the CubeSat cluster were solved, achieving efficient metadata management and load balancing, and improving system performance and reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIHANG UNIV
- Filing Date
- 2023-03-28
- Publication Date
- 2026-04-28
AI Technical Summary
Existing distributed file system designs are unsuitable for CubeSat clusters, resulting in high communication overhead and power consumption. Furthermore, traditional distributed file systems assume high reliability of the communication medium, while CubeSat clusters face a complex communication environment in space with dynamically changing network topology. Existing solutions fail to effectively utilize the Torus network topology characteristics of satellite clouds, leading to insufficient system performance and availability.
A distributed file system for satellite clouds was designed, which adopts a metadata management module based on the TiKV architecture. It combines high availability design of metadata with stateless metadata nodes and uses a topology-aware replica placement strategy to optimize replica placement and communication paths, reduce communication costs and energy consumption, and provide high availability and low latency access.
It achieves efficient metadata management and load balancing in the satellite cloud environment, reduces communication overhead and power consumption, improves data access efficiency and system reliability, and adapts to the dynamic network topology characteristics of the satellite cloud.
Smart Images

Figure CN116401225B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of distributed storage technology, and in particular to a distributed file system for satellite clouds. Background Technology
[0002] After decades of rapid development, modern remote sensing technology and its applications, which originated in the 1960s, have entered a new stage. Remote sensing technology has significant application value in many fields, including crop monitoring and yield estimation, land resource monitoring and surveying, disaster monitoring and assessment, and marine environmental monitoring, and is closely related to national economic development, ecological environmental protection, and national security. Furthermore, remote sensing technology plays a crucial role in scenarios closely related to people's daily lives, such as weather forecasting, map navigation, and air quality monitoring. In recent years, with the continuous development of remote sensing technology, it has acquired new characteristics: high spatial resolution, high temporal resolution, and high spectral resolution, and its application areas have become more extensive, penetrating into all aspects of military and civilian use.
[0003] The current remote sensing satellite mission chain for acquiring information is quite long, mainly including ground mission planning, on-board storage of remote sensing data, satellite-to-ground data transmission, and ground data reception and processing. Furthermore, remote sensing satellites can only transmit data to ground stations during their transits; due to communication duration and bandwidth limitations, the data arrival delay is on the order of days, severely impacting the timeliness of remote sensing satellite operations.
[0004] In recent years, 5G mobile communication technology has entered the stage of large-scale commercial use, and many research institutions and companies have begun to explore and research the next-generation mobile communication technology, 6G. Space-Ground integrated networking is a key technology in 6G, which involves satellites flying at different orbital altitudes, aircraft flying in different airspaces, and ground networks forming a completely new mobile communication network. This space-ground integrated network can achieve on-demand coverage of remote areas, seas, and the air through space-based and air-based systems, while also providing routine network coverage in cities through ground networks. It is highly practical and possesses advantages such as flexible networking and resilience.
[0005] Recently, satellite internet solutions have been validated in space. This approach involves deploying satellites in bulk to form a collaborative computing network, which is expected to serve various fields such as urban development, environmental monitoring, disaster prevention and mitigation, and emergency communications. Furthermore, satellite edge computing can be updated as needed. For example, by using inference from on-orbit AI models, changes in mountain images before and after heavy rain can be compared to detect landslides or other geological disasters in advance, issuing warnings to relevant departments and helping people prepare for emergencies. Therefore, building a cloud computing resource management and service system is of great significance to satellite internet. It can leverage the network, storage, and computing capabilities of collaborative satellites to enable rapid data processing and delivery, overcoming the bottleneck of long data transmission times.
[0006] A key aspect of building a satellite cloud is integrating storage resources. Since remote sensing imagery consists of large files with diverse data types, it is well-suited for storage using distributed file systems. After years of development, distributed file systems have matured, providing a unified view of file resources and offering features such as massive storage, high scalability, redundant replication, and high availability. However, the computing power of nodes in a satellite cloud is limited, and the network topology differs from traditional data center topologies, exhibiting a typical Torus network. This Torus network structure allows nodes to communicate with multiple neighbors, with multiple routing paths between nodes, resulting in better network connectivity, stronger fault tolerance, and robustness. Furthermore, the changing distances between satellites due to their orbital motion present challenges for designing a distributed file system.
[0007] In summary, this invention must take into account the limited core resources of satellite nodes, make full use of the network topology and motion characteristics of satellite clouds, and design a distributed file system for satellite clouds to improve system performance and availability.
[0008] Low Earth Orbit (LEO) satellites are satellites that orbit near the Earth's surface. The altitude of LEO satellites is generally below 2,000 kilometers. Because LEO satellites are closer to the ground, they also move faster. Currently, the vast majority of remote sensing satellites, communication satellites, and space stations use LEO.
[0009] Generally, satellite-to-ground communication uses radio signals, while satellite constellations are interconnected via intersatellite links (ISLs). ISLs connect satellites in different orbits to those in the same orbit. The distance between satellites within the same orbit remains constant throughout the connection, but the distance between satellites in different orbits changes as satellites move. Although satellite movement causes changes in the network topology, established connections must be maintained within the network, with each satellite connected to its successor and predecessor satellites within its orbital plane, as well as to the nearest satellite in each adjacent plane. Furthermore, because light travels 47% faster than fiber optics in a vacuum in space, satellite networks can significantly reduce network latency compared to terrestrial networks.
[0010] Although satellites are highly mobile, the actual topology of the network does not change. It is obvious that a constellation can be modeled as an N×M 2D Torus network, where N is the number of orbits and M is the number of satellites in the orbits.
[0011] With advancements in satellite communication networks, people are exploring further possibilities, such as LEO edge computing. In edge computing, computing and storage resources are embedded within the network and closer to consumers to provide lower access latency, increased throughput, enhanced privacy, and reduced costs. For LEO satellite networks, the network edge is the satellite constellation itself, as the satellites communicate directly with user equipment (ground stations). Therefore, many have proposed using LEO Edge, adding computing and storage resources to LEO satellites to build edge applications such as CDNs or IoT preprocessing.
[0012] Currently, there are very few systems that support large-scale storage on satellites. CubeSat is a pioneer in this field. CubeSat is a small satellite used for space research. It is relatively inexpensive and can accomplish more work in less time. However, CubeSat has limited storage, computing, and communication capabilities, and a single CubeSat cannot complete many tasks on its own. Therefore, it is necessary to build a CubeSat cluster.
[0013] The CubeSat Distributed File System (CDFS) is a distributed storage system designed for CubeSat clusters to support subsequent distributed processing. Existing distributed file systems, such as Google File System, Hadoop Distributed File System, and Lustre, are designed for distributed systems using reliable wired communication media. These file systems are unsuitable for storing data on CubeSat clusters because traditional distributed file systems assume reliable communication media and are designed for systems that organize nodes in a flat or hierarchical manner using racks, where communication costs are not high. For CubeSat clusters in space, traditional distributed file system designs would incur significant communication overhead and consume large amounts of power. CDFS overcomes these problems through a series of solutions, including load balancing, replication during transport, and tree-based routing and aggregation to control message reduction.
[0014] Cloud Constellation plans to establish a cloud infrastructure called Space-Belt, based on low Earth orbit (LEO) satellites. This infrastructure aims to provide secure data storage for large enterprises and government agencies, including internet service providers, telecommunications businesses, and others. To address the pervasive global data insecurity crisis, the system will utilize a combination of LEO-orbiting satellites and secure terrestrial networks, allowing clients to securely store large amounts of sensitive and mission-critical data in space. The advantages of this system are readily apparent, as it protects high-value data by: 1. Completely isolating it from the internet and leased terrestrial lines; 2. Protecting it from cyberattacks and clandestine activities; 3. Protecting it from natural disasters and major events on Earth; 4. Resolving all jurisdictional issues; and 5. Avoiding the risk of violating privacy regulations.
[0015] Existing solutions employ a layered architecture for data storage based on space-based cloud design. In the data storage layer, datasets are stored in a distributed server cluster deployed on LEO satellites. At least one server is embedded in the onboard unit of each LEO satellite. Each server then hosts one or more virtual machines (VMs) that can run popular data storage technologies such as HBase, Hive, and HDFS to provide data management capabilities. Currently, large-scale storage systems entirely based on satellites are relatively rare; however, there are many storage systems that leverage satellites and integrate with terrestrial data centers.
[0016] Distributed file systems designed for terrestrial data centers primarily employ two core technologies: metadata management and multi-replica technology. Currently, there are three main models for metadata services in distributed file systems: centralized metadata management, distributed metadata management, and designs without metadata services.
[0017] Centralized metadata management separates metadata and data storage services. In this architecture, the metadata server is dedicated to responding to client queries and storing metadata. This architecture is simple and clear, and GFS, HDFS, and Lustre all use it. However, a single primary data service presents a single point of failure. To address this, GFS provides fault tolerance for the single primary data service node through operation log files. If a metadata node fails, another metadata node can be quickly started based on the log files, enabling system recovery. Lustre introduces an active / standby architecture. The active metadata node acts as the active node, fulfilling read and write requests from the data nodes, while the standby metadata node acts as a backup node, maintaining state synchronization with the active node. HDFS HA mode employs redundant hot-standby NameNodes, responsible for periodically merging metadata information in the primary NameNode's memory. In the event of a failure, a new primary NameNode is elected via ZooKeeper. In applications with smaller data volumes or specific application environments, a single master data service has shown advantages in reducing the communication costs of metadata access and the overhead of maintaining metadata consistency. However, as cloud computing and big data applications continue to put increasing pressure on the scale, performance, and availability of storage systems, a single master data service architecture is clearly unable to meet the many requirements of big data processing.
[0018] The distributed metadata management model adopts a cluster mode, dividing metadata across different nodes in the cluster. These nodes collaborate to provide metadata services. Compared to a centralized metadata management model, this model allows all nodes in the cluster to share the access load of metadata, effectively solving the problems of single-point performance bottlenecks and single-point failures.
[0019] Distributed metadata management models can be further divided into fully peer-to-peer models and fully distributed models. In a fully peer-to-peer model, each metadata node can provide services externally, and nodes maintain data consistency through consensus protocols such as Gossip and Paxos, periodically synchronizing data and state among nodes. In a fully distributed model, ...
[0020] Nodes in a cluster can only provide partial metadata services. When using it, the services of each node need to be integrated to jointly provide metadata services, as exemplified by Ceph and GPFS. Through its unique CRUSH algorithm and dynamic subtree partitioning strategy, CephFS distributes system tasks and metadata across different nodes, with each node jointly responding to metadata access requests. By distributing the access pressure across multiple nodes, it improves system throughput and eliminates single points of failure. However, distributed metadata management models are more complex to design and implement, and developers must fully consider the inconsistencies and performance overhead caused by distributed metadata management.
[0021] The metadata-free server model is theoretically feasible, as long as a metadata addressing method that doesn't rely on location-based queries is designed. Ideally, the metadata-free node model offers numerous advantages, eliminating single points of failure, performance bottlenecks, and distributed data consistency issues associated with metadata. Therefore, the metadata-free node model allows for linear growth in system concurrency and scalability while also improving system reliability. GlusterFS uses this model, employing a resilient hash algorithm to completely remove metadata nodes and services, resulting in excellent system scalability. GlusterFS simplifies the communication chain for metadata access through hash addressing, significantly improving the efficiency of metadata access. However, the metadata-free service model also faces several challenges. For example, if a node fails, the hash bucket needs to be redistributed, incurring significant data migration costs. Furthermore, due to the lack of a central node, this model is inefficient for operations such as directory traversal.
[0022] Centralized metadata management suffers from single points of failure and performance bottlenecks. Fully distributed models are complex to implement and difficult to maintain consistency. Models without metadata servers lack a central control node and are difficult to control. Furthermore, none of these three models take into account network topology characteristics.
[0023] Regarding multi-replica technology, cloud computing relies on a large number of inexpensive commercial machines, which often experience hardware failures. Multi-replica technology is a common technique in the cloud computing field to improve data reliability. It utilizes redundant copies of data to ensure reliability while improving data access efficiency. A key aspect of replica technology is the replica placement algorithm, which allocates a corresponding set of storage nodes to each file or data block. This directly determines the system's utilization of resources such as disk space, I / O, and energy consumption, directly impacting cluster performance and job execution time. Therefore, optimizing the replica placement algorithm is the most crucial part of optimizing the storage system.
[0024] HDFS employs a rack-aware replica placement strategy. The NameNode can access the network topology of all DataNodes, resulting in a more ideal network environment between nodes on the same rack compared to those on different racks. HDFS stores replicas both within and across racks. Storing replicas within the same rack reduces network transmission overhead, while storing them across different racks improves data security. This strategy improves file read / write efficiency by reducing data transfer between racks, and enhances data reliability and security by storing replicas on different racks. However, because this strategy lacks comprehensive consideration of factors such as node load, node network distance, and node hardware performance, it is prone to load imbalance.
[0025] Wentao Zhao addresses the shortcomings of the default placement strategy by considering the space differences of disks across all nodes, which to some extent makes the data placement strategy on heterogeneous HDFS nodes more optimal. Xiaolong Ye proposes a new replica placement strategy that comprehensively considers the real-time status and disk utilization of each node when selecting nodes for replica placement, thereby achieving the goal of balancing node workload.
[0026] Luo's proposed Replica Placement Strategy (SRPM) is based on the SVM algorithm. SRPM first extracts features from hardware information such as CPU performance, disk speed, and server load of each node in the cluster, as well as the network distance between nodes. Then, it classifies nodes based on the extracted feature values and finally selects suitable nodes for placing replicas from among the better nodes. Zhao Wentao's proposed replica placement strategy is sensitive to node storage space. This strategy calculates the storage space of nodes in the cluster and then selects nodes with more storage space to store replicas. Lin Weiwei's proposed data replica placement strategy comprehensively considers node data load information and network topology distance to make decisions, using a linear weighting method to simplify the multi-objective optimization problem into a single-objective optimization problem. It selects the data storage node with the better decision value to store data replicas. Jing Xu proposed a replica placement strategy based on load and user history information. This strategy not only utilizes user historical replica access characteristics to target replica placement but also considers node load, thus improving system performance to a certain extent. For edge cloud, Li Chunlin proposed a multi-objective optimization model to formulate the optimal replica placement strategy, taking into account various factors such as file unavailability, node load, and network transmission costs. At the same time, in order to solve the problem of file access hotspots caused by sudden requests, he proposed a hotspot replica migration model, which includes the load of data acquisition nodes, hot file selection strategy, and dynamic replica migration.
[0027] The above research focuses on replica placement in traditional data center networks. Different network topologies employ significantly different replica placement algorithms. Compared to wired networks, wireless mesh networks experience reduced throughput on long paths consisting of numerous wireless hops due to wireless channel contention between adjacent mesh nodes and interference from adjacent wireless links. Zakwan proposed a novel heuristic algorithm, MP-DNA, which minimizes the number of hops between the requesting node and the replica server in its cost function, while also considering content popularity. In 2014, Zakwan further proposed an improved replica placement algorithm for WMNs, optimized for scenarios where content is unevenly distributed among nodes. Yaling Tao proposed a caching network scheme constructed from a large number of caching nodes in an Internet edge network to store replicas of sensor data required by end users.
[0028] For replica placement in data center networks, rack-aware strategies reduce data transfer between racks and improve file read / write efficiency, but they do not consider factors such as node load, node hardware performance, and node network distance. Network distance-aware replica placement considers the network distance between nodes, but only applies to data center networks and does not consider load conditions. Load-balancing replica placement strategies comprehensively consider factors such as CPU, storage space, and network I / O, making full use of cluster resources, but lack dynamism and do not fully utilize the Torus network topology characteristics of satellite clouds. Furthermore, most current replica placement solutions for wireless mesh networks are designed for sensor nodes with limited computing and storage resources, and are not suitable for satellite cloud scenarios.
[0029] Distributed file systems for data center networks are already quite mature, but the Torus network topology of satellite clouds differs from that of terrestrial clouds and exhibits time-varying characteristics. This invention addresses satellite cloud scenarios, considering availability and node load, and fully leverages the network topology characteristics of satellite clouds to design and implement a distributed file system for satellite clouds. This provides underlying storage for satellite clouds, ensuring high data availability, reducing data transmission costs, and improving data access efficiency. Summary of the Invention
[0030] To this end, the present invention first proposes a distributed file system for satellite clouds, which consists of five components: a metadata management module, a storage engine, a policy, a client, and a satellite topology service module.
[0031] The metadata management module uses a metadata storage model based on the TiKV architecture for metadata storage. It also adopts a high availability design for metadata by dividing the metadata into different data blocks and designs metadata encoding methods and metadata caching mechanisms to ensure load balancing of the metadata cluster and efficient data access. In addition, it designs a master data service to establish stateless metadata nodes.
[0032] The Metadata Engine is a metadata engine composed of metadata servers. Since the metadata servers only cache the metadata retrieved from the underlying key-value pairs and do not perform persistence operations, they can be re-cached when a crash occurs. Therefore, it is stateless.
[0033] Storage Engine is the data storage module, which consists of all satellites. Each satellite is a block server. It uses a metadata node selection strategy, a replica placement strategy optimized for low-latency access scenarios, and a topology-aware replica placement strategy to store and read / write block data.
[0034] Policy is a strategy module designed for satellite cloud scenarios. It provides the metadata management module with a metadata storage model selection, and provides the Storage Engine with a metadata node selection strategy and a topology-aware replica placement strategy. It also optimizes the replica placement strategy for low-latency access scenarios.
[0035] Users interact with the satellite cloud file system using a client, and the interaction between the client and the satellite cloud file system is achieved through a dedicated file system interface.
[0036] The satellite topology service module is specifically designed for satellite cloud scenarios, providing satellite cloud file systems with information on satellite cloud network topology, inter-satellite link distances, and satellite-to-ground communication intervals, and providing input parameters for mechanisms and strategies.
[0037] The modules communicate with each other using RPC.
[0038] The metadata storage model models metadata information as key-value pairs and stores it using the distributed KV storage system TiKV. It uses two tables, a directory entry table and an index node table, to store directory tree information. The directory entry table stores the content in the directory tree and is converted into KV form in a multi-way tree structure. The key value of the multi-way tree stores the index node ID of the parent directory at the beginning. The index node table stores the index node information in the tree, with the key value being the index node ID. Hard links are implemented using a reference counting method.
[0039] The high-availability design of the metadata uses the replica mechanism of multi-raft-group to divide the data into roughly equal slices according to the key-value range. Each slice has multiple replicas, and one of the replicas is the Leader that provides read and write services. The storage nodes schedule these slices and replicas through the scheduling node to ensure that the data and read / write loads are evenly distributed across all storage nodes.
[0040] The method for establishing the stateless metadata node is as follows: The main data service stores the unified access to TiKV. The block server operates on the metadata through the main data service. The main data service does not actually store data but only caches the data in TiKV. At the same time, the main data service maintains the load conditions of each node in memory. The block server reports its network address when starting, and each block server regularly reports its own load information. The main data service maintains the information of each block server in memory.
[0041] The selection strategy of the metadata node adopts a resource placement algorithm based on the 2D-Torus network. First, regard the network topology as an unweighted graph, that is, make D M = D N , use the hop count as the distance metric, which is called d-Hops placement, where d is the hop count. Use the basic 2D-Torus resource placement algorithm to ensure that the hop count from each node to the metadata node does not exceed d. According to the N×M 2D-Torus network topology of the satellite cloud model, if N and M are both multiples of k, and k = 2d 2 + 2d + 1, the perfect placement position of a k×k 2D-Torus network is:
[0042] [i, 2d 2 i](mod k), i = 0, 1,..., (k - 1)
[0043] Since N and M may not be integer multiples of k, but N = pk + r and M = qk + s, and 0 < s, r < k, thus construct (p + 1)×(q + 1) k×k 2D-Torus sub-blocks, then use the perfect placement algorithm respectively, and finally delete k - r rows and k - s columns. When D M ≠ D N , the network topology is not an unweighted graph. Transform the basic 2D-Torus resource placement algorithm to extend N×M to N×N, where N > M, and then compress it to N×M.
[0044] The topology-aware replica placement strategy selects the optimal replica placement plan through local screening, optimizing for bandwidth, energy, and load balancing, and using the minimum spanning tree algorithm:
[0045] Considering bandwidth performance, after a node receives data, it will delineate an area with that node as the center and a fixed number of hops or distance as the radius. Any selection strategy made afterward can only be considered within this area.
[0046] Considering energy performance, the replication path method is optimized by adopting a chained replication approach. The path is selected with the fewest relay points and the location where the most replicas are placed as the storage relay as the optimization strategy.
[0047] To consider load balancing performance, the load score of all nodes in the region is obtained when placing replicas. The load score is represented by LBF. Input: node storage capacity N. storage Node computing power N cpu Node communication bandwidth N bandwidth Remaining power of node N power Total storage capacity (Total_Storage), Total processing power (Total_Cpu), Total communication bandwidth (Total_Bandwidth), Total power (Total_Power), and storage weight (W). storage Calculate the weight W cpu Communication bandwidth W bandwidth Electricity weight W power The parameters are calculated through the following steps:
[0048] LBF storage =N storage ÷Total_Storage
[0049] LBF cpu =N cpu ÷Ttotal_Cpu
[0050] LBF bandwidth =N bandwidth ÷Total_Bandwidth
[0051] LBF power =N badwidth ÷Total_Power
[0052] LBF = W storage ×LBF storage +W cpu ×LBF cpu +W bandwidth ×LBF bandwidth +W power ×LBF power
[0053] The output LBF is the load score. When filtering the placement of replicas, nodes with load scores below a specified threshold are first excluded. Then, nodes are selected based on their load scores as a probability. In other words, nodes with higher load scores are more likely to be selected.
[0054] Then, combining the three methods, a combined filtering optimization strategy is designed to minimize communication costs, save energy and bandwidth, and still ensure high-reliability data storage. Assuming the total number of nodes is N and the required number of replicas is P, the specific algorithm is described as follows:
[0055] The first node A to receive the write request accesses the Master to obtain the topology location information. It divides a local region centered on the Master based on the specified number of hops or distance, and filters out all nodes outside the region. If the number of nodes in the region is M, then NM candidates are filtered out in this round.
[0056] Then, among these nodes, those below the load score threshold are filtered out, leaving M1 nodes.
[0057] Next, M1 nodes are selected with probability based on their load scores. In this case, nodes with high load scores are highly likely to be selected repeatedly, while nodes with low load scores are less likely to be selected. Then, the selected M1 nodes are deduplicated, leaving M2 nodes. Finally...
[0058] With the remaining M2 nodes, P nodes need to be selected to minimize the communication cost of node A replicating data to these P nodes. This involves constructing an undirected graph with the remaining M2 nodes, where the weights are the number of hops or the distance. Starting from node A, Prim's algorithm is used to construct the minimum spanning tree. The tree stops when the number of nodes is P+1. The P nodes added using Prim's algorithm at this point are the optimal locations for replica placement and replication.
[0059] The optimized replica placement strategy for low-latency access scenarios addresses scenarios where the system tracks and processes the same area, data about hot areas is frequently generated, and hot replicas are frequently accessed. It adopts a topology-aware combined strategy, expanding the locally selected area to the global area. Nodes are not filtered in the area selection step. Load filtering and load score probability selection are used as constraints instead of direct filtering conditions. The optimal placement strategy of the Torus network is used again to reduce the communication cost of accessing replicas.
[0060] The block server specifically includes three parts: block storage, data block verification, and read / write process.
[0061] The block storage structure serializes the data block structure into JSON format and stores it in the block server directory;
[0062] The database verification function is designed as follows: each data block contains several blocks, each block stores a fixed 4MB of data, and the block stores the data checksum. When reading a block, it is necessary to verify whether the stored checksum and the calculated checksum are consistent. If they are inconsistent, it means that the data has been lost or corrupted, and it is necessary to use the data from other copies of the block to correct it.
[0063] The read / write process is implemented through two independent processes: a read process and a write process.
[0064] The read process takes chunk_id as input and outputs error_code, indicating whether the read process was successful. The algorithm first constructs a corresponding ChunkReader based on the chunk_id. Then, it finds the corresponding data block metadata file based on the chunk_id, loads it into memory, and reconstructs the memory image of the data block. Next, it iterates through all blocks contained in the data block. Each iteration reads the data of that block and calculates the checksum. If the checksum verification fails, an error is returned directly; otherwise, the add method of ChunkReader is called. After calling add, ChunkReader temporarily stores the data in a buffer. After all blocks have been traversed, the commit method of ChunkReader is called, and ChunkReader writes a CommitLog entry indicating that the data block was read successfully. If the system crashes during block traversal, the data block read is considered to have failed.
[0065] The write process takes a chunk_id and block data as input. The algorithm first constructs a ChunkWriter based on the chunk_id, then loads the data block's metadata into memory, allocates a new block_id, creates a block, writes the block data, offset, and checksum, and then writes the block to disk. Next, it modifies the data block's metadata, changes the offset, inserts a new block_id, writes the data block information to disk, and finally calls the ChunkWriter's commit method to commit the write. At this point, the block data writing is successful.
[0066] The client is implemented using multi-threaded data block and block transmission and multi-interface access.
[0067] The multi-threaded transmission method involves the client calling MDS to obtain the file's metadata information, first splitting the large file into 64MB data blocks, then submitting the entire data block to the thread pool, then splitting the data block into several blocks, and resubmitting another task to the thread pool. The thread pool will call the write interface provided by Chunkserver to insert into the corresponding data blocks.
[0068] The multi-interface access method is compatible with POSIX, supports Fuse mounting, and establishes a C_Connector to support Fuse mounting, exposing the C++ interface as a C interface, and defining a series of fuse_operations to map the corresponding interface to system-written functions.
[0069] The satellite topology service module provides information related to the satellite network topology, including the satellite's position coordinates and the communication range between the satellite and the ground station. The data comes from data exported in batches by STK Engine. The system of this invention generates source data by calling STK Engine through Python. To avoid frequently starting and calling STK Engine, the satellite topology service will save data for a period of time and provide services externally. If the time slice of the requested content is not in the saved data, it will request STK Engine to obtain the data. The data flow process is as follows: the client calls the satellite topology service module through Brpc. If the requested data is in the database of the satellite topology service module, it will be processed and returned directly. If not, it needs to call STK Engine through Python STKAPI, write the data to the database, and then process and return it.
[0070] The technical effects to be achieved by this invention are as follows:
[0071] 1. High-Availability Distributed Metadata Management Mechanism. Metadata management is the most fundamental and crucial part of a distributed file system, storing information such as file hierarchy and block storage locations. Due to the Torus network topology of satellite clouds, network communication requires multiple hops. Using a single-point metadata server can lead to communication bottlenecks and single points of failure. This invention studies a high-availability distributed metadata management mechanism for satellite cloud Torus networks, which distributes metadata across different metadata servers to ensure efficient and reliable metadata retrieval.
[0072] 2. Topology-Aware Replica Placement Strategy. Replicas ensure system reliability; how replicas are placed determines system reliability and also affects system read efficiency, making it a crucial aspect of distributed file system replication. Satellite cloud payloads differ from powerful terrestrial clouds; their payload storage and computing capabilities are limited, so they must make fuller use of existing resources. Furthermore, inter-satellite link communication bandwidth is also a valuable resource. Typically, data goes through a process of uploading, waiting for processing, scheduling processing, waiting for landing, and landing. Most data remains unchanged after being written and is not frequently accessed. This means that randomly placing global replicas would lead to significant waste of inter-satellite link bandwidth. Therefore, this invention studies a topology-aware replica placement and replication strategy that utilizes the network topology characteristics of satellite clouds to minimize the communication costs and energy consumption associated with replica placement.
[0073] 3. Optimization of Replica Placement Strategy for Low-Latency Access Scenarios. Consider a special scenario where a certain area is frequently photographed by remote sensing satellites over a period of time; this is referred to as a hotspot area in this invention. In this scenario, replicas of hotspot areas are also frequently accessed and updated. In this case, the communication cost during initial replica placement is no longer the primary consideration; the key factor becomes the communication cost during data access and updates. Therefore, this invention studies an optimization strategy for replica placement in low-latency access scenarios to minimize access latency. Attached Figure Description
[0074] Figure 1 System Architecture Diagram
[0075] Figure 2 Metadata directory tree structure
[0076] Figure 3 Distributed KV System Architecture Diagram
[0077] Figure 4 Local copy filtering strategy
[0078] Figure 5 Bandwidth and energy optimization strategies
[0079] Figure 6 Global replica filtering strategy
[0080] Figure 7 Client-side multi-threaded upload process
[0081] Figure 8 Client-side multi-interface access
[0082] Figure 9 Topology service interaction process
[0083] Figure 10 File system interface ls flowchart
[0084] Figure 11 File system interface write flowchart Detailed Implementation
[0085] The following are preferred embodiments of the present invention, which are described in conjunction with the accompanying drawings. However, the present invention is not limited to these embodiments.
[0086] This invention proposes a distributed file system for satellite clouds.
[0087] System Architecture:
[0088] To address the shortcomings of existing technologies, this invention proposes a distributed file system for satellite clouds. The Satellite Cloud File System (SatFS) mainly consists of five components: Metadata Engine, Storage Engine, Policy, Client, and Topo Service. Figure 1 As shown, the Metadata Engine is the metadata management module of the SatFS satellite cloud file system, built on TiKV. It provides highly available and reliable metadata storage, while dividing metadata into different partitions to ensure load balancing and efficient data access within the metadata cluster. The Master node in the Metadata Engine does not persist data and is therefore stateless. The Storage Engine is the data storage module, composed of all satellites. Each satellite is a Chunk Server, storing chunk data and providing read / write chunk functionality. The Policy module of the SatFS satellite cloud file system primarily optimizes various aspects for satellite cloud scenarios, including metadata storage model selection, metadata node selection strategies, topology-aware replica placement strategies, and replica placement strategy optimizations for low-latency access scenarios. Users interact with the SatFS satellite cloud file system using the Client, currently supporting CLI and Fuse mounting. The Topo Service is specifically designed for satellite cloud scenarios, aiming to provide the SatFS satellite cloud file system with information such as the satellite cloud network topology, inter-satellite link distances, and satellite-to-ground communication intervals, providing input parameters for various mechanisms and policies. The modules communicate and coordinate with each other using RPC to form the satellite cloud file system SatFS.
[0089] Metadata Management Module:
[0090] The metadata management module stores metadata using a metadata storage model based on the TiKV architecture. It also employs a high-availability metadata design to divide metadata into different data blocks and designs metadata encoding methods and metadata caching mechanisms to ensure load balancing of the metadata cluster and efficient data access. Additionally, it designs a master data service to establish stateless metadata nodes.
[0091] Metadata storage model;
[0092] Traditional distributed file systems such as HDFS face many problems when using the native NameNode as a metadata server: (1) Metadata information includes the tree-like directory structure of the file system, the mapping relationship between files and data blocks, and the mapping relationship between physical blocks and storage locations. It is stored entirely in memory, and the single-machine carrying capacity is limited. (2) Accessing metadata uses a global read-write lock, resulting in poor read and write throughput performance. (3) As the cluster data scale increases, the restart and recovery time reaches the level of hours, during which no operation can be performed.
[0093] To address the above issues, this invention proposes a key-value (KV) storage model that models metadata information as key-value pairs and stores them using the distributed KV storage system TiKV. TiKV is built on the embedded database RocksDB, providing high-performance read and write access while also avoiding memory invalidation. Figure 2 This is a hypothetical directory tree with 3 directories, 1 symbolic link, and 4 files:
[0094] Table 1 Directory Tree Information Table
[0095]
[0096] The node information of the directory tree is shown in Table 1. This invention uses two tables to store the directory tree information: the Dentry Table and the Inode Table. The Dentry Table is used to store... Figure 2 The contents of the Directory Tree vary depending on the type of Dentry:
[0097] 1. FILE: Stores the file's base name and the Inode ID that the file points to. Based on the Inode ID, an Inode can be uniquely located in the Inode Table. Because the system of this invention supports hard links, information such as the number of file references, permissions, owner, mtime, ctime, and atime are all stored in the Inode Table.
[0098] 2. SYMLINK: Stores the base name of the symbolic link itself and the Inode ID that the symbolic link points to. Based on the Inode ID, an Inode can be uniquely located in the Inode Table. The Inode stores information such as the permissions, owner, mtime, atime, ctime of the symbolic link itself, and the name of the target it points to.
[0099] 3. DIRECTORY: Stores the directory's base name, directory Inode ID, and information such as directory permissions, owner, mtime, atime, and ctime.
[0100] A Dentry Table is essentially a multi-way tree, while the metadata in this invention's system is stored in key-value (KV) format. Therefore, the multi-way tree structure needs to be converted into KV form. The key must contain information about the parent directory for easy indexing. First, the parent directory name is excluded because recording the entire parent directory prefix would result in a large key range, which is inconvenient for storage. Therefore, this invention stores the parent directory's inode ID at the beginning of the key. This utilizes TiKV's prefix scan feature to easily list all child nodes of a directory, reducing the time complexity of the index file. Regarding the key encoding, as shown in Table 3, the first 8 bytes are a fixed length of the parent directory in big-endian order, followed by the current dentry's base name. Figure 2 The contents of the Dentry Table in the directory tree are shown in Table 4.
[0101] Table 3DentryTable storage format
[0102]
[0103] Table 4 Example Dentry Table
[0104]
[0105] Inode Table is used to store Figure 2 The `Inode info` section contains `Inode ID` as the key and reference count, file / symbolic link permissions, owner, `atime`, `ctime`, and `mtime` as the value. The most crucial part is the reference count, which is the core component for implementing hard links. Figure 2 The contents of the Inode Table in the directory tree are shown in Table 5.
[0106] Table 5 Example InodeTable
[0107]
[0108] In summary, representing all system metadata in key-value (KV) format provides a solid foundation for metadata segmentation and horizontal scaling.
[0109] High availability design for metadata:
[0110] Traditional NameNode designs suffer from a single point of failure, and the load on the metadata server increases with the number of files. Therefore, this invention proposes using a distributed key-value (KV) system, such as TIKV, to store metadata. The architecture of a distributed KV system is generally as follows: Figure 3 As shown, unlike traditional full-node backup methods, distributed key-value (KV) systems generally use a multi-raft-group replication mechanism. Data is divided into roughly equal slices (hereinafter referred to as Partitions) according to the range of keys. Each slice has multiple replicas (usually three), one of which is the Leader, providing read and write services. Storage nodes schedule these Partitions and replicas through scheduling nodes to ensure that data and read / write loads are evenly distributed across the storage nodes. This design has many advantages:
[0111] 1. High Availability: The distributed key-value system uses the Raft protocol to maintain data consistency. Any write request can only be made on the Leader, and a write success message will only be returned to the client after a majority of replicas have been written (the default configuration is 3 replicas, meaning all requests must successfully write to at least two replicas). The scheduling node also uses consistency protocols such as Paxos / Raft to avoid single points of failure.
[0112] 2. Load Balancing: The scheduling nodes of the distributed KV system collect information from each storage node through heartbeat packets, and maintain the even distribution of Leaders, the even distribution of storage capacity on each node, and the even distribution of access hotspots according to the scheduling strategy. They also control the speed of Balance, avoid affecting online services, and manage node status.
[0113] 3. Horizontal scaling: Distributed key-value systems generally support horizontal scaling, meaning that new nodes can be added without affecting online services.
[0114] The metadata storage model designed in this invention is based on a distributed key-value storage system. Each MDS is connected to a distributed key-value node, storing data in the underlying distributed key-value store, ensuring high availability and providing load balancing and horizontal scaling capabilities.
[0115] Metadata node selection strategy:
[0116] Different from the flat network structure of a ground data center, the network topology of a satellite cloud is a 2D-Torus network. In a ground environment, the location where the metadata cluster is placed is not a crucial factor because the network of the data center is not a bottleneck. However, in the scenario of a satellite cloud, the strategy for selecting the metadata cluster is of vital importance and affects the latency and throughput capabilities of the entire system. Therefore, based on the resource placement algorithm for a 2D-Torus network, this invention proposes a selection algorithm to meet different QoS requirements.
[0117] First, consider the simplest scenario. We regard the network topology as an unweighted graph, that is, make D M = D N , and use the number of hops as the distance metric. This is called d-Hops placement, where d is the number of hops. In this way, the basic 2D-Torus resource placement algorithm can be used to ensure that the number of hops from each node to the metadata node does not exceed d. According to the N×M 2D-Torus network topology modeled for the satellite cloud, if both N and M are multiples of k, and k = 2d 2 + 2d + 1, then the perfect placement positions for a k×k 2D-Torus network are:
[0118] [i, 2d 2 i](mod k), i = 0, 1, …, (k - 1)
[0119] However, it is not exactly the case that both N and M are integer multiples of k. More generally, N = pk + r and M = qk + s, and 0 < s, r < k. Bae and Bose proposed constructing (p + 1)×(q + 1) k×k 2D-Torus sub-blocks, then using the perfect placement algorithm respectively, and finally deleting k - r rows and k - s columns. The result obtained in this way is not perfect, that is, there may be a node whose distance to the metadata node is slightly greater than d hops, but it does not affect the actual effect. A rather tricky problem is that D M ≠ D N , and the network topology is not an unweighted graph. Therefore, it is necessary to transform the basic 2D-Torus resource placement algorithm, extend N×M to N×N (assuming N > M), and then compress it to N×M to obtain the optimal placement scheme.
[0120] Metadata encoding:
[0121] As mentioned earlier, metadata is mainly stored in four tables: Dentry Table, Inode Table, Fchunk Table, and Dnode Table. In practice, it is only necessary to set different key prefixes for different tables. Table 6 shows the encoding format of the data tables. inode_id and chunk_id are allocated using allocate(). Since the underlying key-value storage stores the maximum value that has been allocated, it can be incremented. Integers are converted into big-endian strings. In addition, to avoid frequent retrieval of inode_id and chunk_id, metadata nodes will retrieve inode_id and chunk_id in batches.
[0122] Table 6 Metadata Encoding Table
[0123]
[0124] Stateless metadata nodes:
[0125] MDS storage is uniformly connected to TiKV. Chunkservers operate on metadata through MDS, but MDS does not actually store data; it only caches data in TiKV. Simultaneously, MDS maintains the load status of each node in memory, including CPU utilization, memory usage, and remaining disk space. This information also does not need to be persisted. Chunkservers report their network addresses upon startup, and each Chunkserver also periodically reports its own load information. MDS maintains information about each Chunkserver in memory, namely the ChunkserverInfo structure, as shown in Table 7, including the Chunkserver's IP address, on / off status, CPU utilization, total memory size, memory usage, total disk space, and remaining disk space. This information is also needed for subsequent decision-making. Therefore, MDS is essentially stateless, thus requiring no additional redundancy mechanism to ensure MDS failure, allowing for easy start-up and shutdown, providing great flexibility. Metadata management separates storage and computation. TiKV is the unified storage for all metadata, providing an infinitely scalable storage platform, while MDS can be viewed as a compute node, performing metadata computation operations. When the metadata access pressure on the entire system increases, either MDS nodes or TiKV nodes can be scaled up. When the access pressure decreases, some MDS nodes can be shut down. When the metadata storage capacity is insufficient, TiKV nodes can be scaled up. Additionally, the capacity and expiration policy of the MDS node cache can be adjusted appropriately based on the business scenario.
[0126] Table 7 ChunkserverInfo Structure Description
[0127]
[0128] Metadata caching mechanism:
[0129] Accessing TiKV from MDS requires network transmission, but satellite cloud network resources are relatively limited and bandwidth is very precious. Therefore, the system of this invention provides a caching mechanism on MDS to cache metadata obtained from TiKV. However, caching inevitably brings consistency issues. This invention chooses to sacrifice some consistency in exchange for precious bandwidth resources.
[0130] This invention employs a caching mechanism based on LRU (Least Recently Used), a commonly used replacement algorithm that selects the least recently used data for eviction. Simultaneously, this invention sets expiration times to ensure that the MDS (Multi-Distributed System) does not access excessively old data, thus preventing disruption to system consistency. If an operation on the cached content fails in the MDS, the cached data will also be evicted, and data will be retrieved again from TiKV.
[0131] Instance placement strategy:
[0132] The solution consists of two parts: a replica placement strategy optimization for low-latency access scenarios and a topology-aware replica placement strategy.
[0133] Topology-aware copy placement strategy:
[0134] Traditional distributed file systems such as GFS and HDFS are designed for computer clusters with flat structures and data centers with hierarchical structures, where nodes are organized into racks and data centers. They also require ample power supplies and reliable networks. However, data distribution and replication operations designed and optimized for this architecture, such as random replica placement, lead to significant communication overhead and energy consumption. To address this, this invention proposes a topology-aware replica placement strategy. Through local filtering, optimization of bandwidth, energy, and load balancing, and the use of a minimum spanning tree algorithm, the optimal replica placement scheme is selected.
[0135] To reduce communication overhead, after a node receives data, it delineates a region centered on that node with a fixed number of hops or distance as the radius. Any subsequent selection strategies can only be considered within this region. For example... Figure 4 As shown, the red dot in the center represents the client. The data written can only be placed within the red dotted line box. This is reasonable and effective for satellite cloud scenarios because after the data is distributed, it will be read by batch processing tasks. When performing calculations, it is only necessary to distribute the operators to each shard. Compared to file data, the bandwidth occupied by operator transmission is minimal, and there are very few random reads. Therefore, local copy placement can effectively reduce the cost of data transmission.
[0136] When selecting replicas, considering the replication path allows for more refined optimizations. Figure 5 For example, suppose A receives file data and wants to make two copies. There are two options: the first is to place the copies at B and C, and the second is to place them at D and E. If we calculate the distances from A to B and C and from A to D and E respectively, the communication cost is 1 + 2 = 3 links. However, in the first option, A can first copy the copy to C, and then C can copy the copy to B, thus only incurring a communication cost of 2 links. Therefore, the first option is better, and this invention calls it "chain copying". Therefore, when selecting the copy placement location, this invention cannot only consider the distance from the source point and each placement point, but also whether this optimization strategy can be used.
[0137] The goal of load balancing is to balance the storage capacity, processing power, communication bandwidth, and power consumption of each node when distributing data across the cluster. During replica placement, the load score of all nodes within the region is obtained. The load score is represented by LBF, and its calculation method is shown in Table 8.
[0138] Table 8. Pseudocode for Load Score Calculation Algorithm
[0139]
[0140] When selecting replica placement locations, nodes with load scores below a specified threshold are first excluded. Then, each node is selected based on its load score as a probability. In other words, nodes with higher load scores are more likely to be selected. Based on this replica selection method, the system tends to be load-balanced.
[0141] Utilizing the three optimization strategies described above, this invention proposes a combined filtering optimization strategy to minimize communication costs, save energy and bandwidth, while still ensuring highly reliable data storage. Assuming the total number of nodes is N and the required number of replicas is P, the specific algorithm is described below:
[0142] 1. Node A that receives the write request accesses the Master to obtain the topology location information, divides a local region centered on it according to the specified number of hops or distance, and filters out all nodes outside the region. If the number of nodes in the region is M, then NM candidates are filtered out in this round.
[0143] 2. Among these nodes, filter out those below the load score threshold, leaving M1 nodes;
[0144] 3. Select M1 duplicate nodes with the probability of each node's load score. In this case, nodes with high load scores are very likely to be selected repeatedly, while nodes with low load scores are less likely to be selected. Then, deduplicate the selected M1 nodes, leaving M2 nodes.
[0145] 4. From the remaining M2 nodes, we need to select P nodes to minimize the communication cost of node A replicating data to these P nodes. Naturally, we can construct an undirected graph with the remaining M2 nodes, using hop count or distance as the weights. Starting from node A, we use Prim's algorithm to construct a minimum spanning tree, stopping when the tree has P+1 nodes. The P nodes added using Prim's algorithm at this point represent the optimal locations for replica placement and replication.
[0146] Optimization of replica placement strategy for low-latency access scenarios:
[0147] When a hotspot area appears during a certain period of time, the system may track and process the same area, and data about the hotspot area is frequently generated and hotspot copies are frequently accessed. In this case, the local copy placement strategy will lead to an increase in global communication cost. This invention proposes a global copy selection strategy for this scenario to reduce the latency of ground base stations accessing hotspot data.
[0148] Similarly, a topology-aware combination strategy is adopted: first, selection is based on region; then, selection is based on load score; and finally, the Prim minimum spanning tree algorithm is used to generate a replica placement scheme. This invention expands the locally selected region to the global region for low-latency access scenarios, meaning that nodes are not filtered in the region selection step. Furthermore, load filtering and load score probability selection are used as constraints, rather than as direct filtering conditions. Likewise, to ensure global access latency, the optimal placement strategy for Torus networks described above is used again to reduce the communication cost of accessing replicas. Since the orbits change periodically, there are theoretically N×M optimal placement schemes, and they are all equivalent.
[0149] like Figure 6 As shown, the location of the black dot is obtained according to the optimal placement strategy of the Torus network. The location obtained by translating the black dot is also an optimal placement scheme, depending on the choice of reference frame. However, these placement schemes are different. If scheme A is chosen, the communication cost when node S replicates the copy is significantly higher, so this scheme is excluded. Then, comparing schemes B and C, it can be seen that the virtual triangles formed by schemes B and C can both enclose S. However, the link communication cost of scheme B is 8, while the link communication cost of scheme C is 9. Therefore, scheme B is better. This invention also selects the replica location based on this principle.
[0150] Block server module
[0151] A block server, also known as a chunk server, primarily performs read and write chunk tasks. Since satellite clouds typically deal with large files such as remote sensing images, the system described in this invention defaults to a maximum chunk size of 64MB. Furthermore, to avoid wasting space, the system treats chunks as logical data, with each chunk containing several blocks, and each block storing a fixed 4MB of data.
[0152] Block storage organization structure:
[0153] Chunks, as a logical storage structure, do not actually store data. As shown in Table 9, `chunk_id` is the unique identifier of a chunk, `bytes_size` represents the actual number of bytes in a chunk, and `blocks` stores the `block_id`s of all blocks in order. A block is a real physical storage unit, as shown in Table 10, which includes `block_id`, the offset of the currently written block, the data, and the checksum.
[0154] Table 9. Description of Chunk Structure
[0155]
[0156] Table 10 Block Structure Description
[0157]
[0158] The chunk structure is serialized into JSON format and stored in the Chunkserver directory. Chunks are essentially metadata, while blocks, which are of fixed size, are also persistently stored in the specified chunk's directory. This method effectively reduces wasted space and improves file read performance.
[0159] Data block verification:
[0160] The verification mechanism of a storage system involves calculating a checksum for user-generated data during storage or network transmission, and transmitting or storing this checksum along with the data. When this data is used again, the checksum is recalculated and compared with the stored checksum to verify data consistency. Network transmission is susceptible to errors due to network conditions; only by verifying each network communication can data be prevented from being lost or corrupted over the network. Although calculating the checksum takes time, it is necessary, and ensuring data accuracy is the primary design principle of a storage system.
[0161] The block stores the data checksum, effectively preventing data loss and corruption due to network transmission and system failures. When reading a block, it is necessary to verify whether the stored checksum and the calculated checksum are consistent. If they are inconsistent, it indicates that the data has been lost or corrupted, and it is necessary to use other copies of the block data to correct it. The system described in this invention uses Google's open-source CRC32 to calculate the checksum, generating a 32-bit checksum, which is stored in the block's checksum.
[0162] Block read / write process:
[0163] Block reading and writing are the core implementation of the block server. The performance of block reading and writing directly affects the read and write performance of the entire distributed file system. At the same time, reading and writing blocks must tolerate scenarios such as network failures and power outages to ensure the consistency of block data. Block reading and writing mainly rely on two classes, ChunkReader and ChunkWriter, which require modification of both Chunk metadata and Block data.
[0164] Table 11 shows the pseudocode for the Chunk read algorithm. The input is `chunk_id`, and the output is `error_code`, indicating whether the read process was successful. The algorithm first constructs a corresponding `ChunkReader` based on the `chunk_id`. Then, it finds the corresponding `Chunk` metadata file based on the `chunk_id`, loads it into memory, and reconstructs the memory image of the `Chunk`. Next, it iterates through all blocks contained in the `Chunk`. Each iteration reads the data of that block and calculates the checksum. If the checksum verification fails, an error is returned directly; otherwise, the `add` method of `ChunkReader` is called. After calling `add`, `ChunkReader` temporarily stores the data in a buffer. After all blocks have been traversed, `ChunkReader`'s `commit` method is called, and `ChunkReader` writes a `CommitLog` to indicate that the `Chunk` was read successfully. This is done to ensure the atomicity of reading the `Chunk`. If the system crashes while traversing blocks, since the `commit` method has not yet been called, the read of the `Chunk` can be considered a failure. The design of `add` and `commit` references Git semantics, providing semantics similar to transaction atomicity.
[0165] Table 11 Reading Chunk Algorithm Pseudocode
[0166]
[0167]
[0168] The block write process is shown in Table 12. The inputs are chunk_id and block data. The algorithm first constructs a ChunkWriter based on the chunk_id, then loads the chunk's metadata into memory, allocates a new block_id, creates a block, writes the block data, offset, and checksum, and then writes the block to disk. Next, it modifies the chunk's metadata, changes the offset, inserts a new block_id, writes the chunk information to disk, and calls the ChunkWriter's commit method to commit the write. At this point, the block data writing is successful.
[0169] Table 12 shows the pseudocode for the Chunk algorithm.
[0170]
[0171] Client module:
[0172] Multi-threaded transmission:
[0173] When transferring large files, transmitting them in the order of Chunks and Blocks would significantly reduce read and write performance, causing latency to be directly proportional to file size, which is unreasonable. Therefore, the client must use multi-threading to transmit Chunks and Blocks. After the client calls MDS to obtain the file's metadata, it can split the file into Chunks, and each Chunk can be further divided into several Blocks. Because the data does not overlap, all transmissions can be performed in parallel. The Chunkserver needs to queue the Chunks and Blocks transmitted in parallel, and logically related data is written in order.
[0174] The system described in this invention uses a thread pool for parallel transmission, wherein the upload processing flow is as follows: Figure 7 As shown, a large file is first split into 64MB chunks, then the entire chunk is submitted to the thread pool. Next, the chunk is split into several blocks, and another task is submitted to the thread pool. The thread pool will call the write interface provided by Chunkserver to insert into the corresponding chunk.
[0175] Multiple interface access:
[0176] To facilitate user operation, this invention system implements multiple interface access methods, such as CLI, Fuse, and HTTP interfaces, see [link to documentation]. Figure 8The CLI interface is a custom interface that is not compatible with POSIX and HDFS, but it is customizable, more flexible and convenient. The CLI is mainly implemented by embedding the Client SDK, in which the Client communicates with MDS and Chunkserver via RPC.
[0177] To ensure compatibility with POSIX, this invention also supports Fuse mounting. Users only need to mount the system to a directory to operate SatFS as if it were a kernel file system. However, since Fuse only supports C, the system needs to create a C_Connector to expose the C++ interface as a C interface. The implementation of Fuse is relatively fixed, requiring the definition of a series of fuse_operations to map the corresponding interfaces to system-written functions. When mount is executed, the operations performed on the file system are translated into system-written functions via VFS, hence the name user-space file system.
[0178] In addition, this system also supports RESTful HTTP interface access. The HTTP interface is implemented through Brpc, so users no longer need a client. However, uploading and downloading files still requires the front end to cooperate.
[0179] Satellite Topology Service Module:
[0180] The satellite topology service node primarily provides information related to the satellite network topology, including satellite position coordinates and communication ranges between satellites and ground stations. Essentially, the satellite topology service is a data service; the data originates from data exported in batches by STKEngine. The system described in this invention generates source data by calling STK Engine via Python. To avoid frequent invocations of STK Engine, the satellite topology service saves data for a period of time before providing it externally. If the time slice of the requested content is not in the saved data, it requests the data from STK Engine. The data flow is roughly as follows... Figure 9 As shown, the Client calls TopoService via Brpc. If the requested data is in the database of TopoService, it processes and returns directly. If not, it needs to call STK Engine via Python STKAPI to write the data to the database before processing and returning.
[0181] File system interface implementation process:
[0182] The interfaces currently supported by the system described in this invention are shown in Table 13. Next, we will analyze two representative processes, ls and write, which involve multiple services. Among them, ls does not involve interaction with Chunkserver.
[0183] Table 13 List of File System Interface Implementations
[0184]
[0185] The Ls process is as follows Figure 10 As shown, MDS first obtains the Inode ID based on the path in `ls`. After obtaining the Inode ID corresponding to the path, it calls the `Get` method of `KV` to retrieve the path information. If the path is a file, it constructs the file's stat information and returns it to the Client. If the path is a directory, it needs to construct a prefix, call `KV`'s `PrefixScan` to perform a prefix lookup, retrieve all files or directories under the directory, construct the stat information, and return it.
[0186] The Write process is as follows: Figure 11 As shown, the Client initiates a Write request to the MDS. The MDS obtains the Inode ID of the file, then calculates the Chunk information of the file, including the Chunk ID and the location of the replica. The Chunk information is written into the KV and returned to the Client. The Client splits the task according to the Chunk information and transmits the file fragments to the corresponding Chunkserver in multiple threads until the entire file is successfully transferred.
Claims
1. A distributed file system for satellite clouds, characterized in that: The system consists of five components: metadata management module, Storage Engine, Policy, client, and satellite topology service module; The metadata management module uses a metadata storage model based on the TiKV architecture for metadata storage. It also adopts a high availability design for metadata, dividing the metadata into different data blocks and designing metadata encoding methods and metadata caching mechanisms to ensure load balancing of the metadata cluster and efficient data access. Additionally, it designs metadata services to establish stateless metadata nodes. The Metadata Engine is a metadata engine composed of metadata servers, which are stateless. Storage Engine is the data storage module, which consists of all satellites. Each satellite is a block server. It uses a metadata node selection strategy, a replica placement strategy optimized for low-latency access scenarios, and a topology-aware replica placement strategy to store and read / write block data. Policy is a strategy module designed for satellite cloud scenarios. It provides the metadata management module with a metadata storage model selection, and provides the Storage Engine with a metadata node selection strategy and a topology-aware replica placement strategy. It also optimizes the replica placement strategy for low-latency access scenarios. Users interact with the satellite cloud file system using a client; The satellite topology service module is specifically designed for satellite cloud scenarios, providing the satellite cloud file system with information on the satellite cloud's network topology, inter-satellite link distances, and satellite-to-ground communication intervals, thus providing input parameters for mechanisms and strategies. The satellite cloud file system needs to select the optimal replica placement strategy based on the network topology, while the system... The modules communicate with each other using RPC, and the system communicates with each other through a dedicated file system interface to enable the client to interact with the satellite cloud file system; The selection strategy for the metadata nodes adopts a resource placement algorithm based on 2D-Torus networks. First, the network topology is treated as an unweighted graph, meaning that... Using hop count as the distance metric, this is called d-Hops placement, where d is the hop count. It uses a basic 2D-Torus resource placement algorithm to ensure that the hop count from each node to the metadata node does not exceed d, based on satellite cloud modeling. 2D-Torus network topology, if and All Multiples of, and ,one The perfect placement for a 2D-Torus network is: because and Not necessarily Not an integer multiple of, but and ,and Therefore, construct indivual The 2D-Torus sub-blocks are then processed using the perfect placement algorithm, and finally deleted. lines and column, when At that time, the network topology is not an unweighted graph. The basic 2D-Torus resource placement algorithm is then used to transform... Extended into ,in Then compress it into .
2. A distributed file system for satellite clouds as described in claim 1, characterized in that: The metadata storage model models metadata information as key-value pairs and stores it using the distributed KV storage system TiKV. It uses two tables, a directory entry table and an index node table, to store directory tree information. The directory entry table stores the content in the directory tree and is converted into KV form in a multi-way tree structure. The key value of the multi-way tree stores the index node ID of the parent directory at the beginning. The index node table stores the index node information in the tree, with the key value being the index node ID. Hard links are implemented using a reference counting method.
3. A distributed file system for satellite clouds as described in claim 2, characterized in that: The metadata high availability design uses a multi-raft-group replication mechanism to divide the data into roughly equal slices according to the range of key values. Each slice has multiple replicas, one of which is the Leader, providing read and write services. Storage nodes schedule these slices and replicas through scheduling nodes to ensure that data and read / write loads are evenly distributed across the storage nodes.
4. A distributed file system for satellite clouds as described in claim 3, characterized in that: The method for establishing stateless metadata nodes is as follows: the master data service stores unified access to TiKV, the block server operates the metadata through the master data service, the master data service does not actually store the data, but only caches the data in TiKV, and at the same time, the master data service maintains the load status of each node in memory, the block server reports its own network address when it starts up, each block server periodically reports its own load information, and the master data service maintains the information of each block server in memory.
5. A distributed file system for satellite clouds as described in claim 4, characterized in that: The topology-aware replica placement strategy selects the optimal replica placement scheme through local filtering, optimization of bandwidth, energy, and load balancing, and the use of the minimum spanning tree algorithm. Considering bandwidth performance, after a node receives data, it will delineate an area with that node as the center and a fixed number of hops or distance as the radius. Any selection strategy made afterward can only be considered within this area. Considering energy performance, the replication path method is optimized by adopting a chained replication approach. The path is selected with the fewest relay points and the location where the most replicas are placed as the storage relay as the optimization strategy. To consider load balancing performance, the load score of all nodes in the region is obtained when placing replicas. Load fraction is represented by LBF. Input: node storage capacity. Node computing power Node communication bandwidth Remaining power at the node Total storage capacity Total processing capacity Total communication bandwidth Total power Storage weight Calculate weights Communication bandwidth Electricity weight The parameters are calculated through the following steps: The output LBF is the load score. When filtering the placement of replicas, nodes with load scores below a specified threshold are first excluded. Then, nodes are selected based on their load scores as a probability. In other words, nodes with higher load scores are more likely to be selected. Then, combining the three methods, a combined filtering optimization strategy is designed to minimize communication costs, save energy and bandwidth, and still ensure high-reliability data storage. Assuming the total number of nodes is N and the required number of replicas is P, the specific algorithm is described as follows: Node A, which first receives the write request, accesses the Master to obtain topology location information. Based on a specified hop count or distance, it divides a local region centered on the Master and filters out all nodes outside this region. If the number of nodes within the region is M, then this round of filtering... One candidate; Then, among these nodes, those below the load score threshold are filtered out, leaving M1 nodes. Next, select M1 duplicate nodes with probability based on the load score of each node. In this case, nodes with high load scores are very likely to be selected repeatedly, while nodes with low load scores are less likely to be selected. Then, deduplicate the selected M1 nodes, leaving M2 nodes. Finally, select P nodes from the remaining M2 nodes to minimize the communication cost of node A replicating data to these P nodes. That is, construct an undirected graph with the remaining M2 nodes, with weights of hop count or distance. Starting from node A, use Prim's algorithm to construct a minimum spanning tree. Stop when the number of nodes in the tree is P + 1. The P nodes added by Prim's algorithm at this point are the optimal replica placement and replication positions.
6. A distributed file system for satellite clouds as described in claim 5, characterized in that: The optimized replica placement strategy for low-latency access scenarios addresses scenarios where the system tracks and processes the same area, data about hot areas is frequently generated, and hot replicas are frequently accessed. It employs a topology-aware combined strategy: first, selection is based on the region; then, selection is based on the load score; and finally, the Prim minimum spanning tree algorithm is used to generate a replica placement scheme. The difference lies in expanding the locally selected region to the global region. Specifically, nodes are not filtered during the region selection step, and load filtering and load score probability selection are used as constraints rather than direct filtering conditions. Furthermore, the Torus network optimal placement strategy is used again to reduce the communication cost of accessing replicas.
7. A distributed file system for satellite clouds as described in claim 6, characterized in that: The block server specifically includes three parts: block storage, data block verification, and read / write process. The block storage structure serializes the data block structure into JSON format and stores it in the block server directory; The data block verification function is designed as follows: each data block contains several blocks, each block stores a fixed 4MB of data, and the block stores the data checksum. When reading a block, it is necessary to verify whether the stored checksum and the calculated checksum are consistent. If they are inconsistent, it means that the data has been lost or corrupted, and it is necessary to use the data from other copies of the block to correct it. The read / write process is implemented through two independent processes: a read process and a write process. The read process takes chunk_id as input and outputs error_code, indicating whether the read process was successful. The algorithm first constructs a corresponding ChunkReader based on chunk_id, then finds the data block metadata file corresponding to the data block based on chunk_id, loads it into memory, and restores the memory image of the data block. Next, it iterates through all blocks contained in the data block. Each iteration requires reading the data of that block and calculating the checksum. If the checksum verification fails, an error is returned directly; otherwise, the add method of ChunkReader is called. After calling the add method, ChunkReader temporarily stores the data in a buffer. After all blocks have been traversed, the commit method of ChunkReader is called. ChunkReader writes a CommitLog to indicate that the data block was read successfully. If the system crashes while traversing the blocks, the data block read is considered to have failed. The write process takes chunk_id and block data as input. The algorithm first constructs a ChunkWriter based on chunk_id, then loads the metadata information of the data block into memory based on chunk_id, allocates a new block_id and creates a block, writes the block data, offset and checksum, and then writes the block to disk. Next, it modifies the data block metadata information, changes the offset, inserts a new block_id, writes the data block information to disk, and calls the commit method of ChunkWriter to commit the write. At this point, the write of the block data is successful.
8. A distributed file system for satellite clouds as described in claim 7, characterized in that: The client is implemented using multi-threaded data block and block transmission and multi-interface access. The multi-threaded transmission method involves the client calling MDS to obtain the file's metadata information, first splitting the large file into 64MB data blocks, then submitting the entire data block to the thread pool, then splitting the data block into several blocks, and resubmitting another task to the thread pool. The thread pool will call the write interface provided by Chunkserver to insert into the corresponding data blocks. The multi-interface access method is compatible with POSIX, supports Fuse mounting, and establishes a C_Connector to support Fuse mounting, exposing the C++ interface as a C interface, and defining a series of fuse_operations to map the corresponding interface to system-written functions.
9. A distributed file system for satellite clouds as described in claim 8, characterized in that: The satellite topology service module provides information related to the satellite network topology, including the satellite's position coordinates and the communication range between the satellite and the ground station. The data comes from data exported in batches by STK Engine. The system generates source data by calling STK Engine through Python. To avoid frequently starting and calling STK Engine, the satellite topology service saves data for a period of time and provides services externally. If the time slice of the requested content is not in the saved data, it requests STK Engine to obtain the data. The data flow process is as follows: the client calls the satellite topology service module through Brpc. If the requested data is in the database of the satellite topology service module, it processes and returns directly. If not, it needs to call STK Engine through the Python STK API, save the data to the database, and then process and return it.
Citation Information
Patent Citations
System-level simulation demonstration verification method facing satellite mobile communication
CN105915304A
Data consistency system based on satellite network environment
CN108111338A