A Network Offloading Method for 6G Satellite-Ground Computing Power Based on Lightweight Blockchain and Hierarchical Reinforcement Learning

By constructing a space-ground computing power network based on lightweight blockchain and hierarchical reinforcement learning, the problems of insufficient computing power resource coordination, poor adaptability of consensus mechanism and lack of security guarantee in the space-ground computing power network are solved. It realizes efficient utilization and real-time offloading of computing power resources across the entire domain, improves security and adaptability, and is suitable for computing offloading services in remote areas and emergency rescue scenarios.

CN122137445APending Publication Date: 2026-06-02NANJING UNIV OF POSTS & TELECOMM

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV OF POSTS & TELECOMM
Filing Date
2026-02-26
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

The space-ground computing network suffers from problems such as insufficient coordination of computing resources, poor adaptability of consensus mechanisms, and lack of security systems. This leads to suboptimal resource allocation, insufficient security, and poor model robustness during the computing offloading process, making it difficult to meet the requirements of optimal global computing power configuration and real-time offloading.

Method used

We employ a lightweight blockchain and layered reinforcement learning approach to construct a three-tier consortium blockchain system consisting of a terminal, a satellite, and a cloud. By combining a node dynamic reputation assessment model, a lightweight satellite-ground practical Byzantine fault-tolerant consensus mechanism, and a smart contract with four-fold verification logic, we optimize the offloading decision and solve the global optimization objective function.

Benefits of technology

It enhances the trust and security of the unloading process and the efficiency of global computing resource utilization, adapts to the dynamic and real-time requirements of satellite-ground networks, reduces communication complexity and consensus latency, ensures data integrity and traceability, and provides efficient computing unloading services suitable for remote areas and emergency rescue scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122137445A_ABST
    Figure CN122137445A_ABST
Patent Text Reader

Abstract

This invention relates to the field of 6G communication technology, and in particular to a 6G satellite-to-ground computing power network offloading method based on lightweight blockchain and hierarchical reinforcement learning. The method includes: S1, constructing a computing power network model based on user local devices, low Earth orbit satellites, and remote cloud data centers; S2, calculating the latency and energy consumption of different communication paths based on the user's offloading decision; S3, establishing a latency model and an energy consumption model by combining the offloading ratio, latency, energy consumption, and blockchain system overhead; S4, solving for the optimal solution of the global optimization objective function to generate the optimal offloading action for each user. This invention solves the technical problems of insufficient collaboration of computing power resources in satellite-to-ground computing power networks, poor adaptability of consensus mechanisms, and lack of security systems, achieving the goal of optimizing the utilization efficiency of computing power resources across the entire domain and adapting to the dynamic and real-time requirements of satellite-to-ground networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of 6G communication technology, and in particular to a method for offloading 6G satellite-to-ground computing power networks based on lightweight blockchain and hierarchical reinforcement learning. Background Technology

[0002] With the rapid development of sixth-generation mobile communication technology (6G) and low-orbit satellite communication technology, Space-Ground Computing Network (SGCPN) has become the core carrier for achieving full-scenario coverage of "air, land, sea" and providing full-domain computing power services. By integrating the coverage advantages of terrestrial cellular networks and satellite networks, it can effectively solve the problems of uneven distribution of ground computing power and limited terminal computing power, and provide efficient computing offloading services for scenarios such as communication in remote areas and emergency rescue. It is currently the focus of research and industrial layout in the field of information and communication.

[0003] However, the inherent characteristics of space-to-ground computing networks, such as spatial heterogeneity, dynamic links, and multi-node distributed nature, lead to three major technical challenges in the computation offloading process: (i) Insufficient coordination of computing resources: The existing strategy adopts a single-level resource allocation method, which does not take into account the latency sensitivity of tasks, the differences in computing power requirements and the dynamic changes of links to achieve cross-level dynamic coordination. Moreover, the resource allocation and offloading decisions are isolated from each other and there is no closed-loop optimization mechanism. This can easily lead to satellite node overload or resource idleness, making it difficult to achieve optimal configuration of computing power across the entire domain. The existing strategy has not established a collaborative scheduling mechanism for the three-level computing power nodes of "terminal-satellite-cloud", and has failed to give full play to the complementary advantages of the computing power of the three types of nodes.

[0004] (ii) Poor adaptability of consensus mechanisms: Existing solutions mostly adopt a single-level flat modeling, which is not adapted to the layered characteristics of a three-level computing power architecture. Moreover, the modeling ignores the security costs such as blockchain consensus latency and verification overhead, and only focuses on performance indicators such as latency and energy consumption, which is out of touch with the actual needs of "security-performance" synergistic optimization. At the same time, it does not consider the dynamic nature of links caused by satellite orbital motion and the temporary offline status of nodes, resulting in insufficient model robustness and poor actual adaptability of the offloading strategy based on it. Furthermore, the traditional distributed consensus mechanism has high communication complexity and cannot adapt to the dynamic access / exit characteristics of satellite-ground network nodes. The consensus latency is too high to meet the real-time offloading requirements.

[0005] (III) Lack of security system: Privacy protection schemes either have excessive computational overhead that is difficult to adapt to the limited computing power of satellites, or ignore the dynamic bandwidth changes of satellite-to-ground links, resulting in uncontrollable latency; the lightweight consensus mechanism does not consider the scenario of dynamic node adjustment, has insufficient fault tolerance performance, and does not design a layered protection strategy based on the different privacy levels of tasks, resulting in low accuracy of node security assessment in the initial stage; the data verification smart contract is not optimized for the characteristics of satellite-to-ground transmission, resulting in low verification efficiency; cross-node data transmission is prone to security threats such as identity forgery and data tampering, and the existing verification methods rely on end-to-end encryption, lack decentralized and trusted verification means, and cannot achieve full lifecycle integrity traceability of data. Summary of the Invention

[0006] To address the shortcomings of the existing technologies, the present invention aims to provide a 6G satellite-to-ground computing power network offloading method based on lightweight blockchain and hierarchical reinforcement learning. This method addresses the technical problems of insufficient collaboration of computing power resources in satellite-to-ground computing power networks, poor adaptability of consensus mechanisms, and lack of security assurance systems. Ultimately, it aims to optimize the utilization efficiency of computing power resources across the entire domain and adapt to the dynamic and real-time requirements of satellite-to-ground networks.

[0007] To achieve the above objectives, the technical solution of the present invention is as follows: A method for network offloading of 6G satellite-to-ground computing power based on lightweight blockchain and hierarchical reinforcement learning includes the following steps: S1. Construct a computing power network model based on user local devices, low Earth orbit satellites and remote cloud data centers, and simultaneously establish a consortium blockchain system model adapted to the computing power network. Then, define the communication method and communication model according to the computing power network model. S2. Based on the user's offloading decision, determine the offloading ratio of user computing tasks to be offloaded to the user's local device, low Earth orbit satellite, and remote cloud data center for computing, calculate the transmission rate of different communication paths, and calculate the latency and energy consumption of different communication paths for processing based on the characteristics of the user computing tasks and communication methods. S3. Design a dynamic reputation evaluation model for nodes in a blockchain system, a lightweight, practical Byzantine fault-tolerant consensus mechanism for satellite and ground systems, and a smart contract for data security verification. Model the overhead of the blockchain system and establish a latency model and an energy consumption model by combining the offloading ratio, latency, energy consumption and the overhead of the blockchain system. S4. With the goal of minimizing the user's computational unloading cost during the interaction between the user and the computing network environment, a global optimization objective function is established. The computational unloading problem is modeled as a hierarchical Markov decision process model, and the optimal solution of the global optimization objective function is solved to generate the optimal unloading action for each user.

[0008] Furthermore, the computing power network model includes N local user devices, a satellite constellation network consisting of M low Earth orbit satellites, and a remote cloud data center; The consortium blockchain system model includes: each satellite node as a lightweight node in the consortium blockchain, and data center nodes as full nodes in the consortium blockchain. Lightweight nodes are equipped with secure data verification smart contracts, generate lightweight new blocks through a consensus mechanism, and collaboratively maintain lightweight blockchain information; full nodes record full blocks containing complete computational offload information and maintain full blockchain information. Communication methods include: user local devices can directly utilize the computing resources of low Earth orbit satellites; low Earth orbit satellites can act as computing nodes to process user computing tasks, or as relay nodes to forward user computing tasks to remote cloud data centers.

[0009] Furthermore, the characteristics of user computing tasks The expression is: ; in: Indicates the user ID. This indicates the amount of data involved in the user's computation task. Indicates completion The amount of computation required by the CPU in the data volume. Indicates completion The amount of computation required by GPUs in the data volume. This indicates the privacy level of user n, with three levels: low, medium, and high. When a user's computing task is offloaded to the user's local device, only the computing latency of the user's computing task is considered. When user computing tasks are offloaded to low Earth orbit satellites, the propagation delay of the satellite-to-ground link, the data transmission delay, and the computing delay of the user computing tasks need to be considered. When user computing tasks are offloaded to a remote cloud data center, the data transmission latency of sending data to the remote cloud data center via low Earth orbit satellite relay needs to be considered.

[0010] Furthermore, the node dynamic reputation assessment mechanism includes: a performance evaluation function for low Earth orbit satellite nodes. And the reputation decay mechanism updates the reputation value of low Earth orbit satellite node m. S3's lightweight, practical, Byzantine fault-tolerant consensus mechanism for space-to-ground applications includes the following steps: S3-1-1. Select the node with the highest current reputation value as the master node, and replace it sequentially in case of failure. The number of consensus nodes is [number missing]. ,and f is the number of Byzantine nodes; S3-1-2. Divide the node reputation value into three categories: low, medium and high. Determine the minimum reputation threshold based on the highest privacy level of the task to be processed. Delete nodes in the consensus node set whose reputation value is lower than the threshold and add high reputation nodes of the corresponding level. S-1-3, The consensus process for deleting and adding consensus nodes is as follows: Initiation phase: The master node broadcasts information about nodes that do not meet the reputation threshold, triggering a dynamic adjustment process for consensus nodes; Deletion phase: After receiving the low-reputation node information broadcast by the master node, each consensus node synchronously verifies whether the reputation value of the node is lower than the reputation threshold corresponding to the current task. If the verification confirms that it does not meet the requirements, each consensus node removes the low-reputation node from its local consensus node set and sends a confirmation message back to the master node. After the master node collects at least 2f+1 valid deletion confirmations, the removal operation of the low-reputation node is completed. Adding phase: Each consensus node initiates an add request to the node with the highest reputation value at the current level, realizing a targeted invitation for the new node; Adding Response Phase: After receiving 2f+1 valid add requests, the new node sends a response message to all consensus nodes; Update phase: Each consensus node broadcasts the updated consensus node set information. When a consensus node receives 2f+1 consistent update confirmation messages, the new node is officially included in the consensus node set, completing the entire dynamic adjustment process.

[0011] Furthermore, the consensus mechanisms for deleting and adding consensus nodes include: Pre-preparation phase: The master node assigns proposal numbers and view numbers, calculates and signs data hashes, and broadcasts pre-preparation messages; Preparation phase: After the consensus node verifies the legality of the message, it broadcasts a preparation message and enters the preparation completed state after collecting at least 2f+1 consistent messages; Submission Phase: Consensus nodes only send submission messages to the master node. The master node collects at least 2f+1 submission messages and then proceeds to the next phase. Broadcast phase: The master node packages and signs the valid commit message and broadcasts the consensus certificate; Response phase: After verifying the credentials, each node stores the information and sends out a response. Lightweight nodes store hashes, while full nodes store complete data.

[0012] Furthermore, S3's data security verification smart contract includes the following steps: S3-2-1 Off-chain preprocessing: Legitimate users submit a unique identifier to the smart contract. and its public key For the raw data and timestamp Calculate hash value and with private key right and Joint signature obtained ,Will Upload to computing power nodes; S3-2-2, On-chain core verification logic: Identity legitimacy verification: Check Whether to register in the blockchain system; if not registered, return failure. S3-2-3, Signature Validity Verification: Passed Obtain the public key ,verify If the statement is invalid, it returns a failure message. S3-2-4, Time Validity Verification: If the current block time is inconsistent with... If the difference exceeds the preset threshold, it is judged as a replay attack and a failure is returned. The preset threshold is five minutes by default. S3-2-5, Data Integrity Verification: Local Computation by Computing Nodes ,like If all validations pass, it returns a failure message; otherwise, it returns a successful data validation message.

[0013] Furthermore, the latency of a blockchain system includes consensus mechanism latency and data integrity verification latency. When the number of Byzantine nodes is... The cost of hash encryption is The cost of a single digital signature generation or verification cycle is [missing information]. The cost of a single message validity verification cycle is... During the cycle, the cost of storage execution by each computing node is... The cycle time and the processing speed of each node are... GHz; When the latency of message transmission across the entire blockchain network or in a targeted manner is The transmission latency required for a single consensus process is consensus mechanism latency The expression is: ; Data integrity verification latency includes off-chain operation latency. Data transmission latency and on-chain operation latency The costs of on-chain authentication and timestamp verification are respectively and Off-chain operations perform hash calculations and generate digital signatures for data, while on-chain operations perform identity verification, signature verification, timestamp verification, and hash calculations. Data integrity verification latency The expression is: ; Total latency of blockchain system The expression is: ; When the average power consumption of the blockchain system is The total energy consumption of the blockchain system The expression is: ; Total latency The expression is: ; Total energy consumption The expression is: .

[0014] Furthermore, the decision to unload The expression is: , , ; in, This indicates the proportion of user computing tasks that are offloaded to the user's local device for computation. This indicates the proportion of user computing tasks that are offloaded to low Earth orbit satellites for computation. This indicates the proportion of user computing tasks that are offloaded to a remote cloud data center for computation. The communication model between the user's local device and the low Earth orbit satellite is as follows: ; The communication model between low Earth orbit satellites and remote cloud data centers is as follows: ; User calculates uninstallation cost The expression is: ; in, and These are the time delay coefficient and the energy consumption coefficient, respectively, and they satisfy... + =1.

[0015] Furthermore, S4 includes the following steps: S4-1. Model the upper-level task scheduling layer, determine the optimal unloading decision for each user task, and define the upper-level state space. It is used to describe three dimensions at time t: the global network state, user task attributes, and the state of the blockchain system, comprehensively depicting the global information required for upper-level decision-making. S4-2, Define the upper-level motion space , used to describe the unloading target decision for each user task at time t; S4-3. Define the upper-level reward function. , used to describe the immediate reward obtained by taking the unloading action in the stated state; S4-4. Model the lower-level node execution layer. Based on the offloading decisions determined by the upper layer, optimize the resource allocation and execution strategies for satellite nodes to improve the single-node task execution efficiency, while satisfying node resource constraints and feeding back the execution results to the upper layer; define the lower-level state space. This is used to describe the local state and task execution details of satellite node m at time t, providing a precise basis for resource allocation decisions. S4-5, Define the lower-level motion space , used to describe the resource allocation of satellite node m to each task in the task queue at time t; S4-6. Define the lower-level reward function. , used to describe the execution efficiency and energy consumption cost of satellite node m at time t.

[0016] Furthermore, the reinforcement learning model is trained using the state space, action space, and reward function of the upper and lower layers to generate the optimal uninstallation action for each user, including steps a~f: a. Initialize the parameters of the upper-layer Actor network Critic network parameters and the target Actor network parameters Target Critic network parameters Initialize the upper-level experience replay pool For each near-Earth orbit satellite, initialize the parameters of the lower-level Actor network. Critic network parameters Target Actor Network Parameters and target Critic network parameters Initialize the lower-level experience replay pool Initialize the number of training epochs, the number of training steps per epoch, the delayed update step size K, the batch size B, and the soft update coefficient. The training steps are the total number of time steps executed in one round of training; b. At the start of each training round, reset the upper-level state. and each lower-level state The initial state of the environment; at each time step, based on the current policy of the Actor network. Select the initial uninstall action. And it operates on the computing network environment, distributing tasks to each satellite node, where... Represents the upper-level state based on time t. And the parameters of the upper-layer policy network Through upper-level strategies Output the corresponding upper-level action; c. For each near-Earth orbit satellite, obtain its local status. According to the current strategy Select lower-level action Execute actions and observe lower-level rewards. and the next state Storage experience Store in the experience replay pool And update the current state space, making ; Observe the upper-level rewards and the next global state Storage experience Store in the experience replay pool And update the current state space, making ; d. Utilize the experience replay mechanism to retrieve data from the upper-level experience replay pool. We randomly sample B data points and optimize the loss function of the upper-layer Critic network by minimizing the loss function. Update the parameters of the upper-layer Critic network ; e. For each near-Earth orbit satellite, from the lower-level experience replay pool We sample B data points and optimize the loss function of the lower-level Critic network by minimizing the loss function. Update the parameters of the lower-level Critic network ; f. If the current training step number is an integer multiple of the delayed update step size K, then update the parameters of the upper-layer Actor network using the policy gradient method. Selecting the optimal unloading action maximizes the computational efficiency of the upper-layer Critic network. value; After each training round, the parameters of the target network in the upper and lower layers are updated based on the soft update formula, so that they gradually approach the parameters of their respective current Actor and Critic networks, achieving smooth adjustment of the target network parameters. The soft update formula is as follows: Once all training rounds are completed, a well-trained reinforcement learning model is obtained.

[0017] Compared with the prior art, the beneficial technical effects of the present invention are as follows: (I) Enhancing the Trust and Security of the Unloading Process: This invention constructs a three-level alliance chain system of "terminal-satellite-cloud", combining a node dynamic reputation assessment model, a lightweight satellite-ground practical Byzantine fault-tolerant consensus mechanism, and a smart contract with four-fold verification logic to achieve decentralized and trustworthy verification of the entire process of node access, data transmission, and unloading execution. This effectively resists security threats such as identity forgery, data tampering, and replay attacks, ensures the integrity and traceability of data throughout its entire lifecycle, and solves the problem of the lack of security protection system in existing technologies.

[0018] (II) Optimize the utilization efficiency of computing power resources across the entire domain: Through hierarchical Markov decision process modeling, the collaborative closed-loop optimization of upper-level task scheduling and lower-level resource allocation is realized. Cross-level resource coordination is carried out in combination with task characteristics, dynamic changes in links and node computing power status to avoid satellite node overload or resource idleness. The complementary advantages of the three-level computing power nodes are fully utilized, significantly improving the utilization rate of computing power resources across the entire domain and solving the pain points of low resource coordination efficiency and insufficient computing power collaboration in existing technologies.

[0019] (III) Adapting to the dynamic and real-time requirements of satellite-ground networks: The lightweight consensus mechanism reduces communication complexity and consensus latency by dynamically adjusting the consensus node set and optimizing the consensus process, adapting to the dynamic access or exit characteristics of satellite nodes; the layered modeling takes into account the requirements of security and performance optimization, incorporates the dynamic nature of links and offline node scenarios, improves the robustness of the model, and makes the offloading strategy more adaptable and more stable in the actual satellite-ground environment, meeting the latency requirements of real-time computation offloading.

[0020] (iv) Balancing security and performance costs: By modeling the overhead of the blockchain system, security costs such as consensus latency and verification overhead are incorporated into the global optimization goal, so as to achieve a dynamic balance between security and performance indicators such as latency and energy consumption. This avoids the performance degradation caused by pursuing security alone and solves the problem of existing technology modeling being out of touch with actual needs and the imbalance between security and performance.

[0021] (V) Expanding the application scenarios of satellite-ground computing power networks: This invention can effectively adapt to scenarios such as communication in remote areas and emergency rescue where there is no ground base station coverage or the link is unstable, providing efficient and reliable computing offloading services for these scenarios, breaking through the adaptation limitations of existing solutions in dynamic and complex scenarios, and helping the full-domain application of 6G satellite-ground converged networks. Attached Figure Description

[0022] Figure 1 This is a flowchart of the 6G satellite-to-ground computing power network offloading method based on lightweight blockchain and hierarchical reinforcement learning, as described in this invention. Figure 2 This is a diagram of the satellite-to-ground computing network architecture of the present invention; Figure 3 This is a diagram of the blockchain system architecture of the satellite-ground computing power network of the present invention; Figure 4 This is a diagram illustrating the consensus node adjustment process of this invention. Figure 5 This is a flowchart illustrating the consensus mechanism process of this invention; Figure 6 This is a diagram illustrating the hierarchical reinforcement learning algorithm framework based on DDPG of this invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the device proposed by this invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The advantages and features of this invention will become clearer from the following description. It should be noted that the drawings are in a very simplified form and use non-precise proportions, only for the purpose of conveniently and clearly illustrating the embodiments of this invention. Please refer to the drawings to make the objectives, features, and advantages of this invention more apparent and understandable. It should be understood that the structures, proportions, sizes, etc., depicted in the accompanying drawings are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed in the specification, and are not intended to limit the implementation conditions of this invention. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportional relationships, or adjustments to the size, without affecting the effects and objectives achieved by this invention, should still fall within the scope of the technical content disclosed in this invention.

[0024] A method for network offloading of 6G satellite-to-ground computing power based on lightweight blockchain and hierarchical reinforcement learning includes the following steps: S1. Construct a computing power network model based on user local devices, low Earth orbit satellites and remote cloud data centers, simultaneously establish a consortium blockchain system model adapted to the computing power network, and define the communication methods between user local devices, remote cloud data centers and low Earth orbit satellites according to the constructed computing power network model. S2. Based on the user's offloading decision, determine the offloading ratio of user computing tasks to be performed on the user's local device, low Earth orbit satellite, and remote cloud data center; establish a communication model between the user's local device, remote cloud data center, and low Earth orbit satellite, and calculate the transmission rate of different communication paths according to Shannon's formula; based on the characteristics of the user's computing tasks and the transmission rate of different communication paths, calculate the latency and energy consumption of the user's local device, low Earth orbit satellite, and remote cloud data center in processing the user's computing tasks, respectively, in combination with the communication method. S3. Design a dynamic reputation evaluation model for nodes, a lightweight space-ground practical Byzantine fault-tolerant consensus mechanism, and a data security verification smart contract in the blockchain system, and model the overhead of the blockchain system; combine the offloading ratio, latency and energy consumption, and the overhead of the blockchain system to establish a latency model and an energy consumption model. S4. To minimize the user's computational unloading cost during the interaction between the user and the computing network environment, a global optimization objective function is established. The computational unloading problem is modeled as a hierarchical Markov decision process model, and the optimal solution of the global optimization objective function is solved to generate the optimal unloading action for each user.

[0025] Specifically, the computing network architecture is as follows: Figure 2 As shown, it includes a satellite constellation network consisting of N local user devices, M low Earth orbit satellites, and a remote cloud data center; Consortium blockchain system models adapted to this computing power network, such as Figure 3 As shown, it includes: each satellite node as a lightweight node in the consortium blockchain, equipped with a secure data verification smart contract, generating lightweight new blocks through a consensus mechanism, and collaboratively maintaining lightweight blockchain information; the data center node as a full node in the consortium blockchain, recording full blocks containing complete computational offloading information, and maintaining full blockchain information.

[0026] The communication methods between user local devices, low Earth orbit satellites, and remote data centers include: user local devices can directly utilize the computing resources of low Earth orbit satellites; low Earth orbit satellites can act as computing nodes to process user computing tasks, or as relay nodes to forward user computing tasks to remote cloud data centers.

[0027] Each user can choose to offload computing tasks to a local terminal, a low-Earth orbit satellite, or a remote cloud data center. The user's offloading decision... Defined as: ; In the formula: This indicates the proportion of user computing tasks that are offloaded to the user's local device for computation. This indicates the proportion of user computing tasks that are offloaded to low Earth orbit satellites for computation. This indicates the proportion of user computing tasks that are offloaded to a remote cloud data center for computation. , , Subject to the following relational constraints: , ; The communication model between a user's local device and a low Earth orbit satellite is defined as follows: ; In the formula: This represents the data transmission rate of the ground-to-space communication between the nth user's local device and the mth low Earth orbit satellite; This indicates the transmission bandwidth between the user's local device and the low Earth orbit satellite; This indicates the number of users who choose to offload their computing tasks to the m-th low Earth orbit satellite. Users offloaded to the same low Earth orbit satellite share bandwidth resources equally. This represents the transfer rate of the local device of the nth user. This represents the channel gain between the local device of the nth user and the mth low Earth orbit satellite; Indicates channel noise power; The communication model between low Earth orbit satellites and remote cloud data centers is defined as follows: ; In the formula: This represents the data transmission rate between the m-th low Earth orbit satellite and the remote cloud data center; Indicates the congestion coefficient; This indicates the transmission bandwidth between low Earth orbit satellites and remote cloud data centers; This indicates the number of users who chose to offload their computing tasks to a remote cloud data center. This represents the transmission rate of the m-th low Earth orbit satellite; This represents the channel gain between the m-th low Earth orbit satellite and the remote cloud data center.

[0028] Characteristics of user computing tasks Defined as: ; In the formula: Indicates the user ID. Indicates the size of the data in the user's computation task, in bytes; Indicates completion The amount of computation required by the CPU in the data volume is measured in CPU cycles. Indicates completion The computational load required by the GPU in the data volume is expressed in the number of GPU floating-point operations. This indicates the privacy level of user n, with three levels: low, medium, and high. Based on the characteristics of user computing tasks and the transmission rates of different communication paths, the latency and energy consumption of user computing tasks processed by local devices, low Earth orbit satellites, and remote cloud data centers are calculated separately, including: when user computing tasks are offloaded to local devices, only the computation latency of user computing tasks is considered. The latency calculation formula is shown below: ; In the formula: This represents the latency required to complete the user's computing task when it is offloaded to the user's local device for computation. This represents the CPU computing power of the local device of the nth user. This represents the GPU computing power of the local device of the nth user. The formula for calculating energy consumption is shown below: ; In the formula: This indicates the energy consumption required for a user's computing task to be completed locally. This represents the capacitance parameters of the local device of the nth user. This represents the GPU power consumption of the nth user's local device; When user computing tasks are offloaded to low Earth orbit satellites, the propagation delay of the satellite-to-ground link, the data transmission delay, and the computing delay of the user computing tasks need to be considered. The formula for calculating the delay is as follows: ; In the formula: This indicates the latency required to complete the user computing task when it is offloaded to a low Earth orbit satellite for computing. This represents the data transmission delay from the nth user's local device to the mth low Earth orbit satellite; Indicates the computation latency of the user's computation task; This represents the CPU computing power allocated to the user's computing task by the m-th low Earth orbit satellite; d represents the GPU computing power allocated to the user's computing task by the m-th low Earth orbit satellite; d represents the distance between the low Earth orbit satellite and the user's local device. Represents the speed of light; This indicates the round-trip propagation delay from Earth to a low Earth orbit satellite; The formula for calculating energy consumption is shown below: In the formula: This represents the energy consumption required to complete the user computing task when the user computing task is offloaded to the m-th low Earth orbit satellite for computing. This represents the transmission power consumption of the local device of the nth user. This represents the capacitance parameter of the m-th low Earth orbit satellite; This represents the GPU power consumption of the m-th low Earth orbit satellite; When user computing tasks are offloaded to a remote cloud data center, the data transmission latency when data is relayed to the remote cloud data center via low Earth orbit satellites needs to be considered. The latency calculation formula is shown below: ; In the formula: This indicates the latency required to complete a user's computing task when it is offloaded to a remote cloud data center. This represents the data transmission latency when data is relayed to a remote cloud data center via the m-th low Earth orbit satellite. This indicates the round-trip propagation delay from Earth via a low Earth orbit satellite to a remote cloud data center; The formula for calculating energy consumption is shown below: ; In the formula: This indicates the energy consumption required to complete a user's computing task when it is offloaded to a remote cloud data center. This represents the transmission energy consumption of the m-th low Earth orbit satellite; The dynamic reputation assessment mechanism for nodes in a blockchain system includes: Define the performance evaluation function for low Earth orbit satellite nodes. for: In the formula, This represents the actual number of tasks completed by low Earth orbit satellite node m. This represents the number of tasks assigned to low Earth orbit satellite node m. The average time delay for low Earth orbit satellite node m to complete its mission. The average mission completion delay for all low Earth orbit satellite nodes. (Used) To indicate satellite reliability, use This indicates the efficiency of low Earth orbit satellite nodes.

[0029] A reputation decay mechanism is introduced to update the reputation value of low Earth orbit satellite node m. Update using the following formula: In the formula: ( >0) is the attenuation coefficient, The larger the size, the faster the reputation declines. It is an exponential decay term of reputation; It is a cumulative reputation item, reflecting the cumulative reputation performance of a node within the time interval [0, t], with recent contributions having a greater impact and long-term contributions having a smaller impact.

[0030] A lightweight, practical, space-to-ground Byzantine fault-tolerant consensus mechanism, including: (1) The number of consensus nodes is ,and (f is the number of Byzantine nodes); (2) Master node selection based on node reputation: Select the node with the highest current reputation value as the master node, and replace it in turn when it fails; (3) Dynamically adjust the consensus node set: Divide the node reputation values ​​into three categories: low, medium, and high. Determine the minimum reputation threshold based on the highest privacy level of the task to be processed. Delete nodes in the consensus node set whose reputation values ​​are lower than the threshold and add high-reputation nodes of the corresponding level. The consensus process for deleting and adding consensus nodes is as follows: Figure 4 As shown, the specific process is as follows: Initiation phase: The master node broadcasts information about nodes that do not meet the reputation threshold (as shown by node LN3 in the figure), triggering the dynamic adjustment process of consensus nodes.

[0031] Deletion Phase: After receiving the low-reputation node information (such as LN3 node) broadcast by the master node, each consensus node synchronously verifies whether the reputation value of the node is lower than the reputation threshold corresponding to the current task. If the verification confirms that it does not meet the requirements, each consensus node removes the low-reputation node from its local consensus node set and sends a confirmation message back to the master node. After the master node collects at least 2f + 1 valid deletion confirmations, the removal operation of the low-reputation node is completed.

[0032] Adding phase: Each consensus node initiates an add request to the node with the highest reputation value at the current level, thereby achieving a targeted invitation for the new node.

[0033] Add Response Phase: After receiving 2f + 1 valid add requests, the new node sends a response message to all consensus nodes.

[0034] Update phase: Each consensus node broadcasts the updated consensus node set information. When a consensus node receives 2f + 1 consistent update confirmation messages, the new node is officially included in the consensus node set, completing the entire dynamic adjustment process.

[0035] (4) The five-stage consensus mechanism process is as follows: Figure 5 As shown, the specific process is as follows: Pre-preparation phase: The master node assigns proposal numbers and view numbers, calculates and signs data hashes, and broadcasts pre-preparation messages; Preparation phase: After verifying the legality of the message, the consensus node broadcasts a preparation message and enters the preparation completed state after collecting at least 2f + 1 consistent messages; Submission Phase: Consensus nodes only send submission messages to the master node. The master node proceeds to the next phase after collecting at least 2f + 1 submission messages. Broadcast phase: The master node packages and signs the valid commit message and broadcasts the consensus certificate; Response Phase: After verifying the credentials, each node stores the information (lightweight nodes store the hash, full nodes store the complete data) and sends out a response. Data security verification of smart contracts includes: (1) Off-chain preprocessing: Legitimate users submit a unique identifier to the smart contract and its public key For the raw data and timestamp Calculate hash value and with private key right and Joint signature obtained ,Will Upload to computing power nodes; (2) On-chain core verification logic: Identity verification: Check Whether to register in the blockchain system; if not registered, return failure. Signature validity verification: Passed Obtain the public key ,verify If the statement is invalid, it returns a failure message. Time validity check: If the current block time is consistent with... If the difference exceeds the preset threshold (default 5 minutes), it is judged as a replay attack and a failure is returned; Data integrity verification: Local computation on computing nodes ,like Then it will return failure; If all validations pass, a data validation success message will be returned. Blockchain system overhead modeling includes: The latency of a blockchain system is mainly divided into two parts: consensus mechanism latency and data integrity verification latency. Assume the number of malicious nodes (Byzantine nodes) is... The cost of hash encryption is The cost of a single digital signature generation or verification cycle is [missing information]. The cost of a single message validity verification cycle is... During the cycle, the cost of storage execution by each computing node is... The cycle time is K GHz, and the processing speed of each node is K GHz.

[0036] (1) Delay in consensus mechanism Master node consensus cost Other consensus node costs It can be represented as: Additionally, assuming the latency of message transmission across the entire blockchain network or in a targeted manner is... The transmission latency required for a single consensus process is Therefore, the consensus mechanism has a latency. It can be represented as: (2) Data integrity verification latency: This part of the latency includes off-chain operation latency, data transmission latency, and on-chain operation latency. The user's data transmission latency has already been included in the latency calculation above, so this part will not be calculated again. Assume the costs of on-chain authentication and timestamp verification are respectively... and Off-chain operations mainly involve hash calculations and digital signature generation of data; therefore, off-chain operation latency is significant. It can be represented as: On-chain operations mainly involve identity verification, signature verification, timestamp verification, and hash calculations; therefore, on-chain operation latency is significant. It can be represented as: Therefore, data integrity verification delay It can be represented as: The total latency of the blockchain system for Assume the average power consumption of the blockchain system is The total energy consumption of the blockchain system for: Latency and energy consumption models, including: Total latency It can be defined as: Total energy consumption It can be defined as: To minimize the user computation offloading cost during the interaction between the user and the computing network environment, a global optimization objective function is established, including: User calculates uninstallation cost Defined as: In the formula, and These are the time delay coefficient and the energy consumption coefficient, respectively, and they satisfy... + = 1. Therefore, the global optimization objective function is to minimize the user's computational unloading cost at each time step. And subject to relevant conditional constraints, it is expressed as follows: Among them, the conditional expression and Ensure that the uninstallation ratio selected by the user meets the requirements, conditionally. and To ensure that the computational load allocated to user n by a low Earth orbit satellite does not exceed its maximum value, the condition is as follows: and Ensure that the allocation of computing resources for each user to low Earth orbit satellites does not exceed the maximum computing resources of the low Earth orbit satellites.

[0037] Solve for the optimal solution to the global optimization objective function and generate the optimal uninstall action for each user, including: modeling the upper-level task scheduling layer, with the goal of determining the optimal uninstall decision for each user's task. Define an upper-level state space to describe three dimensions: the global network state at time t, user task attributes, and the blockchain system state. This space comprehensively characterizes the global information required for upper-level decision-making and is formally represented as follows: Defined as: In the formula, Given the network state at time t, the remaining computing power of each satellite node is used. Load rate of each satellite node and the local computing power of ground terminals To represent, denoted as ; , which represents the core characteristics of all user tasks to be uninstalled at time t; The state of the blockchain system at time t is represented by the reputation value of each node. and consensus node set To represent, denoted as .

[0038] Define an upper-level action space to describe the unloading target decision of each user task at time t. The formal representation is as follows: Defined as: In the formula, , To unload the target satellite.

[0039] Define a higher-level reward function to describe the immediate reward obtained by taking the unloading action in a state, formally represented as follows: Defined as: The goal of modeling the lower-level node execution layer is to optimize resource allocation and execution strategies for satellite nodes based on the offloading decisions determined by the upper layer, thereby improving the task execution efficiency of a single node, while satisfying node resource constraints and feeding back the execution results to the upper layer. Define a lower-level state space to describe the local state and task execution details of satellite node m at time t, providing a precise basis for resource allocation decisions, and formally represent it as follows: Defined as: In the formula, Let be the remaining computing power of the current node at time t. This is a task queue, consisting of all tasks assigned to the current node, containing the core characteristics of each unloading task.

[0040] Define a lower-level action space to describe the resource allocation of satellite node m to each task in the task queue at time t, formally represented as follows: Defined as: In the formula: Where is the length of the task queue, and .

[0041] Define a lower-level reward function to describe the execution efficiency and energy consumption cost of satellite node m at time t, formally expressed as: Defined as: The reinforcement learning model is trained using the state space, action space, and reward function of the upper and lower layers to generate the optimal uninstall action for each user.

[0042] The reinforcement learning model is trained using the state space, action space, and reward function of the upper and lower layers. A hierarchical reinforcement learning algorithm based on DDPG is employed, and the algorithm architecture diagram is shown below. Figure 6 As shown, it specifically includes: Initialize the parameters of the upper-layer Actor network Critic network parameters and the target Actor network parameters Target Critic network parameters Initialize the upper-level experience replay pool. ; For each near-Earth orbit satellite, initialize the parameters of the lower-level Actor network. Critic network parameters and target Actor network parameters Target Critic network parameters Initialize the lower-level experience replay pool. .

[0043] Initialize the number of training epochs, the number of training steps per epoch, the delayed update step size K, the batch size B, and the soft update coefficient. The training steps are the total number of time steps executed in one round of training; At the start of each training round, reset the upper-level state. and each lower-level state This represents the initial state of the environment. At each time step, based on the current Actor network policy Select the initial uninstall action. And it operates on the computing network environment, distributing tasks to each satellite node, where... Represents the upper-level state based on time t. And the parameters of the upper-layer policy network Through upper-level strategies Output the corresponding upper-level action; For each near-Earth orbit satellite, obtain its local status. According to the current strategy Select lower-level action Execute actions and observe lower-level rewards. and the next state Storage experience Store in the experience replay pool And update the current state space, making .

[0044] Observe upper-level rewards and the next global state Storage experience Store in the experience replay pool And update the current state space, making Utilizing the experience replay mechanism to access the upper-level experience replay pool We randomly sample B data points and optimize the loss function of the upper-layer Critic network by minimizing the following loss function. Update the parameters of the upper-layer Critic network : In the formula: This indicates a buffer for replaying upper-level experiences. The sample in the expected value is taken. This represents the Q-value output by the current upper-layer Critic network. Indicates the upper-level time-differential objective. The calculation is as follows: In the formula, This represents the action at the next time step, i.e., time t+1, as output by the target's upper-level policy network. This represents the Q-value output of the target's upper-layer Critic network; Represents the reward at the current moment. Add the state of the next moment The discount values ​​of the target Q-values ​​corresponding to the output actions of the target policy and the target policy together constitute the temporal difference target at the current moment, which is used to guide the update of the upper-layer Critic network.

[0045] For each near-Earth orbit satellite, from the lower-level experience replay pool We sample B data points and optimize the loss function of the lower-level Critic network by minimizing the following loss function. Update the parameters of the lower-level Critic network In the formula: This indicates the lower-level experience replay buffer. The sample in the expected value is taken. This represents the Q-value output by the current lower-level Critic network. Indicates the lower-level temporal difference objective. The calculation is as follows: In the formula: This represents the action at the next time step, i.e., time t+1, as output by the target's lower-level policy network. This represents the Q-value output of the target's lower-level Critic network; Represents the reward at the current moment. Add the state of the next moment The discount values ​​of the target Q-values ​​corresponding to the output actions of the target policy together constitute the temporal difference target at the current moment, which is used to guide the update of the lower-level Critic network.

[0046] If the current training step is an integer multiple of the delayed update step size K, then the parameters of the upper-layer Actor network are updated using the following policy gradient method. Selecting the optimal unloading action maximizes the computational efficiency of the upper-layer Critic network. value: In the formula: The upper-layer policy performance function J represents the effect of the upper-layer Actor network parameters. The gradient; Indicates the distribution of upper-level states The state in Take the expected value; Indicates the action output by the current strategy. At that point, the Q function corresponds to the action. The gradient; Indicates the upper-level Actor strategy Its parameters The gradient; Simultaneously, for each near-Earth orbit satellite, the parameters of the lower-level Actor network are updated using the following strategy gradient method. Selecting the optimal unloading action maximizes the computation of the lower-level Critic network. value: In the formula: The performance function J of the lower-level policy is represented by the parameters of the lower-level Actor network. The gradient; Indicates the distribution of lower-level states The state in Take the expected value; Indicates the action output by the current strategy. At that point, the Q function corresponds to the action. The gradient; Indicates the strategy of the lower-level Actor Its parameters The gradient.

[0047] After each training round, the parameters of the target network in the upper and lower layers are updated based on the following soft update formula, so that they gradually approach the parameters of their respective current Actor network and Critic network, thus achieving smooth adjustment of the target network parameters. Once all training rounds are completed, a trained reinforcement learning model is obtained, and the optimal uninstall action for each user is output.

[0048] In summary, this invention constructs a three-tiered consortium blockchain system consisting of a terminal, satellite, and cloud, combined with a dynamic node reputation assessment model, a lightweight, practical satellite-ground Byzantine fault-tolerant consensus mechanism, and smart contracts with four-fold verification logic. This enables decentralized and trusted verification of the entire process of node access, data transmission, and offloading execution, effectively resisting security threats such as identity forgery, data tampering, and replay attacks, ensuring the integrity and traceability of data throughout its entire lifecycle, and addressing the lack of security assurance systems in existing technologies.

[0049] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0050] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.

Claims

1. A method for network offloading of 6G satellite-to-ground computing power based on lightweight blockchain and hierarchical reinforcement learning, characterized in that, Includes the following steps: S1. Construct a computing power network model based on user local devices, low Earth orbit satellites and remote cloud data centers, and simultaneously establish a consortium blockchain system model adapted to the computing power network. Then, define the communication method and communication model according to the computing power network model. S2. Based on the user's offloading decision, determine the offloading ratio of user computing tasks to be offloaded to the user's local device, low Earth orbit satellite, and remote cloud data center for computing, calculate the transmission rate of different communication paths, and calculate the latency and energy consumption of different communication paths for processing based on the characteristics of the user computing tasks and communication methods. S3. Design a dynamic reputation evaluation model for nodes in a blockchain system, a lightweight, practical Byzantine fault-tolerant consensus mechanism for satellite and ground systems, and a smart contract for data security verification. Model the overhead of the blockchain system and establish a latency model and an energy consumption model by combining the offloading ratio, latency, energy consumption and the overhead of the blockchain system. S4. With the goal of minimizing the user's computational unloading cost during the interaction between the user and the computing network environment, a global optimization objective function is established. The computational unloading problem is modeled as a hierarchical Markov decision process model, and the optimal solution of the global optimization objective function is solved to generate the optimal unloading action for each user.

2. The 6G satellite-to-ground computing power network offloading method based on lightweight blockchain and hierarchical reinforcement learning as described in claim 1, characterized in that: The computing network model includes N local user devices, a satellite constellation network consisting of M low Earth orbit satellites, and a remote cloud data center; The consortium blockchain system model includes: each satellite node as a lightweight node in the consortium blockchain, and data center nodes as full nodes in the consortium blockchain. Lightweight nodes are equipped with secure data verification smart contracts, generate lightweight new blocks through a consensus mechanism, and collaboratively maintain lightweight blockchain information; full nodes record full blocks containing complete computational offload information and maintain full blockchain information. Communication methods include: user local devices can directly utilize the computing resources of low Earth orbit satellites; low Earth orbit satellites can act as computing nodes to process user computing tasks, or as relay nodes to forward user computing tasks to remote cloud data centers.

3. The 6G satellite-to-ground computing power network offloading method based on lightweight blockchain and hierarchical reinforcement learning as described in claim 2, characterized in that: Characteristics of user computing tasks The expression is: ; in: Indicates the user ID. This indicates the amount of data involved in the user's computation task. Indicates completion The amount of computation required by the CPU in the data volume. Indicates completion The amount of computation required by GPUs in the data volume. This indicates the privacy level of user n, with three levels: low, medium, and high. When a user's computing task is offloaded to the user's local device, only the computing latency of the user's computing task is considered. When user computing tasks are offloaded to low Earth orbit satellites, the propagation delay of the satellite-to-ground link, the data transmission delay, and the computing delay of the user computing tasks need to be considered. When user computing tasks are offloaded to a remote cloud data center, the data transmission latency of sending data to the remote cloud data center via low Earth orbit satellite relay needs to be considered.

4. The 6G satellite-to-ground computing power network offloading method based on lightweight blockchain and hierarchical reinforcement learning as described in claim 3, characterized in that, The node dynamic reputation assessment mechanism includes: performance evaluation functions for low Earth orbit satellite nodes. And the reputation decay mechanism updates the reputation value of low Earth orbit satellite node m. S3's lightweight, practical, Byzantine fault-tolerant consensus mechanism for space-to-ground applications includes the following steps: S3-1-1. Select the node with the highest current reputation value as the master node, and replace it sequentially in case of failure. The number of consensus nodes is [number missing]. ,and f is the number of Byzantine nodes; S3-1-2. Divide the node reputation value into three categories: low, medium and high. Determine the minimum reputation threshold based on the highest privacy level of the task to be processed. Delete nodes in the consensus node set whose reputation value is lower than the threshold and add high reputation nodes of the corresponding level. S3-1-3, The consensus process for deleting and adding consensus nodes is as follows: Initiation phase: The master node broadcasts information about nodes that do not meet the reputation threshold, triggering a dynamic adjustment process for consensus nodes; Deletion phase: After receiving the low-reputation node information broadcast by the master node, each consensus node synchronously verifies whether the reputation value of the node is lower than the reputation threshold corresponding to the current task. If the verification confirms that it does not meet the requirements, each consensus node removes the low-reputation node from its local consensus node set and sends a confirmation message back to the master node. After the master node collects at least 2f+1 valid deletion confirmations, the removal operation of the low-reputation node is completed. Adding phase: Each consensus node initiates an add request to the node with the highest reputation value at the current level, realizing a targeted invitation for the new node; Adding Response Phase: After receiving 2f+1 valid add requests, the new node sends a response message to all consensus nodes; Update phase: Each consensus node broadcasts the updated consensus node set information. When a consensus node receives 2f+1 consistent update confirmation messages, the new node is officially included in the consensus node set, completing the entire dynamic adjustment process.

5. The 6G satellite-to-ground computing power network offloading method based on lightweight blockchain and hierarchical reinforcement learning as described in claim 4, characterized in that, The consensus mechanisms for deleting and adding consensus nodes include: Pre-preparation phase: The master node assigns proposal numbers and view numbers, calculates and signs data hashes, and broadcasts pre-preparation messages; Preparation phase: After the consensus node verifies the legality of the message, it broadcasts a preparation message and enters the preparation completed state after collecting at least 2f+1 consistent messages; Submission Phase: Consensus nodes only send submission messages to the master node. The master node collects at least 2f+1 submission messages and then proceeds to the next phase. Broadcast phase: The master node packages and signs the valid commit message and broadcasts the consensus certificate; Response phase: After verifying the credentials, each node stores the information and sends out a response. Lightweight nodes store hashes, while full nodes store complete data.

6. The 6G satellite-to-ground computing power network offloading method based on lightweight blockchain and hierarchical reinforcement learning as described in claim 5, characterized in that, S3's data security verification smart contract includes the following steps: S3-2-1 Off-chain preprocessing: Legitimate users submit a unique identifier to the smart contract. and its public key For the raw data and timestamp Calculate hash value and with private key right and Joint signature obtained ,Will Upload to computing power nodes; S3-2-2, On-chain core verification logic: Identity legitimacy verification: Check Whether to register in the blockchain system; if not registered, return failure. S3-2-3, Signature Validity Verification: Passed Obtain the public key ,verify If the statement is invalid, it returns a failure message. S3-2-4, Time Validity Verification: If the current block time is inconsistent with... If the difference exceeds the preset threshold, it is judged as a replay attack and a failure is returned. The preset threshold is five minutes by default. S3-2-5, Data Integrity Verification: Local Computation by Computing Nodes ,like If all validations pass, it returns a failure message; otherwise, it returns a successful data validation message.

7. The 6G satellite-to-ground computing power network offloading method based on lightweight blockchain and hierarchical reinforcement learning as described in claim 6, characterized in that, The latency of a blockchain system includes consensus mechanism latency and data integrity verification latency. When the number of Byzantine nodes is... The cost of hash encryption is The cost of a single digital signature generation or verification cycle is [missing information]. The cost of a single message validity verification cycle is... During the cycle, the cost of storage execution by each computing node is... The cycle time is 100 GHz, and the processing speed of each node is 100 GHz. When the latency of message transmission across the entire blockchain network or in a targeted manner is The transmission latency required for a single consensus process is consensus mechanism latency The expression is: ; Data integrity verification latency includes off-chain operation latency. Data transmission latency and on-chain operation latency The costs of on-chain authentication and timestamp verification are respectively and Off-chain operations perform hash calculations and generate digital signatures for data, while on-chain operations perform identity verification, signature verification, timestamp verification, and hash calculations. Data integrity verification latency The expression is: ; Total latency of blockchain system The expression is: ; When the average power consumption of the blockchain system is The total energy consumption of the blockchain system The expression is: ; Total latency The expression is: ; Total energy consumption The expression is: 。 8. The 6G satellite-to-ground computing power network offloading method based on lightweight blockchain and hierarchical reinforcement learning as described in claim 7, characterized in that, Uninstallation decision The expression is: , , ; in, This indicates the proportion of user computing tasks that are offloaded to the user's local device for computation. This indicates the proportion of user computing tasks that are offloaded to low Earth orbit satellites for computation. This indicates the proportion of user computing tasks that are offloaded to a remote cloud data center for computation. The communication model between the user's local device and the low Earth orbit satellite is as follows: ; The communication model between low Earth orbit satellites and remote cloud data centers is as follows: ; User calculates uninstallation cost The expression is: ; in, and These are the time delay coefficient and the energy consumption coefficient, respectively, and they satisfy... + =1.

9. The 6G satellite-to-ground computing power network offloading method based on lightweight blockchain and hierarchical reinforcement learning as described in claim 8, characterized in that, S4 includes the following steps: S4-1. Model the upper-level task scheduling layer, determine the optimal unloading decision for each user task, and define the upper-level state space. It is used to describe three dimensions at time t: the global network state, user task attributes, and the state of the blockchain system, comprehensively depicting the global information required for upper-level decision-making. S4-2, Define the upper-level motion space , used to describe the unloading target decision for each user task at time t; S4-3. Define the upper-level reward function. , used to describe the immediate reward obtained by taking the unloading action in the stated state; S4-4. Model the lower-level node execution layer. Based on the offloading decisions determined by the upper layer, optimize the resource allocation and execution strategies for satellite nodes to improve the single-node task execution efficiency, while satisfying node resource constraints and feeding back the execution results to the upper layer; define the lower-level state space. This is used to describe the local state and task execution details of satellite node m at time t, providing a precise basis for resource allocation decisions. S4-5, Define the lower-level motion space , used to describe the resource allocation of satellite node m to each task in the task queue at time t; S4-6. Define the lower-level reward function. , used to describe the execution efficiency and energy consumption cost of satellite node m at time t.

10. The 6G satellite-to-ground computing power network offloading method based on lightweight blockchain and hierarchical reinforcement learning as described in claim 9, characterized in that, The reinforcement learning model is trained using the state space, action space, and reward function of the upper and lower layers to generate the optimal uninstall action for each user, including steps a~f: a. Initialize the parameters of the upper-layer Actor network Critic network parameters and the target Actor network parameters Target Critic network parameters Initialize the upper-level experience replay pool For each near-Earth orbit satellite, initialize the parameters of the lower-level Actor network. Critic network parameters Target Actor Network Parameters and target Critic network parameters Initialize the lower-level experience replay pool Initialize the number of training epochs, the number of training steps per epoch, the delayed update step size K, the batch size B, and the soft update coefficient. The training steps are the total number of time steps executed in one round of training; b. At the start of each training round, reset the upper-level state. and each lower-level state The initial state of the environment; at each time step, based on the current policy of the Actor network. Select the initial uninstall action. And it operates on the computing network environment, distributing tasks to each satellite node, where... Represents the upper-level state based on time t. And the parameters of the upper-layer policy network Through upper-level strategies Output the corresponding upper-level action; c. For each near-Earth orbit satellite, obtain its local status. According to the current strategy Select lower-level action Execute actions and observe lower-level rewards. and the next state Storage experience Store in the experience replay pool And update the current state space, making ; Observe the upper-level rewards and the next global state Storage experience Store in the experience replay pool And update the current state space, making ; d. Utilize the experience replay mechanism to retrieve data from the upper-level experience replay pool. We randomly sample B data points and optimize the loss function of the upper-layer Critic network by minimizing the loss function. Update the parameters of the upper-layer Critic network ; e. For each near-Earth orbit satellite, from the lower-level experience replay pool We sample B data points and optimize the loss function of the lower-level Critic network by minimizing the loss function. Update the parameters of the lower-level Critic network ; f. If the current training step number is an integer multiple of the delayed update step size K, then update the parameters of the upper-layer Actor network using the policy gradient method. Selecting the optimal unloading action maximizes the computational efficiency of the upper-layer Critic network. value; After each training round, the parameters of the target network in the upper and lower layers are updated based on the soft update formula, so that they gradually approach the parameters of their respective current Actor and Critic networks, achieving smooth adjustment of the target network parameters. The soft update formula is as follows: Once all training rounds are completed, a well-trained reinforcement learning model is obtained.