Teaching resource sharing decision optimization method based on hierarchical federal incentive mechanism

By employing a hierarchical federated incentive mechanism and dynamic topology reconfiguration, the problems of low resource storage and routing efficiency, imperfect incentive mechanisms, and lack of dynamic adaptability in the network topology of the teaching resource sharing network are solved. This achieves privacy protection and efficient circulation of teaching resources, and optimizes the fairness of resource allocation and network stability.

CN121961119APending Publication Date: 2026-05-01ZHONGBEI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHONGBEI UNIV
Filing Date
2026-01-19
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing teaching resource sharing networks suffer from low resource storage and routing efficiency, imperfect incentive mechanisms, and a lack of dynamic adaptability in network topology, leading to risks of privacy leaks, resource waste, load imbalance, and overall decline in efficiency.

Method used

A hierarchical federated incentive mechanism is adopted, which determines the storage level and routing strategy by calculating resource entropy, establishes a multi-dimensional reputation weighted model, executes resource scheduling decisions based on dynamic Standberg game, and performs topology reconstruction based on game equilibrium drift rate to achieve dynamic adaptive management of teaching resources.

Benefits of technology

It achieves privacy protection and efficient circulation of teaching resources, optimizes the fairness of resource allocation, ensures that the network maintains a stable Nash equilibrium in a dynamic environment, and improves overall efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121961119A_ABST
    Figure CN121961119A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of Internet education data processing and network communication, and discloses a teaching resource sharing decision optimization method based on a hierarchical federated incentive mechanism, and the method specifically comprises the steps: constructing a hierarchical heterogeneous federated topology network based on resource entropy, determining the storage hierarchy and routing strategy of teaching resources through the calculation result of the resource entropy, and carrying out the decision optimization of the teaching resources. Establishing a multi-dimensional reputation weighted right quantification model, and calculating comprehensive reputation weights of participating nodes; and executing a resource scheduling decision based on the dynamic Stackelberg game, solving a Nash equilibrium solution by using a distributed iterative algorithm, determining an optimal price vector and a resource allocation matrix, monitoring a game equilibrium drift rate to implement dynamic topology reconstruction, and triggering topology splitting or fusion operation according to the Nash equilibrium drift rate and a global utility variance. According to the method, privacy and efficiency are balanced through resource entropy hierarchical storage, pricing is optimized by using a reputation weighted game, and self-adaptive stability of a network architecture is guaranteed by means of topology reconstruction based on a drift rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of internet education data processing and network communication technology, specifically to a method for optimizing teaching resource sharing decisions based on a hierarchical federated incentive mechanism. Background Technology

[0002] With the accelerated advancement of educational informatization, massive amounts of digital teaching resources are generated between various educational institutions and terminal devices. Achieving efficient flow and sharing of heterogeneous teaching resources has become the key to improving educational equity and resource utilization. However, current teaching resource sharing technical solutions generally suffer from rigid resource management architecture, a single incentive mechanism, and a lack of network topology adaptability.

[0003] Existing resource-sharing architectures employ centralized storage or a completely flat peer-to-peer network model. The centralized model is prone to performance bottlenecks when facing massive concurrent requests. It ignores the differences in the sensitivity of teaching resources and stores privacy-related exam data and publicly available courseware materials in the same way. This one-size-fits-all approach increases the risk of privacy leaks and leads to unnecessary waste of cross-domain transmission bandwidth. While the flat model alleviates the single point of failure problem, it is difficult to achieve fine-grained routing and management of resources with different attributes without hierarchical division.

[0004] In terms of incentive mechanisms, existing technical solutions adopt static pricing or simple points exchange strategies. The fixed value assessment methods cannot reflect the dynamic changes in market supply and demand in real time, and lack long-term tracking of the historical behavior of participating nodes and resource quality. Due to the lack of a reputation-based multi-dimensional evaluation system, the network environment is prone to free-riding behavior, that is, some nodes only download resources without contributing, or upload low-quality resources to cheat for points, which leads to the loss of the rights and interests of high-quality resource providers, reduces the willingness of resource providers to continue sharing, and causes a vicious cycle in the sharing ecosystem.

[0005] Furthermore, the existing shared network topology remains fixed after initialization. Actual teaching resource interaction exhibits significant time-varying and tidal effects. When the interaction intensity of local node clusters fluctuates drastically or the game state drifts, the fixed network topology cannot perceive and adapt accordingly. The architecture, lacking dynamic reconstruction capabilities, struggles to maintain the stability of the Nash equilibrium, leading to load imbalance, decreased resource indexing efficiency, and overall network performance degradation over long-term operation. This fails to meet the demands of large-scale, highly dynamic teaching resource sharing. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides a decision optimization method for sharing teaching resources based on a hierarchical federated incentive mechanism, which solves the problems of low resource storage and routing efficiency, imperfect incentive mechanism, and lack of dynamic adaptability of network topology in existing teaching resource sharing networks.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a teaching resource sharing decision optimization method based on a hierarchical federated incentive mechanism, comprising the following steps: S1: Construct a hierarchical heterogeneous federated topology network based on resource entropy. The network architecture is logically divided into a local layer, a federation layer, and a global layer. The storage level and routing strategy of the teaching resources to be shared are determined by calculating the resource entropy of the teaching resources to be shared. The resource entropy is calculated as follows: Identify the set of sensitive features contained in the teaching resources to be shared, obtain the probability that the content of the teaching resources to be shared belongs to each type of sensitivity feature, sum the product of the probability of each type of sensitivity feature with the base-2 logarithm of the probability, and take the negative value of the summation result to obtain the resource entropy.

[0008] Based on the calculated resource entropy, the storage level and routing strategy are determined by comparing the resource entropy with a preset classification threshold. When the resource entropy is greater than the first threshold, the teaching resource to be shared is determined to be a private resource, and the resource metadata is routed to the local layer storage. When the resource entropy is between the second threshold and the first threshold, the teaching resource to be shared is determined to be a regional shared resource, and the resource metadata is routed to the alliance layer storage. When the resource entropy is less than the second threshold, the teaching resource to be shared is determined to be a public resource, and the resource metadata is routed to the global layer storage.

[0009] S2: Establish a multi-dimensional reputation-weighted equity quantification model, maintain a multi-dimensional reputation vector for each participating node in the network, and calculate the comprehensive reputation weight of each participating node based on the multi-dimensional reputation vector. The multi-dimensional reputation vector includes three components: quality reputation, contribution depth, and time-related decay. The time-related decay is updated by multiplying the time-related reputation value of the previous time step by the decay coefficient, and adding the newly added contribution value of the current time step to obtain the time-related decay of the current time step. The decay coefficient is an exponential function value with the natural constant as the base and the negative value of the product of the preset time decay factor and the time interval as the exponent.

[0010] The method for calculating the overall reputation weight of each participating node is to obtain a preset quality weight coefficient, contribution weight coefficient, and timeliness weight coefficient, and then multiply the quality reputation, contribution depth, and timeliness decay by the corresponding weight coefficients respectively, and add the three product results to obtain the overall reputation weight of the participating node.

[0011] S3: Execute resource scheduling decisions based on a dynamic Standberg game, constructing the resource sharing process as a multi-leader, multi-follower game model. Use a distributed iterative algorithm to solve for the Nash equilibrium, obtaining the optimal price vector and resource allocation matrix. In this multi-leader, multi-follower game model, resource providers are defined as leaders. The leader sets the resource incentive price based on its own bandwidth cost and the comprehensive reputation weight to maximize the first utility function. The first utility function consists of resource transaction revenue minus bandwidth congestion cost, plus an incentive subsidy based on the comprehensive reputation weight. Resource requesters are defined as followers. The follower sets the resource demand based on the resource incentive price and network latency cost to maximize the second utility function. The second utility function consists of resource usage value minus payment cost, minus transmission latency loss. The process of solving for the Nash equilibrium includes: The follower calculates the optimal resource demand that maximizes the second utility function based on the received resource incentive unit price and feeds it back to the leader. The leader determines the total load based on the optimal resource demand received from all followers and updates the resource incentive unit price using a gradient ascent algorithm to approach the maximum value of the first utility function. The above process is iteratively executed until the change in the resource incentive unit price is less than a preset error threshold.

[0012] S4: Dynamic topology reconstruction based on game equilibrium drift rate. During the game, the Nash equilibrium drift rate and global utility variance are monitored in real time, and topology splitting or merging operations are triggered according to the monitoring results. Resource scheduling decisions are re-executed under the reconstructed network topology. The value of the Nash equilibrium drift rate is the average of the Euclidean distance between the optimal price vectors of adjacent time steps, and the value of the global utility variance is the variance of the total global utility value within the time sliding window.

[0013] The conditions and specific steps for triggering the topology split operation are as follows: when the monitored Nash equilibrium drift rate is greater than the preset drift threshold, the current alliance layer structure is determined to be unstable. An interaction frequency matrix of nodes in the current alliance is constructed. Based on the interaction frequency matrix, a graph partitioning algorithm is executed to divide the node set of the current alliance layer into two disjoint sub-alliance sets, so that the interaction weight between the sub-alliance sets is minimized and the direct game channel between nodes belonging to different sub-alliance sets is blocked.

[0014] The conditions and specific steps for triggering the topology fusion operation are as follows: when the monitored Nash equilibrium drift rate is less than a preset stability threshold and the average global utility is lower than the maintenance cost threshold, it is determined that the current alliance layer resource liquidity is insufficient. The resource feature centroid vector of the current alliance is calculated. The resource feature centroid vector is composed of the arithmetic mean of the feature vectors of the stored resources of all nodes in the alliance. The feature distance between the current alliance and the adjacent alliances in the network is calculated. The adjacent alliance with the smallest feature distance is selected as the fusion object. A communication bridge connecting the current alliance and the fusion object is established. The distributed ledger data is merged and the resource index is synchronized.

[0015] This invention provides a method for optimizing decision-making in teaching resource sharing based on a hierarchical federated incentive mechanism. It has the following beneficial effects: 1. This invention automatically routes teaching resources to the local layer, alliance layer, or global layer by calculating resource entropy, thus constructing a hierarchical and heterogeneous storage architecture. This content-sensitivity-based hierarchical strategy can accurately identify the private or public attributes of data. While ensuring that highly sensitive teaching data is only visible locally or within a limited scope to protect privacy and security, it reduces the access threshold and cross-domain transmission latency of low-sensitivity public resources, achieving a balance between data privacy protection and resource circulation efficiency.

[0016] 2. This invention establishes a multi-dimensional reputation mechanism that includes time-based decay and combines it with a multi-leader, multi-follower game model for resource scheduling. It directly integrates the comprehensive reputation weight of nodes into the utility function construction process, so that the final resource allocation result and incentive unit price depend not only on the current supply and demand relationship, but also on the historical contribution quality and persistence depth of the nodes. This can suppress malicious nodes and free-riding behavior in the network, incentivize participants to continuously provide high-quality teaching resources, thereby optimizing the resource pricing system and allocation fairness in a distributed environment.

[0017] 3. This invention utilizes the game equilibrium drift rate and global utility variance to monitor the network operation status in real time and trigger topology splitting or merging operations accordingly. The dynamic reconstruction mechanism endows the overall architecture with the ability to adapt and evolve. When the game structure is unstable, it can reduce the interaction complexity by splitting, or introduce external resources by merging when resource liquidity is insufficient. This ensures that the federated network always maintains a stable Nash equilibrium state in a dynamically changing environment, guaranteeing long-term operational stability and overall benefits. Attached Figure Description

[0018] Figure 1 This is the overall flowchart of the present invention; Figure 2 This is a schematic diagram of the hierarchical heterogeneous federated network architecture and resource routing of the present invention; Figure 3 This is a schematic diagram of the multidimensional reputation-weighted equity quantification method of the present invention; Figure 4 This is the resource scheduling decision diagram for the dynamic Stankberg game of the present invention; Figure 5 This is a dynamic topological reconstruction diagram of the game equilibrium drift rate of the present invention. Detailed Implementation

[0019] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] See attached document Figure 1 This invention provides a teaching resource sharing decision optimization method based on a hierarchical federated incentive mechanism, comprising the following steps: S1: Constructing a hierarchical heterogeneous federated topology network based on resource entropy. The first step in constructing this network is to initialize the network architecture. Logically, the network architecture is divided into three layers: the local layer, the consortium layer, and the global layer. The local layer consists of edge computing nodes deployed within each teaching unit, configured to store frequently accessed and latency-sensitive private resources. The consortium layer consists of multiple geographically adjacent or subject-specific local layer nodes connected via a consortium blockchain protocol, configured to maintain a resource-sharing ledger and consensus state within the region. The global layer consists of public cloud nodes, configured to store resource index summary information across the entire network.

[0021] Based on the determined network architecture, the storage and routing strategies for teaching resources are determined by calculating resource entropy. For teaching resources to be shared... Identify the set of sensitive features contained therein, denoted as... ,in This indicates that the resource content belongs to the first... The probability of class-sensitivity features. Calculate resource entropy using the following formula. : ; The calculated resource entropy Compare with the preset classification threshold: when When the value exceeds the first threshold, the resource metadata is routed to local storage; when... When the value is between the second threshold and the first threshold, the resource metadata is routed to the federation layer storage. When the value is less than the second threshold, the resource metadata is routed to the global layer storage.

[0022] S2: Establish a multidimensional reputation-weighted equity quantification model. In the step of establishing the multidimensional reputation-weighted equity quantification model, maintain a multidimensional reputation vector for each participating node in the network. Multidimensional reputation vector It includes three components: quality and reputation. Depth of contribution and time decay Quality and reputation Contribution depth is calculated based on feedback scores from resource recipients after the transaction is completed, and anomaly scores are removed using a weighted algorithm. Based on nodes The depth at which published resources are referenced is calculated. (Time decay is also considered.) Introducing a time decay factor Update according to the following formula: ; in: For time intervals; The new contribution value for the current time step.

[0023] Further calculate the nodes based on the following formula. Comprehensive reputation weight : ; compute nodes Comprehensive reputation weight ,in: , These are the weight coefficients for each dimension, and their sum is 1.

[0024] S3: Execute resource scheduling decisions based on a dynamic Standborg game. In this step, the resource sharing process is constructed as a multi-leader, multi-follower game model. The resource provider, as the leader, sets the resource incentive price based on its bandwidth cost and reputation weight. The goal is to maximize the first utility function. The first utility function includes resource transaction revenue, bandwidth congestion costs, and reputation-based weights. Incentive subsidies. The resource requester, as a follower, receives subsidies based on the resource incentive unit price. Determine resource requirements based on network latency costs. The goal is to maximize the second utility function, which includes the value of resource use, payment costs, and transmission delay losses. A distributed iterative algorithm is used to solve the Nash equilibrium of a multi-leader, multi-follower game model, yielding the optimal price vector and resource allocation matrix for the current time window.

[0025] S4: Dynamic topology reconstruction based on game equilibrium drift rate. In the dynamic topology reconstruction step based on game equilibrium drift rate, the Nash equilibrium drift rate is monitored in real time during the game. and global utility variance Nash equilibrium drift rate This represents the magnitude of change in the optimal price vector within a continuous time window. Global utility variance. This indicates the fluctuation of the sum of the utility of all nodes within the alliance.

[0026] When the monitored Nash equilibrium drift rate If the drift exceeds a preset threshold and the duration exceeds a preset time limit, the current alliance layer structure is deemed unstable, triggering a topology split operation. The topology split operation uses a spectral clustering algorithm to divide the current alliance layer's node set into two independent sub-alliance sets based on the historical interaction frequency and utility similarity between nodes, and then re-initializes the consensus parameters of each sub-alliance.

[0027] When the monitored Nash equilibrium drift rate When the liquidity of resources in the current consortium layer falls below a preset stability threshold and the average global utility is lower than the maintenance cost threshold, it is determined that the current consortium layer resource liquidity is insufficient, triggering a topology fusion operation. The topology fusion operation identifies the logically closest adjacent consortium layer nodes or isolated nodes, merges the reputation ledger through a cross-chain protocol, generates a new consortium layer structure, and updates the logical distance table between nodes. Through the aforementioned topology splitting or fusion operation, the resource scheduling decision steps are re-executed under the new network topology.

[0028] See attached document Figure 2 The teaching resource sharing decision optimization method based on the hierarchical federation incentive mechanism runs in a network environment composed of multiple distributed computing nodes, and logically divides the network environment into a local edge computing layer, a regional alliance consensus layer, and a global index layer.

[0029] The local edge computing layer consists of edge servers distributed across various independent teaching units. These edge servers are equipped with storage and computing modules to handle high-frequency local access requests and store private teaching resources with resource entropy greater than a first threshold. The regional consortium consensus layer comprises multiple edge server nodes connected via a communication network. These nodes run consortium blockchain protocols to maintain the ledger state of shared resources and store regional teaching resources with resource entropy between the first and second thresholds. The global index layer consists of public cloud servers or public blockchain nodes and stores public resource index summaries with resource entropy less than the second threshold; it does not store resource entity files.

[0030] The nodes at each level are connected via network communication interfaces, and the decision-making optimization for sharing teaching resources is performed according to the following process: First, the resource feature extraction and routing distribution process is executed. When a resource provider terminal uploads teaching resources to be shared, the computing nodes extract a set of sensitive features of the teaching resources to be shared, and calculate the resource entropy based on the probability distribution. According to the calculated resource entropy value, the resource metadata and entity files are transmitted to the corresponding storage locations in the local edge computing layer, regional alliance consensus layer, or global index layer. The resource feature extraction and routing distribution process establishes the initial distribution state of resources in the physical network.

[0031] Secondly, a multi-dimensional reputation status update process is executed. Historical interaction data of each node in the network is recorded, including the recipient's score after a resource transaction, the number of times the resource is cited, and the timestamp of the contribution. Based on this data, the comprehensive reputation weight of each node is periodically updated according to preset weight coefficients. This comprehensive reputation weight, as a key parameter in subsequent resource scheduling decisions, is stored in the distributed ledger of the regional alliance consensus layer, ensuring the immutability and consistency of the data.

[0032] Subsequently, a game theory-based resource scheduling decision-making process is executed. This process is automatically triggered between nodes in the regional alliance consensus layer via smart contracts. The current resource-sharing environment is instantiated as a multi-leader, multi-follower Standborg game model. In this model, resource providers act as leaders, calculating and broadcasting resource incentive prices based on their bandwidth load and overall reputation weight. Resource requesters, acting as followers, receive the incentive prices and, considering their own resource value assessment and network latency costs, calculate and report their optimal resource demand. Distributed alternating direction multipliers are used to conduct multiple rounds of information exchange and iterative calculations between leaders and followers until the strategies of all participants no longer change significantly, thus obtaining the Nash equilibrium solution for the current time window—the determined transaction price and allocation scheme.

[0033] Simultaneously, a network topology status monitoring and reconstruction process is executed. During the aforementioned game iteration process, the optimal price vector and global utility value for each round are collected in real time. The Nash equilibrium drift rate is calculated to characterize the stability of the game solution, and the global utility variance is calculated to characterize the equilibrium of resource allocation.

[0034] When monitoring data shows that the Nash equilibrium drift rate exceeds a preset drift threshold, it is determined that the node structure of the current regional alliance consensus layer is hindering game convergence, and a topology split command is executed. Based on the historical interaction frequency matrix between nodes, the node set in the original regional alliance consensus layer is divided into two logically independent sub-alliances, restricting direct game interactions between high-conflict nodes.

[0035] When monitoring data shows that the Nash equilibrium drift rate is lower than a preset stability threshold and the average global utility is lower than a preset cost threshold, it is determined that the current regional alliance consensus layer has insufficient resource liquidity, and a topology fusion instruction is executed. The nearest neighboring alliance or isolated node in the logical network is identified, a cross-chain communication channel is established, the reputation ledger and resource index are merged, and a new regional alliance consensus layer structure is formed.

[0036] By cyclically executing the above process, the network topology can be dynamically evolved according to the equilibrium state of the game, ensuring efficient matching and transmission of teaching resources between different levels and nodes.

[0037] See attached document Figure 2 The constructed network architecture is logically divided into a local layer, a federation layer, and a global layer, specifically including: The local layer (L1) consists of a set of edge computing nodes, denoted as set. Each edge computing node is deployed within an independent physical management domain, which corresponds to the local area network environment of a single school or teaching unit. Each edge computing node is configured with local storage and a local processor. The local storage is used to store private teaching resource entity files generated by the management domain to which the node belongs that are not publicly available to the external network. The local processor is used to process resource access requests from terminals within this management domain and to perform index updates for local resources.

[0038] The alliance layer (L2) consists of a set of alliance nodes, denoted as set. The consortium nodes are connected via wide area network (WAN) communication links, forming a peer-to-peer (P2P) topology. Consortium nodes are selected from edge computing nodes in the local layer that possess public IP addresses or pending gateway permissions. All nodes in the consortium layer run a permissioned blockchain consensus protocol and jointly maintain a distributed shared ledger. This distributed shared ledger records metadata of regional shared resources, resource transaction records, and node reputation status data. The consortium layer defines the trust boundaries for resource sharing, allowing only authenticated nodes to join.

[0039] The global layer (L3) consists of a set of public service nodes, denoted as a set. Public service nodes are deployed on public cloud infrastructure or public blockchain networks. The global layer communicates with consortium layer nodes via standard internet protocols. The global layer's storage space only stores the digital fingerprints (hash values) of all resources on the network and global routing indexes; it does not store the physical files of teaching resources. The global layer is used to respond to cross-consortium resource discovery requests; global layer nodes redirect requests to consortium layer or local layer nodes that own the resource entity.

[0040] There are clear data flow and control relationships between the above-mentioned layers. Local layer nodes connect to their respective consortium layer nodes through an internal gateway and upload index digests of local resources. Consortium layer nodes periodically synchronize the root hash of the resource state within their consortium to the global layer nodes to ensure the consistency of the index data stored in the global layer. The network topology supports dynamic reconfiguration based on control commands, and the consortium layer node set... It can split into multiple non-overlapping subsets according to instructions, or merge with neighboring sets of nodes.

[0041] See attached document Figure 2 First, receive the teaching resources to be processed. and the teaching resources to be processed The metadata and content are scanned to extract the teaching resources to be processed. The set of sensitive features included. The set of sensitive features is denoted as... ,in This represents the total number of predefined sensitive feature categories, including copyright restriction features, privacy data features, and large file features. For each sensitive feature category... Calculate the teaching resources to be processed Belongs to the sensitive feature classification probability Probability It is a normalized value obtained by analyzing the tag matching degree in resource metadata and the keyword density in content, satisfying... .

[0042] Based on the extracted probability distribution, the information entropy formula is used to calculate the teaching resources to be processed. resource entropy Resource entropy Used to quantify teaching resources to be processed The formula for calculating the uncertainty and sensitivity in storage location selection is as follows: ; in: Indicates teaching resources The resource entropy value; It represents a logarithmic operation with base 2.

[0043] The calculated resource entropy Input to route mapping function According to resource entropy With preset high sensitivity threshold and low sensitivity threshold Based on the size relationship, determine the target storage level of the teaching resource metadata. Routing mapping function The logical definition is as follows: ; in: Represents the set of local layer nodes; Represents the set of nodes in the alliance layer; This represents the set of nodes in the global layer.

[0044] when At that time, determine the teaching resources to be processed. For highly sensitive private resources, execute local storage instructions to store the teaching resources to be processed. The entity files are encrypted and stored in the private storage area of ​​the local edge computing node, and the teaching resources to be processed are only recorded in the index table of the local layer. Metadata is not broadcast to the federation layer or global layer of the teaching resources to be processed. The evidence storage information.

[0045] when At that time, determine the teaching resources to be processed. To facilitate regional resource sharing, execute consortium storage instructions. This will transfer the teaching resources to be processed. The entity files are uploaded to the distributed file system maintained by the alliance layer nodes, and the teaching resources to be processed are... The hash fingerprint and metadata are written into the shared ledger of the consortium blockchain, enabling the processing of teaching resources. exist Nodes within the set are visible and accessible to each other.

[0046] when At that time, determine the teaching resources to be processed. For public resources, execute global indexing commands. Preserve the storage state of the resource entity file in the source node, and simultaneously generate the teaching resources to be processed. The system generates a global unified resource identifier and a digital fingerprint, and publishes the global unified resource identifier and digital fingerprint to the public index database of the global layer nodes for inspection and discovery by all nodes in the network.

[0047] See attached document Figure 3 A periodic reputation assessment process is executed to quantify the behavioral characteristics of each participating node in the network, providing a basis for evaluating the performance of each node in the network. Maintain a dynamically updated reputation vector Reputation Vector In terms of time It consists of three independently calculated components: quality reputation component Contribution depth component and time-degradation components .

[0048] First calculate the nodes. Quality, reputation, weight Quality, reputation, and weight Reflecting nodes As a resource provider, you receive feedback and evaluations from historical transactions. Retrieve data from the distributed ledger and nodes. Extract the set of users who have requested and used node resources from all associated historical transaction records. For sets Every user in Get its node rating Introducing users Self-reputation weight As a weighting factor, the quality reputation component is calculated using a weighted average algorithm, and the calculation formula is as follows: ; in: Indicates user The rating score submitted after the transaction is completed; Indicates user The overall reputation weight at the evaluation time is used to reduce the impact of low-reputation node scores on the results; the summation symbol traverses the set. All users who provided reviews.

[0049] Next, compute nodes Contribution depth component Contribution depth component Reflecting nodes The breadth of dissemination and utilization value of published resources on the network. Traversing nodes. Collection of resources that have been published and stored in the resource index repository For sets Each resource in Statistical Resources Total number of times referenced by other nodes Reference behavior includes direct downloading of resources, being referenced as metadata links, or being used as the parent of derived resources. A logarithmic function is used to smooth the number of references, calculated as follows: ; in: Represents a single resource Reference count; Represents the natural logarithm operation; the constant 1 is used to prevent calculation errors due to zero reference count; the summation symbol traverses nodes. All published resources.

[0050] Last computed node Time-degradation component Age decay component Reflecting nodes The contribution behavior is considered in terms of temporal continuity, with recent contributions given higher weight and older contributions given lower weight. (Read node) In the previous time step Timeliness and reputation value and get the current time interval. internal nodes The newly generated contribution value The component is updated using an exponential decay model, calculated as follows: ; in: This represents the preset time decay coefficient, used to control the rate at which historical reputation is forgotten. It is a natural constant; This represents the time difference between the current update time and the previous update time. The value is obtained by linearly weighting the number of transactions completed and the number of resources uploaded by the node within the current time window. The node is generated through the calculation of these three components. Complete reputation vector at the current moment .

[0051] See attached document Figure 3 This maps multi-dimensional reputation components to a single scalar weight to facilitate quantification in subsequent game theory models. First, nodes are read from a distributed ledger or local database. At the current time step Multidimensional reputation vector data. Multidimensional reputation vector data includes quality reputation components generated in the pre-computation steps. Contribution depth component and time-degradation components .

[0052] Then, a preset set of weight coefficients is obtained. This set of weight coefficients, including quality weight coefficients, is stored in configuration parameters or a smart contract. Contribution weighting coefficient and timeliness weighting coefficient All three weighting coefficients are non-negative real numbers and satisfy the normalization constraint: ; in: Used to adjust the degree of impact of user ratings on overall benefits; Used to adjust the impact of resource citation on overall equity; Used to adjust the degree of impact of continuous contribution behavior on overall equity.

[0053] Based on the linear weighted model, the read reputation components are multiplied and summed with their corresponding weight coefficients to generate nodes. At the current time step Comprehensive reputation weight The calculation formula is as follows: ; in: Represents a node The overall equity weighting value, the overall equity weighting value It is a scalar; This indicates the numerical value of the quality and reputation component; Indicates the numerical value of the depth component; This indicates the value of the time-degradation component.

[0054] The calculated comprehensive credit weight value Write node The overall reputation weight value is shown in the status attribute table. These parameters are directly input into the subsequent Standberg game model to determine the nodes. Incentive subsidy amount when acting as a resource provider. Overall reputation weight. The higher the value, the stronger the node. The higher the overall contribution in a resource-sharing network, the greater the utility gain it obtains during the game.

[0055] See attached document Figure 4 This paper abstracts the node interaction process in a resource-sharing network into a master-slave game structure, first defining the set of game participants. The set of leaders is defined as all nodes in the network that provide teaching resources and have remaining bandwidth, denoted as […]. ,in This represents the total number of resource provider nodes. All nodes in the network that issue resource acquisition requests are defined as the follower set, denoted as . ,in This represents the total number of resource requesting nodes. Any physical node, within different time windows, belongs to a set based on its behavior (uploading or downloading). or set .

[0056] Define the policy space for the leader node. For a set Each leader node in Its decision variable is the resource incentive unit price. Resource incentive unit price This represents the number of payment vouchers or points required to transmit a unit of teaching resources. Leader node. The strategy space is a non-negative real number field, that is... In the first phase of the game, the leader node... Proactively develop and broadcast pricing strategies The aim is to control its own resource output and maximize its first utility function by adjusting prices.

[0057] Define the policy space for follower nodes. For a set... Each follower node in Follower nodes The decision variables are the resource demand vectors for each leader node. : ; Where: Components Represents follower nodes To the leader node The size of the resource data requested to be transferred. Follower node. The strategy space of a follower node is limited by its local storage capacity and processing power. In the second phase of the game, the follower node... After observing the price vectors published by all leader nodes Subsequently, the resource demand was passively adjusted. The aim is to maximize the second utility function under budget and performance constraints.

[0058] Constructing the execution sequence and information exchange mechanism of the game. The multi-leader, multi-follower game model follows the Stankelberg game protocol, i.e., a dynamic game with complete information and a sequential action order. The game process consists of two related decision-making phases: The first phase is the price-setting phase, with the leader node... Based on the prediction of the follower's reaction function, the optimal price is determined first. The second phase is the demand response phase, with follower nodes. Based on known prices Determine the optimal demand quantity .

[0059] The leader node's decision-making relies on the inverse induction of the optimal response of the follower node. Through this hierarchical decision-making structure, the supply and demand balance in the resource market is simulated and optimized.

[0060] See attached document Figure 4 First, construct the resource requester (follower node). The second utility function Second utility function Used to quantify the net gain of follower nodes under a specific resource allocation scheme. Follower node The goal is to adjust the resource demand vector , making The second utility function consists of three parts: the value of acquiring resources, the cost of paying for resources, and the cost of network latency. Its mathematical expression is: ; in: Represents follower nodes Total utility value; Represents the set of leader nodes; Represents follower nodes To the leader node The amount of resources requested; Represents follower nodes The resource value coefficient reflects the degree of preference for teaching content; The logarithmic revenue function represents the diminishing marginal utility of resource acquisition; Represents the leader node The set unit resource price; This represents the time-delay sensitivity coefficient, used to convert time costs into utility losses; Indicates the current network topology Below, follower nodes With leader node The logical communication distance between them.

[0061] Build a resource provider (leader node) The first utility function First utility function Used to quantify the net profit of the leader node in providing resource services. Leader node The goal is to anticipate that followers will make the best response. Under the premise of setting resource unit prices , making The first utility function consists of three parts: resource transaction revenue, bandwidth congestion cost, and reputation incentive subsidy. Its mathematical expression is: ; in: Represents the leader node Total utility value; Represents the set of follower nodes; Represents follower nodes Regarding price The optimal demand response quantity achieved; Represents the leader node The current total load.

[0062] In the formula The bandwidth congestion cost function, expressed in quadratic form, reflects the non-linear cost increase caused by increased load, and is specifically defined as follows: ; in: The bandwidth congestion coefficient reflects the degree of limitation on the physical bandwidth of a node.

[0063] In the formula The item represents additional incentive subsidies. Among them: These are preset incentive factors used to adjust the level of subsidies; The nodes calculated in the preceding steps The overall reputation weight is determined by the additional incentive subsidy, which directly couples the node's reputation status to the game payoff. This allows nodes with high reputation to obtain higher utility returns under the same load, thus gaining an advantageous position in the game equilibrium.

[0064] See attached document Figure 4 Perform the initialization steps. This occurs within the time window at the start of the game. Set the iteration counter Each resource provider node acting as a leader Initialize its resource incentive unit price This is the preset benchmark price or the equilibrium price of the previous time window. Each resource requester node acting as a follower... Initialize its resource demand vector This is a zero vector. Simultaneously, a learning rate parameter is set to control the iteration step size. and the error threshold used to determine convergence .

[0065] This then enters the follower strategy update phase. In the... In this iteration, each resource requester node Receive the current price vector broadcast by all leader nodes. Resource requester node According to the second utility function First-order optimality conditions are calculated for each leader node. Optimal resource demand Based on the properties of partial derivatives of the utility function, the optimal resource demand can be calculated directly using the following analytical formula: ; in: Indicates the first Nodes in the next iteration To the node The amount of resources requested; This is the resource value coefficient; For nodes In the The price set in the next iteration; Transmission delay cost per unit of resource; function This is used to ensure the non-negativity of resource requirements. After calculation, all resource requester nodes will update the required quantities. Feedback is sent to the corresponding leader node. .

[0066] Next, the leader strategy update phase begins. After receiving feedback from all resource requester nodes, each resource provider node... Calculate its current total load requirements Resource provider nodes The gradient ascent algorithm is used to update the resource incentive unit price in order to approximate the first utility function. The maximum value point. The price update formula is defined as follows: ; in: For the first The pricing strategy for the next iteration; The preset gradient step size satisfies This represents the projection operation onto the non-negative real number field; The first utility function is expressed in terms of price. The gradient value of the first utility function with respect to price. The gradient value is determined by the node It is calculated locally based on the current marginal revenue, marginal congestion cost, and the marginal rate of change of reputation incentives.

[0067] Finally, the convergence determination step is performed. The Euclidean distance between the price strategies of all leader nodes in two adjacent iterations is calculated. The formula is: ; Will With error threshold Compare. If The game has not yet reached equilibrium, so let And return to the follower policy update phase. If Once the game has converged to a Nash equilibrium, the iteration process is terminated, and the current strategy combination is changed. Once the optimal resource scheduling scheme for the current time window is confirmed, the actual resource transmission and payment settlement will be executed according to the optimal resource scheduling scheme.

[0068] See attached document Figure 5 First, create a structure with a length of A time-sliding window is used to record historical game data. At each discrete time step... The Nash equilibrium solution obtained from the aforementioned steps, i.e., the optimal price vector, is... The utility values ​​of each participating node are stored in a time-sliding window. The data set within the time-sliding window is denoted as: ; in: Indicates the current time step The historical set of game states under the given conditions; Indicates at time step The optimal price vector of Nash equilibrium obtained through game theory; Indicates at time step The total global utility of the network under the following conditions; Indicates the current time step index; This indicates the preset size of the time-based sliding window; This represents the historical time variable that is within the time sliding window.

[0069] Calculate the Nash equilibrium drift rate based on data within a time sliding window. The Nash equilibrium drift rate is used to characterize the volatility of resource trading price strategies over a continuous time period. It is obtained by calculating the average Euclidean distance between the optimal price vectors of adjacent time steps. The calculation formula is as follows: ; in: Indicates time The Nash equilibrium drift rate; Indicates time The Nash equilibrium price vector; The L2 norm (Euclidean norm) of a vector is used to represent the vector. The length of the sliding window. The larger the value, the more unstable the game equilibrium point in the current network, and the more frequent the strategy adjustments between nodes.

[0070] Simultaneously calculate the total global utility value at the current time step. The total global utility is defined as the sum of the utilities of all resource-providing nodes (leaders) and resource-requesting nodes (followers) in a Nash equilibrium state. The calculation formula is as follows: ; in: Indicates time The total global utility; For the leaders to gather; For the followers to gather; For the first The utility value of a leader in equilibrium; For the first The utility value of each follower in equilibrium. Further, based on the calculated global utility sequence, the global utility variance is calculated. Global utility variance is used to characterize the oscillations in overall utility during the sliding window. First, the arithmetic mean of global utility within the window is calculated. The variance is then calculated using the following formula: ; in: Indicates the current time step Global utility variance under; This indicates the preset size of the time-based sliding window; Indicates the current time step index; This represents the historical time variable that is within the time sliding window; Indicates at time step The total global utility of the network under the following conditions; Indicates the current time step The average global utility of the network within the time sliding window described below.

[0071] The calculated Nash equilibrium drift rate and global utility variance The state parameter is output to the topology reconstruction decision logic. Nash equilibrium drift rate. and global utility variance Together, they constitute a quantitative basis for judging whether the current alliance layer structure is suitable for the current resource supply and demand relationship, and are used to trigger subsequent topology splitting or topology merging operations.

[0072] See attached document Figure 5 Based on the monitoring indicators calculated in the aforementioned steps, the node composition and connection relationships of the alliance layer are changed by executing topology splitting or merging commands. First, the calculated Nash equilibrium drift rate is... With the preset drift threshold Compare. Drift threshold The maximum allowable range of fluctuations in game strategies is defined.

[0073] When detected If a node group with severe policy conflicts is identified within the current regional alliance consensus layer, hindering the convergence of the global game, a topology splitting process is triggered. An interaction frequency matrix of nodes within the current alliance is then constructed. The matrix dimension is , of which elements Indicates a sliding window internal nodes With nodes The total number of resource requests and responses that occurred between them.

[0074] Based on the interaction frequency matrix Perform the minimum cut graph partitioning algorithm to partition the current set of nodes in the alliance. Divide into two disjoint subsets and The objective function for partitioning is defined as: ; in , This represents the set of two complementary sub-community nodes formed after a network topology split; Indicates belonging to a set Any participating node in the process; Indicates belonging to a set Any participating node in the process; Represents a node With nodes Weights on the strength of the interaction coupling between them.

[0075] The constraints are satisfied: and .

[0076] Based on the calculated subset and Logically, this blocks direct game channels between nodes belonging to different subsets, and respectively and Instantiate independent distributed ledgers to decouple the physical network topology.

[0077] Simultaneously, the calculated Nash equilibrium drift rate will be... With the preset stability threshold Compare and average global utility With the preset cost threshold Compare the stability thresholds. Cost threshold used to determine whether a game has reached a steady state. Used to determine whether the utility of the current resource allocation meets the minimum economic requirements.

[0078] When detected and When the current regional alliance is determined to be in an inefficient local equilibrium state, i.e., a deadlock state, the topology fusion process is triggered. The current alliance is then calculated. resource feature centroid vector resource feature centroid vector It consists of the arithmetic mean of the feature vectors of the storage resources of all nodes in the alliance.

[0079] Traverse the set of neighboring alliances with the closest logical distance in the network Calculate the characteristic distance between the current alliance and each of its neighboring alliances. The formula for calculating feature distance is as follows: ; in: Indicates the current federal community and the first Feature Euclidean distance between candidate external communities; This represents the centroid vector representing the resource distribution characteristics of the current federal community; Indicates the first Centroid vectors representing the resource distribution characteristics of each candidate external community; This represents the calculation operation of the L2 norm (or Euclidean distance) of a vector.

[0080] Select feature distance Minimum Adjacency Alliance As a fusion object, it executes the cross-chain handshake protocol.

[0081] Establish connections to the current alliance Alliance with the target The communication bridge merges the distributed ledger data from both sides and synchronizes the resource index tables. The resulting new set of nodes... In the next step, they will jointly participate in the game, thereby expanding the resource search space and introducing external liquidity.

[0082] After performing a topology split or topology merge operation, update the global index layer. The routing mapping table ensures that subsequent resource discovery requests can be correctly routed to the reconstructed alliance layer nodes.

[0083] See attached document Figure 1 This process is executed collaboratively by all public service nodes across the network. After performing topology splitting or merging operations, it verifies the convergence of the network state and establishes the final steady-state structure. It first initiates a recursive game-theoretic solution program. At the time step after the topology reconstruction operation is completed... For the new alliance set formed after the restructuring Reinitialize the Stankberg game model. Force all nodes to operate based on the new network connectivity. Then, re-execute the aforementioned distributed game equilibrium solution algorithm to calculate the Nash equilibrium solution under the new topology. And the corresponding global utility value.

[0084] Constructing global situation functions The global situation function is used to quantitatively evaluate the global effectiveness of the current network topology. It consists of two parts: the cumulative global utility of the entire network and the network maintenance communication overhead. Its mathematical expression is defined as: ; in: Indicates time step The overall network situation value; This represents the total number of alliances currently existing in the network; Indicates the first The total internal global utility of each alliance under the current Nash equilibrium; This is a preset overhead penalty coefficient used to balance utility gains and communication costs; This represents the set of active links in the current network topology; Represents a node With nodes The communication bandwidth cost or heartbeat overhead required to maintain the connection.

[0085] Perform the potential gain determination step. Calculate the difference in the global state function before and after topology reconstruction: ; in: This represents the change in the system's potential energy function before and after topological reconstruction; This represents the global situational energy function used to evaluate the state of the network topology; Indicates at time step The updated federated network topology will be reconstructed. Indicates at time step The current federal network topology.

[0086] when When it is determined that the current topology reconfiguration operation has improved overall performance, the reconfiguration result is accepted and the global routing table is updated. If the reconstruction is deemed invalid, a rollback operation is performed to restore the network topology to the previous time step. The system will then implement a temporary cooling-off lock on the node that triggered the reconstruction, prohibiting it from operating within a preset cooling-off period. The topology change request was initiated again.

[0087] Perform a global convergence check. Continuously monitor the rate of change of the global situation function, and define the convergence criteria as follows: ; in: Indicates the current time step The overall situational energy function value of the network topology is shown below. Indicates a step in the future time. The overall situational energy function value of the network topology is shown below. This indicates the preset convergence threshold (or relative error tolerance); Indicates the length of the observation time window used for convergence testing; This represents the time offset variable within the observation time window; This indicates the current time step index.

[0088] After confirming global steady state, the current alliance layer structure and global index mapping relationship are locked. At this point, resource requesting nodes (followers) determine the optimal resource demand. Establish a peer-to-peer data transmission channel with the resource provider node (leader). The resource provider node, based on the aforementioned resource entropy... Upon determining the result, the encrypted resource entity data block is read from local storage or the consortium distributed file system and transmitted to the resource requester node through the established channel. After receiving the data, the resource requester node automatically pays the corresponding calculated resource incentive voucher to the resource provider node according to the smart contract, completing the physical closed loop of resource sharing.

Claims

1. A teaching resource sharing decision optimization method based on a hierarchical federated incentive mechanism, characterized in that, Includes the following steps: S1: Construct a hierarchical heterogeneous federated topology network based on resource entropy, logically divide the network architecture into local layer, federation layer and global layer, and determine the storage level and routing strategy of teaching resources by calculating the resource entropy of the teaching resources to be shared. S2: Establish a multi-dimensional reputation-weighted equity quantification model, maintain a multi-dimensional reputation vector for each participating node in the network, and calculate the comprehensive reputation weight of each participating node based on the multi-dimensional reputation vector; S3: Execute resource scheduling decisions based on dynamic Stankberg game, construct the resource sharing process as a multi-leader multi-follower game model, use distributed iterative algorithm to solve the Nash equilibrium solution, and obtain the optimal price vector and resource allocation matrix; S4: Dynamic topology reconstruction based on game equilibrium drift rate. During the game, the Nash equilibrium drift rate and global utility variance are monitored in real time, and topology splitting or merging operations are triggered according to the monitoring results. Resource scheduling decisions are re-executed under the reconstructed network topology.

2. The teaching resource sharing decision optimization method based on a hierarchical federated incentive mechanism according to claim 1, characterized in that, The resource entropy is calculated by identifying the set of sensitive features contained in the teaching resources to be shared, obtaining the probability that the content of the teaching resources to be shared belongs to each type of sensitivity feature, summing the product of the probability of each type of sensitivity feature with the base-2 logarithm of the probability, and taking the negative value of the summation result to obtain the resource entropy.

3. The teaching resource sharing decision optimization method based on a hierarchical federated incentive mechanism according to claim 2, characterized in that, The specific method for determining the storage level and routing strategy of the teaching resources is to compare the calculated resource entropy with a preset classification threshold; When the resource entropy is greater than the first threshold, the teaching resource to be shared is determined to be a private resource, and the resource metadata is routed to the local layer storage. When the resource entropy is between the second threshold and the first threshold, the teaching resource to be shared is determined to be a regional shared resource, and the resource metadata is routed to the alliance layer storage. When the resource entropy is less than the second threshold, the teaching resource to be shared is determined to be a public resource, and the resource metadata is routed to the global layer storage.

4. The teaching resource sharing decision optimization method based on a hierarchical federated incentive mechanism according to claim 1, characterized in that, The multidimensional reputation vector includes three components: quality reputation, contribution depth, and time-related decay. The update method for the timeliness decay is to multiply the timeliness reputation value of the previous time step by the decay coefficient, and add the new contribution value of the current time step to obtain the timeliness decay of the current time step. The decay coefficient is an exponential function value with the natural constant as the base and the negative value of the product of the preset time decay factor and the time interval as the exponent.

5. The teaching resource sharing decision optimization method based on a hierarchical federated incentive mechanism according to claim 4, characterized in that, The specific method for calculating the overall reputation weight of each participating node is as follows: Obtain the preset quality weight coefficient, contribution weight coefficient, and timeliness weight coefficient. Multiply the quality reputation, contribution depth, and timeliness decay by the corresponding weight coefficients, and add the multiplication results to obtain the comprehensive reputation weight of the participating node.

6. The teaching resource sharing decision optimization method based on a hierarchical federated incentive mechanism according to claim 1, characterized in that, The multi-leader, multi-follower game model is constructed as follows: The resource provider is defined as the leader. The leader sets the resource incentive unit price based on its own bandwidth cost and the comprehensive reputation weight in order to maximize the first utility function. The first utility function consists of the resource transaction revenue minus the bandwidth congestion cost, plus the incentive subsidy based on the comprehensive reputation weight. The resource requester is defined as a follower. The follower determines the resource demand based on the resource incentive unit price and network latency cost in order to maximize the second utility function. The second utility function is composed of the resource use value minus the payment cost and then minus the transmission latency loss.

7. The teaching resource sharing decision optimization method based on a hierarchical federated incentive mechanism according to claim 6, characterized in that, The specific process of solving for Nash equilibrium using a distributed iterative algorithm includes: The follower calculates the optimal resource requirement that maximizes the second utility function based on the received resource incentive unit price and feeds it back to the leader. The leader determines the total load based on the optimal resource requirements of all followers received, and updates the resource incentive unit price using a gradient ascent algorithm to approximate the maximum value of the first utility function. The above process is executed iteratively until the change in the unit price of the resource incentive meets the preset convergence condition.

8. The teaching resource sharing decision optimization method based on a hierarchical federated incentive mechanism according to claim 1, characterized in that, The definitions of the Nash equilibrium drift rate and the global utility variance are as follows: The Nash equilibrium drift rate is the average of the Euclidean distances between the optimal price vectors at adjacent time steps. The value of the global utility variance is the variance of the total global utility value within the time sliding window.

9. The teaching resource sharing decision optimization method based on a hierarchical federated incentive mechanism according to claim 8, characterized in that, The specific steps of the topology splitting operation are as follows: When the monitored Nash equilibrium drift rate is greater than the preset drift threshold, the current alliance layer structure is determined to be unstable, and the interaction frequency matrix of the nodes in the current alliance is constructed. Based on the interaction frequency matrix, a graph partitioning algorithm is executed to divide the node set of the current alliance layer into two disjoint sub-alliance sets, thereby minimizing the interaction weight between the sub-alliance sets and blocking the direct game channel between nodes belonging to different sub-alliance sets.

10. The teaching resource sharing decision optimization method based on a hierarchical federated incentive mechanism according to claim 8, characterized in that, The specific steps of the topology fusion operation are as follows: When the monitored Nash equilibrium drift rate is less than the preset stability threshold and the average global utility is lower than the maintenance cost threshold, it is determined that the current alliance layer resource liquidity is insufficient. Calculate the centroid vector of the resource features of the current alliance, which reflects the feature distribution center of the storage resources of all nodes in the alliance; Calculate the characteristic distance between the current alliance and its neighboring alliances in the network, select the neighboring alliance with the smallest characteristic distance as the fusion object, establish a communication bridge connecting the current alliance and the fusion object, merge the distributed ledger data and synchronize the resource index.