Mountainous biomass collection, storage and transportation collaborative scheduling method and system based on multi-agent game

The collaborative scheduling method for biomass collection, storage, and transportation in mountainous and hilly areas, based on multi-agent game theory, solves the problems of high transportation costs and unbalanced benefit distribution caused by the dispersion of resource points and complex terrain in mountainous and hilly areas. It achieves increased farmer income, reduced transportation costs, and improved resource collection coverage, thus forming a sustainable business ecosystem.

CN122453014APending Publication Date: 2026-07-24TIANFU YONGXING LAB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANFU YONGXING LAB
Filing Date
2026-04-23
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Biomass collection, storage and transportation in mountainous and hilly areas face problems such as scattered resource points, complex terrain leading to high transportation costs, and unbalanced distribution of benefits. Existing technologies cannot dynamically adjust and flexibly switch operating models, resulting in insufficient participation from farmers, high moral hazard among intermediaries, and high procurement costs for end users.

Method used

A collaborative scheduling method for biomass collection, storage, and transportation in mountainous and hilly areas is adopted, which uses multi-agent game theory to calculate weighted topological centrality by acquiring static and dynamic data, evaluates the functional status level of nodes and the weight of benefit allocation, dynamically adjusts the collaborative mode, and combines the MARL framework and the improved Shapley value method for path planning and benefit allocation.

Benefits of technology

It has increased farmers' income by 15-25%, reduced transportation costs by 20-30%, increased resource collection coverage to 90%, eliminated monitoring costs and default risks, and formed a sustainable business ecosystem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122453014A_ABST
    Figure CN122453014A_ABST
Patent Text Reader

Abstract

The application discloses a mountain and hilly biomass collection, storage and transportation collaborative scheduling method and system based on multi-agent game, relates to the technical field of wisdom agriculture and supply chain collaboration, and aims to solve the problems of unbalanced benefit distribution, ignored terrain constraints and rigid collaborative mechanism in the prior art. The application first acquires static attribute data, dynamic operation data of the collection and storage nodes, and geographical distribution and resource quantity of all resource points of the biomass to be collected and transported; determines the dispersion degree of the geographical distribution of all resource points, the balance degree of the resource quantity, and the difference degree of the appeals of the benefit related parties; takes the dispersion degree, the balance degree, the static attribute data of each collection and storage node and the difference degree of the appeals as judgment indexes, determines the collaborative mode of the benefit related parties according to a preset judgment mechanism, executes collaborative planning, and completes dynamic planning of the collection path and the corresponding resources; and finally realizes dynamic allocation of resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart agriculture and supply chain collaboration technology, specifically to a collaborative scheduling method and system for biomass collection, storage and transportation in mountainous and hilly areas based on multi-agent game theory. Background Technology

[0002] Biomass energy, as an important renewable energy source, plays a vital role in optimizing the energy structure and promoting rural economic development. However, the collection, storage, and transportation of biomass in the mountainous and hilly areas of southern my country (such as Sichuan, Yunnan, Chongqing, and Hunan) face unique challenges: resource points are highly dispersed and small in scale, complex terrain leads to high transportation costs, and the traditional multi-level chain of "farmers-middlemen-purchasing stations-terminals" suffers from a severe imbalance in profit distribution, resulting in low income for farmers, cumbersome middlemen, and high procurement costs for end users, creating a predicament of "unable to collect, unable to transport, and dissatisfied with all parties."

[0003] Existing technologies suffer from three main shortcomings: First, the profit distribution mechanism is rigid, often employing fixed acquisition methods or simple profit sharing, failing to dynamically adjust based on market fluctuations, transportation difficulties, and quality differences, leading to insufficient farmer participation and high moral hazard among intermediaries. Second, route planning neglects terrain constraints and multi-party game theory; traditional VRP models assume plain transportation and only optimize the logistics base, failing to consider the impact of mountain slopes on fuel consumption, and not incorporating economic behaviors such as price games and alliance formation into the optimization. Third, there is a lack of adaptive coordination mechanisms; when resource distribution exhibits a "small concentration, large dispersion" characteristic, it cannot flexibly switch between centralized and distributed operation modes, making it difficult to balance economies of scale and response speed.

[0004] In recent years, multi-agent reinforcement learning (MARL) has shown potential in distributed logistics coordination, and game theory has been increasingly applied in the distribution of benefits in the supply chain. However, a method to integrate the two and systematically design for the collection, storage and transportation of biomass in mountainous and hilly areas has not yet been studied. Summary of the Invention

[0005] To address the aforementioned problems, this invention aims to provide a collaborative scheduling method and system for biomass collection, storage, and transportation in mountainous and hilly areas based on multi-agent game theory, thereby solving the problems of unbalanced benefit distribution, insufficient terrain adaptability, and rigid collaborative mechanisms in existing technologies.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A collaborative scheduling method and system for biomass collection, storage, and transportation in mountainous and hilly areas based on multi-agent game theory includes the following steps: Obtain static attribute data, dynamic operation data, and the geographical distribution and resource quantity of all resource points of biomass to be collected and transported within the mountainous and hilly network. Based on the static attribute data of each storage node, the weighted topological centrality index of each storage node is calculated. By combining weighted topological centrality index and dynamic operational data, the real-time functional status level and benefit distribution weight of each storage node are evaluated. Determine the dispersion of the geographical distribution of all resource points, the balance of resource quantity, and the differences in the demands of stakeholders; Using dispersion, balance, static attribute data of each storage node, and demand differences as judgment indicators, the collaboration mode of stakeholders is determined according to the preset judgment mechanism, and collaboration planning is executed to complete the dynamic planning of the collection path and its corresponding resources; among them, the collaboration mode includes multi-agent parallel collaboration mode and centralized collaboration mode. Based on the real-time functional status level and benefit allocation weight of each collection and storage node, the resources corresponding to the planned collection path are dynamically allocated to the collection and storage nodes with matching functional status.

[0007] Through the above technical solutions, the following further information is provided: static attribute data includes geographic coordinates, total storage capacity, terrain slope level, and topological connection relationship in the transportation network; dynamic operational data includes real-time storage capacity, task load, historical revenue sharing ratio, and current bargaining power index; resource points include farmers, intermediaries, and primary processing points.

[0008] Based on the above technical solution, the formula for calculating the weighted topological centrality index of each storage node is as follows: ; in: The weighted topological centrality index for each storage node. To achieve proximity centrality based on road network distance, This represents the normalized value of the average terrain slope around the node. The distance to the end market is the inverse of the equivalent distance; α, β, γ are weighting coefficients, and satisfy α+β+γ=1.

[0009] Based on the above technical solutions, the preset judgment mechanism is as follows: When the dispersion, balance, static attribute data of each storage node, and demand difference all simultaneously meet their respective preset parallel operation thresholds, the current task is considered to have high decoupling capability, and the multi-agent parallel collaboration mode is activated; conversely, if any indicator fails to reach the corresponding threshold, the task is considered to have low decoupling capability, and the centralized collaboration mode is activated.

[0010] Furthermore, through the above technical solutions, in the multi-agent parallel collaboration mode, all resource points are divided into multiple non-overlapping sub-clusters according to the principle of geographical proximity, and an independent agent is assigned to each sub-cluster. Each agent, based on the MARL framework, learns the optimal collection path and the benefit distribution scheme within the sub-cluster through centralized training and distributed execution.

[0011] Furthermore, based on the above technical solutions, in step S500, the MARL framework adopts the CTDE architecture, which includes: Centralized training phase: The central Critic network acquires global state information, including the location, load, path planning, and revenue distribution status of all agents, and calculates the global advantage function; In the distributed execution phase, each agent makes independent decisions based on local observations, including the distribution of resource points within the corresponding sub-cluster, terrain conditions, and local market prices. The agent's action space includes path selection, acquisition pricing, and sending collaboration requests to other agents. Reward function design: The single-step reward equals the first weight coefficient multiplied by the local profit increment, plus the second weight coefficient multiplied by the global profit increment, and minus the third weight coefficient multiplied by the penalty for path overlap or vicious price competition.

[0012] Furthermore, through the above technical solutions, in the centralized collaborative mode, a centralized cyclic collection path is uniformly constructed for all resource points, and the improved Shapley value method or Nash negotiation model is used for scheduling and allocation on a global scale.

[0013] This invention also proposes a benefit equilibrium system for decentralized collection, storage, and transportation of biomass based on multi-agent cooperation and dynamic game theory, characterized in that it includes the following components for executing the aforementioned method: The data acquisition module is used to acquire static attribute data, dynamic operational data, and geographical distribution and resource quantity information of each storage node; The mountain adaptation assessment module is used to calculate the weighted topological centrality based on the static attribute data of each storage node, and to evaluate the real-time functional status level and benefit allocation weight of the node based on dynamic operation data. The task mode parsing module is used to determine the decoupling capability of tasks based on the dispersion of geographical distribution, the balance of resource quantity, static attribute data, and the differences in the demands of stakeholders. The path planning module is used to execute a multi-agent parallel collaboration mode or a centralized interest coordination mode based on the decoupling determination result, so as to realize the path planning of resource points. The profit distribution module is used to dynamically allocate resources and determine the profit-sharing ratio based on the real-time functional status level and the game equilibrium solution.

[0014] Based on the above technical solutions, further: the mountain adaptability assessment module includes: The weighted topological centrality calculation submodule is used to calculate the weighted topological centrality based on the static attribute data of each storage node; The real-time functional status level calculation submodule is used to calculate the real-time functional status level of nodes based on dynamic operational data. The weight allocation submodule is used to allocate node weights based on weighted topological centrality and dynamic operational data.

[0015] Based on the above technical solutions, the task mode parsing module further includes: The geographic distribution dispersion calculation submodule is used to calculate the geographic distribution dispersion based on the geographic distribution of resource points; The resource balance calculation submodule is used to calculate the resource balance based on the resource quantity of resource points. The interest claim difference calculation submodule is used to calculate the interest claim difference among stakeholders. The task decoupling determination submodule is used to determine the decoupling capability of a task based on the dispersion of geographical distribution, the balance of resource quantity, static attribute data, and the differences in the demands of stakeholders. The determination result is either a highly decoupling task or a low decoupling task; a multi-agent parallel collaboration mode or a centralized interest coordination mode. The path planning module includes a multi-agent parallel collaboration mode submodule and a centralized interest coordination mode submodule. The multi-agent parallel collaboration mode submodule is used to execute the multi-agent parallel collaboration mode when the result is a highly decoupling task; the centralized interest coordination mode submodule is used to execute the centralized interest coordination mode when the result is a low decoupling task.

[0016] The beneficial effects of this invention are: 1. By using a hierarchical game model and a dynamic profit-sharing mechanism, we ensure that farmers' income is 15-25% higher than that of the traditional model, and that the reasonable profits of intermediaries are guaranteed, thus forming a sustainable business ecosystem.

[0017] 2. The weighted topological centrality index incorporates terrain slope into network evaluation. MARL path planning automatically avoids inefficient road sections with high slopes, which can reduce transportation costs by 20-30% compared to traditional planar distance optimization models.

[0018] 3. The dual-mode switching mechanism for high / low decoupled tasks enables the system to handle large-scale collection in "small centralized" areas and adapt to distributed collaboration in "large decentralized" areas, increasing resource collection coverage to over 90%.

[0019] 4. Combining blockchain-based evidence storage with IoT-triggered mechanisms eliminates the monitoring costs and default risks associated with traditional contracts, improving contract execution efficiency by over 40%. Attached Figure Description

[0020] Figure 1 The flowchart is a method for balancing the benefits of decentralized collection, storage and transportation of biomass based on multi-agent collaboration and dynamic game theory. Figure 2 A schematic diagram of the system structure of this invention; Figure 3 This is a schematic diagram of the Agent coordination mechanism under the MARL-CTDE architecture. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0022] Example 1 See Figures 1-3 This application discloses a collaborative scheduling method for biomass collection, storage, and transportation in mountainous and hilly areas based on multi-agent game theory, including the following steps: S100. Obtain static attribute data, dynamic operation data, and geographical distribution and resource quantity information of biomass resource points to be collected and transported within the mountainous and hilly network for each collection and storage node. S200. Based on the static attribute data of each storage node, the weighted topological centrality index of each storage node is calculated. S300, and comprehensively evaluate the real-time functional status level and benefit distribution weight of each storage node based on dynamic operational data; S400: The dispersion of the geographical distribution of all resource points, the balance of resource quantity, the static attribute data of each storage node, and the difference in the interests of stakeholders are used as judgment indicators. When all of the above indicators simultaneously meet their respective preset parallel operation thresholds, the current task is deemed to have high decoupling capability, and the multi-agent parallel collaboration mode is activated. Conversely, if any indicator fails to reach the corresponding threshold, the task is deemed to have low decoupling capability, and the centralized collaboration mode is activated. S500, based on the collaborative mode determined by S400, executes collaborative planning, where: In the multi-agent parallel collaborative mode, all resource points are divided into multiple non-overlapping sub-clusters according to the principle of geographical proximity, and an independent agent is assigned to each sub-cluster. Each agent is based on the MARL framework and learns the optimal collection path and the interest distribution scheme within the sub-cluster through centralized training and distributed execution. At the same time, agents share intent information through message passing mechanism to avoid path conflicts and bidding competition. In the centralized collaborative model, a centralized cyclical collection path is uniformly constructed for all resource points, and an improved Shapley value negotiation model is adopted to distribute benefits globally. Both models ensure that the gains of each participant are no less than their retained gains under non-cooperative game conditions, and the total residual increment of the supply chain is dynamically allocated according to the contribution of each party. S600: Based on the real-time functional status level and benefit allocation weight of each collection and storage node, dynamically allocate the resources corresponding to the planned collection path to the collection and storage node with the matching functional status.

[0023] In S100, multi-source heterogeneous data is acquired; the static attribute data includes geographic coordinates, total storage capacity, terrain slope level, and topological connection relationship in the transportation network; the dynamic operation data includes real-time storage capacity, task load, historical revenue sharing ratio, and current bargaining power index; the resource points include farmers, intermediaries, and primary processing points.

[0024] The terrain slope grade is calculated using DEM (Digital Elevation Model) data, as shown in Table 1. It is divided into four levels according to the slope angle: plain (<5°), gentle slope (5°-15°), steep slope (15°-25°), and dangerous slope (>25°). Different levels correspond to different transportation cost coefficients. The unit distance transportation cost in dangerous slope areas can be 2.5-3 times that in plain areas.

[0025] Table 1. Terrain Slope Grade Table ; Benchmark weight =0.25 (when four nodes are equally distributed), actual weight .

[0026] ; ; in Let v be the shortest path distance from node v to all other nodes, calculated using Dijkstra's algorithm.

[0027] The transportation network is modeled as an undirected graph G=(V,E), where V is the set of road intersections and access nodes, and E is the set of passable road segments. Topological connections are represented using... express, Represents a node and Direct connection; simultaneously record road segment attribute vectors These represent the road segment length, average gradient, and traffic capacity level, respectively.

[0028] In S200, a topological centrality evaluation model adapted to mountainous and hilly terrain is constructed. Based on the static attribute data of the storage nodes and combined with the influence coefficient of terrain slope level on transportation costs, the weighted topological centrality index is analyzed.

[0029] The weighted topological centrality index takes into account road network connectivity, terrain accessibility, and market proximity, overcoming the shortcomings of traditional centrality indices that only consider topological distance and ignore geographical constraints.

[0030] The formula for calculating the weighted topological centrality index of each storage node is as follows: ; in: The weighted topological centrality index for each storage node. To achieve proximity centrality based on road network distance, This represents the normalized value of the average terrain slope around the node. The distance to the end market is the inverse of the equivalent distance; α, β, γ are weighting coefficients, and satisfy α+β+γ=1.

[0031] Proximity centrality Calculation based on road network diagram G structure: ,in Let v be the shortest path distance from node v to all other nodes, calculated using Dijkstra's algorithm.

[0032] , so that the value range is [0,1].

[0033] Slope normalization Extract DEM data within a 2km buffer zone surrounding node v and calculate the average slope angle. Normalized according to four levels of classification: Plains: Gentle slope: ;steep slope: Dangerous slope: The formula uses 1- That is, the smaller the slope, the greater the contribution of centrality.

[0034] The principles for setting the weighting coefficients α, β, γ are as follows: Default values: α=0.5 (road network connectivity priority), β=0.3 (terrain adaptability), γ=0.2 (market proximity).

[0035] Dynamic adjustment rules: Adaptively adjusted based on regional characteristics. When the average slope of the region is >15°, β is increased by 0.1 (topographic constraints are strengthened); when the region has only one end market, γ is decreased by 0.1 (market factors are weakened); when the road network density is <1.0km / km², α is increased by 0.1 (connectivity is prioritized). The adjustment must maintain α+β+γ=1.

[0036] Adjustment mechanism: set by the system administrator during initialization, or automatically optimized through reinforcement learning based on historical scheduling results (adjusted quarterly based on transportation cost feedback).

[0037] In step S300, the real-time functional status level and benefit distribution weight of the nodes are evaluated by combining the weighted topological centrality index and the revenue sharing history and bargaining power in the dynamic operation data.

[0038] Real-time functional status level determination is based on a constructed multi-dimensional status scoring function, the expression of which is: .

[0039] in: This represents the real-time warehouse capacity utilization rate (the more empty space, the higher the score). Task load rate (the lower the load, the higher the score). This is a normalized value for the historical revenue sharing ratio. This represents the normalized value of the current bargaining power index. , , , These are the weighting coefficients, and their sum is 1.

[0040] Level mapping: Score(v)∈[0.8,1.0] is level 1, [0.6,0.8) is level 2, [0.4,0.6) is level 3, and [0,0.4) is level 4.

[0041] Benefit distribution weight mapping: Where λ∈[0.5,0.7] is the topological weight percentage. This is the weighted topological centrality normalized value. The higher the level and the greater the centrality, the greater the assigned weight, and the more revenue share is obtained when undertaking cross-regional scheduling tasks.

[0042] In S400, the task decoupling capability is determined and a collaborative mode is selected. When the geographical dispersion, resource balance, road network connectivity, and interest demand difference meet the corresponding preset parallel conditions, it is determined to be a highly decoupling task and triggers the multi-agent parallel collaborative mode; otherwise, it is determined to be a lowly decoupling task and triggers the centralized interest coordination mode.

[0043] Geographical distribution dispersion: Calculate the standard deviation of the road network distances for all resource points. .

[0044] in Let be the road network distance from the i-th resource point to its nearest storage node. Average distance. Threshold. =3.5km (can be dynamically adjusted according to the area, such as when the area is >1000km2, the threshold is relaxed to 4.5km).

[0045] Resource balance: using the coefficient of variation .

[0046] .

[0047] in The standard deviation of resource quantity in each sub-cluster. This represents the average resource quantity. Threshold. =0.4.

[0048] Road network connectivity: Road network density , The total length of roads within the area. Area of ​​the region. Threshold. .

[0049] The introduction of differences in interests is one of the key innovations of this invention. When the differences in the expected distribution of benefits among the parties are too large (such as farmers expecting a minimum guaranteed price higher than the price that intermediaries can afford), forced parallel coordination will lead to the breakdown of the alliance. At this time, it is necessary to switch to centralized coordination, with the core enterprise (storage station or terminal) taking the lead in formulating a unified pricing and distribution plan.

[0050] The method for quantifying the difference in interests is as follows:

[0051] in, Let be the utility function value of the i-th stakeholder. For average utility, The degree of difference in interests among stakeholders.

[0052] Farmers: ,in For the benefit of farmers, Actual acquisition price, For sales volume, For unit labor cost, Government subsidies; Middlemen: ,in For the benefit of middlemen, For resale price, For the acquisition price, For transportation costs, For warehousing costs; Storage and collection stations: ,in For the benefit of the storage station, For shipping rates, For the quantity of goods received, This is a long-term contract price. For outbound volume, For operating costs, Depreciation cost; End users: ,in For end-user utility, Let the energy conversion efficiency function be... For processing costs.

[0053] Threshold for difference in interests: This is a fixed empirical value, which can be dynamically adjusted according to market fluctuation cycles (e.g., adjusted to 0.3 during peak season).

[0054] In S500, collaborative planning is performed based on the collaborative mode determined by S400.

[0055] Multi-agent parallel collaborative mode: The resource point cluster is decoupled into several geographically adjacent sub-clusters, and each sub-cluster corresponds to an independent agent; each agent learns the optimal collection path and local benefit allocation strategy through the MARL framework.

[0056] MARL adopts the CTDE architecture, which includes the Actor network and the Central Critic network, as well as a communication module.

[0057] Actor Network (each agent is independent): Input dimension = local observation dimension (4 × number of sub-cluster resource points + terrain feature dimension), hidden layer = [256, 128, 64], output dimension = action space size (path selection: number of sub-cluster resource points + 1 integration point; pricing: 10 discrete price levels; cooperation request: 2-dimensional binary). Activation functions: ReLU (hidden layer), Softmax (output layer, path selection), Sigmoid (output layer, cooperation request).

[0058] Central Critic Network: Input dimension = global state concatenation (concatenation of all agent observations + global resource distribution matrix), hidden layer = [512, 256, 128], output dimension = 1 (global state value V). Employs a QMIX hybrid network with monotonicity constraints. , where f is a nonlinear hybrid function with positive weights.

[0059] Communication module: inter-agent message encoder (GRU network, 64 hidden units), decoder (fully connected layer, output intent vector 32-dimensional).

[0060] The above structure employs a centralized training phase where a central Critic acquires global information, while in the distributed execution phase, each Agent makes independent decisions based on local observations. Agents coordinate through a message passing mechanism to avoid path overlap and price conflicts.

[0061] Specifically, the implementation path of the MARL algorithm is as follows: Observation coding: This involves encoding the coordinates of resource points within the sub-cluster. Topographic slope Estimated resource quantity Local market prices Encoded as a state vector st∈R4n (n is the number of resource points in the sub-cluster).

[0062] Actor Network: A two-layer fully connected network (hidden layer 256→128 units, ReLU activation) is used to output the probability distribution of actions. The actions include: the next access node number, the local purchase price (discretized into 10 levels within the range of [160,220] yuan / ton), and the cooperation request flag.

[0063] Critic Network: The central Critic receives the global state concatenation vector (states of all agents + global resource distribution) and outputs the global state value. .

[0064] Training process: The PPO algorithm is used, with each episode length being the number of resource points in the sub-cluster × 2 steps, training 10,000 episodes, with a learning rate of 3 × 10⁻⁴ and a discount factor γ = 0.95.

[0065] The reward function design reflects the balance between local and global interests: the single-step reward expression is: ; in, For single-step rewards, To increase local profits, This is the first weighting coefficient, with a value of 0.4. To increase overall profits, This is the second weighting coefficient, with a value of 0.5. As a penalty item, This is the third weighting coefficient, with a value of 0.1.

[0066] Penalties include path overlap penalties. Punishment for vicious price competition The sum of all additions.

[0067] Path overlap penalty: .

[0068] in Agent and Path overlap length, Agent Total path length, The conflict weight is 0.8 for adjacent agents and 0.2 for non-adjacent agents.

[0069] Penalties for Malicious Price Competition: , in Provide quotes for two agents. =15 yuan / ton is the allowable price difference threshold. =5 km is the geographical distance threshold. This is an indicator function. A penalty is triggered when the distance between two agents is less than 5km and the price difference exceeds 15 yuan.

[0070] This design prevents agents from sacrificing overall efficiency in pursuit of maximizing local profits (such as malicious bidding to snatch resources).

[0071] Centralized Collaborative Model: Construct a centralized cyclical collection path and use an improved Shapley value method for global benefit distribution.

[0072] Traditional Shapley value:

[0073] The allocation value after introducing the correction factor:

[0074] The correction term is calculated as follows: Resource input adjustment: , For the first Resource input To average input, =0.4 Risk-taking revision: , It is a risk coefficient (weighted by transportation risk, market risk, and natural risk). =0.35 Adjustment of bargaining power: , This is a bargaining power index. =0.25 Normalization constraints: (Total revenue of the major leagues), if the total is not equal after adjustment, it will be scaled proportionally.

[0075] The S600 includes dynamic resource scheduling and contract execution. Based on the real-time functional status level and benefit allocation weight of the collection nodes, the resources corresponding to the planned collection path are dynamically allocated to the collection nodes with matching functional status; and smart contracts are generated for phased execution.

[0076] The smart contract incorporates a three-layered protection mechanism: a bottom layer (ensuring farmers' minimum income and stabilizing supply expectations), an incentive layer (distributing excess profits based on contribution to incentivize quality improvement), and a constraint layer (automatic execution triggered by IoT data to prevent default). Contract terms are stored on a blockchain using a consortium blockchain architecture (such as Hyperledger Fabric). Participating nodes include: storage stations (endorsing nodes), end users (ordering nodes), and regulatory authorities (anchoring nodes). Contract data is stored in a private data collection; only transaction hashes are recorded on the blockchain to ensure privacy, immutability, and automatic execution.

[0077] A biomass decentralized collection, storage and transportation benefit equilibrium system based on multi-agent collaboration and dynamic game theory includes: a data acquisition module, used to acquire static attribute data, dynamic operation data and geographical distribution and resource quantity information of each collection and storage node; The mountain adaptation assessment module is used to calculate the weighted topological centrality based on the static attribute data of each storage node, and to evaluate the real-time functional status level and benefit allocation weight of the node based on dynamic operation data. The task mode parsing module is used to determine the decoupling capability of tasks based on the dispersion of geographical distribution, the balance of resource quantity, static attribute data, and the differences in the demands of stakeholders. The path planning module is used to execute a multi-agent parallel collaboration mode or a centralized interest coordination mode based on the decoupling determination result, so as to realize the path planning of resource points. The profit distribution module is used to dynamically allocate resources and determine the profit-sharing ratio based on the real-time functional status level and the game equilibrium solution.

[0078] The mountain adaptability assessment module includes: The weighted topological centrality calculation submodule is used to calculate the weighted topological centrality based on the static attribute data of each storage node; The real-time functional status level calculation submodule is used to calculate the real-time functional status level of nodes based on dynamic operational data. The weight allocation submodule is used to allocate node weights based on weighted topological centrality and dynamic operational data.

[0079] The task mode parsing module includes: The geographic distribution dispersion calculation submodule is used to calculate the geographic distribution dispersion based on the geographic distribution of resource points; The resource balance calculation submodule is used to calculate the resource balance based on the resource quantity of resource points. The interest claim difference calculation submodule is used to calculate the interest claim difference among stakeholders. The task decoupling determination submodule is used to determine the decoupling capability of a task based on the dispersion of geographical distribution, the balance of resource quantity, static attribute data, and the differences in the demands of stakeholders. The determination result is either a highly decoupling task or a low decoupling task; a multi-agent parallel collaboration mode or a centralized interest coordination mode. The path planning module includes a multi-agent parallel collaboration mode submodule and a centralized interest coordination mode submodule. The multi-agent parallel collaboration mode submodule is used to execute the multi-agent parallel collaboration mode when the result is a highly decoupling task; the centralized interest coordination mode submodule is used to execute the centralized interest coordination mode when the result is a low decoupling task.

[0080] The MARL path planning module further includes: an observation encoding unit (encoding geographic, topographic, and market information into state vectors), a communication coordination unit (realizing message passing and intent sharing among agents), a policy network unit (outputting path and pricing decisions using an Actor-Critic architecture), and a value decomposition unit (using the QMIX algorithm to solve the multi-agent credit allocation problem).

[0081] Application Case 1: Rice Straw Collection, Storage, and Transportation in the Hilly Area on the Edge of the Sichuan Basin Setting: A county with a jurisdiction area of ​​1200 km² 2 The terrain is mainly hilly (average slope 12°), encompassing 15 townships and 180 administrative villages. There are approximately 2,400 rice straw resource points (average planting area of ​​3-5 mu per household), with an annual rice straw collection capacity of about 80,000 tons. There are currently 5 collection and storage nodes (S1-S5), and the end users are 1 biomass briquette fuel processing plant (annual demand of 60,000 tons) and 2 biomass gasification power plants (combined annual demand of 50,000 tons).

[0082] Step S100: Data Acquisition Static attributes: S1 coordinate (0,0), storage capacity 800 tons, average slope around 8°; S2 (15,12), storage capacity 600 tons, slope 15°; S3 (-8,-10), storage capacity 500 tons, slope 18°; S4 (20,18), storage capacity 400 tons, slope 6°; S5 (-5,15), storage capacity 350 tons, slope 20°.

[0083] Dynamic data: S1 currently occupies 320 tons, with 4 tasks in the current task queue, and the historical average profit sharing ratio is 45:25:30 for farmers: middlemen: storage stations; S2 occupies 480 tons, with 7 tasks in the task queue, and a historical dispute rate of 12% (negotiation conflicts).

[0084] Resource points: Through remote sensing identification and village-level reporting, the GPS coordinates, estimated rice straw volume (400 kg per mu), and current holders (farmers / small vendors) of 2,400 resource points are obtained.

[0085] Step S200-S300 Mountain Adaptability Assessment: Calculate the weighted topological centrality (weights α=0.5, β=0.3, γ=0.2): S1: =0.85, =0.2 (8° normalized), =0.9 (15km from the terminal), =0.5×0.85+0.3×(1__0.2)+0.2×0.9=0.845 S2: =0.72, =0.5, =0.7, =0.665 S3: =0.65, =0.6, =0.6, =0.595 S4: =0.68, =0.15, =0.85, =0.695 S5: =0.55, =0.75, =0.5, =0.525 Real-time functional status levels: S1 (Level 1, Strategic Hub), S4 (Level 2, Regional), S2 (Level 3, Local Node), S3, S5 (Level 4, Supplementary Site).

[0086] Step S400: Task decoupling determination: Geographic dispersion (standard deviation of road network distance): Taking 2400 resource points as an example, calculate the shortest road network distance from each resource point to the nearest storage node (S1-S5). (Equivalent distance based on Dijkstra's algorithm, considering slope correction factor): Distance distribution: mean =3.8km, standard deviation km, maximum distance km.

[0087] Threshold setting basis: Based on service radius level classification, the upper limit of the service radius for level 2 nodes is 25km, and for level 3 nodes it is 15km. Taking 1 / 4 of the service radius of level 3 nodes as the dispersion threshold: 15km × 0.23 ≈ 3.5km (the empirical coefficient 0.23 is derived from historical scheduling data: when...) When the radius is less than 3.5km, the probability that the load balancing coefficient (CV) of each agent after sub-cluster partitioning is less than 0.4 is greater than 85%.

[0088] determination: =4.2km>3.5 km, but considering the resource balance degree CV=0.35 and the road network density 1.8, it is still judged that the high decoupling condition is met. Resource balance (coefficient of variation, CV): 0.35 (threshold 0.4, satisfied); Road network connectivity (density): 1.8 km / km 2 (Threshold 1.5, satisfied); Differences in interests : 0.18 (threshold 0.25, satisfied).

[0089] Judgment result: Highly decoupled task, initiate multi-agent parallel cooperative mode.

[0090] Step S500: Path planning and benefit distribution: Sub-cluster partitioning: The DBSCAN algorithm (neighborhood radius 3km, minimum number of points 5) was used to divide the 2400 resource points into 6 sub-clusters: C1 (420 points, 1680 tons), C2 (380 points, 1520 tons), C3 (510 points, 2040 tons), C4 (290 points, 1160 tons), C5 (450 points, 1800 tons), and C6 (350 points, 1400 tons).

[0091] Agent configuration: One collection agent is deployed in each sub-cluster (6 in total). The agent action space includes: path node selection, local purchase price setting ([160,220] yuan / ton range), and collaboration request sending.

[0092] Training process: The central Critic network collects global state (agent location, load, and price) and outputs a global value estimate; each AgentActor network outputs action probabilities based on local observations (resource distribution within the sub-cluster, terrain slope, and competitor prices). The training converges after 10,000 episodes.

[0093] Execution result: C1-Agent path: P1-P5-P12-...-M1 (integration point), path length 28km, collection time 4.2h, local price 198 yuan / ton; C2-Agent and C3-Agent detected a price conflict in the border area (C2 quoted 195 yuan, C3 quoted 205 yuan). Through message communication and coordination, C2 focused on the northern area and C3 focused on the southern area to avoid malicious bidding. Each agent executes in parallel, reducing the total collection cycle from 14 days in the traditional model to 6 days.

[0094] Step S600 Smart Contract Execution: Generate a smart contract for the C1 cluster: a guaranteed minimum price of 160 yuan / ton (subsidies will be automatically triggered if the price is lower than this). The excess profit (terminal price 380 - purchase cost 198 - transportation cost 45 - other costs 25 = 112 yuan / ton) is distributed according to contribution: farmers 60% (67.2 yuan), middlemen 25% (28 yuan), and storage stations 15% (16.8 yuan).

[0095] IoT trigger: When the GPS track shows that the transport vehicle has arrived at the integration point M1 and the weighbridge data is uploaded to the blockchain, the payment instruction is automatically executed.

[0096] Application Case 2: Low Decoupling Task Scenarios (Severe Conflicts of Interest) Scenario change: If, due to market price fluctuations, farmers expect the guaranteed price to rise to 190 yuan / ton, while the highest purchase price that intermediaries can afford is only 185 yuan / ton, the difference in interests is ΔU=0.32 (exceeding the threshold of 0.25).

[0097] Mode switching: If the task is determined to be of low decoupling potential, switch to a centralized interest coordination mode. Led by end users (biomass power plants), a unified purchase price of 190 yuan / ton was set for the entire county, with a fixed commission of 25 yuan / ton for intermediaries and a direct subsidy of 5 yuan / ton for farmers from the power plants. An improved Shapley value method was used to allocate the total surplus in the supply chain: considering farmers' resource input (weight 0.4), intermediary transportation risk (weight 0.35), and power plant scale effect (weight 0.25), the final allocation ratio was farmers:intermediaries:power plants = 48:27:25; a centralized TSP algorithm was used for route planning, aiming to minimize total transportation cost, and five large transport vehicles were uniformly dispatched for collection. A comparison with the traditional fixed-price acquisition model is shown in Table 2. Table 2: Comparison of the present invention with the traditional fixed-price purchase model ; Application Case 1 and Application Case 2 are merely examples; in actual applications, the parameters can be adjusted according to the specific regional characteristics. This invention achieves a win-win situation and efficient operation for the collection, storage, and transportation of biomass in mountainous and hilly areas through the deep integration of multi-agent collaboration and dynamic game theory.

[0098] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A collaborative scheduling method for biomass collection, storage, and transportation in mountainous and hilly areas based on multi-agent game theory, characterized in that, Includes the following steps: Obtain static attribute data, dynamic operation data, and the geographical distribution and resource quantity of all resource points of biomass to be collected and transported within the mountainous and hilly network. Based on the static attribute data of each storage node, the weighted topological centrality index of each storage node is calculated. By combining weighted topological centrality index and dynamic operational data, the real-time functional status level and benefit distribution weight of each storage node are evaluated. Determine the dispersion of the geographical distribution of all resource points, the balance of resource quantity, and the differences in the demands of stakeholders; Using dispersion, balance, static attribute data of each storage node, and demand differences as judgment indicators, the collaboration mode of stakeholders is determined according to the preset judgment mechanism, and collaboration planning is executed to complete the dynamic planning of the collection path and its corresponding resources; among them, the collaboration mode includes multi-agent parallel collaboration mode and centralized collaboration mode. Based on the real-time functional status level and benefit allocation weight of each collection and storage node, the resources corresponding to the planned collection path are dynamically allocated to the collection and storage nodes with matching functional status.

2. The method according to claim 1, characterized in that: Static attribute data includes geographic coordinates, total storage capacity, terrain slope grade, and topological connections in the transportation network; dynamic operational data includes real-time storage capacity, task load, historical revenue sharing ratio, and current bargaining power index; resource points include farmers, intermediaries, and primary processing points.

3. The method according to claim 2, characterized in that: The formula for calculating the weighted topological centrality index of each storage node is as follows: ; in: The weighted topological centrality index for each storage node. To achieve proximity centrality based on road network distance, This represents the normalized value of the average terrain slope around the node. The distance to the end market is the inverse of the equivalent distance; α, β, γ are weighting coefficients, and satisfy α+β+γ=1.

4. The method according to claim 3, characterized in that: The preset judgment mechanism is as follows: When the dispersion, balance, static attribute data of each storage node, and demand difference all simultaneously meet their respective preset parallel operation thresholds, the current task is considered to have high decoupling capability, and the multi-agent parallel collaboration mode is activated; conversely, if any indicator fails to reach the corresponding threshold, the task is considered to have low decoupling capability, and the centralized collaboration mode is activated.

5. The method according to claim 4, characterized in that: In the multi-agent parallel collaboration mode, all resource points are divided into multiple non-overlapping sub-clusters according to the principle of geographical proximity, and an independent agent is assigned to each sub-cluster. Each agent, based on the MARL framework, learns the optimal collection path and the benefit distribution scheme within the sub-cluster through centralized training and distributed execution.

6. The method according to claim 5, characterized in that: The MARL framework uses the CTDE architecture, which includes: Centralized training phase: The central Critic network acquires global state information, including the location, load, path planning, and revenue distribution status of all agents, and calculates the global advantage function; In the distributed execution phase, each agent makes independent decisions based on local observations, including the distribution of resource points within the corresponding sub-cluster, terrain conditions, and local market prices. The agent's action space includes path selection, acquisition pricing, and sending collaboration requests to other agents. Reward function design: The single-step reward equals the first weight coefficient multiplied by the local profit increment, plus the second weight coefficient multiplied by the global profit increment, and then minus the third weight coefficient multiplied by the penalty term.

7. The method according to claim 6, characterized in that: In the centralized collaborative mode, a centralized cyclic collection path is uniformly constructed for all resource points, and the improved Shapley value method or Nash negotiation model is used for scheduling and allocation on a global scale.

8. A benefit equilibrium system for decentralized collection, storage, and transportation of biomass based on multi-agent cooperation and dynamic game theory, characterized in that, For performing the method according to any one of claims 1 to 7, comprising: The data acquisition module is used to acquire static attribute data, dynamic operational data, and geographical distribution and resource quantity information of each storage node; The mountain adaptation assessment module is used to calculate the weighted topological centrality based on the static attribute data of each storage node, and to evaluate the real-time functional status level and benefit allocation weight of the node based on dynamic operation data. The task mode parsing module is used to determine the decoupling capability of tasks based on the dispersion of geographical distribution, the balance of resource quantity, static attribute data, and the differences in the demands of stakeholders. The path planning module is used to execute a multi-agent parallel collaboration mode or a centralized interest coordination mode based on the decoupling determination result, so as to realize the path planning of resource points. The profit distribution module is used to dynamically allocate resources and determine the profit-sharing ratio based on the real-time functional status level and the game equilibrium solution.

9. The system according to claim 8, characterized in that, The mountain adaptability assessment module includes: The weighted topological centrality calculation submodule is used to calculate the weighted topological centrality based on the static attribute data of each storage node; The real-time functional status level calculation submodule is used to calculate the real-time functional status level of nodes based on dynamic operational data. The weight allocation submodule is used to allocate node weights based on weighted topological centrality and dynamic operational data.

10. The system according to claim 9, characterized in that, The task mode parsing module includes: The geographic distribution dispersion calculation submodule is used to calculate the geographic distribution dispersion based on the geographic distribution of resource points; The resource balance calculation submodule is used to calculate the resource balance based on the resource quantity of resource points. The interest claim difference calculation submodule is used to calculate the interest claim difference among stakeholders. The task decoupling determination submodule is used to determine the decoupling capability of a task based on the dispersion of geographical distribution, the balance of resource quantity, static attribute data, and the differences in the demands of stakeholders. The determination result is either a highly decoupling task or a low decoupling task; a multi-agent parallel collaboration mode or a centralized interest coordination mode. The path planning module includes a multi-agent parallel collaboration mode submodule and a centralized interest coordination mode submodule. The multi-agent parallel collaboration mode submodule is used to execute the multi-agent parallel collaboration mode when the result is a highly decoupling task; the centralized interest coordination mode submodule is used to execute the centralized interest coordination mode when the result is a low decoupling task.