Industrial chain demand matching method

By modeling the demander and provider of the industrial chain into two-part graphs, and using graph neural network and attention mechanism for dynamic matching, the problem of traditional methods being unable to respond to market changes and inefficiency in real time is solved, and efficient and accurate demand matching is achieved.

CN120146505APending Publication Date: 2025-06-13SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510247239.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The traditional industrial chain demand matching method cannot adapt to market changes in real time, resulting in lagging matching strategies and being inefficient when processing massive supply and demand data, making it difficult to meet the needs of immediate decision-making.

Method used

The demander and providers in the industrial chain are modeled as two-part graphs, and the nodes and edges are extracted and iteratively updated through the graph neural network to generate optimal nodes and edge embedding vectors, and the candidate actions are selected in combination with the attention mechanism, and the optimal matching actions are output through the policy network.

Benefits of technology

It realizes dynamic adjustment of the weight of supply and demand relationships, responds to market changes in real time, improves the accuracy and efficiency of matching, reduces the computational complexity, and enhances the dynamic adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146505A_ABST
    Figure CN120146505A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial chain demand matching method, which is applied to the technical field of intelligent optimization and resource scheduling, and comprises the following steps: modeling a demand side and a provider in an industrial chain into a bipartite graph; respectively carrying out feature extraction on nodes and edges of the bipartite graph to obtain a node embedding vector and an edge embedding vector; iteratively updating the node embedding vector and the edge embedding vector until the node embedding vector and the edge embedding vector converge to a preset threshold value, and generating an optimal node embedding vector and an optimal edge embedding vector; forming a current environment state based on the optimal node embedding vector and the optimal edge embedding vector; screening a candidate action set according to a current environment state and an attention mechanism, and outputting an optimal matching action through a strategy network; according to the method, the real-time performance, fairness and economy of industrial chain resource matching are remarkably improved, and the method has wide practical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent optimization and resource scheduling, and particularly to a method for matching the demands of an industrial chain. Background Art

[0002] With the accelerating advancement of global economic integration, the scale and complexity of modern industrial chains have reached unprecedented heights. The rapid changes in market demands, the significant shortening of product life cycles, and the increasing number of supply chain links have imposed higher requirements on the accuracy and response speed of resource supply-demand matching in industrial chains. Traditional demand matching models, such as static methods based on the Hungarian algorithm, have shown obvious limitations in dealing with these dynamic changes. They are unable to adapt to market fluctuations such as price changes and resource quantity increases or decreases in real time, resulting in matching strategies lagging behind the changes in actual demands. In addition, in the face of a large amount of supply-demand data, the time complexity of traditional algorithms is high, and they are inefficient in dealing with the global optimal matching problem of nodes at the ten-thousand level, making it difficult to meet the needs of immediate decision-making.

[0003] To overcome these defects, the present application proposes a method for matching the demands of an industrial chain. Summary of the Invention

[0004] The purpose of the present application is to provide a method for matching the demands of an industrial chain, aiming to solve the above problems.

[0005] To achieve the above purpose, the present application provides the following technical solutions:

[0006] The present application provides a method for matching the demands of an industrial chain, including:

[0007] Modeling the demand side and the supply side in the industrial chain as a bipartite graph;

[0008] Performing feature extraction on the nodes and edges of the bipartite graph respectively to obtain node embedding vectors and edge embedding vectors;

[0009] Iteratively updating the node embedding vectors and the edge embedding vectors until convergence to a preset threshold to generate optimal node embedding vectors and optimal edge embedding vectors;

[0010] Forming the current environmental state based on the optimal node embedding vectors and optimal edge embedding vectors;

[0011] Screening a candidate action set according to the current environmental state and the attention mechanism, and outputting an optimal matching action through a policy network.

[0012] The present application provides a method for matching the demands of an industrial chain, which has the following beneficial effects:

[0013] (1) This application ingeniously transforms the complex supply - demand relationship in the industrial chain into a bipartite graph. The edges are determined based on the supply - demand quantity and price conditions, which is then transformed into a maximum - matching problem of the bipartite graph. Considering the resource quantity and price matching degree comprehensively, it provides a solid mathematical foundation for the matching algorithm. It adopts a dynamic bipartite - graph modeling and weight - quantization mechanism to update the weights of the supply - demand relationship in real - time to adapt to market changes. Through this mechanism, the weights of the supply - demand relationship can be dynamically adjusted to ensure that the matching strategy is always consistent with the current market conditions, thereby improving the accuracy and efficiency of matching. This dynamic adjustment ability is not available in traditional static matching methods and can better cope with market uncertainties and fluctuations.

[0014] (2) Use the message - passing mechanism of the graph neural network (GNN) to extract features of the bipartite - graph nodes, effectively mine the hidden information in the graph, and generate embedding vectors that comprehensively represent the graph - structure features and node attributes.

[0015] (3) By introducing attention - driven action pruning, the number of candidate actions is reduced, and the computational complexity is lowered. Through the attention mechanism, key candidate actions are automatically selected, and unimportant actions are ignored, thus reducing the number of candidate actions, lowering the computational complexity, and improving the decision - making efficiency. This action - pruning method enables the model to focus more on key decisions, improving the accuracy and speed of matching.

[0016] (4) Design a multi - objective reward function to guide the global optimal solution, which can better balance exploration and exploitation during the training process, and finally find the global - optimal matching strategy to maximize the overall benefits of the industrial chain. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic flowchart of a method for matching industrial - chain demands in Embodiment 1 of this application;

[0018] Figure 2 It is a flowchart of step S5 in Embodiment 1 of this application;

[0019] Figure 3 It is a schematic diagram of the experience replay buffer in Embodiment 1 of this application;

[0020] Figure 4 It is a schematic diagram of the structure of the policy network in Embodiment 1 of this application;

[0021] Figure 5 It is a schematic flowchart of the alternating training of the policy network and the value network in Embodiment 1 of this application;

[0022] Figure 6 It is a flowchart of a method for matching industrial - chain demands in Embodiment 1 of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] It should be understood that the specific embodiments described herein are merely for explaining the present application and are not intended to limit the present application.

[0024] The following analyzes the solutions in the prior art in combination with relevant technologies.

[0025] With the development of big data and artificial intelligence technologies, graph neural networks (GNNs) and reinforcement learning (RL) provide new ideas and tools. GNNs can effectively process graph-structured data, capture the complex relationships between the supply and demand sides, and improve the quality of matching decisions through embedded representation learning. RL, on the other hand, provides a flexible framework that enables the system to optimize strategies through trial-and-error learning in an uncertain environment, achieving balanced resource allocation from local to global. This combination can not only significantly improve computational efficiency but also enhance the system's dynamic adaptability, enabling it to perceive and respond to market changes in real time and ensuring the effective allocation of resources. By integrating the advantages of GNNs and RL, a supply-demand matching system with a global perspective can be constructed.

[0026] Application No. 202310420383.7 discloses a bipartite graph matching task offloading method based on social perception under cache constraints, focusing on task offloading in mobile edge computing networks. First, it calculates the unit popularity of tasks of remote devices and sorts the cache, then constructs a weighted bipartite graph based on device social relationships and D2D links, and uses the KM matching algorithm to complete the matching of tasks and device nodes. Application No. 202210851325.5 discloses a user access control method based on bipartite graph matching in mobile edge computing networks. For mobile edge computing networks, it establishes an optimization problem model for user access control and transforms it into a bipartite graph matching problem, and solves it using the Hungarian algorithm. By optimizing user device access, data upload volume, and channel allocation, it maximizes the number of user devices that can complete computing tasks before the deadline.

[0027] Although traditional algorithms such as the Hungarian algorithm and the Hopcroft-Karp algorithm can obtain theoretically optimal solutions in static bipartite graphs, their core defect lies in adopting a static modeling mechanism. This means that they cannot respond in real time to changes in a dynamic environment, such as increases or decreases in supply and demand resources, fluctuations in price thresholds, etc. Since these algorithms rely on a fixed input data set for one-time calculation, once external conditions change, the original matching strategy may lag or even completely fail, resulting in unreasonable resource allocation or low efficiency. The present application aims to enable the intelligent agent to adjust the matching strategy in real time according to the changing market demands and industrial chain environment by utilizing the dynamic learning characteristics of reinforcement learning, thereby significantly improving the timeliness and accuracy of demand matching.

[0028] In addition, for large-scale application scenarios, the time complexity of traditional algorithms is relatively high. Especially when dealing with a large number of nodes, the action space grows exponentially with the number of nodes, which directly limits its feasibility in practical applications. When using reinforcement learning (RL) to solve this problem, due to the need to process a huge state space, the training efficiency is low and it is difficult to converge, making it impractical to directly apply RL methods.

[0029] Existing matching methods often focus on maximizing the single-step immediate reward, such as the revenue brought by a single match, while ignoring the importance of global goals, such as the balanced allocation of total resources and the fairness of price differences. This method is prone to falling into local optimal solutions, that is, although certain specific matches seem beneficial, they may not be the best choice from an overall perspective. Therefore, it is crucial to develop a matching strategy that can balance local interests and global benefits to avoid uneven overall resource allocation caused by excessive focus on short-term gains.

[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.

[0031] Embodiment 1

[0032] Please refer to Figure 1 , which is a schematic flowchart of a method for matching industrial chain demands in Embodiment 1 of the present application; the steps include:

[0033] S1: Model the demand side and the supply side in the industrial chain as a bipartite graph.

[0034] In this embodiment, the industrial chain is abstracted as a bipartite graph, which contains two independent node sets: the resource supply side and the demand side. Specifically, it is divided into resource demand sheets and resource supply sheets, which respectively write the quantity of resources needed and provided, the highest price acceptable to the resource demand side, and the lowest price acceptable to the resource supply side. In actual operation, it is necessary to build multi-dimensional data collection channels. Through online platforms, such as specially developed industrial chain supply and demand information interaction websites, or by using existing industry vertical platforms, a large amount of data can be quickly collected. At the same time, the collected data needs to be preliminarily sorted and cleaned to remove duplicate, incorrect or inconsistent information, providing a high-quality data basis for subsequent steps such as bipartite graph modeling.

[0035] Each node represents an entity in the industrial chain, and each edge represents potential transaction possibilities or benefit values. The bipartite graph not only captures the basic connection relationships among entities in the industrial chain but also reflects the transaction possibilities or benefits through edge weights, transforming the problem into finding the maximum matching problem. Specifically, the set of demand-side vertices U = {u 1 , u 2 , …, u n} has attributes d i which is the resource demand quantity of the demand-side node, and which is the maximum acceptable price of the demand-side node. The set of supplier-side vertices V = {v 1 , v 2 , …, v m} has attributes s j which is the resource supply quantity of the supplier-side node, and which is the minimum acceptable price of the supplier-side node.

[0036] Among them, the weight w ij of the edge is expressed as:

[0037]

[0038] In the above formula, is the quantification of the matching degree of supply and demand quantities; is the quantification of the overlapping degree of price ranges; α, β ∈ [0, 1] are hyperparameters that respectively control the weights of the matching degree of supply and demand quantities and the price overlapping degree. When a new demand order arrives, a new demand node u n+1 is added, and the edge weights between it and the supplier-side nodes are calculated. If the resource supply quantity s j of the supplier v j decreases, the weight of the associated edge e ij is updated.

[0039] S2: Feature extraction is respectively performed on the nodes and edges of the bipartite graph to obtain node embedding vectors and edge embedding vectors.

[0040] In this embodiment, according to the demand order and supplier order information, feature vectors are respectively constructed for the demand-side nodes and the supplier-side nodes, including resource quantity and price information. First, the demand-side nodes u ∈ U and the supplier-side nodes v ∈ V are initialized to obtain the node embedding vectors and Each edge (u, v) is initialized to obtain the edge embedding vector

[0041] The initialization method is a linear transformation based on node features, and the formula is expressed as:

[0042]

[0043] Among them, x u and x v are the feature vectors of the demand-side node u and the supply-side node v respectively; w ij is the weight of the corresponding edge; Linear is a linear transformation.

[0044] S3: Iteratively update the node embedding vector and the edge embedding vector until convergence to a preset threshold, generating the optimal node embedding vector and the optimal edge embedding vector.

[0045] In this embodiment, in order to more precisely understand and predict the interaction patterns among entities in the industrial chain, a graph neural network (GNN) is adopted. The GNN can iteratively update the representation vector of each node according to the node and its neighbor information, thereby generating an embedding vector that can not only reflect the characteristics of the node itself but also reflect the information of its local network structure. In this process, each node collects information from its direct neighbors and updates its own state representation by combining its own characteristics. This mechanism makes the GNN particularly suitable for processing data sets with complex topological structures, such as the resource supply and demand network in the industrial chain.

[0046] During the message passing process, each node updates the node embedding vector and the edge embedding vector according to the obtained neighbor information; the message passing is expressed as:

[0047]

[0048] Among them, and are the neighbor node sets of the demand-side node u and the supply-side node v respectively, and φ is a non-linear message function; and are the message vectors received by the demand-side node u and the supply-side node v at the k-th layer respectively; and are the node embedding vectors of the demand-side node u and the supply-side node v at the (k - 1)-th layer respectively; is the edge embedding vector of the edge between the demand-side node u and the supply-side node v at the (k - 1)-th layer; k is the number of layers of message passing.

[0049] Aggregate the obtained neighbor information, and the updated node embedding vector and the edge embedding vector are respectively expressed as:

[0050]

[0051] Among them, ψ is a non-linear aggregation function; χ is a non-linear update function that takes the node, the updated message, and the previous-round embedding vector of the edge as inputs and outputs the updated edge feature embedding vector, enabling the edge features to be continuously optimized as the node states are updated.

[0052] Repeat the message passing and aggregation steps. When the change in the embedding vector is less than the preset constant ∈, it is determined that the training has converged.

[0053] S4: Based on the optimal node embedding vector and the optimal edge embedding vector, form the current environmental state.

[0054] In this embodiment, by fusing the features of the optimal node embedding vector h u , h v and the optimal edge embedding vector w uv , the state encoding formula is obtained:

[0055]

[0056] Among them, M t ∈ [0, 1] is the matching state matrix of the matching situation at time t, and the matrix size is ∣U∣×∣V∣; CONCAT is used to concatenate several vectors and matrices in a preset dimension; Flatten is used to convert multi-dimensional vectors and matrices into one-dimensional vectors and remove their multi-dimensional structure information.

[0057] S5: According to the current environmental state and the attention mechanism, screen the candidate action set, and output the optimal matching action through the policy network.

[0058] In this embodiment, each action is defined as a triple (d i , s j , q ij ), indicating that the demander d i is matched with the provider s j , and the resource allocation amount q ij is allocated; among them, the calculation formula for the resource allocation amount is and are respectively the remaining quantity of resource demand and the remaining quantity of resource supply. By calculating the minimum value of the remaining quantities of both parties, over-allocation is avoided.

[0059] Through the action mask mechanism, the probability of invalid actions is set to zero, and the probability distribution is re-normalized based on the Softmax function. If the remaining quantity of the demander d i is 0, then all actions containing d i are masked; if the remaining quantity of the provider s j is 0, then all actions containing s jActions; If multiple actions are selected simultaneously, it is necessary to ensure that the sum of the allocated quantities does not exceed the provider's inventory or the demander's remaining quantity.

[0060] To reduce the computational complexity and improve the decision-making quality, before selecting an action, a candidate action set is filtered through an attention mechanism, which not only improves the decision-making speed but also enhances the decision-making accuracy. To achieve attention-driven action pruning, it is necessary to combine a cross-graph attention mechanism with a hierarchical screening strategy to dynamically filter out low-value matching pairs, thereby reducing the number of candidate actions and improving the decision-making efficiency.

[0061] Specifically, calculate the attention score score(d i and the provider node s j respectively, and the attention weight a i , s j ), and the formula is expressed as: ij

[0062]

[0063] where LeakyReLU is the activation function; W is the learnable weight matrix; is the Softmax function;

[0064] Narrow the candidate action set through rough screening: The goal of rough screening is to quickly filter out obviously unreasonable matching pairs. Considering price compatibility and inventory feasibility, the highest price of the demander ≥ the lowest price of the provider or the remaining inventory of the provider ≥ 10% of the remaining demand of the demander. Then, perform fine screening according to the attention weight to screen out the key candidate actions and reduce the number of actions. Sort by score and retain the top K actions with the highest scores; dynamically adjust K according to the preset supply and demand tightness. When the inventory is sufficient, reduce K to reduce the computational amount, and vice versa, increase K to avoid missing high-quality matches.

[0065] In the policy network, design a multi-objective reward function, calculate the reward according to the action result, and give a high reward if the decision reduces the overall cost of the industrial chain, increases the profit or meets more market demands, and give a low reward otherwise. The multi-objective reward function includes: local revenue r profit , demand satisfaction rate r satisfaction , resource balance reward r fairness , price fairness reward r price_fairness .

[0066] Local revenue: The reward for maximizing profit is r profit = ∑ i,j a ij ·(p d - p s ), where p d is the highest price of the demander, p sFor the provider's lowest price, a ij is the attention weight.

[0067] The reward for demand satisfaction rate is the proportion of the matching quantity to the total demand.

[0068] Global benefit: The resource balance reward measures the distribution balance among suppliers through the Gini coefficient, that is, r fairness = -Gini(supplier allocation quantity).

[0069] The price fairness reward measures the fairness of price differences through the standard deviation, determined by r price_fairness = -std(matching price difference).

[0070] In summary, the multi-objective reward function is expressed as:

[0071] r t = αr profit + βr satisfaction + γr fairness + δr price_fairness ,

[0072] where the weight coefficients α, β, γ, δ are dynamically adjusted through learning.

[0073] Please refer to Figure 3 , which is a schematic diagram of the experience replay buffer in Embodiment 1 of this application.

[0074] The experience replay buffer is used to store interaction experiences and support offline learning and data reuse. The state, action, reward, and next state of each interaction are composed into an experience tuple (s t , a t , r t , s t+1 ), which is stored in the experience replay buffer. Uniform sampling is used in the initial training stage to ensure data diversity. Priority sampling is adopted in the later stage of training to focus on learning high-error samples and accelerate convergence. When the buffer is full, new experiences overwrite old experiences, high-priority samples are preferentially retained, and low-priority samples are regularly cleared to reduce storage overhead. Priority sampling is determined by p i = |δ i | + ∈ represents the priority of the i-th sample, which is used to measure the importance of the sample for the agent's learning. The higher the priority of the sample, the greater the probability of being sampled; δ i is the temporal difference error, and η, ∈ are hyperparameters.

[0075] Please refer to Figure 4, which is a schematic structural diagram of the policy network in Embodiment 1 of this application. Based on the policy network, the probability distribution of candidate actions is output according to the current environmental state, and actions are selected according to the probability distribution. The probability distribution is expressed as:

[0076] π θ (a t |s t ) = Softmax(MLP policy (s t ))

[0077] Among them, the policy network π θ is a multi-layer perceptron MLP; Softmax converts the output of the MLP into a probability distribution, so that the output value of each category can be interpreted as a probability.

[0078] The policy network is used to map the state to the action probability distribution and needs to balance exploration and exploitation. Sample data from the experience replay buffer, use the policy gradient algorithm, and combine the advantage function to update the policy network parameters to maximize the expected cumulative reward.

[0079] Among them, the objective function is the clipped surrogate objective function, whose purpose is to limit the difference between the new policy and the old policy when updating the policy network parameters, and avoid unstable performance caused by too large an update amplitude. It is expressed as:

[0080]

[0081] In the above formula, is the importance sampling ratio, which is used to balance the difference between the new policy and the old policy. π θ (a t ∣s t ) is the probability that the new policy selects action a t under state s t ; is the probability that the old policy selects the same action under the same state; clip(r t (θ), 1 - ∈, 1 + ∈) is the clipping operation, which clips the value of r t (θ) to the interval (1 - ∈, 1 + ∈); E t is the expectation operator, which means taking the expectation of the data after time step t; is the advantage function, γ is the discount factor, which is used to measure the importance of future rewards. The value range is generally between [0, 1]. The closer the value is to 0, the more it focuses on immediate rewards, and the closer it is to 1, the higher the degree of emphasis on future rewards; λ is a parameter used to control the contribution of different time steps in the advantage function estimation; l is the offset of the time step, which is used to traverse and sum different time steps when calculating the advantage function; δ t = r t+γV(s t+1 ) - V(s t ), V(s t ) is the state value function, representing the expected cumulative reward that can be obtained in the future under state s t ;

[0082] In the objective function, a policy entropy term is added to measure the uncertainty of the policy. The larger the entropy value, the stronger the randomness of the policy; the smaller the entropy value, the more certain the policy. In the update of the policy network, an entropy regularization term is introduced to encourage the policy to maintain a certain degree of exploration. The formula is expressed as:

[0083]

[0084] In the above formula, L total is the total loss function, which is used to comprehensively measure the loss during the update of the policy network, combining the clipped surrogate objective function and the policy entropy term; η is the entropy regularization coefficient, which is used to control the weight of the policy entropy term in the total objective function and adjust the degree of policy exploration; is the entropy of policy π under state s t , measuring the uncertainty of the policy in the current state; is the general expression of the entropy of policy π, calculating the entropy of the probability distribution over all actions; ∑ a π(a|s t ) log π(a|s t ) represents the sum of the product of the logarithm of the probability π(a|s t ) of each action a selected by policy π under state s t ) and its own probability, which is used to calculate the policy entropy.

[0085] The calculation formula of the policy gradient algorithm is:

[0086]

[0087] In the above formula, represents the policy gradient, which is used to update the policy network parameter θ; represents the gradient of the logarithmic probability of policy π θ selecting action a t under state s t ; A(s t , a t ) represents the advantage function of taking action a t under state s t .

[0088] In addition, this application also proposes a value network for estimating the state value V(s t), guiding the optimization direction of the policy network. Specifically, the gradient descent algorithm is adopted to update the parameters of the value network by minimizing the mean square error between the value estimated by the value network and the actual value. The value network V φ The formula for estimating the state value is expressed as: V φ (s t ) = MLP value (s t ). To consider optimizing multiple objectives such as profit and fairness, a multi-headed value network is designed.

[0089] The target value is calculated by the Bootstrapping method, and the calculation formula is:

[0090]

[0091] where y t is the target value at time step t; r t is the immediate reward at time step t; γ is the discount factor; s t+1 is the state at time t + 1; is the output of the target network. To reduce the overestimation problem, the double Q-learning technique is used.

[0092] The mean square error (MSE) is used to calculate the gap between the predicted value and the target value:

[0093] L(φ) = E t [(V φ (s t ) - y t ) 2 ,

[0094] In the above formula, L(φ) is the loss function of the value network; V φ (s t ) is the output of the value network; y t is the target value at time step t.

[0095] Soft update of the target network: φ target ← τφ + (1 - τ)φ target , τ ≈ 0.005

[0096] In the above formula, φ target is the parameter of the target network; τ is the soft update coefficient.

[0097] In summary, in this Embodiment 1, the relationship between the resource providers and demanders in the industrial chain is modeled as a bipartite graph, and a graph neural network (GNN) is used to extract the features of nodes and edges. The current environmental state, which is integrated by the node embedding vectors and edge information of the bipartite graph, is used to make optimal decisions. Additionally, attention-driven action pruning is introduced to reduce the number of candidate actions, lower the computational complexity, and improve the decision-making efficiency. A multi-objective reward function is also designed, which combines profit, demand satisfaction rate, resource balance, and price fairness to guide the agent to explore the global optimal solution. By integrating the advantages of reinforcement learning and graph neural networks, real-time response and precise matching to the dynamic changes in the industrial chain demand are achieved, providing technical support for the efficient, dynamic, and balanced matching of industrial chain resource supply and demand.

[0098] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article or method comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article or method. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, apparatus, article or method comprising such element.

[0099] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

[0100] Although the embodiments of the present application have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application. The scope of the present application is defined by the appended claims and their equivalents.

[0101] Certainly, the present invention can also have other various implementation manners. Based on this implementation manner, other implementation manners obtained by those of ordinary skill in the art without any creative work belong to the scope protected by the present invention.

Claims

1. A method for matching industrial chain demand, characterized in that: include: Model the demand side and the supply side in the industrial chain as a bipartite graph; Performing feature extraction on the nodes and edges of the bipartite graph respectively to obtain a node embedding vector and an edge embedding vector; Iteratively updating the node embedding vector and the edge embedding vector until they converge to a preset threshold, thereby generating an optimal node embedding vector and an optimal edge embedding vector; Based on the optimal node embedding vector and the optimal edge embedding vector, forming a current environment state; The candidate action set is screened according to the current environment state and the attention mechanism, and the optimal matching action is output through the policy network.

2. The industrial chain demand matching method according to claim 1, characterized in that: The step of modeling the demand side and the supply side in the industrial chain as a bipartite graph specifically includes the following steps: The demand side vertex set U={u1,u2,…,u n The properties of} are d i is the resource demand quantity of the demand-side node, The highest price accepted by the demand side node; The provider vertex set V = {v1,v2,…,v m The properties of} are s j The number of resources provided by the provider node. The lowest accepted price of the provider node; The weight of the edge w ij It is expressed as: in, To quantify the matching degree of supply and demand; is to quantify the overlap of price ranges; α, β∈[0,1] are hyperparameters, which control the weights of supply-demand quantity matching and price overlap respectively; When a new demand order arrives, a new demand node u is added n+1 , calculate the edge weight between it and the provider node; if the provider v j Resource quantity j If it decreases, update the associated edge e ij The weight of .

3. The industrial chain demand matching method according to claim 2, characterized in that: The step of extracting features from the nodes and edges of the bipartite graph to obtain node embedding vectors and edge embedding vectors specifically includes the following steps: Initialize the demand side node u∈U and the provider side node v∈V to get the node embedding vector and Initialize each edge (u, v) to get the edge embedding vector The initialization method is a linear transformation based on node features, and the formula is expressed as: Among them, x u and x v are the feature vectors of demand node u and provider node v respectively; w ij is the weight of the corresponding edge; Linear is a linear transformation.

4. The industrial chain demand matching method according to claim 3 is characterized in that: The step of iteratively updating the node embedding vector and the edge embedding vector until they converge to a preset threshold and generating an optimal node embedding vector and an optimal edge embedding vector specifically includes the following steps: In the process of message passing, the node embedding vector and edge embedding vector are updated according to the obtained neighbor information; the message passing is expressed as: in, and are the neighbor node sets of the demand side node u and the provider side node v respectively, φ is the nonlinear message function; and are the message vectors received by the demand node u and the provider node v at the kth layer respectively; and are the node embedding vectors of the demand side node u and the provider node v at the k-1 layer respectively; is the edge embedding vector of the edge between the demand node u and the provider node v at the k-1 layer; k is the number of layers of message transmission; Aggregate the obtained neighbor information and update the node embedding vector and edge embedding vector Respectively expressed as: Among them, ψ is a nonlinear aggregation function; X is a nonlinear update function; Repeat message passing and aggregation, and when the change in the embedding vector is less than a preset constant ∈, the training is judged to have converged.

5. The industrial chain demand matching method according to claim 4 is characterized in that: The step of forming the current environment state based on the optimal node embedding vector and the optimal edge embedding vector specifically includes the following steps: By feature fusion, the optimal node embedding vector h u ,h v and the optimal edge embedding vector w uv , and get the state encoding formula: Among them, M t ∈0,1 is the matching state matrix of the matching situation at time t, and the matrix size is |U|×|V|; CONCAT is used to connect several vectors and matrices in a preset dimension; Flatten is used to convert multi-dimensional vectors and matrices into one-dimensional vectors.

6. The industrial chain demand matching method according to claim 5, characterized in that: The step of screening the candidate action set according to the current environment state and the attention mechanism and outputting the best matching action through the strategy network specifically includes the following steps: Each action is defined as a triple (d i ,s j ,q ij ), indicating that the demand side d i With providers j Matching, allocating resources q ij ; The calculation formula for the amount of allocated resources is: and The remaining quantity of resource requirements and the remaining quantity of resources provided are respectively; The probability of invalid actions is set to zero through the action mask mechanism, and the probability distribution is renormalized based on the Softmax function; if the demand side d i If the remaining number is 0, then all the i action; if the provider s j If the remaining number is 0, then all the j Actions; Calculate the demand side node d separately i With the provider node j The attention score score(d i ,s j ) and attention weight α ij , the formula is: Among them, LeakyReLU is the activation function; W is the learnable weight matrix; is the Softmax function; Narrow the set of candidate actions through coarse screening; perform fine screening based on attention weights to screen out key candidate actions; sort by score and retain the top K actions with the highest scores; dynamically adjust K based on the preset supply and demand tension.

7. The industrial chain demand matching method according to claim 6, characterized in that: In the strategy network, a multi-objective reward function is designed; the multi-objective reward function includes: local benefit r profit , demand satisfaction rate r satisfaction , Resource Balance Reward fairness , price fairness reward price_fairness ; The multi-objective reward function is expressed as: r t =αr profit +βr satisfaction +γr fairness +δr price_fairness , Among them, r profit =∑ i,j a ij ·(p d -p s ), p d is the highest price on the demand side, p s is the lowest price provided by the provider, a ij is the attention weight; r fairness =-Gini (supplier allocation); r price_fairness =-std(matching price difference); weight coefficients α, β, γ, δ are dynamically adjusted through learning. The policy network parameters and value network parameters are updated by the policy gradient algorithm to optimize the global matching strategy; wherein the value network is used to estimate the state value V(s t ), guiding the optimization direction of the policy network.

Citation Information

Patent Citations

  • User access control method based on bipartite graph matching in mobile edge computing network

    CN115278692A

  • Bipartite graph matching task unloading method based on social perception under cache constraint

    CN116431244A