Supply chain problem solving method and electronic equipment

By generating the first node features of the node and combining the information of the node and its neighboring nodes, the problem that the solution to the supply chain problem in the existing technology does not meet the target is solved, and a more efficient solution effect is achieved.

CN121961376APending Publication Date: 2026-05-01LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LENOVO (BEIJING) LTD
Filing Date
2026-01-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies, when solving supply chain problems, only consider the information of each graph node itself, resulting in solutions that do not meet the solution objectives.

Method used

By generating the first node characteristics of the node, and combining the variables and constraints of the node and its nearest neighbors, the variable values ​​that satisfy the constraints and the solution objective are determined.

Benefits of technology

It improves the accuracy and efficiency of solving supply chain problems, and can better meet the solution objectives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121961376A_ABST
    Figure CN121961376A_ABST
Patent Text Reader

Abstract

The invention discloses a supply chain problem solving method and electronic equipment, and the method comprises the steps: obtaining a to-be-processed map corresponding to a supply chain problem, the supply chain problem comprising variables related in a plurality of supply chain problems, at least one constraint condition limiting the variables, and a solving target of the supply chain problem, nodes contained in the to-be-processed map comprise variable nodes representing variables and constraint nodes representing constraint conditions, edges contained in the to-be-processed map represent the relationship between the variables and the constraint conditions corresponding to the nodes, and for at least one node in the to-be-processed map, determining neighbor nodes of the nodes in the to-be-processed map, the first node feature of the node is generated according to variables and / or constraint conditions represented by the node and neighbor nodes of the node, the neighbor nodes of the node comprise nodes which are communicated with the node and have the distance to the node smaller than a distance threshold value, and the first node feature of at least one node is generated according to the first node feature of the at least one node; and determining a variable value which meets the constraint condition and the solving target and corresponds to each variable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method for solving supply chain problems and an electronic device. Background Technology

[0002] In related technologies, in order to solve supply chain problems, the supply chain problem is usually converted into a corresponding graph, and the solution to the supply chain problem is obtained by processing the nodes representing variables and constraints in the graph.

[0003] Related techniques only consider the information of each node in each graph when processing graphs, so the solutions obtained sometimes cannot meet the solution objectives. Summary of the Invention

[0004] Therefore, this application discloses the following technical solution:

[0005] The first aspect of this application provides a method for solving supply chain problems, including:

[0006] Obtain the unprocessed graph corresponding to the supply chain problem. The supply chain problem includes multiple variables involved in the supply chain problem, at least one constraint condition restricting the variables, and the solution objective of the supply chain problem. The nodes contained in the unprocessed graph include variable nodes representing the variables and constraint nodes representing the constraints. The edges contained in the graph represent the relationship between the variables and constraints corresponding to the nodes.

[0007] For at least one node in the graph to be processed, the nearest neighbor nodes of the node in the graph to be processed are determined, so as to generate a first node feature of the node based on the variables and / or constraints represented by the node and its nearest neighbor nodes. The nearest neighbor nodes of the node include nodes that are connected to the node and whose distance to the node is less than a distance threshold.

[0008] Based on the first node characteristics of at least one of the nodes, determine the variable values ​​corresponding to each of the variables that satisfy the constraints and the solution objective.

[0009] A second aspect of this application provides an electronic device, including a memory and a processor;

[0010] The memory is used to store computer programs;

[0011] The processor is used to execute the computer program to perform:

[0012] Obtain the unprocessed graph corresponding to the supply chain problem. The supply chain problem includes multiple variables involved in the supply chain problem, at least one constraint condition restricting the variables, and the solution objective of the supply chain problem. The nodes contained in the unprocessed graph include variable nodes representing the variables and constraint nodes representing the constraints. The edges contained in the graph represent the relationship between the variables and constraints corresponding to the nodes.

[0013] For at least one node in the graph to be processed, the nearest neighbor nodes of the node in the graph to be processed are determined, so as to generate a first node feature of the node based on the variables and / or constraints represented by the node and its nearest neighbor nodes. The nearest neighbor nodes of the node include nodes that are connected to the node and whose distance to the node is less than a distance threshold.

[0014] Based on the first node characteristics of at least one of the nodes, determine the variable values ​​corresponding to each of the variables that satisfy the constraints and the solution objective. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0016] Figure 1 This is a flowchart of a supply chain problem-solving method provided in an embodiment of this application;

[0017] Figure 2 This is a schematic diagram of a spectrum to be processed provided in an embodiment of this application;

[0018] Figure 3 This is a schematic diagram illustrating how a spectrum to be processed is divided into sub-species, as provided in an embodiment of this application.

[0019] Figure 4 This is a schematic diagram of a supply chain problem-solving method provided in an embodiment of this application;

[0020] Figure 5 This is a schematic diagram of a training process provided in an embodiment of this application;

[0021] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] This embodiment provides a method for solving supply chain problems. Please refer to [link to relevant documentation]. Figure 1 The method may include the following steps.

[0024] S101, obtain the graph to be processed corresponding to the supply chain problem. The supply chain problem includes multiple variables involved in the supply chain problem, at least one constraint condition of the limiting variables, and the solution objective of the supply chain problem. The nodes contained in the graph to be processed include variable nodes representing variables and constraint nodes representing constraints. The edges contained in the graph represent the relationship between the variables and constraints corresponding to the nodes.

[0025] S102, for at least one node in the graph to be processed, determine the nearest neighbor nodes of the node in the graph to be processed, so as to generate the first node feature of the node according to the variables and / or constraints represented by the node and its nearest neighbor nodes. The nearest neighbor nodes of the node include nodes that are connected to the node and whose distance to the node is less than a distance threshold.

[0026] S103, based on the first node characteristics of at least one node, determine the variable values ​​corresponding to each variable that satisfy the constraints and the solution objective.

[0027] The beneficial effects of this embodiment are as follows:

[0028] When solving supply chain problems, a graph corresponding to the supply chain problem is obtained. For at least one node in the graph, the first node feature of the node is determined by combining the variables or constraints represented by the node and the variables and / or constraints represented by its nearest neighbor nodes. In this way, the information of the nearest neighbor nodes of a node can be integrated into the first node feature of the node. Thus, in the process of solving based on the first node feature, the information and relationship between the node and its nearest neighbor nodes are obtained to obtain variable values ​​that are more in line with the solution objective.

[0029] The supply chain problem in S101 can be any multivariate programming problem related to the supply chain field. For example, the supply chain problem can be a production scheduling problem, a spare parts problem, or other supply chain-related planning problems, without limitation.

[0030] The production scheduling problem refers to a situation where, for a given set of workpieces to be processed, there are a total of Sum1 pieces (e.g., Sum1 equals 100, meaning there are 100 workpieces to be processed). It is known that each workpiece needs to be processed sequentially through a specified number of operations (e.g., 5 operations). Each operation of each workpiece can be processed on one of several processing machines. The processing time for different operations varies among different processing machines. In order to minimize the total time required to process all workpieces, it is necessary to determine which processing machine to use for each operation of each workpiece.

[0031] In production scheduling problems, the variables that need to be solved can be represented by x. ijk Indicates that x ijk The variable value can be 1 or 0. If it is 1, it means that the j-th process of the i-th workpiece is processed on the k-th processing machine. If it is 0, it means that the j-th process of the i-th workpiece is not processed on the k-th processing machine.

[0032] The objective can be to minimize the target value, which can be the total processing time from the first operation of the first workpiece to the last operation of the last workpiece.

[0033] Constraints may include: Condition 1, each process of a workpiece can only be processed by one machine; Condition 2, the completion time of a workpiece on a machine should be greater than or equal to the sum of its start time on that machine and the processing time of that machine; Condition 3, a workpiece can only be processed in the next process after the previous process is completed; Condition 4, a machine can only be used for another process after the current process is completed.

[0034] The spare parts problem refers to determining, based on a variety of raw materials and a given quantity of each material, which products to produce and the quantity of each product to produce, in order to maximize revenue from the sale of the products. Different products have different selling prices, resulting in different revenues per unit sold, and different products require different quantities of different raw materials. For example, given three raw materials and a given inventory level for each material, three different products can be produced. The spare parts problem then requires determining how many units of each of these three products need to be produced to maximize revenue with the given inventory levels.

[0035] In the spare parts problem, the variable that needs to be solved can be A. i Let represent the planned production quantity of product i, and constraints may include: the inventory quantity S of raw material k. kThe total consumption of raw material k for all planned production products should be less than or equal to the inventory level S. k The objective can be to maximize the target value, which can be the total revenue that can be obtained after selling all the products planned to be produced, assuming that all the products produced can be sold.

[0036] In step S101, each variable (hereinafter referred to as variable) and each constraint in the supply chain problem can be represented by a corresponding node in a graph. Based on the relationship between the variables and constraints, edges are connected between the nodes representing the variables and the nodes representing the constraints. These edges represent the constraints imposed on the variables, thus obtaining the graph to be processed corresponding to the supply chain problem. In this graph, the nodes representing the variables are not connected to each other, and the nodes representing the constraints are not connected to each other. Therefore, the graph to be processed corresponding to the supply chain problem is equivalent to a bipartite graph.

[0037] For a node representing a variable, this node can include at least one variable attribute information for the corresponding variable. The at least one variable attribute information includes, but is not limited to: type, variable value, and value range. The type indicates whether the variable is continuous (i.e., the variable value can be any real number), integer (i.e., the variable value can only be an integer), or binary (i.e., the variable value is 1 or 0); the variable value is the current value of the variable.

[0038] For a node representing a constraint, this node can include at least one constraint attribute that describes the corresponding constraint. The constraint attribute may include, but is not limited to, constraint values ​​and constraint relationships. The constraint value is a numerical value used to restrict the values ​​of each variable, and the constraint relationship can be any one of greater than, greater than or equal to, less than, less than or equal to, and equal to.

[0039] For any edge, this edge can represent the constraint condition of the node connected by the edge. The variables of the node connected by the edge have a constraint effect, that is, the variables at one end of the edge are constrained by the constraint condition at the other end of the edge.

[0040] Based on the previous example of the spare parts problem, we know there are 3 types of raw materials, and their inventory quantities are denoted as S. k The value of k ranges from 1 to 3. There are three products that can be processed: product 1, product 2, and product 3. The three corresponding variables are denoted as A. i The value of i ranges from 1 to 3, and the variable value represents the planned production quantity of the corresponding product. For example, A1 equals 100, indicating a planned production of 100 units of product 1. Producing product 1 consumes raw materials 1 and 2, producing product 2 consumes raw materials 1 and 3, and producing product 3 consumes raw materials 1, 2, and 3. Based on the above example, we can determine... Figure 2The diagram shows the unprocessed graph corresponding to the spare parts problem. Edge L1 represents the planned production quantity A1 of product 1, which is limited by the constraint of raw material 1. Maximum production quantity 1 corresponds to the aforementioned value range, representing the maximum quantity that can be produced when all raw materials are used to produce product 1. The same applies to maximum production quantities 2 and 3.

[0041] Optionally, each edge in the graph to be processed may or may not have a corresponding edge weight. The edge weight represents the degree to which the variable value is constrained by the corresponding constraint condition. Combined with... Figure 2 In the example of the spare parts problem, the edge weight of edge L1 can be equal to the amount of raw material 1 2 consumed in producing one product 1.

[0042] In step S102, a corresponding first node feature can be generated for each node in the graph to be processed, or the first node feature can be generated only for some nodes in the graph to be processed. In the latter case, the first node feature can be generated only for the nodes representing variables (denoted as variable nodes) in the graph to be processed. The information of the nodes representing constraints (denoted as constraint nodes) has already been incorporated into the first node feature of the variable nodes, so the first node feature of the constraint nodes does not need to be generated.

[0043] Generating only the first node features of some nodes can reduce the computational load of steps S102 and S103, thereby improving the processing efficiency of the method in this embodiment.

[0044] Before executing S102, a distance threshold N can be determined. This threshold can be a fixed value determined empirically, or a dynamic value determined based on the graph to be processed. That is, different graphs may have different applicable distance thresholds. For example, it can be determined based on the total number of nodes or edges in the graph; the larger the total number of nodes or edges, the larger the distance threshold. Alternatively, it can be determined based on the type of supply chain problem corresponding to the graph. For instance, the distance threshold for a spare parts problem may differ from that for a production scheduling problem.

[0045] After determining the distance threshold N, for any node in the graph to be processed, the nearest neighbor nodes can be determined based on the N-hop sampling method in related technologies.

[0046] For any two nodes in the graph to be processed, if there exists at least one path consisting of multiple interconnected edges, with these two nodes at both ends of the path, then the two nodes are connected. The distance of the shortest path between these two nodes is the distance between them. The distance of a path can be defined as the number of edges contained in the path.

[0047] by Figure 2For example, the node of product 1 is connected to the node of product 2, and the distance between them is 2. The node of product 1 is connected to the node of raw material 1, and the distance between them is 1.

[0048] For any node X in the graph to be processed, we can start from node X and search for every node connected to X. For each connected node found, we determine if the distance from this connected node to X is less than a distance threshold N. If it is less, this connected node is determined to be a nearest neighbor of X, thus obtaining all the nearest neighbors of X. Alternatively, we can define connected nodes whose distance is less than or equal to the distance threshold as nearest neighbors.

[0049] For any node X in the graph to be processed, node X can be used as a seed node. The information represented by the seed node and the information represented by all the nearest neighbor nodes of the seed node (whether variables or constraints) are fused to obtain the first node feature of the seed node.

[0050] There are no restrictions on the method for generating the first node feature. The following uses the seed node as an example to illustrate several optional methods for generating the first node feature.

[0051] One method for generating the first node feature is as follows: For any node, if the node represents a variable, use a pre-built convolutional network model for extracting node features to process the variable attribute information contained in the node to obtain the feature vector of the node, which is denoted as the initial node feature; if the node represents a constraint condition, use a pre-built convolutional network model for extracting node features to process the constraint attribute information contained in the node to obtain the initial node feature of the node; finally, for the seed node, add the initial node feature of the seed node and the initial node features of all the nearest neighbor nodes of the seed node and take the average, and the result is used as the first node feature of the seed node.

[0052] One method for generating the first node feature is to, for each nearest neighbor node of the seed node, assign a distance weight to the nearest neighbor node based on the distance from the nearest neighbor node to the seed node. The smaller the distance, the larger the distance weight, and the larger the distance, the smaller the distance weight. Then, sum the initial node features of all the nearest neighbors according to the corresponding distance weights, and concatenate the weighted sum with the initial node features of the seed node to form the first node feature of the seed node. Alternatively, the first node feature of the seed node can be obtained by adding the weighted sum with the initial node features of the seed node.

[0053] An alternative method for generating the first node features could be:

[0054] Based on the initial node characteristics of a node and the initial node characteristics of its nearest neighbors, multiple aggregation processes are performed to obtain the first node characteristics of the node. The initial node characteristics of a node are determined based on the variables or constraints represented by the node.

[0055] The process of one aggregation process includes:

[0056] For each node to be aggregated, the previous aggregation feature of the node to be aggregated and the previous aggregation feature of the neighboring nodes directly connected to the node to be aggregated are fused to obtain the current aggregation feature of the node to be aggregated.

[0057] The nodes to be aggregated include the nodes used for aggregation processing and the nodes' nearest neighbor nodes. The previous aggregation feature of a node refers to the aggregation feature obtained after the previous aggregation processing. In the first aggregation processing, the previous aggregation feature of a node is the corresponding initial node feature, and the last aggregation feature of a node is the node's first node feature.

[0058] For any node, its neighboring nodes can be defined as nodes that are connected to this node and whose distance is equal to 1, for example... Figure 2 The node representing product 1 and the node representing raw material 1 are each other's neighbor nodes.

[0059] In the above method for generating the first node feature, the process of the kth aggregation process can be represented by the following formulas (1) and (2).

[0060]

[0061]

[0062] In formulas (1) and (2), h v,k The h represents the k-th aggregation feature of the node v to be aggregated, which is equivalent to the aggregation feature mentioned above; v,k-1 This represents the (k-1)th aggregation feature of the node v to be aggregated, which is equivalent to the previous aggregation feature mentioned above; h u,k-1 The feature represents the (k-1)th aggregation of node u; node u belongs to set Nv, which is the set of all neighboring nodes of the node to be aggregated; h Nv,k This represents the neighbor node fusion feature obtained during the k-th aggregation process.

[0063] `Aggregate()` is the fusion function, which merges all aggregated features within the parentheses. The fusion method can be summation, averaging, weighted averaging, or other methods, without limitation. For example, a weighted average can be calculated based on the edge weights of the connecting nodes. `Concat()` is the concatenation function, which concatenates the aggregated features within the parentheses with the fusion features of neighboring nodes to form a concatenated feature. `W` is a preset transformation coefficient matrix used to convert the concatenated feature output by `Concat()` into a feature with the same dimension as the aggregated features. `Sigma()` represents the activation function, which can be the hyperbolic tangent function `tanh()` or other commonly used activation functions in related fields, without limitation.

[0064] The transformation coefficient matrices used in the (k-1)th and kth aggregation processes can be the same or different.

[0065] During the first aggregation process, when k equals 1, h v,0 h is equal to the initial node characteristics of the node to be aggregated, v; u,0 It equals the initial node characteristics of node u.

[0066] Where k ranges from 1 to the distance threshold N. For example, if the distance threshold is 5, then in order to obtain the first node feature of the seed node, the above aggregation process can be performed 5 times on the seed node and all its nearest neighbor nodes. After the 5th aggregation process, the 5th aggregation feature of the seed node is used as the first node feature of the seed node.

[0067] The method for obtaining the first node features described above is equivalent to:

[0068] For the seed node and all its nearest neighbor nodes, each of them is treated as node v, and the initial node features are calculated according to formulas (1) and (2) to obtain the first aggregated features of the seed node and the first aggregated features of each nearest neighbor node, thus completing the first aggregated process; thereafter, each aggregated process is executed in sequence until N aggregated processes are completed.

[0069] It can be seen that for a seed node, each time the aggregation process is performed, the aggregation feature of the seed node can incorporate the information of the nearest neighbor nodes at a greater distance. For example, the first aggregation process can incorporate the information of the nearest neighbor node at a distance of 1, the second aggregation process can incorporate the information of the nearest neighbor node at a distance of 2, and so on. After performing N aggregation processes, the first node feature of the seed node can incorporate the information of all the nearest neighbor nodes of the seed node.

[0070] Obtaining the first node feature using the third method described above ensures that the first node feature of the seed node can be integrated with the information of the seed node and each of its nearest neighbors.

[0071] Optionally, step S103 can be implemented as follows:

[0072] Based on the solution objective, constraints, and the first node characteristics of at least one node, the solution process is executed in multiple rounds to obtain the variable values ​​corresponding to each variable that satisfy the constraints and the solution objective.

[0073] The solution process in one round includes:

[0074] By combining the variable values ​​of each variable from each previous round with the first node features of at least one node, the second node features of at least one node are obtained;

[0075] Based on the solution objective, constraints, and the characteristics of the second node of at least one node, the variable values ​​of each variable in the current round are determined. The variable values ​​obtained in different rounds are different. In the solution process of the first round, the initial values ​​of each variable are randomly determined based on the constraints. In the last round, the variable values ​​of each variable are the variable values ​​corresponding to each variable that satisfy the constraints and the solution objective.

[0076] In each round of the solution process, after obtaining the second node features, the variable values ​​of each variable in the previous round can be adjusted according to the solution objective, constraints, and the second node features of at least one node, so as to obtain the variable values ​​of each variable in the current round.

[0077] The second node features can be obtained in the following ways:

[0078] The first node features of at least one node are combined into a feature matrix, denoted as the graph representation of the graph to be processed. The graph representation of the graph to be processed and the solution results of each previous round are input into a pre-constructed feature model that can process time-series information. After calculation based on the feature model, a feature matrix is ​​obtained from the feature layer output of the feature model, denoted as the output feature matrix. Each row of the output feature matrix is ​​equivalent to the second node feature of a node.

[0079] For example, if S102 obtains the first node features of M nodes, then the graph representation of the graph to be processed is equivalent to an M-row feature matrix, where each row is equal to the first node feature of a node. The feature model is used to calculate the graph representation of the graph to be processed and the solution results of each previous round to obtain an M-row output feature matrix of the feature layer output of the feature model. The M row vectors contained therein are the second node features of the aforementioned M nodes.

[0080] The solution results for each round include: the solution results for each round that has been executed since the first round; and the solution result for a round, which refers to the variable values ​​of each variable obtained at the end of the solution process for a round.

[0081] The above feature model can be any neural network model with any structure that can be used to process time-series information in the relevant technical field. This embodiment does not limit its structure and working principle.

[0082] When performing a solution process in multiple rounds, a solution termination condition can be set. If, after the solution process in a certain round, the variable values ​​of each variable in the current round are determined to meet the solution termination condition, then the solution is terminated, and the variable values ​​of each variable in the last round are output as the variable values ​​corresponding to each variable that meet the constraints and the solution objective.

[0083] There are no restrictions on the termination condition. Below are some examples of optional termination conditions. One termination condition is that the number of rounds in the solution process reaches the maximum number of rounds. For example, if the maximum number of rounds is 50, then after completing the 50th round of the solution process, it can be determined that the termination condition is met.

[0084] The solution termination condition can also be that the cumulative time of the solution process reaches or exceeds the solution time limit. For example, if the solution time limit is set to 5 minutes, and after completing one round of the solution process, it is found that the cumulative time of the solution process has reached 5 minutes and 10 seconds, then the solution termination condition can be determined to be met.

[0085] The solution termination condition can also be that the target value calculated based on the variable values ​​of each variable in the current round is greater than or equal to (or less than or equal to) the set target value convergence threshold. Taking the spare parts problem as an example, if the total revenue calculated based on the variable values ​​of each variable in the current round is greater than the set revenue convergence threshold, then the solution termination condition can be determined to be satisfied.

[0086] Each round of the solution process yields a batch of variable values ​​for that round (i.e., the values ​​of the aforementioned variables), and the variable attribute information included in the variable nodes also contains these variable values. Therefore, after each round of the solution process, the variable attribute information of the corresponding variable nodes in the graph to be processed can be updated based on the batch of variable values ​​from the current round. Then, based on the updated graph, a new first node feature for at least one node is determined, denoted as the first node feature for the current round. The first node feature used in the next round of the solution process can be the first node feature from the previous round.

[0087] The first node feature used in the first round of the solution process is the initial first node feature determined by setting the variable values ​​in the graph to be processed to their initial values. In other words, the first round of the solution process integrates the initial values ​​of each variable, randomly determined based on the constraints. Figure 2 For example, the initial values ​​of A1, A2, and A3 can all be set to 1 or 0 to ensure that the constraints of the corresponding raw materials are met.

[0088] The advantage of obtaining the variable values ​​in the above manner is that, during the solution process, not only is the information of each node in the graph to be processed considered, but the temporal information in the solution process (i.e., the variable values ​​of each variable in each previous round) is also further integrated, which is conducive to obtaining variable values ​​that are more in line with the solution objective more quickly. Taking the aforementioned spare parts problem as an example, by integrating temporal information, the total benefit corresponding to the variable values ​​obtained may be greater, or the solution termination condition may be met more quickly.

[0089] Optionally, the variable values ​​for each variable in the current round are determined based on the solution objective, constraints, and the second node characteristics of at least one node, including:

[0090] The solution strategy for the current round is determined based on the second node characteristics of at least one node. The solution strategy for the current round is used to indicate the variables whose values ​​need to be adjusted during the solution process of the current round.

[0091] Based on the solution objective and constraints, adjust the variable values ​​of each variable from the previous round to the variable values ​​indicated by the solution strategy of the current round, in order to obtain the variable values ​​of each variable in the current round.

[0092] In this embodiment, the second node features of at least one node can be processed according to the decision model to obtain the solution strategy for the current round. The solution strategy for the current round can be a T-bit binary number, where each bit corresponds to a variable node in the graph to be processed, or to a variable in the supply chain problem in S101. T equals the total number of variables in the supply chain problem. A value of 1 for a bit indicates that the variable value corresponding to this bit needs to be adjusted in the solution process of the current round. A value of 0 for a bit indicates that the variable value corresponding to this bit does not need to be adjusted in the solution process of the current round, and the original variable value of this variable before the solution process of the current round remains unchanged.

[0093] After obtaining the solution strategy for the current round, the solution objective, constraints, variable values ​​from the previous round, and solution strategy for the current round can be input into the solver to obtain the variable values ​​for the current round.

[0094] During the current round of problem-solving, the solver obtains the variable values ​​for each variable through two phases: destruction and repair. Specifically, the solver executes the destruction phase based on the predictions of the decision model for the current round. The destruction phase involves removing the variable values ​​of the variables to be adjusted from the previous round's variable values. Then, the solver executes the repair phase. The repair phase involves keeping the values ​​of the non-adjustable variables unchanged within a given time limit, searching for variable values ​​of the variables to be adjusted that best meet the solution objectives based on constraints and the solution objective. After the search time reaches the time limit, the solver outputs the variable values ​​of the non-adjustable variables and the searched variable values ​​of the adjusted variables as the variable values ​​for each variable in the current round.

[0095] During the repair phase, the solver can search for the variable values ​​of the variable to be adjusted based on a greedy algorithm or other commonly used optimization algorithms in related technical fields, which will not be elaborated here.

[0096] It should be noted that the solver mentioned above can be understood as a program that searches for variable values ​​based on a specific optimization algorithm and the aforementioned solution strategy, and does not include model parameters that need to be determined through training.

[0097] The variable to be adjusted refers to the variable that needs to be adjusted according to the solution strategy of the current round, while the variable not to be adjusted refers to the variable other than the variable to be adjusted.

[0098] The decision-making model can be a pre-built neural network model, and its model structure can refer to relevant technologies without limitation.

[0099] The advantage of using the above method to solve for the variable values ​​of each variable in the current round is that:

[0100] On the one hand, by removing the variable values ​​of the variables to be adjusted in the previous round, we can avoid these variable values ​​limiting the search range of the solution process in the current round, and prevent the variable values ​​obtained from getting trapped in local optima and failing to reach or approach the global optima.

[0101] On the other hand, by determining the solution strategy for each round based on the second node characteristics of at least one node, the solution strategy can be dynamically determined based on the node characteristics of the graph to be processed and the real-time time series information during the solution process, which is conducive to obtaining variable values ​​of each variable that are more in line with the solution objective.

[0102] Optionally, if the solution process in each round is based on a solution strategy, then the method for obtaining the second node features in the aforementioned embodiments can be:

[0103] By integrating the variable values ​​of each variable from each previous round, the solution strategy of each previous round, and the first node features of at least one node, the second node features of at least one node are obtained.

[0104] In the first round of the solution process, the solution strategy for each previous round is a randomly determined initial solution strategy.

[0105] In this embodiment, the method for obtaining the second node feature can be:

[0106] The first node features of at least one node are combined into a feature matrix, which is denoted as the graph representation of the graph to be processed. The graph representation of the graph to be processed, the solution results of each previous round, and the solution strategy of each previous round are all input into the feature model. After calculation based on the feature model, a feature matrix is ​​obtained from the feature layer output of the feature model, which is denoted as the output feature matrix. Each row of the output feature matrix is ​​equivalent to the second node feature of a node.

[0107] The solution strategy for each round includes the solution strategy for each round that has been executed since the first round.

[0108] The beneficial effect of this embodiment is that by further combining the solution strategies of each previous round when determining the features of the second node, the features of the second node can more comprehensively reflect the solution process of each previous round, thereby improving the efficiency and accuracy of the solution process.

[0109] Optionally, the variable values ​​of each variable from each previous round and the first node features of at least one node are fused to obtain the second node features of at least one node, including:

[0110] For each subgraph in multiple subgraphs, the variable values ​​of each variable in each previous round and the first node features of the nodes contained in the subgraph are fused to obtain the sub-node features of the nodes contained in the subgraph.

[0111] Based on the sub-node features of the nodes contained in each sub-graph, the second node features of at least one node in the graph to be processed are determined, and multiple sub-graphs are obtained by splitting the graph to be processed.

[0112] In this embodiment, the map to be processed can be first split into multiple sub-maps. The method of splitting is not limited.

[0113] For example, you can set the maximum number of nodes that each subgraph can contain, and then split the graph to be processed into multiple subgraphs based on the edge segmentation method. When splitting based on the edge segmentation method, each edge contained in the graph to be processed can be retained in at least one subgraph. This can preserve the information of the graph to be processed to the maximum extent and avoid the loss of the original information of the graph to be processed when splitting into subgraphs.

[0114] For each subgraph, the graph representation of the subgraph to be processed can be formed by the first node features of the nodes contained in the subgraph according to the method of forming the graph representation of the subgraph in the previous embodiment. Then, the graph representation of the subgraph and the variable values ​​of each variable in each previous round are input into the aforementioned feature model to obtain the subgraph feature matrix output by the feature layer of the feature model. Each row of the subgraph feature matrix is ​​equivalent to the child node feature of a node contained in the subgraph.

[0115] Optionally, if a solution strategy is determined during the solution process, the sub-node features of the nodes contained in the sub-graph can be obtained according to the method of the aforementioned embodiment, based on the variable values ​​of each variable in each previous round, the solution strategy in each previous round, and the first node features of the nodes contained in the sub-graph.

[0116] In this embodiment, the feature model may further include an output head connected after the feature layer. After obtaining the sub-node features of the nodes contained in each sub-map using the above method, the output head can process the sub-node features of the nodes contained in each sub-map to obtain the second node features of the nodes in the map to be processed.

[0117] Optionally, the method for determining the second node features of at least one node in the graph to be processed based on the child node features of the nodes contained in each sub-graph can be:

[0118] For at least one node in the graph to be processed, if only one subgraph contains the node, the child node features of the node contained in the subgraph are determined as the second node features of the node in the graph to be processed.

[0119] For at least one node in the graph to be processed, if multiple subgraphs contain the node, the sub-node features of the nodes contained in the multiple subgraphs are fused to obtain the second node feature of the node in the graph to be processed.

[0120] After splitting the graph to be processed into sub-graphs, some nodes in the graph to be processed may appear repeatedly in multiple sub-graphs.

[0121] by Figure 3 For example, Figure 3 The map to be processed shown in (1) is split into two sub-maps, as follows: Figure 3 The first sub-map shown in (2) and Figure 3 The second subgraph of (3). It can be seen that in order to retain each edge in the graph to be processed, the three nodes of product 2, raw material 1 and raw material 2 appear repeatedly in the two subgraphs, that is, both subgraphs contain product 2, raw material 1 and raw material 2.

[0122] To address this issue, when processing the child node features of nodes within each sub-graph, the output head can directly determine the child node features of nodes that do not appear repeatedly as the corresponding second node features. For nodes that appear repeatedly, the child node features of these nodes in each sub-graph can be fused to obtain the corresponding second node features. The fusion method can be summation, averaging, weighted averaging, or other methods, without limitation.

[0123] Combination Figure 3 For example, in the graph to be processed, the second node feature of product 1 is equal to the sub-node feature of product 1 in the first sub-graph; the second node feature of product 3 and the second node feature of raw material 3 are equal to the sub-node features of product 3 and raw material 3 in the second sub-graph; the second node feature of product 2 is equal to the average of the sub-node features of product 2 in the first sub-graph and the sub-node features of product 2 in the second sub-graph; the second node features of raw material 1 and raw material 2 can be obtained in the same way.

[0124] Optionally, the output head may also include one or more fully connected layers (MLPs) for linear transformation. For each node, if only one subgraph contains this node, the output head can calculate the child node features of this node in the subgraph based on the model parameters contained in the fully connected layers to obtain the second node feature of this node in the graph to be processed. If multiple subgraphs contain this node, the child node features of this node in the multiple subgraphs are fused to obtain a fused result, and then the fused result is calculated based on the model parameters contained in the fully connected layers to obtain the second node feature of this node in the graph to be processed. The values ​​of these model parameters contained in the fully connected layers can be determined during the training phase.

[0125] Optionally, whether to split the graph to be processed into sub-graphs can be determined based on the size of the graph to be processed. For example, if the number of nodes in the graph to be processed is greater than a certain threshold, it is split into sub-graphs, and the second node features are obtained based on the sub-graphs according to the method of this embodiment. If the number of nodes in the graph to be processed is less than or equal to this threshold, it is not split into sub-graphs. In this case, the method of the aforementioned embodiment can be used to directly take each row of the feature matrix output by the feature layer of the feature model as the second node feature of the corresponding node in the graph to be processed.

[0126] The above method splits the graph to be processed into sub-graphs and then calculates the features of each sub-node separately. Its purpose is that when the graph to be processed contains too many nodes, inputting the first node features of all nodes into the feature model at the same time may exceed the processing capacity of the feature model due to the large amount of input data, which may cause the processing to stall. Furthermore, the large amount of input data may cause the feature model to not pay enough attention to the first node features of each node, resulting in inaccurate output second node features.

[0127] By splitting the graph into sub-graphs using the above method, the computational load when calculating sub-node features in each feature model's feature layer can be reduced, ensuring smooth execution of the computation process and avoiding inaccurate results due to excessive data.

[0128] See last for reference. Figure 4 The implementation process of the supply chain problem-solving method in this embodiment will be explained in conjunction with the aforementioned examples.

[0129] Based on the variables and constraints of the spare parts problem, Figure 2 The graph to be processed shown randomly determines the initial solution strategy and the initial values ​​of each variable that satisfy the constraints. Then, the first round of the solution process is executed. The first node feature, the variable values ​​of each variable in the 0th round (i.e., the initial values), and the solution strategy of the 0th round (i.e., the initial solution strategy) are input into the feature model to obtain the second node feature output by the feature model. The decision model processes the second node feature to obtain the solution strategy of the 1st round, i.e., the solution strategy of round 1 shown in the figure. Based on the solution strategy of round 1, the solver combines the constraints, the solution objective, and the variable values ​​of each variable in the 0th round to obtain the variable values ​​of each variable in the 1st round.

[0130] Next, in the second round of solving, the first node features, the variable values ​​of each variable in rounds 0 to 1, and the solution strategy of rounds 0 to 1 are input into the feature model to obtain the second node features output by the feature model. The decision model processes the second node features to obtain the solution strategy for round 2. Based on the solution strategy for round 2, the solver combines the constraints, the solution objective, and the variable values ​​of each variable in round 1 to solve for the variable values ​​of each variable in round 2.

[0131] This process continues in each subsequent round until the Nth round. The first node feature, the variable values ​​of each variable in rounds 0 to N-1, and the solution strategy for rounds 0 to N-1 are input into the feature model. The feature model, decision model, and solver process the data sequentially to obtain the variable values ​​of each variable in round N. At this point, it is determined that the solution termination condition is met. Therefore, the variable values ​​of each variable in round N are output as the variable values ​​corresponding to each variable that satisfy the constraints and the solution objective.

[0132] Optionally, the models used in the foregoing embodiments can be pre-trained based on the sample maps. Additionally, to train these models, a pre-tuned neural network model, called a discriminant model, can be introduced during training to evaluate the solution process in each round.

[0133] Based on the discriminant model, the training methods for the feature model, decision model, and discriminant model in this embodiment may include:

[0134] 1. Obtain the sample map corresponding to the sample supply chain problem;

[0135] 2. For at least one sample node in the sample graph, determine the nearest neighbor sample nodes of the sample node in the sample graph, so as to generate the first sample node feature of the sample node based on the sample variables and / or sample constraints represented by the sample node and its nearest neighbor sample nodes.

[0136] 3. Based on the sample solution objective and sample constraints contained in the sample supply chain problem, as well as the first sample node characteristics of at least one sample node, execute multiple rounds of solution process to obtain the variable values ​​of each sample variable that satisfy the sample constraints and sample solution objective;

[0137] 4. Determine the reinforcement learning loss based on the multiple reward values ​​corresponding to the solution process of multiple rounds. The reward value of each round is determined based on the variable values ​​of each sample variable obtained in the solution process of the corresponding round.

[0138] 5. Update the model parameters of the feature model, decision model, and discriminant model based on the reinforcement learning loss;

[0139] Among them, the feature model is used to obtain the second node features of the nodes in the graph to be processed, the decision model is used to determine the solution strategy for each round, the solver is used to obtain the variable values ​​of each variable in each round, the discriminant model is used to calculate the reward value for each round according to the solution strategy for each round, and the reward value is used to calculate the reinforcement learning loss.

[0140] The sample supply chain problem can be a real supply chain problem or a virtual supply chain problem constructed according to training requirements. The variables, constraints and solution objectives of the sample supply chain problem are denoted as sample variables, sample constraints and sample solution objectives.

[0141] The implementation method of step 1 is the same as that of obtaining the map to be processed, and will not be described in detail.

[0142] For steps 2 and 3, see [link / reference]. Figure 5Starting from the first round of the solution process, in each round, the first sample node features, the sample variable values ​​from each previous round, and the sample solution strategy can be input into the feature model. After being processed sequentially by the feature model, decision model, and solver, the variable values ​​of each variable in the current round are obtained. Here, the sample solution strategy refers to the solution strategy obtained based on the sample graph, and the sample variable values ​​refer to the variable values ​​of the sample variables. The input to the decision model can be the second sample node features obtained after processing by the feature model. The method of obtaining the second sample node features is the same as that in the previous embodiment. Secondly, in the first round of the solution process, the preset initial sample variable values ​​and the initial sample solution strategy are input into the feature model. The above specific implementation method is consistent with the method of obtaining the first node features of the nodes of the graph to be processed and performing multiple rounds of solution process based on the graph to be processed in the previous embodiment, and will not be repeated.

[0143] Unlike the solution process for the graph to be processed, see [link to relevant documentation]. Figure 5 When processing sample maps, after each round of the solution process, the reward value for the current round can be obtained based on the discriminative model. This reward value is then used to derive the reinforcement learning loss for the current round. For example, after completing the solution process for the first round, the reinforcement learning loss L1 for the first round can be obtained; after completing the solution process for the Nth round, the reinforcement learning loss L2 for the first round can be obtained. N The solution process for multiple rounds of the sample map can also be stopped if the aforementioned termination conditions are met, which will not be elaborated further.

[0144] The discriminant model is a pre-trained neural network model used to evaluate the variable values ​​of each variable in any round and the solution strategy for that round in order to obtain the reward value of the corresponding round. This reward value reflects how much the target value calculated based on the sample variable values ​​of this round (such as the total revenue of the spare parts problem) is different from the expected target. The larger the gap, the smaller the reward value, and the smaller the gap, the larger the reward value.

[0145] In this scenario, it is assumed that the solution process for the sample map is performed in N rounds. For any round t+1, where t ranges from 1 to N, the reinforcement learning loss L in round t+1 is... t+1 It can be calculated using the following formulas (3) to (7).

[0146]

[0147]

[0148]

[0149]

[0150]

[0151] In formulas (3) to (7), E[] represents the expected function for calculating the expected value of the data within the parentheses. The specific calculation method can be found in the relevant technology, which will not be elaborated here.

[0152] Y is a preset coefficient, s t+1 Let a represent the sample variable value in round t+1. t+1 Let s represent the sample solving strategy in round t+1. t and a t Let represent the sample variable value and sample solution strategy in round t, respectively;

[0153] Q() represents the reward value for the corresponding round calculated by the discriminative model based on the sample variable values ​​within parentheses and the solution strategy. For example, Q(s) t+1 a t+1 () represents the reward value in round t+1;

[0154] n represents the total number of alternative strategies determined by the decision model in the (t+1)th round of the solution process, PAI i This represents the probability of the i-th alternative strategy determined by the decision model. The decision model outputs the solution strategy by selecting the alternative strategy with the highest probability from these alternative strategies. log() represents the logarithm of the value in parentheses with base 10.

[0155] X t+1 The target value is calculated based on the sample variable values ​​in round t+1, for example, for the spare parts problem X. t+1是 The total return X is calculated based on the sample variable values ​​from round t+1. t The target value of the sample is calculated based on the sample variable values ​​in round t.

[0156] Norm() represents a normalization function that performs normalization operations on the values ​​within parentheses. For methods of normalization operations, please refer to relevant techniques.

[0157] In step 5, the model parameters of the feature model, decision model, and discriminant model can be updated based on the reinforcement learning loss of each round. After the reinforcement learning loss of each round is obtained, the model parameters of the feature model, decision model, and discriminant model can be updated based on the reinforcement learning loss of that round. After the update is completed, the solution process of the next round is executed for the sample map until the solution termination condition is met.

[0158] In step 5, updating the model parameters of the feature model, decision model, and discriminator model based on the reinforcement learning loss can also be achieved by continuously executing the solution process for multiple rounds until the solution termination condition is met. After the solution process for multiple rounds is completed, the reinforcement learning loss of each round is fused to obtain the total reinforcement learning loss. The model parameters of the feature model, decision model, and discriminator model are then updated based on this total reinforcement learning loss.

[0159] Optionally, the two methods for updating model parameters mentioned above can be combined to update the model parameters of the feature model, decision model, and discriminant model.

[0160] Optionally, as explained above, for cases where the graph to be processed is split into sub-graphs, the feature model may include an output head and at least one feature layer.

[0161] For cases involving both feature layers and output heads, methods for updating the model parameters of the feature model, decision model, and discriminant model based on reinforcement learning loss can include:

[0162] The model parameters of at least one feature layer of the feature model, the decision model, and the discriminant model are updated based on the reinforcement learning loss.

[0163] If at least one feature layer, decision model, and discriminant model meet the convergence condition, the model parameters of the output head of the feature model are updated according to the reinforcement learning loss. Here, at least one feature layer is used to obtain the sub-node features of the nodes contained in each sub-graph in the graph to be processed, and the output head is used to determine the second node features of at least one node in the graph to be processed based on the sub-node features of the nodes contained in each sub-graph.

[0164] In other words, for the case where the feature model includes a feature layer and an output head, the training process can be divided into two stages. In the first stage, only at least one feature layer, decision model, and discriminant model of the feature model are trained. After the feature layer, decision model, and discriminant model are trained (i.e., the convergence condition is met), the second stage begins. In the second stage, the feature layer and other models are frozen, that is, the trained model parameters are kept unchanged, and only the output head of the feature model is trained according to the reinforcement learning loss.

[0165] In the first stage of training, the sample map used for training can be a sample map with a small number of nodes that does not need to be split, or it can be a sub-sample map that has already been split. In this stage, the output head of the feature model can be deactivated, and the output of the feature layer can be directly used as the second sample node feature input into the decision model. Then, the feature layer, decision model and discriminant model are updated based on reinforcement learning loss.

[0166] In the second stage of training, the feature layer of the feature model and other models are frozen, and the output head of the feature model is enabled. At this time, the sample map used for training can contain a large number of nodes. Accordingly, the sample map can be split into multiple sub-sample maps according to the aforementioned method of splitting the map to be processed. The output of the feature layer is used as the sub-node features of the nodes contained in each sub-sample map. The sub-node features of these sub-sample maps are input into the output head. After processing by the output head, the second sample node features corresponding to the sample map are obtained. Then, the sample variable values ​​and reinforcement learning loss are calculated based on the second sample node features. The model parameters contained in the output head are updated according to the reinforcement learning loss.

[0167] The purpose of training the feature layer and output head separately in the two-stage method described above is to reduce the number of model parameters updated each time, and to prevent the training process from failing to converge due to too many model parameters being updated, or from causing overfitting problems after convergence.

[0168] In any embodiment of this application, the convergence condition may include, but is not limited to: the cumulative number of updates being greater than or equal to the upper limit of the number of updates; the most recently calculated reinforcement learning loss being less than or equal to a preset first convergence threshold; and the difference between the most recently calculated reinforcement learning loss and the previously calculated reinforcement learning loss being less than or equal to a preset second convergence threshold. The cumulative number of updates refers to the number of times the model parameters are updated. For example, updating the model parameters of the feature model, decision model, and discriminant model once based on a reinforcement learning loss is denoted as updating the model parameters once.

[0169] This embodiment also provides an electronic device; please refer to [link / reference]. Figure 6 This includes a memory 601 and a processor 602;

[0170] Memory 601 is used to store computer programs;

[0171] Processor 602 is used to execute computer programs to perform:

[0172] Obtain the graph to be processed corresponding to the supply chain problem. The supply chain problem includes multiple variables involved in the supply chain problem, at least one constraint condition of the limiting variables, and the solution objective of the supply chain problem. The nodes contained in the graph to be processed include variable nodes representing variables and constraint nodes representing constraints. The edges contained in the graph represent the relationship between the variables and constraints corresponding to the nodes.

[0173] For at least one node in the graph to be processed, determine the nearest neighbor nodes of the node in the graph to be processed, so as to generate the first node feature of the node based on the variables and / or constraints represented by the node and its nearest neighbor nodes. The nearest neighbor nodes of the node include nodes that are connected to the node and whose distance to the node is less than a distance threshold.

[0174] Based on the first node characteristics of at least one node, determine the variable values ​​corresponding to each variable that satisfy the constraints and the solution objective.

[0175] The working principle of the electronic device in this embodiment can be found in the relevant steps of the supply chain problem solving method in the foregoing embodiments, and will not be repeated here.

[0176] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0177] For ease of description, the above systems or devices are described separately as various modules or units based on their functions. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware components.

[0178] In this document, relational terms such as first, second, third, and fourth are used merely to distinguish one entity or operation from another, without necessarily requiring or implying any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0179] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for solving supply chain problems, comprising: Obtain the unprocessed graph corresponding to the supply chain problem. The supply chain problem includes multiple variables involved in the supply chain problem, at least one constraint condition restricting the variables, and the solution objective of the supply chain problem. The nodes contained in the unprocessed graph include variable nodes representing the variables and constraint nodes representing the constraints. The edges contained in the graph represent the relationship between the variables and constraints corresponding to the nodes. For at least one node in the graph to be processed, the nearest neighbor nodes of the node in the graph to be processed are determined, so as to generate a first node feature of the node based on the variables and / or constraints represented by the node and its nearest neighbor nodes. The nearest neighbor nodes of the node include nodes that are connected to the node and whose distance to the node is less than a distance threshold. Based on the first node characteristics of at least one of the nodes, determine the variable values ​​corresponding to each of the variables that satisfy the constraints and the solution objective.

2. The method according to claim 1, wherein determining the variable values ​​corresponding to each of the variables that satisfy the constraints and the solution objective based on the first node characteristics of at least one of the nodes includes: Based on the solution objective, the constraints, and the first node characteristics of at least one of the nodes, a solution process is executed in multiple rounds to obtain the variable values ​​corresponding to each of the variables that satisfy the constraints and the solution objective. The solution process in one round includes: By fusing the variable values ​​of each variable from each previous round with the first node feature of at least one node, a second node feature of at least one node is obtained; Based on the solution objective, the constraints and the second node features of at least one of the nodes determine the variable values ​​of each variable in the current round. The variable values ​​obtained in different rounds are different. In the solution process of the first round, the initial values ​​of each variable are randomly determined based on the constraints. In the last round, the variable values ​​of each variable are the variable values ​​corresponding to each variable that satisfy the constraints and the solution objective.

3. The method according to claim 2, wherein determining the variable values ​​of each of the variables in the current round based on the solution objective, the constraints, and the second node characteristics of at least one of the nodes includes: The solution strategy for the current round is determined based on the second node characteristics of at least one of the nodes, and the solution strategy for the current round is used to indicate the variables whose values ​​need to be adjusted during the solution process of the current round; Based on the solution objective and the constraints, adjust the variable values ​​of the variables in the previous round to the variable values ​​indicated by the solution strategy in the current round, so as to obtain the variable values ​​of the variables in the current round.

4. The method according to claim 3, wherein fusing the variable values ​​of each of the variables from each previous round and the first node feature of at least one of the nodes to obtain the second node feature of at least one of the nodes comprises: By integrating the variable values ​​of each variable from each previous round, the solution strategy from each previous round, and the first node feature of at least one node, a second node feature of at least one node is obtained; In the first round of the solution process, the solution strategy for each previous round is a randomly determined initial solution strategy.

5. The method according to claim 2, wherein fusing the variable values ​​of each of the variables from each previous round and the first node feature of at least one of the nodes to obtain the second node feature of at least one of the nodes comprises: For each of the multiple subgraphs, the variable values ​​of each variable from each previous round and the first node features of the nodes contained in the subgraph are fused to obtain the subnode features of the nodes contained in the subgraph. Based on the sub-node features of each node contained in the sub-graph, the second node feature of at least one node in the graph to be processed is determined, and the multiple sub-graphs are obtained by splitting the graph to be processed.

6. The method according to claim 5, wherein determining the second node feature of at least one node in the graph to be processed based on the child node features of the nodes contained in each of the sub-graphs comprises: For at least one node in the graph to be processed, if only one subgraph contains the node, the child node features of the node contained in the subgraph are determined as the second node features of the node in the graph to be processed. For at least one node in the graph to be processed, if multiple subgraphs contain the node, the sub-node features of the node contained in the multiple subgraphs are fused to obtain the second node feature of the node in the graph to be processed.

7. The method according to claim 1, wherein generating the first node feature of the node based on the variables and / or constraints represented by the node and its nearest neighbor nodes comprises: Based on the initial node characteristics of the node and the initial node characteristics of the node's nearest neighbors, multiple aggregation processes are performed to obtain the first node characteristics of the node. The initial node characteristics of the node are determined according to the variables or constraints represented by the node. The process of one aggregation process includes: For each node to be aggregated, the previous aggregation feature of the node to be aggregated and the previous aggregation feature of the neighboring nodes directly connected to the node to be aggregated are fused to obtain the current aggregation feature of the node to be aggregated. The nodes to be aggregated include the node used for aggregation processing and the node's nearest neighbor nodes. The previous aggregation feature of the node refers to the aggregation feature obtained after the previous aggregation processing. In the first aggregation processing, the previous aggregation feature of the node is the corresponding initial node feature, and the last aggregation feature of the node is the first node feature of the node.

8. The method according to claim 3, further comprising: Obtain the sample map corresponding to the sample supply chain problem; For at least one sample node in the sample map, determine the nearest neighbor sample nodes of the sample node in the sample map, so as to generate the first sample node feature of the sample node based on the sample variables and / or sample constraints represented by the sample node and its nearest neighbor sample nodes; Based on the sample solution objective and sample constraints contained in the sample supply chain problem, and the first sample node characteristics of at least one of the sample nodes, multiple rounds of solution process are executed to obtain the variable values ​​of each of the sample variables that satisfy the sample constraints and the sample solution objective; The reinforcement learning loss is determined based on multiple reward values ​​corresponding to multiple rounds of the solution process. The reward value of each round is determined based on the variable values ​​of each sample variable obtained in the corresponding round of the solution process. The model parameters of the feature model, decision model, and discriminant model are updated based on the reinforcement learning loss. The feature model is used to obtain the second node features of the nodes in the graph to be processed, the decision model is used to determine the solution strategy for each round, the discriminant model is used to calculate the reward value for each round according to the solution strategy for each round, and the reward value is used to calculate the reinforcement learning loss.

9. The method according to claim 8, wherein the feature model comprises an output head and at least one feature layer; The step of updating the model parameters of the feature model, decision model, and discriminant model based on the reinforcement learning loss includes: The model parameters of at least one feature layer, the decision model, and the discriminant model of the feature model are updated according to the reinforcement learning loss. When the at least one feature layer, the decision model, and the discriminant model satisfy the convergence condition, the model parameters of the output head of the feature model are updated according to the reinforcement learning loss. The at least one feature layer is used to obtain the sub-node features of the nodes contained in each sub-graph in the graph to be processed, and the output head is used to determine the second node features of at least one node in the graph to be processed based on the sub-node features of the nodes contained in each sub-graph.

10. An electronic device, comprising a memory and a processor; The memory is used to store computer programs; The processor is used to execute the computer program to perform: Obtain the unprocessed graph corresponding to the supply chain problem. The supply chain problem includes multiple variables involved in the supply chain problem, at least one constraint condition restricting the variables, and the solution objective of the supply chain problem. The nodes contained in the unprocessed graph include variable nodes representing the variables and constraint nodes representing the constraints. The edges contained in the graph represent the relationship between the variables and constraints corresponding to the nodes. For at least one node in the graph to be processed, the nearest neighbor nodes of the node in the graph to be processed are determined, so as to generate a first node feature of the node based on the variables and / or constraints represented by the node and its nearest neighbor nodes. The nearest neighbor nodes of the node include nodes that are connected to the node and whose distance to the node is less than a distance threshold. Based on the first node characteristics of at least one of the nodes, determine the variable values ​​corresponding to each of the variables that satisfy the constraints and the solution objective.