Closed-loop supply chain network optimization method for improving non-dominated sorting genetic algorithm
By improving the non-dominant sorting genetic algorithm and introducing the adaptive control mechanism of reinforcement learning algorithm, the problems of low efficiency, insufficient diversity and complex parameter adjustment in closed-loop supply chain network optimization are solved, and the diversity improvement of efficient cost and sustainability balance reconciliation is achieved.
Patent Information
- Application Number
- CN202510035262.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-13
AI Technical Summary
The existing technology is less efficient when optimizing the closed-loop supply chain network, lacks population diversity, slow convergence speed, and complex parameter adjustment, making it difficult to effectively solve the complex problems of large-scale supply chain networks.
Improve the non-dominant sorting genetic algorithm (NSGA-II), combines the adaptive control mechanism of reinforcement learning algorithms, and improves the efficiency and diversity of the algorithm through mixed initialization strategies and adaptive adjustment of the cross rate and variability rate.
The efficient balance between supply chain cost and sustainability is achieved, the diversity and quality of understanding is improved, local optimal risks are reduced, and the convergence speed and computing efficiency of the algorithm are significantly improved.
Smart Images

Figure CN119990799A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computing methods, and in particular to a closed-loop supply chain network optimization method using an improved non-dominated sorting genetic algorithm. Background Art
[0002] The closed-loop supply chain network is different from the traditional supply chain system. It not only covers the forward process from raw material procurement to production and storage, but also includes the reverse process such as recycling, disassembly and remanufacturing of defective products. In the optimization of closed-loop supply chain networks, precise methods include: ε constraint method, goal programming and branch and bound method. In solving the closed-loop supply chain network, the existing technology uses multi-objective meta-heuristic algorithms (such as non-dominated sorting genetic algorithm and multi-objective particle swarm algorithm) to find the optimal balance point between multiple objectives when dealing with large-scale supply chain networks by simulating natural selection and genetic operations. This parallel search and optimization capability enables the algorithm to explore multiple possible solution sets in complex decision spaces.
[0003] However, although existing optimization methods can provide high-quality solutions, they are relatively inefficient when dealing with highly complex and large-scale problems, and face significant challenges in computing time and memory requirements; the challenges faced by existing solution methods include: Insufficient population diversity: As the number of iterations increases, the population diversity may decrease rapidly, and the algorithm is prone to fall into local optimality. Slow convergence: In a high-dimensional search space, it may take more iterations for the algorithm to converge to a satisfactory solution set, which increases the computational cost. Complex parameter adjustment: Genetic algorithm parameters such as crossover probability and mutation probability have a significant impact on algorithm performance, but the appropriate parameter settings often rely on experience or repeated trials. Summary of the invention
[0004] In order to solve the above problems, the present invention provides a closed-loop supply chain network optimization method based on an improved non-dominated sorting genetic algorithm.
[0005] To implement the above technical solution, the details are as follows:
[0006] S1. Establish the objective function and constraints of the closed-loop supply chain network optimization model with cost and carbon emissions as optimization targets;
[0007] The objective functions include: cost objective function and second objective function;
[0008] The expression of the cost objective function is as follows:
[0009]
[0010] Where y1 represents the minimum total cost of the closed-loop supply chain network; t *represents the cycle; T represents the total cycle; v represents the vth supplier; V represents the total number of suppliers; m represents the product; M1 represents the part; M2 represents the subassembly; the number of parts purchased from the vth supplier is expressed as The corresponding purchase price is C m,v ; The number of subcomponents is expressed as The corresponding purchase price is SC m,v ; When selecting a supplier, the first parameter Second parameter is 1, otherwise 0; j M represents the set of production units; j W represents the warehouse unit set; j R Represents a collection of recycling units; j m Represents the production unit; j w Indicates warehouse; r Represents the recovery unit; M(j m ) represents the product produced by the production unit; M(j w ) represents the products stored in the warehouse; M(j r ) indicates products recycled by a recycling unit; Indicates the quantity of products produced by the production unit; represents the production cost; Indicates the number of products stored in the warehouse; represents the storage cost; Indicates the quantity of products recycled by the recycling unit; represents the recovery cost; Indicates the transportation quantity of products between units and between units and warehouses; TC m,j,j' represents the transportation cost; l represents the emission reduction technology level; L represents the set of emission reduction technology levels; Indicates whether production unit j m The investment level is l for emission reduction technology; h represents the investment cost coefficient of emission reduction technology;
[0011] The expression of the carbon emission objective function is as follows:
[0012]
[0013] Where y2 represents the total carbon emissions of the minimized closed-loop supply chain network, including carbon emissions in the production process, carbon emissions in the recycling process, and carbon emissions in the transportation process; Q M Indicates the carbon emissions generated by producing a unit of product; Q R It represents the carbon emissions generated by recycling a unit product. For the calculation of carbon emissions during transportation, not only the weight of the unit product wg is considered. m , and also considers the transportation distance between units and between units and warehouses.j,j' , where j represents a unit or warehouse, and j' represents (transported to) another unit or warehouse, so as to calculate the total carbon emissions during transportation; Q T That is, the carbon emissions generated by transporting a unit weight (kg) of product per kilometer;
[0014] The constraints include: supplier selection, maximum supply capacity, production capacity, recycling capacity, transportation process, input and output material balance of each unit, warehouse inventory balance and demand satisfaction constraints.
[0015] S2, setting parameters for the closed-loop supply chain network optimization model and initializing the parameters of the genetic algorithm;
[0016] Specifically, the information is loaded into the S1 model; in which the finished product demand in each of the 14 production cycles is randomly generated in the range of 700 to 1000; the information includes: various cost values, number of cycles and demand;
[0017] Initialize the genetic algorithm (NSGA-II) parameters, including the number of iterations, population size N, variable dimension d, initial crossover probability Pc, and mutation probability Pm;
[0018] The closed-loop supply chain network optimization model is solved using the NSGA-II algorithm;
[0019] S3, generate the initial population by mixing random initialization and Latin hypercube sampling method;
[0020] The present invention is a mixed integer model, each individual in the population represents a decision scheme, the decision scheme is a solution, and each individual consists of a set of integer variables and binary decision variables (0 / 1);
[0021] Integer variables include: number of products purchased, number of products produced, number of products shipped, number of products recycled, and number of products in stock;
[0022] Binary decision variables include: supplier selection, whether the production unit invests in emission reduction technology;
[0023] Furthermore, for each integer variable, a hybrid generation method of Latin hypercube sampling and random initialization is used;
[0024] Furthermore, for each binary decision variable, a random initialization method is used to generate a number of 0 or 1;
[0025] That is, the present invention adopts a hybrid initialization strategy to generate an initial population through random initialization and Latin hypercube sampling method, thereby ensuring uniform distribution of sample points in the decision space and improving population diversity;
[0026] The steps to generate the initial population using a mixture of random initialization and Latin hypercube sampling are as follows:
[0027] S3.1. According to the total number of individuals in the population N and the individual dimension d, set the variable of a certain dimension in the individual to X z , then the value range of the variable is Where z = 1, 2, ..., d;
[0028] S3.2, divide the obtained dimension into equally spaced intervals with the same number as the total number of individuals N, and randomly select a point in each interval, where non-integer numbers need to be rounded off;
[0029] The partition interval expression is as follows:
[0030]
[0031] S3.3, let the selection result of each interval be Then the sample points can be expressed as All sample points are merged into the initial group, and the value of the sampling point is X LHS ={X1,X2,...,X N};
[0032] S3.4, based on the Latin hypercube sampling process, the present invention uses the Latin hypercube sampling method to generate N / 2 individuals for integer variables to ensure that each integer variable is evenly distributed in the multidimensional space; the other N / 2 individuals are generated by random initialization to avoid the aggregation of samples while introducing some randomness, which helps the algorithm to jump out of potential local optimality and ensure the extensiveness of the search;
[0033] S4, setting the inverse of the objective function as the fitness function, and calculating the fitness function; performing fast non-dominated sorting on the initial population, calculating the crowding distance and constructing the Pareto frontier;
[0034] The fitness function includes: the inverse of the cost objective function and the inverse of the second objective function
[0035] S5, setting the current population in S4 as the parent population, generating a new population through the parent population and calculating the best fitness function, the mean crowding distance and the number of non-dominated solutions of the new population;
[0036] Specifically, individuals with lower non-dominated ranks are preferentially selected from the parent population; if the ranks are the same, individuals with larger crowding distances are selected for crossover and mutation; then, after performing crossover and mutation operations, the offspring population is generated and merged with the parent population to form a new population;
[0037] Furthermore, for the new population, the fitness function value is calculated again Calculate the mean of the fitness function of the population Calculate the optimal fitness function value. The present invention selects the maximum value as the optimal fitness function value, that is, Perform fast non-dominated sorting and calculate the crowding distance d of each individual t (i) The mean crowding distance of all individuals in the population and the number of non-dominated solutions N t .
[0038] S6, determine whether the current generation is greater than the maximum number of iterations, if yes, execute step S10, if not, execute step S7;
[0039] S7, calculating the comprehensive status value according to the result of S5;
[0040] The expression of the comprehensive state value S is as follows:
[0041] S=w1·d+w2·n+w3·b+w4·p
[0042] Wherein, w1 is the weight of crowding improvement; w2 is the weight of non-fragmentation improvement; w3 is the weight of optimal fitness improvement; w4 is the weight of population average fitness improvement; d represents crowding improvement; n represents non-fragmentation improvement; b represents optimal fitness improvement; p represents population average fitness improvement;
[0043] d represents the expression of congestion improvement as follows:
[0044]
[0045] In the formula, d t (i) represents the crowding distance of the i-th individual in the current generation population; represents the average crowding distance of all individuals in the current generation population; d1(i) represents the crowding distance of the i-th individual in the initial generation population; Represents the average crowding distance of all individuals in the initial generation population;
[0046] The expression for the improvement in the non-fragmentation quantity is as follows:
[0047]
[0048] Where Nt is the number of non-dominated solutions of the contemporary population; N1 is the number of non-dominated solutions of the initial generation population;
[0049] The expression for the best fitness improvement is as follows:
[0050] b=0.5·b f1 +0.5 bf2
[0051] Where b f1 , b f2 It represents the ratio of the optimal fitness value of the two fitness functions in the current generation to the optimal fitness value in the initial generation, that is, the ratio of the maximum values:
[0052] The expression for the improvement of the average fitness of the population is as follows:
[0053] p=0.5·p f1 +0.5·p f2
[0054] In the formula, p f1 、p f2 It represents the ratio of the mean of the two fitness functions in the current generation to the mean of the fitness functions in the initial generation. The expression is as follows:
[0055]
[0056] In the formula, p is the state value, which indicates the overall improvement of the fitness of the current generation population. This indicator reflects the comprehensive improvement of the fitness of the population after several generations.
[0057] S8, input the result of S7 into the Q-learning algorithm, construct the reward function, calculate the Q value, and update the Q value table;
[0058] The Q value update formula is as follows:
[0059]
[0060] Where s is the current state, a is the current action, r is the immediate reward, α is the learning rate, γ is the discount factor, and s′ is the next state;
[0061] The reward function is constructed as follows:
[0062] R=R1+R2
[0063] In the formula, the setting of R1 is based on
[0064] If n=1, R1=0,
[0065] If n<1, R1=-1,
[0066] If n>1, R1=1;
[0067] R2 setting basis
[0068] If d=1, R2=0,
[0069] If d<1, R2=-1,
[0070] If d>1, R2=1;
[0071] S9, such as Figure 3 As shown, according to the Q value table and ε-greedy strategy learned in S8, actions are selected, and the crossover rate and mutation rate of the algorithm are adjusted; and S4 to S9 are repeated until the maximum number of generations is reached;
[0072] That is, the ε-greedy strategy is run based on the initially set genetic algorithm (NSGA-II) parameters, and actions are randomly selected under probability ε to promote exploration, and the action with the largest Q value is selected under probability 1-ε. Then, action operations are randomly performed based on a set of initialization Pc and Pm values set during initialization. One action is performed in one iteration, and each action corresponds to a parameter adjustment. The design of the action space is as follows:
[0073] Action 1: Pc increases by 0.1;
[0074] Action 2: Pc decreases by 0.1;
[0075] Action 3: Keep Pc unchanged;
[0076] Action 4: Pm increases by 0.01;
[0077] Action 5: Pm decreases by 0.01;
[0078] Action 6: Keep Pm unchanged;
[0079] Ensure that Pc and Pm are always between 0 and 1. If the adjusted value is out of range, it will be clamped to the boundary value;
[0080] S10, the algorithm iteration ends and the final optimization result is output.
[0081] Beneficial effects of the present invention:
[0082] The present invention provides a more advanced and efficient optimization solution in optimizing the closed-loop supply chain network by introducing and improving the non-dominated sorting genetic algorithm (NSGA-II) and combining it with the adaptive control mechanism of the reinforcement learning algorithm.
[0083] The present invention successfully achieves an efficient balance between the dual objectives of supply chain cost and sustainability through an improved non-dominated sorting genetic algorithm and reinforcement learning mechanism. Compared with the existing technology, it demonstrates excellent application value in terms of rapid convergence and rich solution set selection in complex environments.
[0084] The present invention improves the diversity and quality of solutions, that is, through the comprehensive evaluation of congestion, the ratio of the number of non-dominated solutions and fitness, the algorithm can more comprehensively judge the quality of solutions. The improved NSGA-II algorithm can more effectively approach the Pareto frontier and maintain better diversity in the solution space, thereby providing a high-quality solution set that better meets actual needs.
[0085] The present invention reduces the risk of local optimal solutions. Traditional optimization methods often fall into local optimal solutions, while the present invention adopts a hybrid initialization strategy to avoid sample aggregation while introducing some randomness, which helps the algorithm to jump out of potential local optimal solutions, ensure the breadth of the search, and help find a more global optimization solution.
[0086] The performance of an algorithm usually depends on empirical parameter settings. The present invention enhances the adaptability and dynamic response capability of the algorithm by introducing a reinforcement learning mechanism, and dynamically adjusts its parameters according to the real-time status of the genetic algorithm. Ultimately, the convergence speed of the present invention is significantly lower than that of other metaheuristic algorithms. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Figure 1 is a flow chart of the method of the present invention;
[0088] Figure 2 A schematic diagram of the electronic assembly closed-loop supply chain network of the present invention;
[0089] Figure 3 It is the algorithm flow chart of the present invention;
[0090] Figure 4 This is a comparison image of the Pareto front solution set of the present invention and other algorithms. DETAILED DESCRIPTION
[0091] The present invention is further described in detail below in conjunction with specific embodiments.
[0092] like Figure 1 As shown, a closed-loop supply chain network optimization method based on an improved non-dominated sorting genetic algorithm includes the following steps:
[0093] S1. Establish the objective function and constraints of the closed-loop supply chain network optimization model with cost and carbon emissions as optimization targets;
[0094] In this embodiment, a fourteen-day production cycle is considered, wherein the duration of each cycle is one day.
[0095] like Figure 2 As shown in the figure, the closed-loop supply chain network optimization model includes: 9 raw materials (6 parts and 3 sub-assemblies), 3 intermediate products, 3 semi-finished products and 3 finished products;
[0096] Specifically, the forward production process consists of suppliers, production units and warehouses. In forward logistics, three external suppliers provide the materials that need to be processed: parts and sub-assemblies. The procurement cost of each supplier is different. Different types of parts are transported to production units A~C and processed into different intermediate products. At the same time, different types of sub-assemblies are transported to production unit D and assembled with the intermediate products to make semi-finished products. In the final production stage, semi-finished products are transported to production units E~G and processed into finished products. Finally, qualified finished products will be transported to product warehouses and finally sold to customers.
[0097] Reverse recycling consists of disassembly units, remanufacturing, reprocessing and disposal units. For unqualified semi-finished products and finished products, the disassembly unit will break them down into parts, sub-assemblies and intermediate products. The remanufacturing unit will process unqualified intermediate products and sub-assemblies. The reprocessing unit repairs unqualified parts. All waste generated during the recycling process will be sent to the disposal unit for processing. Usable products will be sent to the corresponding warehouse for reuse;
[0098] The objective functions include: cost objective function and second objective function;
[0099] The expression of the cost objective function is as follows:
[0100]
[0101] Where y1 represents the minimum total cost of the closed-loop supply chain network; t * represents the cycle; T represents the total cycle; v represents the vth supplier; V represents the total number of suppliers; m represents the product; M1 represents the part; M2 represents the subassembly; the number of parts purchased from the vth supplier is expressed as The corresponding purchase price is C m,v ; The number of subcomponents is expressed as The corresponding purchase price is SC m,v ; When selecting a supplier, the first parameter Second parameter is 1, otherwise 0; j M represents the set of production units; j W represents the warehouse unit set; j R Represents a collection of recycling units; j m Represents the production unit; j w Indicates warehouse; r Represents the recovery unit; M(j m ) represents the product produced by the production unit; M(j w ) represents the products stored in the warehouse; M(j r ) indicates products recycled by a recycling unit; Indicates the quantity of products produced by the production unit; represents the production cost; Indicates the number of products stored in the warehouse; represents the storage cost; Indicates the quantity of products recycled by the recycling unit; represents the recovery cost; Indicates the transportation quantity of products between units and between units and warehouses; TC m,j,j' represents the transportation cost; l represents the emission reduction technology level; L represents the set of emission reduction technology levels; Indicates whether production unit j m The investment level is l for emission reduction technology; h represents the investment cost coefficient of emission reduction technology;
[0102] j M is the set of production units of product m; j W is the warehouse set of product m; j R is the collection of recycling units of product m; j is the total collection (including production units, warehouses, and recycling units); Represents the total purchase cost of parts and subassemblies; Represents the transportation quantity of products between units and between units and warehouses, TC m,j,j' For transportation costs;
[0103] Furthermore, the logistics flows considered in the transportation process include: product transportation from production units to warehouses, product transportation from warehouses to production units, transportation of defective products from warehouses to recycling units, product transportation between recycling units, and return of recycled materials from recycling units to warehouses; in addition, the production units are technologically upgraded and renovated, and each production unit is invested in emission reduction technologies before the supply chain is started. the total cost of investing in emission reduction technologies;
[0104] The expression of the carbon emission objective function is as follows:
[0105]
[0106] Where y2 represents the total carbon emissions of the minimized closed-loop supply chain network, including carbon emissions in the production process, carbon emissions in the recycling process, and carbon emissions in the transportation process; Q M Indicates the carbon emissions generated by producing a unit of product; Q R It represents the carbon emissions generated by recycling a unit product. For the calculation of carbon emissions during transportation, not only the weight of the unit product wg is considered. m , and also considers the transportation distance between units and between units and warehouses. j,j', where j represents a unit or warehouse, and j' represents (transported to) another unit or warehouse, so as to calculate the total carbon emissions during transportation; Q T That is, the carbon emissions generated by transporting a unit weight (kg) of product per kilometer;
[0107] The constraints include: supplier selection, maximum supply capacity, production capacity, recycling capacity, transportation process, input and output material balance of each unit, warehouse inventory balance and demand satisfaction constraints.
[0108] S2, setting parameters for the closed-loop supply chain network optimization model and initializing the parameters of the genetic algorithm;
[0109] Specifically, information including: various cost values, cycle numbers and demand quantities are loaded into the S1 model; in which the finished product demand quantity of each cycle in the 14 production cycles is randomly generated in the range of 700 to 1000;
[0110] In this embodiment, one cycle is set to 1 day, and the finished product demand for a total of 14 cycles is calculated;
[0111] Initialize the genetic algorithm (NSGA-II) parameters with the number of iterations = 200, the population size (total number of individuals) N = 80, the dimension of the variable d = 2779, the initial crossover probability Pc is set to 0.6, and the mutation probability Pm is set to 0.05;
[0112] The closed-loop supply chain network optimization model is solved using the NSGA-II algorithm;
[0113] S3, generate the initial population by mixing random initialization and Latin hypercube sampling method;
[0114] The present invention is a mixed integer model, each individual in the population represents a decision scheme, the decision scheme is a solution, and each individual consists of a set of integer variables and binary decision variables (0 / 1);
[0115] Integer variables include: number of products purchased, number of products produced, number of products shipped, number of products recycled, and number of products in stock;
[0116] Binary decision variables include: supplier selection, whether the production unit invests in emission reduction technology;
[0117] Furthermore, for each integer variable, a hybrid generation method of Latin hypercube sampling and random initialization is used;
[0118] Furthermore, for each binary decision variable, a random initialization method is used to generate a number of 0 or 1;
[0119] That is, the present invention adopts a hybrid initialization strategy to generate an initial population through random initialization and Latin hypercube sampling method, thereby ensuring uniform distribution of sample points in the decision space and improving population diversity;
[0120] The steps to generate the initial population using a mixture of random initialization and Latin hypercube sampling are as follows:
[0121] S3.1. According to the total number of individuals in the population N and the individual dimension d, set the variable of a certain dimension in the individual to X z , then the value range of the variable is Where z = 1, 2, ..., d;
[0122] S3.2, divide the obtained dimension into equally spaced intervals with the same number as the total number of individuals N, and randomly select a point in each interval, where non-integer numbers need to be rounded off;
[0123] The partition interval expression is as follows:
[0124]
[0125] S3.3, let the selection result of each interval be (represents a point randomly selected in the uth interval), then the sample point can be expressed as All sample points are merged into the initial group, and the value of the sampling point is X LHS ={X1,X2,...,X N};
[0126] S3.4, based on the Latin hypercube sampling process, the present invention uses the Latin hypercube sampling method to generate N / 2 individuals for integer variables to ensure that each integer variable is evenly distributed in the multidimensional space; the other N / 2 individuals are generated by random initialization to avoid the aggregation of samples while introducing some randomness, which helps the algorithm to jump out of potential local optimality and ensure the extensiveness of the search;
[0127] S4, setting the inverse of the objective function as the fitness function, and calculating the fitness function; performing fast non-dominated sorting on the initial population, calculating the crowding distance and constructing the Pareto frontier;
[0128] The fitness function includes: the inverse of the cost objective function and the inverse of the second objective function
[0129] The fast non-dominated sorting of the initial population, calculation of the crowding distance and construction of the Pareto front are automatically generated by the computer;
[0130] S5, setting the current population in S4 as the parent population, generating a new population through the parent population and calculating the best fitness function, the mean crowding distance and the number of non-dominated solutions of the new population;
[0131] Specifically, individuals with lower non-dominated ranks are preferentially selected from the parent population; if the ranks are the same, individuals with larger crowding distances are selected for crossover and mutation; then, after performing crossover and mutation operations, the offspring population is generated and merged with the parent population to form a new population;
[0132] Furthermore, for the new population, the fitness function value is calculated again Calculate the mean of the fitness function of the population Calculate the optimal fitness function value. The present invention selects the maximum value as the optimal fitness function value, that is, Perform fast non-dominated sorting and calculate the crowding distance d of each individual t (i) The mean crowding distance of all individuals in the population and the number of non-dominated solutions N t .
[0133] S6, determine whether the current generation is greater than the maximum number of iterations, if yes, execute step S10, if not, execute step S7;
[0134] The judgment process is automatically identified by the computer;
[0135] S7, calculating the comprehensive status value according to the result of S5;
[0136] The expression of the comprehensive state value S is as follows:
[0137] S=w1·d+w2·n+w3·b+w4·p
[0138] Wherein, w1 is the weight of crowding improvement; w2 is the weight of non-fragmentation improvement; w3 is the weight of optimal fitness improvement; w4 is the weight of population average fitness improvement; d represents crowding improvement; n represents non-fragmentation improvement; b represents optimal fitness improvement; p represents population average fitness improvement; in this embodiment, w1=w2=w3=w4=0.25;
[0139] Furthermore, d represents the improvement in crowding. By comparing the crowding distance deviation between the current generation and the initial generation, the improvement in the diversity of the population between generations is measured, which can help the optimization algorithm better evaluate and adjust the distribution and diversity of the population in the multi-objective space. Diversity is crucial to avoid premature convergence, and the improvement in crowding can show the ability of the population to effectively explore in the search space. The expression is as follows:
[0140]
[0141] In the formula, d t (i) represents the crowding distance of the i-th individual in the current generation population; represents the average crowding distance of all individuals in the current generation population; d1(i) represents the crowding distance of the i-th individual in the initial generation population; Represents the average crowding distance of all individuals in the initial generation population;
[0142] Furthermore, the expression for the improvement of the non-fragmentation quantity is as follows:
[0143]
[0144] Where Nt is the number of non-dominated solutions of the contemporary population; N1 is the number of non-dominated solutions of the initial generation population;
[0145] This design evaluates the ability of the population to obtain excellent solutions. The increase in the number of non-dominated solutions indicates the progress of the population on the Pareto frontier, which can be used as an important indicator to evaluate the effectiveness of the algorithm;
[0146] Furthermore, the expression of the optimal fitness improvement is as follows:
[0147] b=0.5·b f1 +0.5 b f2
[0148] Where b f1 、b f2 It represents the ratio of the optimal fitness value of the two fitness functions in the current generation to the optimal fitness value in the initial generation, that is, the ratio of the maximum values:
[0149] b is a comprehensive indicator used to measure the degree of improvement in the fitness of the best individual of the current generation compared with the best individual of the initial generation during the multi-objective optimization process. This indicator reflects the relative progress of the best individual of each generation in terms of the two objectives and emphasizes the evolution of the best individual of the population. It can quickly determine whether the optimization process is effective, because the improvement of the best individual is usually the direct goal of optimization.
[0150] Furthermore, the expression for the improvement of the average fitness of the population is as follows:
[0151] p=0.5·p f1 +0.5·p f2
[0152] In the formula, p f1 、p f2 It represents the ratio of the mean of the two fitness functions in the current generation to the mean of the fitness functions in the initial generation. The expression is as follows:
[0153]
[0154] In the formula, p is the state value, which indicates the overall improvement of the fitness of the current generation population. This indicator reflects the comprehensive improvement of the fitness of the population after several generations. In this way, the optimization effect of the algorithm on different objectives can be evaluated.
[0155] S8, input the result of S7 into the Q-learning algorithm, construct the reward function, calculate the Q value, and update the Q value table;
[0156] The Q value update formula is as follows:
[0157]
[0158] Where s is the current state, a is the current action, r is the immediate reward, α is the learning rate, γ is the discount factor, and s′ is the next state;
[0159] The reward function is constructed as follows:
[0160] R=R1+R2
[0161] In the formula, the setting of R1 is based on
[0162] If n=1, R1=0,
[0163] If n<1, R1=-1,
[0164] If n>1, R1=1;
[0165] R2 setting basis
[0166] If d = 1, R2 = 0,
[0167] If d<1, R2=-1,
[0168] If d>1, R2=1;
[0169] This reward function is designed to improve both the quantity and distribution quality of solutions, guide the algorithm to better explore and utilize the solution space in multi-objective optimization, ensure that the solution is not only optimal but also evenly distributed on the Pareto frontier, and strike a balance between exploration (finding diversified solutions) and utilization (optimizing the quality of solutions);
[0170] S9, such as Figure 3 As shown, according to the Q value table and ε-greedy strategy learned in S8, actions are selected, and the crossover rate and mutation rate of the algorithm are adjusted; and S4 to S9 are repeated until the maximum number of generations is reached;
[0171] That is, the ε-greedy strategy is run based on the initially set genetic algorithm (NSGA-II) parameters, and actions are randomly selected under probability ε (set to 0.2 in this embodiment) to promote exploration, and the action with the largest Q value is selected under probability 1-ε, and then the action operation is randomly performed based on a set of initialization Pc and Pm values set during initialization. One action is performed in one iteration, and each action corresponds to a parameter adjustment. The design of the action space is as follows:
[0172] Action 1: Pc increases by 0.1;
[0173] Action 2: Pc decreases by 0.1;
[0174] Action 3: Keep Pc unchanged;
[0175] Action 4: Pm increases by 0.01;
[0176] Action 5: Pm decreases by 0.01;
[0177] Action 6: Keep Pm unchanged;
[0178] Ensure that Pc and Pm are always between 0 and 1. If the adjusted value is out of range, it will be clamped to the boundary value;
[0179] S10, the algorithm iteration ends, and the final optimization result is output; that is, the optimal solution and the Pareto frontier optimal solution set image, such as Figure 4 As shown;
[0180] In order to find a solution that strikes a balance between cost and carbon emissions from the Pareto frontier optimal solution set, a simple linear weighted method is used to evaluate the comprehensive score of each solution obtained by the present invention and the NSGA-II algorithm, and the optimization results are compared by comparing the high and low comprehensive scores;
[0181] The expression of the comprehensive score is as follows:
[0182] Score = W5 × cost + W6 × carbon emissions
[0183] In the formula, W5 represents the weight of cost; W6 represents the weight of carbon emissions;
[0184] In this example, weight W5 = W6 = 0.5, indicating that cost and carbon emissions are equally important to the comprehensive score;
[0185] In this embodiment, all calculations are generated by computer code.
[0186] Finally, the optimal solution results are shown in Table 1 and Table 2;
[0187] It should be noted that in practical applications, the choice of weights may affect the results, and the weights need to be adjusted according to specific decision-making needs.
[0188] Table 1 Comparison of objective function 1 (total cost and other partial costs)
[0189]
[0190] Table 2 Comparison of objective function 2 (total carbon emissions and other carbon emissions)
[0191]
[0192] In this study, by analyzing the data in Table 1 and Table 2, we can clearly see the significant advantages of the present invention in closed-loop supply chain network optimization. According to Table 1, the total cost of the method of the present invention is 141,957,126 yuan, which saves 114,750 yuan compared to the 142,071,876 yuan of the traditional NSGA-II method. This cost difference is mainly due to the significant reduction in procurement costs, production costs, and transportation costs, which were reduced by 28,055 yuan, 33,286 yuan, and 84,230 yuan, respectively. This shows that by optimizing resource allocation and improving production efficiency, the method of the present invention effectively reduces the cost of raw materials and production links and achieves significant cost savings. Although the recovery cost is slightly higher than that of the NSGA-II algorithm, this does not weaken the overall cost-effectiveness, but instead highlights the importance of the present invention in environmental protection and sustainable resource utilization.
[0193] In Table 2, the total carbon emissions of the method of the present invention are 86753.5Kg, while the traditional NSGA-II method is 86810.5Kg, a reduction of 57Kg. Although this difference seems small, any reduction is worthy of attention in the context of environmental protection, especially in large-scale production and industrial activities, where the cumulative effect may bring significant environmental impacts. This balance not only reflects cost savings in economic benefits, but also demonstrates its contribution to sustainable development in environmental protection.
[0194] Table 3 shows the comparison results of the present invention and other multi-objective algorithms for different performance indicators. All algorithms are run 10 times under the same hardware and software environment, and the average is taken. For algorithms involving crossover and mutation operations, such as the present invention and NSGA-II, the initial crossover rate is set to 0.6 and the mutation rate is 0.05. The additional parameters of other algorithms use the default values.
[0195] Table 3 Comparison of performance indicators of the present invention and other multi-objective algorithms
[0196]
[0197] In Table 3, IGD (Inverse Generational Distance): This indicator is used to evaluate the average distance between the non-dominated solution set obtained by the algorithm and the true Pareto optimal solution set. The smaller the IGD value, the closer the solution set generated by the algorithm is to the true Pareto frontier, and the better the optimization effect. The IGD value of the present invention is 0.1427, which is better than other algorithms such as NSGA-II (0.1913) and MOEAD (0.1642), indicating that the diversity and convergence of the present invention are better, showing better Pareto frontier coverage and optimization solution quality.
[0198] SM (spatial evaluation method): This index is used to measure the distribution degree of non-dominated solutions, that is, the diversity and uniformity of solutions. The smaller the SM value, the more uniform the distribution of the solution set and the better the coverage of the decision space. The SM value of the present invention is 0.5094, which is better than the SM values of other algorithms, indicating that the solutions generated by it are more evenly distributed, which helps decision makers obtain diversified choices.
[0199] NPS (Number of Non-dominated Solutions): refers to the number of non-dominated solutions (i.e., Pareto frontier solutions) generated by the algorithm. The number of solutions generated by the present invention is 23, indicating that it has a strong balance between global search capability and local search capability. This helps the algorithm find high-quality solutions in complex optimization problems. It provides decision makers with more options and improves decision-making flexibility.
[0200] CPU time: This indicator is the running time of the algorithm for 100 iterations. The running time of the present invention is 125.8960 seconds, which is faster than the running time of other algorithms, indicating that the present invention has higher computational efficiency and faster convergence speed.
[0201] In summary, the present invention is aimed at the closed-loop supply chain network optimization problem. The traditional metaheuristic algorithm does not fully explore the solution space and is prone to fall into local optimality. Its key parameters cannot be effectively adjusted dynamically during the calculation process to cope with the complexity of the problem. The present invention constructs an electronic assembly closed-loop supply chain optimization network and provides an improved non-dominated sorting genetic algorithm, which can adaptively adjust key parameters, enhance the comprehensive exploration capability of the solution space, and improve the quality and diversity of solutions. This algorithm not only improves the optimization efficiency, but also ensures the balance and coordination between different objectives in the solution set, and ultimately shows excellent performance and broad application potential in closed-loop supply chain network optimization.
Claims
1. A closed-loop supply chain network optimization method based on an improved non-dominated sorting genetic algorithm, characterized by: The following steps are involved: S1. Establish the objective function and constraints of the closed-loop supply chain network optimization model with cost and carbon emissions as optimization targets; S2, setting parameters for the closed-loop supply chain network optimization model and initializing the parameters of the genetic algorithm; S3, generate the initial population by mixing random initialization and Latin hypercube sampling method; S4, setting the inverse of the objective function as the fitness function, and calculating the fitness function; performing fast non-dominated sorting on the initial population, calculating the crowding distance and constructing the Pareto frontier; S5, setting the current population in S4 as the parent population, generating a new population through the parent population and calculating the best fitness function, the mean crowding distance and the number of non-dominated solutions of the new population; S6, determine whether the current generation is greater than the maximum number of iterations, if yes, execute step S10, if not, execute step S7; S7, calculating the comprehensive status value according to the result of S5; S8, input the result of S7 into the Q-learning algorithm, construct the reward function, calculate the Q value, and update the Q value table; S9, select actions according to the Q value table and ε-greedy strategy learned in S8, and adjust the crossover rate and mutation rate of the algorithm; and repeat S4 to S9 until the maximum number of generations is reached; S10, the algorithm iteration ends and the final optimization result is output.
2. The closed-loop supply chain network optimization method of the improved non-dominated sorting genetic algorithm according to claim 1 is characterized by: The objective functions of the closed-loop supply chain network optimization model include: a cost objective function and a carbon emission objective function; The expression of the cost objective function is as follows: Where y1 represents the minimum total cost of the closed-loop supply chain network; t * represents the cycle; T represents the total cycle; v represents the vth supplier; V represents the total number of suppliers; m represents the product; M1 represents the part; M2 represents the subassembly; the number of parts purchased from the vth supplier is expressed as The corresponding purchase price is C m,v ; The number of subcomponents is expressed as The corresponding purchase price is SC m,v ; When selecting a supplier, the first parameter Second parameter is 1, otherwise 0; j M represents the set of production units; j W represents the warehouse unit set; j R Represents a collection of recycling units; j m Represents the production unit; j w Indicates warehouse; r Represents the recovery unit; M(j m ) represents the product produced by the production unit; M(j w ) represents the products stored in the warehouse; M(j r ) indicates products recycled by a recycling unit; Indicates the quantity of products produced by the production unit; represents the production cost; Indicates the number of products stored in the warehouse; represents the storage cost; Indicates the quantity of products recycled by the recycling unit; represents the recovery cost; Indicates the transportation quantity of products between units and between units and warehouses; TC m,j,j' represents the transportation cost; l represents the emission reduction technology level; L represents the set of emission reduction technology levels; Indicates whether production unit j m The investment level is l for emission reduction technology; h represents the investment cost coefficient of emission reduction technology; The expression of the carbon emission objective function is as follows: Where y2 represents the total carbon emissions of the minimized closed-loop supply chain network, including carbon emissions in the production process, carbon emissions in the recycling process, and carbon emissions in the transportation process; Q M Indicates the carbon emissions generated by producing a unit of product; Q R It represents the carbon emissions generated by recycling a unit product. For the calculation of carbon emissions during transportation, not only the weight of the unit product wg is considered. m , and also considers the transportation distance between units and between units and warehouses. j,j' , where j represents a unit or warehouse, and j' represents (transported to) another unit or warehouse, so as to calculate the total carbon emissions during transportation; Q T That is, the carbon emissions generated by transporting a unit weight (kg) of product per kilometer; The constraints include: supplier selection, maximum supply capacity, production capacity, recycling capacity, transportation process, input and output material balance of each unit, warehouse inventory balance and demand satisfaction constraints.
3. The closed-loop supply chain network optimization method of the improved non-dominated sorting genetic algorithm according to claim 1 is characterized by: In the initial population generated by a mixture of random initialization and Latin hypercube sampling, each individual in the population represents a decision plan, and each individual consists of a set of integer variables and binary decision variables; Integer variables include: the number of purchased products, the number of produced products, the number of shipped products, the number of recycled products, and the number of product inventory; for each integer variable, Latin hypercube sampling and random initialization methods are used to generate the variable; The binary decision variables include: the choice of supplier, whether the production unit invests in emission reduction technology; a random initialization method is used to generate a number of 0 or 1; The steps to generate the initial population using a mixture of random initialization and Latin hypercube sampling are as follows: S3.
1. According to the total number of individuals in the population and the individual dimension, set the variable of a certain dimension in the individual to X z , then the value range of the variable is Where z = 1, 2, ..., d; S3.2, divide the obtained dimension into equally spaced intervals with the same number as the total number of individuals, and randomly select a point in each interval, where non-integer numbers need to be rounded off; The expression for partitioning intervals is as follows: S3.3, let the selection result of each interval be The sample points are represented as All sample points are merged into the initial group, and the value of the sampling point is X LHS ={X1,X2,…,X N }; S3.
4. Based on the Latin hypercube sampling process, for integer variables, the Latin hypercube sampling method is used to generate N / 2 individuals; the other N / 2 individuals are generated by random initialization.
4. The closed-loop supply chain network optimization method of the improved non-dominated sorting genetic algorithm according to claim 1 is characterized by: The expression for calculating the comprehensive status value through the result of S5 is as follows: S=w1·d+w2·n+w3·b+w4·p In the formula, w1 is the weight of crowding improvement; w2 is the weight of non-fragmentation improvement; w3 is the weight of optimal fitness improvement; w4 is the weight of population average fitness improvement; d represents crowding improvement; n represents non-fragmentation improvement; b represents optimal fitness improvement; and p represents population average fitness improvement.
5. The closed-loop supply chain network optimization method of the improved non-dominated sorting genetic algorithm according to claim 1 is characterized by: The result of S7 is input into the Q-learning algorithm, a reward function is constructed, the Q value is calculated, and the Q value table is updated. The Q value update formula is as follows: Where s is the current state, a is the current action, r is the immediate reward, α is the learning rate, γ is the discount factor, and s′ is the next state; The reward function is constructed as follows: R=R1+R2 In the formula, the setting of R1 is based on If n=1, R1=0, If n<1, R1=-1, If n>1, R1=1; R2 setting basis If d=1, R2=0, If d<1, R2=-1, If d>1, R2=1.
Citation Information
Cited By
Method for automatically optimizing continuous intelligent microreactor platform
CN120493571A
Flexible cable assembly parameter optimization method and device based on Q-Learning genetic algorithm and medium
CN120805732A
Flexible cable assembly parameter optimization method, device and medium based on q-learning genetic algorithm
CN120805732B