New energy automobile supply chain multi-objective combination optimization method, device, equipment and medium
By adopting a multi-objective combination optimization method based on deep reinforcement learning in the new energy vehicle supply chain, combined with the NSGA-III algorithm, the balance of economic costs, carbon emissions and risk losses in the supply chain is solved, and more efficient optimization and better solution quality are achieved.
Patent Information
- Application Number
- CN202510096969.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the multi-target optimization of the new energy vehicle supply chain, it is difficult to effectively balance economic costs, carbon emissions and transportation risk losses, especially in the problem of site selection and distribution, the impact of carbon emissions and recycling during transportation cannot be fully considered.
The multi-objective combination optimization method of the new energy vehicle supply chain based on deep reinforcement learning is adopted. By building a multi-objective mathematical model, combining the NSGA-III algorithm and the deep reinforcement learning model, the site selection and allocation plan are optimized to achieve comprehensive optimization of economic costs, carbon emissions and risk losses.
A more comprehensive optimization effect is achieved, the quality of optimization efficiency and solutions is improved, and the global optimal solution can be found in a short time, avoiding the occurrence of local optimal solutions.
Smart Images

Figure CN120069399A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of supply chain management, and particularly to a multi-objective combined optimization method, device, equipment and medium for a new energy vehicle supply chain. Background Art
[0002] With the wide promotion of the application of new energy vehicles globally, in the process of optimizing the supply chain of new energy vehicles, problems in many aspects such as site selection, allocation, and route planning are involved. And these problems usually not only involve economic costs, but also multi-objective optimization tasks such as carbon emissions, environmental benefits, and transportation accidents. Most traditional supply chain optimization methods mainly focus on a single objective, such as minimizing transportation costs and improving distribution efficiency. However, in practical applications, especially in the new energy vehicle supply chain, see Figure 3 , it is often necessary to consider multiple objectives simultaneously, including environmental benefits (such as carbon emissions), economic costs, transportation accident risks, etc. Therefore, such problems are usually modeled as multi-objective optimization problems.
[0003] Currently, there are already some multi-objective algorithms for supply chain optimization. Taking the "Multi-Supply-Point Emergency Material Site Selection Optimization Method Considering Timeliness and Fairness" in the patent application number CN202210584497.0 as an example, this method constructs a multi-objective emergency material distribution model, focuses on timeliness and fairness, and realizes the optimal allocation of emergency materials by optimizing route planning. This method adopts a strategy of combining a mutation operator based on an initialized individual strategy with a local search operator, which improves the optimization effect of the distribution plan. However, existing methods mostly focus on problems of timeliness or fairness, and there are still certain limitations for multi-objective problems that comprehensively consider risk losses and carbon emissions.
[0004] Generally speaking, although the current technical solutions can solve the optimization problems of some specific objectives (such as timeliness, cost, etc.), in multi-objective optimization, how to balance risk losses, carbon emissions and economic costs, especially in the context of new energy vehicle supply chain site selection and allocation, is still an urgent problem to be solved. The existing technologies fail to fully consider the impact of carbon emissions during the transportation process and carbon emissions in the recycling link on the entire supply chain, and there is a large uncertainty in the quantification of risk losses. Especially when considering multiple objectives and complex constraints, the existing optimization algorithms are difficult to achieve a good balance between the quality of solutions and computational efficiency. Summary of the Invention
[0005] To at least partly solve one of the technical problems existing in the prior art, an object of the present invention is to provide a multi-objective combined optimization method, device, equipment and medium for a new energy vehicle supply chain based on deep reinforcement learning.
[0006] The first technical solution adopted by the present invention is as follows:
[0007] A multi-objective combined optimization method for a new energy vehicle supply chain, comprising the following steps:
[0008] Obtain the geographical location coordinates of cities, and construct a multi-objective mathematical model for the location-allocation problem considering transportation risks and carbon emissions;
[0009] Determine the constraint conditions and objective functions of the multi-objective mathematical model;
[0010] Use NSGA-III to optimize and solve the multi-objective mathematical model to obtain an initial solution set of the new energy vehicle location and distribution plan;
[0011] Build a deep reinforcement learning model based on the crossover operator and mutation operator of NSGA-III;
[0012] Use the initial solution set to train the deep reinforcement learning model to improve the solving performance of NSGA-III. After completion of training, obtain a combined optimization model based on deep reinforcement learning;
[0013] Apply the combined optimization model based on deep reinforcement learning to solve the actual location-allocation scenario to obtain an optimal new energy vehicle location and distribution plan.
[0014] Furthermore, there are three objective functions of the multi-objective mathematical model, which are respectively: first, the minimum economic cost, including supply cost, location cost, distribution cost and recycling cost; second, the minimum carbon emissions during transportation under different scenarios; third, the minimum risk loss during transportation under different scenarios;
[0015] The constraint conditions of the multi-objective mathematical model include: facility location cost constraint, supplier supply constraint, factory production capacity constraint, distribution center storage capacity constraint, retailer demand constraint, recycling center processing capacity constraint.
[0016] Furthermore, the working mode of the combined optimization model based on deep reinforcement learning is as follows:
[0017] A1. Chromosome encoding: Divide the chromosome into three segments for encoding. The first segment is the candidate distribution center, and each gene corresponds to the supply factory of each distribution center; the second segment is the retailer, and each gene corresponds to the distribution center of each retailer; the third segment is the recycling center, and each gene corresponds to the recycling center of the corresponding retailer;
[0018] A2. Initialize the population: According to the chromosome encoding in step A1, randomly generate a set of initial solutions to obtain an initialized population;
[0019] A3. Non-dominated sorting: Perform non-dominated sorting on the individuals in the population and assign different non-dominated ranks;
[0020] A4. Selective Optimization: Select according to the non-dominated sorting result and crowding degree, and retain individuals with better diversity.
[0021] A5. Crossover and Mutation: According to the virtual fitness value, replicate the population after non-dominated sorting and perform selection operations; given the crossover probability, perform crossover operations; randomly select mutation genes with a given mutation probability; in the crossover operation, exchange some decision variables of two parent individuals to generate new offspring individuals; in the mutation operation, randomly change the location of a certain factory or adjust the transportation volume of a certain transportation route.
[0022] A6. Main Process: Combine the initial population with its offspring population to form a new population, perform non-dominated sorting on it and generate a series of non-dominated sets, and calculate the crowding degree; gradually put the newly generated individuals into the new parent population until the size of the parent population exceeds the preset population size; perform crowding degree sorting on the individuals in the combined new population, and select the dominant solutions on the Pareto front; return to execute steps A3 to A6 to form a new offspring population until the set iteration threshold is reached.
[0023] Further, the selective optimization specifically includes:
[0024] Calculate the objective components according to the objective function.
[0025] Find all non-dominated individuals in the population according to the objective components, and assign them a shared virtual fitness value to form the first-level non-dominated individual set; continue to classify the other individuals in the population according to the non-dominated relationship and assign new virtual fitness values to form higher-level non-dominated individual sets until all individuals are classified.
[0026] Further, during the crossover process, adjust the crossover probability and method according to the objective differences of risk, carbon emissions, and cost and the distribution of the current solution set; through more refined adjustment, make the crossover effectively explore the multi-objective solution space, ensure that more diverse candidate solutions can be generated during optimization, and maintain the uniform distribution of the Pareto front.
[0027] Further, during the mutation process, based on the positive correlation and loss difference between the carbon emissions and risk losses of the current solution set, make adjustments based on the density information of the space on the basis of random mutation.
[0028] Further, the construction of the deep reinforcement learning model with NSGA-III-based crossover operator and mutation operator includes:
[0029] Accelerate and optimize the initial solution set by combining deep learning, and continuously improve and optimize the crossover operator and selection operator in NSGA-III through deep learning algorithms to improve the solution efficiency and accuracy.
[0030] The second technical solution adopted by the present invention is:
[0031] A multi-objective combined optimization device for a new energy vehicle supply chain, comprising:
[0032] A model construction module, configured to obtain the geographical location coordinates of a city and construct a multi-objective mathematical model for the location-allocation problem considering transportation risks and carbon emissions;
[0033] A target determination module, configured to determine the constraint conditions and objective function of the multi-objective mathematical model;
[0034] An initial solution module, configured to use NSGA-III to optimize the multi-objective mathematical model and obtain an initial solution set of the new energy vehicle location and distribution plan;
[0035] A reinforcement learning module, configured to build a deep reinforcement learning model based on the crossover operator and mutation operator of NSGA-III;
[0036] A model training module, configured to train the deep reinforcement learning model using the initial solution set, and obtain a combined optimization model based on deep reinforcement learning after completion of training;
[0037] A model application module, configured to use the combined optimization model based on deep reinforcement learning to solve the actual location-allocation scenario and obtain an optimal new energy vehicle location and distribution plan.
[0038] The third technical solution adopted by the present invention is:
[0039] An electronic device, the electronic device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above-mentioned multi-objective combined optimization method for a new energy vehicle supply chain.
[0040] The fourth technical solution adopted by the present invention is:
[0041] A computer-readable storage medium, and at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the above-mentioned multi-objective combined optimization method for a new energy vehicle supply chain.
[0042] The fifth technical solution adopted by the present invention is:
[0043] A computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute the above-mentioned method.
[0044] The beneficial effects of the present invention are as follows: The present invention adopts a multi-objective optimization model, which can simultaneously consider economic costs, carbon emissions, and transportation risk losses, bringing a more comprehensive optimization effect. The present invention adopts an optimization framework combining DRL and NSGA-III, bringing higher optimization efficiency and solution quality. Through the selection, crossover, and mutation processes in DRL-MOA, it has the characteristic of significantly accelerating the convergence speed of the algorithm, avoiding the emergence of local optimal solutions, and thus being able to find the global optimal solution in a relatively short time. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the related technical solution drawings in the embodiments of the present invention or the prior art. It should be understood that the drawings introduced below are only for conveniently and clearly expressing some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0046] Figure 1 It is a combined optimization framework diagram based on deep reinforcement learning in an embodiment of the present invention;
[0047] Figure 2 It is a Pareto optimal solution set diagram of NSGA-III in an embodiment of the present invention;
[0048] Figure 3 It is a new energy vehicle supply chain structure diagram;
[0049] Figure 4 It is a step flowchart of the Pareto optimal solution set diagram of NSGA-III in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0051] In the description of the present invention, it should be understood that for the orientation description, such as up, down, front, back, left, right, etc., the orientation or positional relationship indicated is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0052] In the description of the present invention, the meaning of "several" is one or more, the meaning of "multiple" is two or more, and understandings such as "greater than", "less than", "exceeding", etc. do not include the recited number, and understandings such as "above", "below", "within", etc. include the recited number. If there is a description of "first" and "second", it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.
[0053] In the description of the present invention, unless otherwise clearly defined, words such as "set", "installed", "connected", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meaning of the above words in the present invention in combination with the specific content of the technical solution.
[0054] Term Explanation:
[0055] DRL: Deep Reinforcement Learning, deep reinforcement learning.
[0056] NSGA-III: The third generation of non-dominated sorting genetic algorithm.
[0057] In view of the existing technical problems, the present invention proposes a combinatorial optimization framework based on deep reinforcement learning. First, NSGA-III is used to perform a preliminary solution on the multi-objective model to obtain a non-dominated initial solution set; then, DRL is combined to accelerate the optimization of the initial solution set. After continuously improving and optimizing the crossover and selection operators in NSGA-III, they are encapsulated into the DRL-MOA algorithm model to improve the overall solution efficiency. The specific steps include: constructing an initial individual population strategy, designing a mutation operator and a crossover operator, constructing a selection strategy and an acceptance criterion, outputting the optimal operator and evaluating it, so as to obtain the optimal new energy vehicle location and distribution plan. The three-objective multi-level new energy vehicle location and distribution multi-objective model proposed by the present invention considers the impact of transportation accidents on the path, defines the concept of risk loss, and at the same time, to reduce transportation carbon emissions, introduces the concept of carbon emission reduction, and details the relevant calculation methods. Through this multi-objective model, while ensuring the distribution efficiency, it can optimize the economic cost, reduce carbon emissions and reduce risk losses.
[0058] Embodiment 1
[0059] As Figure 4 shown, this embodiment provides a multi-objective combinatorial optimization method for a new energy vehicle supply chain, including the following steps:
[0060] S1. Obtain the geographical location coordinates of the city and construct a multi-objective mathematical model for the location allocation problem considering transportation risk and carbon emissions;
[0061] S2. Determine the constraint conditions and objective functions of the multi-objective mathematical model;
[0062] S3. Use NSGA-III to perform an optimal solution on the multi-objective mathematical model to obtain an initial solution set of the new energy vehicle location and distribution plan;
[0063] S4. Build a deep reinforcement learning model based on the crossover operator and mutation operator of NSGA-III;
[0064] S5. Use the initial solution set to train the deep reinforcement learning model to improve the solution performance of NSGA-III. After completion of training, a combinatorial optimization model based on deep reinforcement learning is obtained;
[0065] S6. Apply the combinatorial optimization model based on deep reinforcement learning to solve the actual location allocation scenario to obtain the optimal new energy vehicle location and distribution plan.
[0066] The method of this embodiment comprehensively considers economic costs, transportation risk losses, and carbon emissions. In the problem of selecting a distribution route, it is necessary to consider both the transportation costs and transportation accident risks of different routes, as well as the additional costs brought about by carbon emissions. Vehicles of different power types have different carbon emission coefficients per kilometer during transportation, which leads to differences in carbon emissions and transportation distances for different routes. At the same time, under the same scenario, there may be significant differences in the probability of a transportation accident occurring and the expected distance of an accident. In a certain scenario, if a route with a greater risk and a longer transportation distance is selected, when an accident occurs, the congested distance and time will also be longer, resulting in higher carbon emissions. Therefore, the selection of a transportation route needs to consider multiple factors such as transportation costs, carbon emissions, and risk losses simultaneously.
[0067] In addition, the selection of facility locations has an important impact on the optimization of distribution routes. The load capacity, energy type of the vehicle, and their related risk coefficients will also affect the number of transportation trips, carbon emissions, and the economic costs per unit product or per unit transportation distance. At the same time, for the location-allocation problem, traditional optimization methods have many limitations. The location-routing problem usually involves a huge solution space and complex constraint conditions, and it is necessary to balance multiple factors such as economic costs and transportation distances. Although NSGA-III, as a global optimization method, can search for the optimal solution in a relatively large solution space by simulating the process of natural selection and evolution. However, when solving the location-allocation problem, NSGA-III still faces the problems of slow convergence speed and high computational cost, especially in large-scale problems, where the convergence speed is slow and a long computational time is required.
[0068] Based on this, this embodiment proposes a hybrid optimization method (DRL-MOA) that combines DRL and NSGA-III. Through the interaction between the agent and the environment, DRL continuously optimizes the decision-making strategy according to the set reward signal, with a fast convergence speed and the ability to avoid the problem of local optimal solutions. Combining DRL with NSGA-III can effectively enhance the evolutionary process of NSGA-III, accelerate the convergence of the algorithm, avoid falling into local optimal solutions, and thus improve the quality of the solution. In this embodiment, DRL guides NSGA-III to perform more intelligent selection, crossover, and mutation operations, thereby accelerating the generation and optimization of solutions and improving the solving efficiency and accuracy.
[0069] In addition, the improvement points and technical points of this embodiment for NSGA-III are as follows: The core idea of NSGA-III is to select Pareto front solutions based on non-dominated sorting and crowding distance. This patent improves the crossover and mutation strategies in NSGA-III to enhance its performance in dealing with complex problems.
[0070] 1) Improvement of the crossover strategy: An adaptive crossover strategy is adopted to adjust the crossover method according to the distribution of the solution set, enhancing the diversity and uniformity of solutions and avoiding over-concentration in certain regions. In traditional NSGA-III, the crossover operation generates offspring individuals by exchanging information between two parent individuals. However, in multi-objective optimization, it may not effectively maintain the diversity of solutions. Especially when the solution space of large-scale problems is relatively complex, there may be local search or over-concentration in certain regions, losing the ability to comprehensively explore the solution space. The improvement in this embodiment is to adjust the crossover probability and method during the crossover process according to the objective differences of risk, carbon emissions, and cost, as well as the distribution of the current solution set. Through more refined adjustment, the crossover can effectively explore the multi-objective solution space, ensuring that more diverse candidate solutions can be generated during optimization and maintaining the uniform distribution of the Pareto front. The innovation of this adaptive crossover strategy makes NSGA-III more flexible in the face of complex problems, avoiding the limitations of traditional methods in certain scenarios and enhancing the diversity and distribution of solutions.
[0071] 2) Improvement of the mutation strategy: Optimize the mutation process based on the density information in the objective space, increasing the search for sparse regions by reducing the variation amplitude in known dense regions, and enhancing the optimization effect and computational efficiency. In traditional mutation operations, individual genes are randomly changed within a large solution space range to avoid falling into local optimal solutions. In multi-objective optimization, random mutation may generate unnecessary solutions due to overly random search, increasing the computational cost and unable to fully guide the search towards the target direction. In this embodiment, during the mutation process, based on the positive correlation and loss difference between carbon emissions and risk losses in the current solution set, adjustments are made based on the density information of the space on the basis of random mutation. For example, when the solutions of carbon emissions and risk are dense, the mutation process will reduce the variation amplitude in this region, thereby increasing the search for sparse regions and avoiding excessive exploration of known high-density solution regions. The innovation of this strategy lies in dynamically adjusting the mutation strategy through the relationship between objectives, making the search more intelligent and directional, while improving the diversity and uniformity of solutions, and thus enhancing the overall optimization effect.
[0072] The following combines the attached Figure 1 to explain the above method in detail.
[0073] See Figure 1, the technical solution proposed in this embodiment mainly includes: 1) A multi-objective optimization method combining deep reinforcement learning and non-dominated sorting algorithm is proposed, which can optimize the selection of facility location, distribution route and distribution vehicle on the premise of comprehensively considering transportation cost, accident loss and carbon emission. 2) In the DRL-MOA algorithm model, DRL intelligently guides the selection, crossover and mutation processes in NSGA-III, thereby accelerating the convergence speed, avoiding local optimal solutions and improving the quality of the final solution. During the distribution process, considering the carbon emission differences and transportation risks of vehicles with different power energy sources, a reasonable objective function is formulated to achieve the comprehensive optimization of minimizing risk loss, carbon emission and transportation cost. 3) Through DRL-MOA, with the help of the strategy optimization ability of reinforcement learning, the solution efficiency of the location-allocation problem is further improved, ensuring the high efficiency and operability of the solution.
[0074] (a) Model design
[0075] Select a suitable optimization algorithm according to the characteristics of the problem in the combinatorial optimization framework. In the model design stage of the location-allocation problem, use NSGA-III to find the initial non-dominated solution set. In the model construction, the symbols shown in Table 1, Table 2 and Table 3 represent sets, parameters and decision variables.
[0076] Table 1 Explanation of set variable table
[0077]
[0078]
[0079] Table 2 Parameter variable table
[0080]
[0081] Table 3 Decision variable table
[0082]
[0083] Objective function: This embodiment considers the carbon emissions and risk losses in the recycling center and transportation process, and promotes its application in the transportation field by optimizing the location, route and vehicle combination. The model includes three optimization objectives: one is to minimize the economic cost, including supply cost, location cost, distribution cost and recycling cost; the second is to minimize the carbon emissions during transportation under different scenarios; the third is to minimize the risk losses during transportation under different scenarios. The expressions are as follows:
[0084]
[0085] Constraints: They are facility location cost constraint, supplier supply constraint, factory production capacity constraint, distribution center storage capacity constraint, retailer demand constraint, and recycling center processing capacity constraint respectively. The expressions are as follows:
[0086]
[0087] Among them, single-source supply constraint:
[0088]
[0089] Penalty constraint for transportation accidents caused by shortages of raw materials, products and services:
[0090]
[0091] Transportation flow balance constraint:
[0092]
[0093] Variable 0-1 constraint and non-negativity constraint:
[0094]
[0095] (b) Model training
[0096] The DRL method is used to optimize the location-allocation scheme obtained by NSGA-III to further improve the quality of the solution. First, multiple location problem instances are randomly generated, and NSGA-III is used to calculate the solutions of each instance and obtain the corresponding objective function values. Then, DRL is used to optimize the selection, mutation, and crossover operator operations in NSGA-III. After further improving the quality of the solution, it is stored as DRL-MOA. Among them, the key steps of using DRL for the location-allocation scheme of NSGA-III are as follows:
[0097] First, obtain the geographical location coordinates at the district level from Baidu Maps, calculate the distance using spherical geographical coordinates, substitute parameters such as the actual carbon emission data from enterprise research and example data from cited literature into the model, and obtain the preliminary location-allocation scheme of the new energy vehicle supply chain in Guangzhou, Shenzhen, Foshan and other places in Guangdong Province through NSGA-III, corresponding to the economic cost, risk loss, and carbon emission target values. To meet the demand for large-scale real-time solution, the economic cost of each candidate facility construction is further optimized, as well as the location-allocation scheme of the transportation risk coefficient and carbon emission coefficient of different vehicle models between facility entities as an appropriate state space; the crossover probability (0.3 - 0.8) of the NSGA-III crossover operator and the mutation probability (0.02 - 0.06) of the mutation operator are intelligently adjusted, and the strategy combinations of the mutation probability and crossover rate are continuously given as the action space; the reward function is designed according to the carbon emission and risk loss targets in the system scheme.
[0098] Secondly, during the training process of introducing the DRL model, whenever the agents in the DRL adjust the crossover operator and mutation operator, resulting in a reduction in carbon emissions and a decrease in risk losses, and the cost is lower than the threshold, the reward is greater; if the objective function deteriorates after the adjustment, a negative reward is given. For the scenarios where the production capacity of the vehicle factory is limited and the budgets of the factory, recycling center, and distribution center exceed the standard, penalties are given. The system will calculate the new reward and return it to the agent to further attempt in the unexplored area, adaptively challenging the parameters of the crossover operator and mutation operator, thereby optimizing the NSGA-III search process. Due to the constraints of large-scale problems, especially in aspects such as construction budget, production capacity limitation, and damage during transportation, which are very complex, the DRL algorithm is used to flexibly handle complex constraints for the designed penalty function, ensuring that while always meeting the constraints, it gradually learns how to adjust the selection operator and crossover operator, and automatically adjusts to select more potential solutions to improve the diversity and balance of the solutions.
[0099] Finally, through multiple iterations and feedback of the DRL agent, the strategy is gradually improved to obtain a near-optimal site allocation solution. The system will continuously compare the optimization results of each round with the preliminary solution obtained by NSGA-III, and analyze the improvement in carbon emissions and risk losses. If the DRL method effectively improves the initial solution, the training model is saved as the DRL-MOA model; if the DRL method cannot improve the initial solution, its parameter model is further optimized, and the average value of multiple trainings is continuously updated. Finally, within the allowable error probability and iteration time, the average objective optimization solution of the parameter-optimal model is selected, and then the next combined optimization solution is formed.
[0100] (c) Form the combined optimization solution
[0101] After the DRL-MOA model is established, a combined optimization solution is generated based on the parameters of the training model. Among them, in randomly generating a certain scale of site allocation problem instances, the NSGA-III model is used to calculate the solution of the site distribution problem, calculate the objective function value of this solution, and select the trained DRL model to update the operator of the NSGA-III model according to the characteristics of the problem to be optimized, obtaining a combined optimization solution at the second level. Among them, in this embodiment, the improved NSGA-III is used to solve the model to obtain the Pareto solution set. The specific steps include:
[0102] 1) Chromosome encoding. The chromosome is divided into three segments for encoding. The first segment is the candidate distribution center, and each gene corresponds to the supply factory of each distribution center; the second segment is the retailer, and each gene corresponds to the distribution center of each retailer. The third segment is the recycling center, and each gene corresponds to the recycling center of the corresponding retailer. For the randomly generated chromosome above, its validity needs to be verified.
[0103] 2) Initialize the population. According to the chromosome coding in the previous step, randomly generate a set of initial solutions to obtain the initialized population.
[0104] 3) Non-dominated sorting. Perform non-dominated sorting on the individuals in the population and assign different non-dominated ranks.
[0105] 4) Elite selection. Select according to the non-dominated sorting result and crowding distance, and retain individuals with better diversity. First, calculate the objective component f(s) + Q(s) according to the objective function formula.
[0106] Where is the set of objective functions, Q(s) = ε t (∑max(0, ∑z ij -P i )) + ε 2 (∑max(0, ∑z ij -S i )) + ε 3 (∑max(0, ∑z ij -D i )) + ε 4 (∑max(0, ∑w rc -R rc )) is the total penalty function for violating the capacity constraints, including the penalty functions for the factory supply capacity v 1 (s), the distribution center supply capacity v 2 (s), the retailer demand v 3 (s) and the recycling center processing capacity v 4 (s). Then, find all non-dominated individuals in the population according to the three objective function components, and assign them a shared virtual fitness value to form the first-level non-dominated individual set. In addition, continue to classify the other individuals in the population according to the non-dominated relationship and assign new virtual fitness values to form higher-level non-dominated individual sets until all individuals are classified.
[0107] 5) Crossover and mutation. According to the virtual fitness value, copy the population after non-dominated sorting and perform selection operations. Then, given the crossover probability, perform crossover operations. Subsequently, randomly select mutation genes given the mutation probability. In the crossover operation, exchange some decision variables of the two parent individuals to generate new offspring individuals. In the mutation operation, the location of a certain factory can be randomly changed or the transportation volume of a certain transportation route can be adjusted.
[0108] 6) Main process. Combine the initial population with its offspring population to form a new population, perform non-dominated sorting on it and generate a series of non-dominated sets, and calculate the crowding degree. Gradually put the newly generated individuals into the new parent population until the size of the parent population exceeds the preset population size. Sort the individuals in the combined new population according to the crowding degree and select the dominating solutions on the Pareto front. Then, execute Steps 3 to 6 to form a new offspring population, and repeat the above process until the set iteration threshold is reached.
[0109] The following combines the attached Figure 2 drawings and specific embodiments to supplement and explain the method of this embodiment.
[0110] This embodiment takes the new energy vehicle supply chain network constructed with vehicle manufacturers in Guangdong Province as the center as an example, and its carbon emission data is provided by a leading automobile group. The distance between facilities uses the map navigation distance. Considering that the carbon emission coefficients per kilometer of vehicles of different energy types are different, it is assumed that the risk coefficients of pure electric vehicles, hybrid vehicles, and fuel vehicles are 1.2, 1.0, and 0.8 respectively, and the carbon emission coefficients per kilometer are 0.5, 1.2, and 1.5 respectively. The retailer's demand is obtained from the annual average demand based on the population in the jurisdiction. Set the fixed site selection cost according to the market average price of factories, distribution centers, and recycling centers in the region in 2024. Set the capacity they possess according to the site selection scale of factories and distribution centers, and input the basic parameters of candidate factories and site selection distribution centers. According to the research on vehicle speed, load, and carbon emissions, combined with the data of the carbon emission coefficients per kilometer of different vehicle models of a leading automobile group in the investigation, input the emissions per unit distance and the cost per unit distance of different vehicle models in normal state and in the event of a transportation accident. Next, input the unit transportation cost of candidate factories, candidate distribution centers, candidate recycling centers, and each transportation path respectively, and give the transportation risk coefficients of vehicle models between different supply chain entities.
[0111] As Figure 2 shown, compared with the Pareto optimal solution under tight constraint conditions, this algorithm obtains the globally unique optimal solution.
[0112] Furthermore, to analyze the impact of the risk coefficient, whether there is a recycling center, and the vehicle energy type on the model results in the event of a transportation accident, based on the selected corresponding vehicle models, solve the model for whether there is a recycling center and whether there is a risk loss, and obtain Figure 2 the Pareto solution set. Subsequently, select a mixed vehicle model to transport products from the factory to the distribution center and retailers, a single vehicle model from the distribution center to retailers, and from retailers to the recycling center, and solve the model for the above situations to give the Pareto solution set, as shown in Table 4.
[0113] Table 4 Pareto solution set and decision-making scheme for the multi-objective optimization problem of the supply chain under vehicles of different energy types
[0114]
[0115] The group meeting optimization method based on deep reinforcement learning (DRL-MOA) proposed in this embodiment is experimentally compared with the NSGA-III and MOEA / D multi-objective evolutionary algorithms commonly used in existing patents. The maximum number of iterations of NSGA-III and MOEA / D are set to 500, 1000, 4000, and 6000 respectively. The population sizes of NSGA-III and MOEA / D are set to 100. The number of sub-problems of DRL-MOA is also set to 100. In addition, only non-dominated solutions are retained in the final solution set of the problem.
[0116] (1) In terms of comparing the advantages and disadvantages of various solution schemes. The iteration times and running times of different algorithms are evaluated for the solution scheme to select the optimal scheme. As shown in Table 5, the Hypervolume (HV, hypervolume) performance indicators and solution times of each algorithm in the table are obtained by independently executing the algorithm five times and taking the average value for all instance results. DRL-MOA performs the best in all experimental results, and the running time of DRL-MOA is much lower than that of other algorithms. Increasing the number of iterations can improve the optimization performance of the algorithm, but it will increase the solution time. NSGA-III requires 28 seconds for 6000 iterations, and MOEA / D (traditional multi-objective evolutionary algorithm) requires 89 seconds for 6000 iterations. The DRL method can achieve the best optimization level in the shortest solution time and realize the fast and accurate solution of the multi-objective location-allocation problem.
[0117] Table 5 Iteration times and running times of different algorithms
[0118]
[0119] Note: In traditional methods, usually 500 - 1000 iterations are used, and more than 1000 iterations belong to relatively large numbers of iterations and complex solutions.
[0120] (2) In terms of solution performance and speed. Compared with other algorithms, DRL-MOA can obtain the optimal HV index value, indicating that the method proposed in the present invention has a faster solution speed. This experiment proves the effectiveness of DRL-MOA. At the same time, in solving large-scale multi-objective combinatorial optimization problems such as the location-allocation of new energy vehicle supply chains, NSGA-III and MOEA / D cannot achieve algorithm convergence in a short time. Compared with traditional multi-objective evolutionary algorithms, the method proposed in the present invention has a better Pareto front in terms of diversity.
[0121] The group meeting optimization model designed in this embodiment can be effectively extended to the implementation of multi-objective combinatorial optimization problems with different numbers of facility entities. For example, it shows good performance in experiments with 100 and 150 facility entity scales. Especially during the deep reinforcement learning process, the three objective function values of risk, carbon emissions, and cost are used as return values to update the policy. For large-scale problems, the temporal difference method in neural networks can also be adopted to improve the training efficiency of the distribution path sub-problem. In addition, the combinatorial framework can be further adjusted according to the diversity and complexity of the location-allocation problem. For example, advanced technologies such as the attention mechanism and graph neural network in deep reinforcement learning can be introduced to improve the quality of the solution and the optimization speed. Guided by deep reinforcement learning, NSGA-III can find better solutions in large-scale location problems and improve the adaptability and efficiency of the algorithm.
[0122] In summary, compared with the prior art, this embodiment has at least the following advantages and beneficial effects:
[0123] (1) Since the present invention adopts a multi-objective optimization model, it can simultaneously consider economic costs, carbon emissions, and transportation risk losses, bringing a more comprehensive optimization effect. Traditional optimization methods often only focus on a single objective (such as cost minimization or path optimization), while the present invention comprehensively considers multiple objectives (such as transportation costs, carbon emissions, and accident risks), making the combinatorial optimization results more in line with actual needs. When solving supply chain problems, it can not only reduce economic costs, but also effectively reduce environmental impacts and improve safety, meeting the requirements of green development and sustainable development.
[0124] (2) Since the present invention adopts an optimization framework combining DRL and NSGA-III, it has higher optimization efficiency and solution quality. By integrating the selection, crossover, and mutation processes in DRL-MOA, it has the characteristic of significantly accelerating the convergence speed of the algorithm, avoiding the emergence of local optimal solutions, and thus being able to find the global optimal solution in a shorter time. Compared with traditional genetic algorithms, this method has stronger adaptability and higher solution efficiency, and is particularly suitable for large-scale complex combinatorial optimization problems.
[0125] (3) Since the present invention adopts the policy optimization mechanism of reinforcement learning, it accelerates the improvement of the model's adaptability and solution efficiency. Then, based on NSGA-III improving the diversity and stability of the optimization solution, it can effectively generate a Pareto front solution set with a good distribution, avoiding the problems of single solutions or unstable solutions that may occur in traditional methods.
[0126] Embodiment 2
[0127] This embodiment provides a multi-objective combinatorial optimization device for a new energy vehicle supply chain, including:
[0128] A model construction module, configured to obtain the geographical location coordinates of a city and construct a multi-objective mathematical model for the location-allocation problem considering transportation risks and carbon emissions;
[0129] An objective determination module, configured to determine the constraint conditions and objective function of the multi-objective mathematical model;
[0130] An initial solution module, configured to perform optimal solution of the multi-objective mathematical model using NSGA-III to obtain an initial solution set of the new energy vehicle location-allocation plan;
[0131] A reinforcement learning module, configured to build a deep reinforcement learning model based on the crossover operator and mutation operator of NSGA-III;
[0132] A model training module, configured to train the deep reinforcement learning model using the initial solution set, and obtain a combined optimization model based on deep reinforcement learning after the training is completed;
[0133] A model application module, configured to use the combined optimization model based on deep reinforcement learning to solve the actual location-allocation scenario and obtain an optimal new energy vehicle location-allocation plan.
[0134] Since this device is a multi-objective combined optimization device for a new energy vehicle supply chain in an embodiment of the present invention, and the principle of solving problems by this device is similar to that of this method, the implementation of this device can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.
[0135] Embodiment 3
[0136] An embodiment of the present invention further provides an electronic device, the electronic device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement a multi-objective combined optimization method for a new energy vehicle supply chain as Figure 4 shown.
[0137] It can be understood that the memory may include a Random Access Memory (RAM) and may also include a Read-Only Memory. Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, codes, code sets or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function, instructions for implementing the above various method embodiments, etc.; the data storage area can store data created according to the use of the server, etc.
[0138] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect various parts within the entire server. By running or executing instructions, programs, code sets or instruction sets stored in the memory, and by calling data stored in the memory, it executes various functions of the server and processes data. Optionally, the processor may be implemented in at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor may integrate a combination of one or several of a Central Processing Unit (CPU) and a modem, etc. Among them, the CPU mainly processes the operating system and application programs, etc.; the modem is used to process wireless communications. It can be understood that the above modem may not be integrated into the processor and may be implemented separately through a single chip.
[0139] Since this electronic device is an electronic device corresponding to a multi-objective combinatorial optimization method for a new energy vehicle supply chain in an embodiment of the present invention, and the principle of how this electronic device solves problems is similar to that of this method, the implementation of this electronic device can refer to the implementation process of the above method embodiment, and repeated parts will not be elaborated.
[0140] Embodiment 4
[0141] An embodiment of the present invention further provides a computer-readable storage medium, in which at least one instruction, at least one segment of program, code set or instruction set is stored, and the at least one instruction, the at least one segment of program, the code set or instruction set is loaded and executed by a processor to implement Figure 4 a multi-objective combinatorial optimization method for a new energy vehicle supply chain as shown.
[0142] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disc memories, tape memories, or any other computer-readable medium capable of carrying or storing data.
[0143] Since this storage medium is the storage medium corresponding to a multi-objective combined optimization method for a new energy vehicle supply chain in an embodiment of the present invention, and the principle of solving problems by this storage medium is similar to that of this method, the implementation of this storage medium can refer to the implementation process of the above method embodiment, and repeated parts will not be elaborated.
[0144] Embodiment 5
[0145] In some possible implementation manners, each aspect of the method in the embodiment of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps of a multi-objective combined optimization method for a new energy vehicle supply chain according to various exemplary implementation manners described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in high-level programming languages such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0146] It should be understood that each part of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0147] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0148] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable those of ordinary skill in the art to understand the content of the present invention and implement it accordingly, and it cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered by the protection scope of the present invention.
Claims
1. A multi-objective combination optimization method for a new energy vehicle supply chain, characterized in that: The following steps are involved: Obtain the geographical coordinates of the city and construct a multi-objective mathematical model for site selection and allocation considering transportation risks and carbon emissions; Determining the constraints and objective function of the multi-objective mathematical model; NSGA-III is used to optimize the multi-objective mathematical model to obtain the initial solution set of the new energy vehicle site selection and distribution plan; Build a deep reinforcement learning model based on the crossover operator and mutation operator of NSGA-III; The deep reinforcement learning model is trained using the initial solution set, and after the training is completed, a combined optimization model based on deep reinforcement learning is obtained; The combinatorial optimization model based on deep reinforcement learning is used to solve actual site selection and allocation scenarios to obtain the optimal site selection and distribution plan for new energy vehicles.
2. The multi-objective combined optimization method for the new energy vehicle supply chain according to claim 1 is characterized in that: The objective functions of the multi-objective mathematical model are as follows: first, the economic cost is minimized, including supply cost, site selection cost, distribution cost and recycling cost; second, the carbon emission in the transportation process is minimized under different scenarios; third, the risk loss in the transportation process is minimized under different scenarios; The constraints of the multi-objective mathematical model include: facility location cost constraints, supplier supply constraints, factory production capacity constraints, distribution center storage capacity constraints, retailer demand constraints, and recycling center processing capacity constraints.
3. The multi-objective combined optimization method for a new energy vehicle supply chain according to claim 1 is characterized in that: The combined optimization model based on deep reinforcement learning works as follows: A1. Chromosome coding: The chromosome is divided into three sections. The first section is the candidate distribution center, and each gene corresponds to the supply factory of each distribution center; the second section is the retailer, and each gene corresponds to the distribution center of each retailer; the third section is the recycling center, and each gene corresponds to the recycling center of the retailer; A2. Initialize the population: According to the chromosome encoding in step A1, randomly generate a set of initial solutions to obtain the initialized population; A3. Non-dominated sorting: non-dominated sorting of individuals in the population, assigning different non-dominated levels; A4. Best selection: select based on non-dominated sorting results and crowding degree, and retain individuals with better diversity; A5. Crossover and mutation: According to the virtual fitness value, the non-dominated sorted population is replicated and the selection operation is performed; given the crossover probability, the crossover operation is performed; given the mutation probability, the mutant gene is randomly selected; In the crossover operation, some decision variables of two parent individuals are exchanged to generate new offspring individuals; in the mutation operation, the location of a factory is randomly changed or the transportation volume of a transportation route is adjusted; A6. Main process: merge the initial population and its offspring population to form a new population, perform non-dominated sorting on it and generate non-dominated sets, and calculate the crowding degree; gradually put the newly generated individuals into the new parent population until the parent size exceeds the preset population size; sort the individuals in the merged new population by crowding degree, and select the dominant solution on the Pareto frontier; Return to execute step A3 to step A6 to form a new offspring population until the set iteration threshold is reached.
4. The multi-objective combined optimization method for the new energy vehicle supply chain according to claim 3 is characterized in that: The preferred selection specifically includes: Calculate the target component according to the target function; According to the target component, all non-dominated individuals in the population are found and assigned a shared virtual fitness value to form a first-level non-dominated individual set; other individuals in the population are further graded according to the non-dominated relationship and new virtual fitness values are assigned to form a higher-level non-dominated individual set until all individuals are graded.
5. The multi-objective combined optimization method for the new energy vehicle supply chain according to claim 3 is characterized in that: During the crossover process, the probability and method of crossover are adjusted according to the target differences in risk, carbon emissions, and cost, as well as the distribution of the current solution set. Through more refined adjustments, crossover can effectively explore the multi-objective solution space, ensure that more diverse candidate solutions can be generated during optimization, and maintain a uniform distribution of the Pareto frontier.
6. The multi-objective combined optimization method for the new energy vehicle supply chain according to claim 3 is characterized in that: During the mutation process, according to the positive correlation and loss difference between carbon emissions and risk losses in the current solution set, adjustments are made based on the spatial density information on the basis of random mutation.
7. The multi-objective combined optimization method for a new energy vehicle supply chain according to claim 1 is characterized in that: The deep reinforcement learning model based on the crossover operator and mutation operator of NSGA-III is constructed, including: The initial solution set is accelerated and optimized by combining deep learning. The crossover operator and selection operator in NSGA-III are continuously improved and optimized through deep learning algorithms to improve the solution efficiency and accuracy.
8. A multi-objective combination optimization device for a new energy vehicle supply chain, characterized in that: include: The model building module is used to obtain the geographical coordinates of the city and build a multi-objective mathematical model for site allocation problems considering transportation risks and carbon emissions; A target determination module, used to determine the constraint conditions and target function of the multi-target mathematical model; An initial solution module, used to use NSGA-III to optimize the multi-objective mathematical model and obtain an initial solution set of the new energy vehicle site selection and distribution plan; Reinforcement learning module, used to build deep reinforcement learning models based on crossover operator and mutation operator of NSGA-III; The model training module is used to train the deep reinforcement learning model using the initial solution set, and obtain a combined optimization model based on deep reinforcement learning after the training is completed; The model application module is used to apply the combinatorial optimization model based on deep reinforcement learning to solve actual site selection and allocation scenarios and obtain the optimal site selection and distribution plan for new energy vehicles.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-supply-point emergency material site selection optimization method considering timeliness and fairness
CN115330288A
Cited By
Suspension system parameter determination method and device, equipment, medium and program product
CN121211996A