Electric shovel production scheduling method based on joint evolution exploration optimization algorithm
Through the joint evolution exploration optimization algorithm, the multi-objective conflict problem in the electric shovel production scheduling is solved, efficient production scheduling and resource utilization are achieved, and it is suitable for complex electric shovel production scenarios.
Patent Information
- Application Number
- CN202510126108.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-06-27
AI Technical Summary
There are multiple target conflicts in electric shovel production scheduling, and existing optimization methods are difficult to effectively solve, resulting in low production efficiency, high cost and low resource utilization.
The electric shovel production scheduling method based on joint evolution exploration optimization algorithm is adopted. By building a multi-objective optimization model, using neural networks and collaborative optimization population mechanisms, the balance between global search and local search is achieved, and a variety of high-quality production scheduling solutions are generated.
It realizes a power shovel production scheduling solution that performs excellently on multiple optimization goals, improves production scheduling efficiency and resource utilization, and is suitable for complex power shovel production scheduling scenarios.
Smart Images

Figure CN120218639A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an electric shovel production scheduling method based on a joint evolution search optimization algorithm to solve the electric shovel production scheduling problem and can optimize the multi-objective sequence optimization tasks in the electric shovel production process. Background Art
[0002] (I) Importance and Complexity of Electric Shovel Production
[0003] Widespread application and key role of electric shovels:
[0004] As heavy engineering machinery and equipment, electric shovels play a crucial role in fields such as mining, construction, and infrastructure, and are mainly used to handle materials such as earth and rock, and ore. Their working efficiency and production capacity directly affect the progress of large open-pit mines and major construction projects.
[0005] With the rapid development of the global mining and construction industries, the demand for electric shovels has gradually increased, which has prompted the design and production of electric shovels to become more complex and precise. The manufacturing of electric shovels involves the processing and assembly of numerous components, including core components such as buckets, boom assemblies, upper mechanisms, and propulsion systems, and each component requires precise processing and careful assembly through multiple processes.
[0006] Challenges faced in production scheduling:
[0007] In the electric shovel production process, the processing procedures and process routes of each component vary significantly due to differences in structure and function. The processing of each component may involve multiple operations, such as cutting, milling, drilling, welding, and heat treatment, and in each operation, different machines can usually be selected for processing, which constitutes a typical multi-process, multi-machine scheduling problem.
[0008] Different from the traditional assembly line production mode, the processing technology of electric shovel components has a high degree of flexibility and parallelism. Therefore, how to determine a reasonable processing sequence for each component and select the most suitable machine for processing has become the core problem in production scheduling. This not only requires considering the processing requirements of the components but also taking into account the performance differences of the machines, because the rationality of the scheduling scheme directly affects key indicators such as production efficiency, processing cost, and equipment utilization rate.
[0009] (II) Limitations of Existing Optimization Methods
[0010] Dilemma of traditional optimization methods:
[0011] The production scheduling problem of electric shovels usually involves multiple conflicting objectives, such as minimizing the total production duration, reducing machine idle time, balancing the workload, and improving resource utilization. These objectives restrict each other and are difficult to achieve simultaneously through simple optimization methods. Therefore, traditional optimization methods often show obvious limitations when solving such complex scheduling problems and cannot effectively meet the needs of actual production.
[0012] Deficiencies of existing evolutionary algorithms:
[0013] Although many evolutionary algorithms have been developed in the past few decades to solve sequence optimization problems, the existing evolutionary algorithms for optimizing neural networks still have many problems when dealing with the production scheduling problem of electric shovels. For example, when dealing with a large number of variables, the efficiency of these algorithms is low and the convergence speed is slow, which severely limits their application in industrial scenarios sensitive to computing resources. In addition, the performance of evolutionary algorithms is sensitive to parameter settings and requires fine-tuning according to specific problems, which undoubtedly increases the complexity and difficulty of using the algorithms and makes them face many challenges in practical applications. Summary of the Invention
[0014] In view of the deficiencies of the existing technology, the present invention provides a production scheduling method for electric shovels based on joint evolution to explore optimization algorithms, which can effectively balance global search and local search to obtain a set of production scheduling schemes for electric shovels that perform excellently on multiple optimization objectives and meet the constraint conditions. It can be seen that the present invention can obtain multiple high-quality solutions in one run to meet different production scheduling requirements of electric shovels, improve production scheduling efficiency and resource utilization, and show significant advantages in practical applications, and is applicable to complex production scheduling scenarios of electric shovels.
[0015] To achieve the above technical objectives, the present invention will adopt the following technical solutions:
[0016] A production scheduling method for electric shovels based on joint evolution to explore optimization includes the following steps:
[0017] Step 1: Construct an optimization model for the production scheduling of electric shovels:
[0018] Determine the set of parts to be processed, their operation sequences, and processing times, where: the set of parts to be processed is used to determine the scheduling objects, the operation sequences are used to determine the processing order of each part to be processed, and the processing times are used to calculate the completion duration of the corresponding parts to be processed;
[0019] Set the objective function and constraint conditions to minimize the maximum completion time among all parts to be processed; the objective function is constructed based on the completion duration of the parts to be processed, and when seeking optimization, it is used to evaluate whether the solutions of the operations of each part to be processed meet the objectives and constraints;
[0020] Use a coding scheme to represent the solution containing the operations of each component to be processed as a real vector;
[0021] Step 2. Search for the optimal electric shovel production scheduling plan:
[0022] Use joint evolution to search for and optimize the Pareto optimal solution of the electric shovel production scheduling optimization model constructed in Step 1, so as to obtain the optimal electric shovel production scheduling plan. This optimal electric shovel production scheduling plan performs excellently in multiple optimization objectives and meets the constraint conditions set in Step 1.
[0023] Preferably, in Step 2, the joint evolution search and optimization algorithm includes a neural network Net and two co-optimized populations;
[0024] Each population contains N individuals, and each individual represents a possible electric shovel production scheduling plan; the electric shovel production scheduling plan includes the operation sequence of each component of the electric shovel and the machine allocation plan;
[0025] In the iterative process, select parents from the two populations respectively, generate offspring based on the mutation granularity and using specific crossover and mutation operators, and update the two populations respectively through environmental selection and truncation selection based on the objective function;
[0026] The mutation granularity is adaptively adjusted through the neural network Net;
[0027] When the iterative process of the joint evolution search and optimization algorithm reaches the maximum evaluation number, terminate the iteration;
[0028] At this time, return the population as the final output result of the joint evolution search and optimization algorithm; the individuals in the population represent the relatively optimal electric shovel production scheduling plans obtained after multiple iterations of optimization, and each individual includes the operation sequence of the components and the machine allocation plan.
[0029] Preferably, in Step 1, the construction of the electric shovel production scheduling optimization model specifically includes the following steps:
[0030] Step 1.1. Determine the set of components to be processed J = {J1, J2, …, J L}, where L is the total number of components to be processed; any component to be processed J l has a predefined operation sequence O k1 , O k2 , …, O lj ; l represents the component number, and j represents the operation number in the operation sequence; given the set of machines W = {W1, W2, …, W M}, where M is the total number of machines; clarify the processing time T lj of the operation O k on the machine W ljk ;
[0031] Step 1.2: Calculate the J of the part to be processed l Completion time Among them: ET ljk is the idle waiting time between operations; a l J is the part to be processed l The number of operations;
[0032] Step 1.3: Set the objective function and constraints:
[0033] The objective function includes the first and second objective functions, which are as follows:
[0034] f1=max l=1,…,L C l ;
[0035]
[0036] E l =max(0,d l1 -C l );
[0037] T l =max(0,C l -d l2 );
[0038] Where: f1 represents the first objective function; C l Represents the part to be processed J i completion time; f2 represents the second objective function; E l Penalty for early completion; T l Penalty for delayed completion;d l1 J is the part to be processed l The earliest completion time, d l2 J is the part to be processed l The latest completion time; α and β are the weight coefficients of early and late penalties respectively;
[0039] The constraint condition is: the second objective function f2 is less than the given constraint value SC;
[0040] Step 1.4: Use an encoding scheme to represent the solution involving multiple component operations as a real vector.
[0041] Preferably, in step 1.4, the encoding scheme is specifically as follows: for a solution containing multiple component operations to be processed, an operation sequence is obtained by sorting and decoding the real vector elements, wherein each dimension corresponds to an ordered operation, and the execution order of the component operations is determined according to the order of the elements, and the level of the component operations is predefined and cannot be modified.
[0042] Preferably, an optimal production scheduling scheme for electric shovels is found by implementing a joint evolution search optimization algorithm, which specifically includes the following steps:
[0043] Step 2.1: Initialize the neural network Net and two populations Population1 and Population2, and set the current evaluation count FE to be equal to the population size;
[0044] Step 2.2: Enter the iterative loop until the maximum evaluation count FEmax is reached;
[0045] Step 2.2.1: Randomly select N parent individuals from Population1 and Population2 respectively;
[0046] Step 2.2.2: Generate offspring individuals by using simulated binary crossover and polynomial mutation operations based on the mutation granularity action. The specific operations are as follows;
[0047] For the crossover operation, given two parent solutions x 1 and x 2 , the generation formulas for the offspring o 1 and o 2 are:
[0048]
[0049] where: i = 1, …, d, d is the dimension of the decision variable, and the parameter β is the weight coefficient of the delay penalty, which is calculated by the following formula:
[0050]
[0051] In the formula: μ ∈ [0, 1] is a random number, and n represents the number of iterations;
[0052] For the mutation operation, each offspring individual o mutates according to the comparison between the random number r and the mutation probability prob. The formula is:
[0053]
[0054] where i = 1, …, d, r ∈ [0, 1] is a random number, prob is the mutation probability, u i = 1 is the upper limit, and l i = 0 is the lower limit; the parameter δ is calculated by the following formula:
[0055]
[0056] In the formula, η is a parameter;
[0057] Step 2.2.3: Merge the generated offspring individuals into Population1 and Population2 respectively;
[0058] Step 2.2.4: Perform environmental selection operation on Population1, and retain N solutions. The selection basis is the constrained Pareto dominance relationship;
[0059] Step 2.2.5: Update the experience memory pool M, record the reward and state information, train the neural network Net, and adaptively adjust the mutation granularity action;
[0060] Step 2.3: After the iteration is completed, return the final Population1 as the optimization result.
[0061] Preferably, in Step 2.2.5, an adaptive mutation mechanism is introduced, and a deep Q-network is used to select the mutation granularity of the joint evolution exploration optimization algorithm; specifically, it includes the following steps:
[0062] Step 3.1: Define five candidate values of mutation granularity 1 / d, and 5 / d, where d is the dimension of the decision variable; and use reinforcement learning to determine the optimal mutation granularity;
[0063] Step 3.2: Define the population state s diversity and feasibility to define the population state s t , where i represents the number of objective functions used for traversal, f i (x) is the objective function value, m is the number of objective functions, obj i is the average value of the objective functions, and CV(x) is the constraint violation degree;
[0064] Step 3.3: Use the hypervolume indicator as the reward signal r t , and use a deep Q-network to select the mutation granularity;
[0065] The hypervolume indicator is calculated based on the population state;
[0066] The deep Q-network architecture includes an input layer, two hidden layers, and an output layer; the input layer receives the population state information, the hidden layers process the information through non-linear activation functions, and the output layer provides the Q value of a specific action-state pair, so that the deep Q-network can learn an effective mapping from the state to the action value;
[0067] The deep Q-network is based on the current state s t , action a, the obtained reward r t and the new state s t+1Perform training updates, where the current state reflects the system situation and serves as the basis for decision-making; the action refers to the selection of mutation granularity, which affects the exploration direction of the scheduling scheme; the reward is the feedback evaluation of the action, which helps to learn favorable mutation granularity; the new state is the state of the system after taking the action, providing information for the next decision-making, and jointly contributing to the optimization of the electric shovel production scheduling scheme. The agent selects the mutation granularity according to the update rule Q(s,a;θ)←Q(s,a;θ)+α(y - Q(s,a;θ)), where a is the learning rate, θ is the network parameter, and the target y = e + γmax a' Q(s',a';θ - ), where e is the reward obtained after executing the action a, s' is the next state, θ - is the target network parameter; a' is; γ is the discount factor;
[0068] Calculate the evaluation usage ratio λ, and the calculation formula is:
[0069] λ = FE / FEmax;
[0070] In the formula: FE represents the current evaluation quantity, FE = FE + |Q|, |Q| is the total number of generated offspring individuals; FEmax represents the maximum evaluation quantity;
[0071] If λ < 0.25 or r < 0.3λ, randomly select one from five candidate mutation granularities as the mutation granularity for the next iteration; otherwise, select the mutation granularity with the highest Q value as the mutation granularity for the next iteration according to the output of the current deep Q network.
[0072] Preferably, in step 2.2.5, the training process of the deep Q network is specifically as follows:
[0073] First, calculate the evaluation usage ratio λ;
[0074] Then, compare the evaluation usage ratio λ with 0.25: If λ < 0.25, skip this training; if λ = 0.25, use all samples in the experience memory pool to train the deep Q network; if λ > 0.25, update the deep Q network with the ten nearest samples in the experience memory pool every ten iterations;
[0075] During the training process, adjust the parameters of the neural network through the backpropagation algorithm according to the difference between the prediction result of the deep Q network and the actual reward signal.
[0076] Preferably, the non-linear activation function adopted by the hidden layer is the RelU function.
[0077] Preferably, in step 2.2.4, the specific process of Population1 performing environmental selection includes:
[0078] First, calculate the constraint violation degree CV(x) of each solution x, and the calculation formula is:
[0079] CV(x) = max{0, F2(x) - SC};
[0080] Where: f2(x) is the second objective function value, and SC is the constraint value;
[0081] Then, determine the dominance relationship between solutions by comparing the constraint violation degree and the objective function value, and select N excellent individuals to be retained in Population1 through non-dominated sorting and crowding distance; specifically, if two individuals x and y satisfy:
[0082] CV(x) < CV(y);
[0083] Or
[0084] Then it is said that x dominates y, select the non-dominated individuals and further screen out N individuals according to the crowding distance to be retained in Population1.
[0085] Preferably, in step 2.1, the initialization method of the population individuals adopts the method of randomly generating operation sequences and randomly allocating machines
[0086] Compared with the prior art, the present invention has many significant beneficial effects:
[0087] 1. Improve model characteristics: It can effectively generate a sparse neural network model, reduce the model complexity, reduce the consumption of computing resources, and enhance the model generalization ability.
[0088] 2. Improve detection accuracy: Optimize the neural network weights, improve the fault detection accuracy and robustness, and accurately identify the fault status of the electric shovel.
[0089] 3. Accelerate the optimization speed: Rely on dynamic variable clustering and improved genetic operators to improve the search efficiency of the algorithm and reduce the calculation time.
[0090] 4. Reduce resource requirements: Generate a sparse model, reduce the consumption of computing resources for model training and inference, and adapt to resource-constrained industrial scenarios. Description of the Drawings
[0091] Figure 1 It is the framework flowchart of the electric shovel production scheduling method based on joint evolution exploration and optimization described in the present invention. Detailed Embodiments
[0092] As Figure 1 shown, the electric shovel production scheduling method based on the joint evolution exploration and optimization algorithm described in the present invention includes the following steps:
[0093] Step 1. Data Preparation and Initialization:
[0094] Step 1.1. Collect the production data of the electric shovel:
[0095] Record in detail the relevant information of each component involved in the production process of the electric shovel, including the operations required for each component, the machines on which each operation can be performed, and the corresponding processing time. At the same time, determine the earliest completion time d i1 and the latest completion time d i2 of each component for subsequent calculation of the penalty cost.
[0096] Specifically, determine the set of parts to be processed J = {J1, J2,..., J L}, where L is the total number of parts to be processed; any part to be processed J l has a predefined operation sequence O k1 , O k2 ,..., O lj ; l represents the serial number of the part, and j represents the serial number of the operation in the operation sequence; given the set of machines W = {W1, W2,..., W M}, where M is the total number of machines; clarify the processing time T lj of the operation O k on the machine W ljk .
[0097] Based on the collected production data of the electric shovel, calculate the completion time l of any part J where: ET ljk is the idle waiting time between operations; a l is the number of operations of the part to be processed J l .
[0098] Set the optimization objective function, including:
[0099] Take f1 = max l=1,...,L C l as the first objective function to minimize the maximum completion time among all parts.
[0100] Define E l = max(0, d l1 - C l ) as the penalty for early completion, and T l = max(0, C l - d l2 ) as the penalty for late completion, where d l1 is the earliest completion time of part J l , and d l2 is the latest completion time of part J l .
[0101] Let α and β be the weight coefficients of the early and late penalties respectively, and use it as the second objective function, and require that f2 is less than the constraint value SC given by the user, to construct a constrained multi-objective optimization model.
[0102] The core objective of the optimization objective function is to minimize the maximum completion time among all components, that is, the first objective function f1 = max l=1,…,L C l ; this helps to ensure the optimization of the entire production cycle, so that the most time-consuming component can be processed as soon as possible. At the same time, considering that each component should be completed within a specific time window, an early completion penalty E l = max(0, d l1 - C l ) and a late completion penalty T l = max(0, C l - d l2 ) are introduced, where d l1 and d l2 are the earliest completion time and the latest completion time of component J l respectively. The second objective function is Here, α and β are the weights of the early and late penalties respectively, which are used to flexibly balance the costs of deviating from the predetermined delivery date. Due to strict time management requirements, the second objective function must satisfy f2 < SC (SC is the constraint value specified by the user), thus making this problem a typical constrained multi-objective optimization problem.
[0103] To facilitate the optimization of the production scheduling sequence problem using various algorithms, a coding scheme representing the solution with real vectors is proposed. Specifically: for a solution containing multiple part operations, the operation sequence is obtained by decoding the sorting of real vector elements, where each dimension corresponds to an ordered operation, and the execution order of part operations is determined according to the element order, and the levels of part operations are predefined and cannot be modified.
[0104] Taking a processing task with three components (assuming component 1 has three operations, component 2 has three operations, and component 3 has four operations) as an example, its solution can be represented as a real-coded vector, such as (0.10, 0.42, 0.58, 0.15, 0.29, 0.81, 0.23, 0.36, 0.77, 0.93). By sorting the real elements of this vector in ascending order, the obtained permutation (such as (1, 4, 7, 5, 8, 2, 3, 9, 6, 10)) can be converted into an operation sequence, where elements 1, 2, 3 correspond to the three operations of component 1, elements 4, 5, 6 correspond to the three operations of component 2, and elements 7, 8, 9, 10 correspond to the four operations of component 3. It should be noted that the order of all operations of each component is predefined, so there is no need to associate elements with specific operations.
[0105] When calculating the completion time, operations will be carried out sequentially on a specific machine in order. If there are other operations being carried out on the same machine, waiting is required. Through this coding scheme, the sequence optimization problem is transformed into a continuous constrained multi-objective optimization problem, which can theoretically be processed by many constrained multi-objective evolutionary algorithms. However, due to the conflict of objective functions and the strictness of constraints, as well as the high discretization of the solution space caused by the conversion from real vectors to discrete sequences, it is still challenging for many existing algorithms to find feasible Pareto optimal solutions. For this reason, the present invention provides a joint evolutionary exploration optimization algorithm.
[0106] Step 1.2, Initialize algorithm parameters:
[0107] Randomly initialize a deep neural network Net, whose structure includes an input layer (receiving population state information), two hidden layers (each containing 10 nodes, using the RelU function as the activation function to enhance learning ability), and an output layer (the output is the Q value of a specific action-state pair).
[0108] Empty the experience memory pool (M); the experience memory pool is used to store information such as states, actions, and rewards during the operation of the algorithm.
[0109] Randomly initialize two populations Population1 and Population2, each population contains N individuals, and each individual represents a possible production scheduling scheme for electric shovels (i.e., operation sequence and machine allocation scheme). The initialization method of individuals can adopt the method of randomly generating operation sequences and randomly allocating machines to ensure the diversity of the initial population.
[0110] Randomly initialize the action of the agent, which is used to determine the mutation granularity, and is initialized to one of the five candidate mutation granularities (1 / d, 5 / d), where d is the dimension of the decision variable.
[0111] Set the current evaluation number (FE) to the size of the initial population, i.e., |P|, and set the maximum evaluation number, which is determined according to the complexity of the problem and the limitation of computing resources. For example, it is set to 10000.
[0112] Step 2. Iterative optimization process
[0113] Adopt the joint evolution search optimization algorithm for iterative optimization to find the Pareto optimal solution of the electric shovel production scheduling optimization model, which specifically includes the following steps:
[0114] Step 2.1. Parent selection:
[0115] In each iteration, randomly select N individuals from Population1 and Population2 respectively as parents. This random selection method helps to maintain the diversity of the population, avoid premature convergence to local optimal solutions, and at the same time provides the algorithm with the opportunity to explore different solution space regions.
[0116] Step 2.2. Offspring generation:
[0117] Step 2.2.1. Based on the current mutation granularity action, perform genetic operations on the selected parent individuals to generate offspring individuals.
[0118] Specifically, for the parent individuals in Population1 and Population2, the following operations are performed respectively:
[0119] Step 2.2.2. Simulated binary crossover:
[0120] Given two parent solutions x 1 and x 2 , first randomly generate a number μ in the range [0, 1], and calculate β according to the value of μ (when μ ≤ 0.5, when μ > 0.5, where η is a parameter, and the parameter η can be set to η = 20). Then, generate two offspring solutions o (i = 1, …, d, where d is the dimension of the decision variable) according to the formula 1 and o 2 .
[0121] Step 2.2.3. Polynomial mutation:
[0122] For each offspring solution o, first randomly generate a number r and a mutation probability prob = 1 / d in the range [0, 1]. Then, mutate each offspring individual o according to the comparison between the random number r and the mutation probability prob. The formula is:
[0123]
[0124] where \(i = 1,\ldots,d\), \(r\in[0,1]\) is a random number, prob is the mutation probability, \(u\) i \(= 1\) is the upper limit, and \(l\) i \(= 0\) is the lower limit.
[0125]
[0126] Mutation operation is performed on each dimension \(i\).
[0127] Step 2.2.4: Add the generated \(N\) offspring individuals to Population1 and Population2 respectively to complete the update of the population.
[0128] Step 2.3: Environmental selection and population update:
[0129] Step 2.3.1: Perform environmental selection on Population1:
[0130] The environmental selection of Population1 is based on the constrained Pareto dominance relationship, and \(N\) solutions are selected through non-dominated sorting and crowding distance.
[0131] Specifically, first calculate the constraint violation degree \(CV(x)=\max\{0,f_2(x)-SC\}\) of each solution \(x\), where \(f_2(x)\) is the second objective function value and \(SC\) is the constraint value.
[0132] The second objective function value \(f_2(x)\) is expressed as:
[0133] \(f_2(x)=\sum_{i = 1}\) n (\(\alpha E\) i +\(\beta T\) i );
[0134] \(E\) i \(=\max(0,d\) i1 \(-C\) i );
[0135] \(T\) i \(=\max(0,C\) i \(-d\) i2 );
[0136]
[0137] In the above formula, \(T\) ij is the processing time of operation \(O\) ij on machine \(W\) k , and \(ET\) ij is the idle waiting time between adjacent operations.
[0138] If the constraint violation degree CV(x) of solution x is smaller, it indicates that the degree of constraint violation of solution x is smaller. When CV(x) = 0, solution x is feasible.
[0139] Then, the dominance relationship between solutions is determined by comparing the constraint violation degree and the objective function value, and N excellent individuals are selected and retained in Population1 through non-dominated sorting and crowding distance. Specifically, if two individuals x and y satisfy:
[0140] CV(x) < CV(y);
[0141] Or
[0142] Then x is said to dominate y, and non-dominated individuals are selected and further screened out N individuals according to the crowding distance.
[0143] Step 2.3.2. Perform a truncation operation on Population2:
[0144] Only retain N individuals in Population2 that perform better on the second objective function f2, that is, select individuals with smaller f2 values.
[0145] Step 3. Experience recording and neural network update:
[0146] Step 3.1. Determine the reward and state:
[0147] The population state is defined by convergence, diversity, and feasibility metrics. Convergence is evaluated by calculating the average performance of the population on each objective function, that is where f i (x) represents the value of the i-th objective function, and m is the number of objective functions. Diversity is measured by calculating the distribution of the population on each objective function, and the formula is where obj i is the average value of the i-th objective function. Feasibility is used to measure the degree to which the population satisfies the problem constraints and is defined as where CV(x) is the constraint violation degree of solution x. These convergence, diversity, and feasibility elements together constitute the population state s t , providing key information for the reinforcement learning agent to determine the ideal mutation granularity.
[0148] The hypervolume metric is used as a reward signal to evaluate the effectiveness of resource allocation, and measures convergence and diversity by evaluating the volume enclosed by the solution set. Each training entry includes the current state, action, obtained reward, and new state, and this information is stored in the experience memory pool to continuously enhance the decision-making ability of the agent.
[0149] Calculate the state of the current population, including convergence Diversity and feasibility
[0150] Calculate the hypervolume indicator (HV) as a reward signal based on the state. This reward signal is used to evaluate the effectiveness of the current resource allocation and reflects the balance degree of the algorithm between exploration and exploitation. Combine the current state, the action taken (mutation granularity), the obtained reward, and the new population state into a record and insert it into the experience memory pool (M).
[0151] Step 3.2, Train the neural network:
[0152] Train the neural network according to the data in the experience memory pool (M).
[0153] First, calculate the evaluation utilization ratio (λ = FE / FEmax).
[0154] If λ < 0.25, skip this training; if λ = 0.25, use all samples in the experience memory pool to train the neural network; if λ > 0.25, update the neural network with the ten most recent samples in the experience memory pool every ten iterations. During the training process, adjust the parameters of the neural network through the backpropagation algorithm according to the difference between the prediction result of the neural network and the actual reward signal to improve the prediction accuracy of the neural network for the optimal mutation granularity action. Step 3.3, Adaptive mutation granularity selection: Select the mutation granularity action for the next iteration through the adaptive mutation mechanism.
[0155] First, randomly generate a number r between [0, 1]. If λ < 0.25 or r < 0.3λ, randomly select one from the five candidate mutation granularities (1 / d,[[]] and 5 / d) as the mutation granularity for the next iteration; otherwise, select the mutation granularity with the highest Q value according to the output of the current neural network (Net) as the mutation granularity for the next iteration.
[0156] The adaptive mutation mechanism for selecting the mutation granularity for the next iteration specifically includes the following steps:
[0157] Step 3.3.1, Define five mutation granularity candidate values 1 / d,[[]] and 5 / d (d is the decision variable dimension), and use reinforcement learning to determine the optimal mutation granularity;
[0158] Step 3.3.2, Define the population state s by calculating the convergence Diversity and feasibility to define the population state st , where f i (x) is the objective function value, m is the number of objective functions, and obj i is the average value of the objective functions, and CV(x) is the constraint violation degree;
[0159] Step 3.3.3: Use the hypervolume indicator as the reward signal r t to construct a deep Q-network to select the mutation granularity. The deep Q-network architecture includes an input layer, two hidden layers, and an output layer. The deep Q-network is trained and updated based on the current state s t , action a, obtained reward r t and new state s t+1 . The agent selects the mutation granularity according to the update rule Q(s,a;θ)←Q(s,a;θ)+α(y - Q(s,a;θ)), where α is the learning rate, θ is the network parameter, and the target y = r + γmax a' Q(s',a';θ - ), where r is the reward after action a, s' is the next state, and θ - is the target network parameter, and γ is the discount factor.
[0160] The deep Q-network architecture includes an input layer with 4 nodes, two hidden layers, and an output layer. The input layer receives a 4-dimensional state vector. The hidden layers process information through non-linear activation functions. The output layer provides the Q value for a specific action-state pair, enabling the network to learn an effective mapping from states to action values.
[0161] Step 3.4: Evaluation quantity update:
[0162] Update the current evaluation quantity by increasing the number of offspring individuals generated in this iteration, i.e., FE = FE + |Q|, where |Q| is the total number of offspring individuals generated (2N in the present invention).
[0163] Then, determine whether FE is less than FEmax. If so, continue with the next iteration; otherwise, enter the termination phase.
[0164] Step 4: Termination and result output:
[0165] When the iteration process reaches the maximum number of evaluations (FEmax), the algorithm terminates the iteration. At this time, Population1 is returned as the final output result of the algorithm. The individuals in Population1 represent the relatively optimal solutions to the electric shovel production scheduling problem obtained after multiple iterations of optimization. Each individual contains information such as the operation sequence of components and the machine allocation scheme. These solutions can be further analyzed and applied according to actual needs to extract common optimization features to guide the scheduling decision-making in the electric shovel production process, improve production efficiency, reduce costs, balance the workload, and ensure meeting constraints such as time windows.
Claims
1. A method for electric shovel production scheduling based on joint evolutionary exploration optimization, characterized in that: The following steps are involved: Step 1: Construct an optimization model for electric shovel production scheduling: Determine the set of parts to be processed and their operation sequence and processing time, where: the set of parts to be processed is used to determine the scheduling object, the operation sequence is used to determine the processing order of each part to be processed, and the processing time is used to calculate the completion time of the corresponding part to be processed; Set the objective function and constraints to minimize the maximum completion time among all parts to be processed; The objective function is constructed based on the completion time of the parts to be processed, and is used to evaluate whether the solutions for the operations of each part to be processed meet the objectives and constraints during optimization. A coding scheme is used to represent the solution including the operations of each component to be processed as a real vector; Step 2: Find the optimal shovel production scheduling plan: The joint evolutionary exploration optimization is used to find the Pareto optimal solution of the electric shovel production scheduling optimization model constructed in step one, so as to obtain the optimal electric shovel production scheduling plan. The optimal electric shovel production scheduling plan performs well in multiple optimization objectives and meets the constraints set in step one.
2. The electric shovel production scheduling method based on joint evolutionary exploration optimization according to claim 1 is characterized in that: In step 2, the joint evolutionary exploration optimization algorithm includes a neural network Net and two collaboratively optimized populations; Each population contains N individuals, each of which represents a possible shovel production scheduling plan. The shovel production scheduling plan includes the operation sequence of each shovel component and the machine allocation plan. In the iterative process, parents are selected from the two populations, and offspring are generated based on the mutation granularity and using specific crossover and mutation operators. The two populations are updated through environmental selection and truncation selection based on the objective function. The mutation granularity is adaptively adjusted through the neural network Net; When the iteration process of the joint evolutionary exploration optimization algorithm reaches the maximum number of evaluations, the iteration is terminated; At this time, the population is returned and used as the final output result of the joint evolutionary exploration optimization algorithm; the individuals in the population represent the optimal electric shovel production scheduling plan obtained after multiple iterative optimizations, and each individual contains the component operation sequence and machine allocation plan.
3. The electric shovel production scheduling method based on joint evolutionary exploration optimization according to claim 2 is characterized in that: In step 1, the construction of the shovel production scheduling optimization model specifically includes the following steps: Step 1.1, determine the set of parts to be processed J = {J1, J2, ..., J L }, L is the total number of parts to be processed; any part to be processed J l With predefined operation sequence O k1 ,O k2 ,…,O lj ; l represents the serial number of the component, j represents the serial number of the operation in the operation sequence; given a machine set W = {W1, W2, ..., W M }, M represents the total number of machines; clear operation O lj On machine W k Processing time T ljk ; Step 1.2: Calculate the J of the part to be processed l Completion time Among them: ET ljk is the idle waiting time between operations; a l J is the part to be processed l The number of operations; Step 1.3: Set the objective function and constraints: The objective function includes the first and second objective functions, which are as follows: f1=max l=1,…,L C l ; E l =max(0,d l1 -C l ); T l =max(0,C l -d l2 ); Where: f1 represents the first objective function; C l Represents the part to be processed J i completion time; f2 represents the second objective function; E l Penalty for early completion; T l Penalty for delayed completion;d l1 J is the part to be processed l The earliest completion time, d l2 J is the part to be processed l The latest completion time; α and β are the weight coefficients of early and late penalties respectively; The constraint condition is: the second objective function f2 is less than the given constraint value SC; Step 1.4: Use an encoding scheme to represent the solution involving multiple component operations as a real vector.
4. The electric shovel production scheduling method based on joint evolutionary exploration optimization according to claim 3 is characterized in that: In step 1.4, the encoding scheme is specifically as follows: for solutions containing multiple component operations to be processed, an operation sequence is obtained by sorting and decoding the real vector elements, where each dimension corresponds to an ordered operation, and the execution order of the component operations is determined according to the order of the elements, and the level of the component operations is predefined and cannot be modified.
5. The electric shovel production scheduling method based on joint evolutionary exploration optimization according to claim 3 is characterized in that: The optimal shovel production scheduling solution is found by implementing the joint evolutionary exploration optimization algorithm, which includes the following steps: Step 2.1, initialize the neural network Net and two populations Population1 and Population2, and set the current evaluation quantity FE equal to the population size; Step 2.2, enter the iterative loop until the maximum evaluation number FEmax is reached; Step 2.2.1, randomly select N parent individuals from Population1 and Population2 respectively; Step 2.2.2, based on the mutation granularity action, use simulated binary crossover and polynomial mutation operations to generate offspring individuals. The specific operations are as follows; For the crossover operation, given two parent solutions x 1 and x 2 , offspring o 1 and 2 The formula for generating is: Where: i = 1, ..., d, d is the decision variable dimension, and parameter β is the weight coefficient of delay penalty calculated by the following formula: Where: μ∈ [ 0,1 ] is a random number, n represents the number of iterations; For the mutation operation, each offspring individual o mutates according to the comparison between the random number r and the mutation probability prob, and the formula is: Where i=1,…,d,r∈[0,1] is a random number, prob is the mutation probability, u i =1 is the upper limit, l i =0 is the lower limit; the parameter δ is calculated by the following formula: Where η is a parameter; Step 2.2.3, merge the generated offspring individuals into Population1 and Population2 respectively; Step 2.2.4, perform the environment selection operation on Population1, retain N solutions, and select based on the constrained Pareto dominance relationship; Step 2.2.5, update the experience memory pool M, record the reward and state information, train the neural network Net, and adaptively adjust the mutation granularity action; Step 2.3: After the iteration is completed, return the final population1 as the optimization result.
6. The electric shovel production scheduling method based on joint evolutionary exploration optimization according to claim 5 is characterized in that: In step 2.2.5, an adaptive mutation mechanism is introduced, and a deep Q network is used to select the mutation granularity of the joint evolutionary exploration optimization algorithm; specifically, the following steps are included: Step 3.1, define five candidate values of variation granularity 1 / d, and 5 / d, where d is the dimension of decision variables; and reinforcement learning is used to determine the optimal mutation granularity; Step 3.2: By calculating the convergence of the population Diversity and feasibility To define the population state s t , where i represents the number of iterations of the objective function, and f i (x) is the objective function value, m is the number of objective functions, obj i is the average value of the objective function, CV(x) is the constraint violation degree; x represents the population; represents the solution; Step 3.3: Use the super volume index as the reward signal r t , a deep Q network is used to select the mutation granularity; The super volume index is calculated based on the population status; The deep Q network architecture consists of an input layer, two hidden layers, and an output layer; the input layer receives population state information, the hidden layer processes the information through a nonlinear activation function, and the output layer provides the Q value of a specific action-state pair, thereby enabling the deep Q network to learn an effective mapping from state to action value; The deep Q network is based on the current state s t , action a, reward r t and the new state s t+1 Training updates are performed, where the current state reflects the system situation and is the basis for decision-making; the action refers to the selection of mutation granularity, which affects the exploration direction of the scheduling plan; the reward is the feedback evaluation of the action, which helps to learn the favorable mutation granularity; the new state is the state of the system after the action is taken, which provides information for the next decision and jointly helps optimize the production scheduling plan of the electric shovel. The agent selects the mutation granularity according to the update rule Q(s,a;θ)←Q(s,a;θ)+α(yQ(s,a;θ)), where a is the learning rate, θ is the network parameter, and the target y=e+γmax a' Q(s',a';θ - ), where e is the reward obtained after executing action a, s' is the next state, and θ - is the target network parameter; a' is; γ represents the discount factor; Calculate the evaluation usage ratio λ, the calculation formula is: λ=FE / FEmax; Where: FE represents the current evaluation quantity, FE = FE + |Q|, |Q| is the total number of offspring individuals generated; FEmax represents the maximum evaluation quantity; If λ<0.25 or r<0.3λ, one of the five candidate mutation granules is randomly selected as the mutation granule for the next iteration; otherwise, the mutation granule with the highest Q value is selected as the mutation granule for the next iteration according to the output of the current deep Q network.
7. The electric shovel production scheduling method based on joint evolutionary exploration optimization according to claim 6 is characterized in that: In step 2.2.5, the training process of the deep Q network is as follows: First, the evaluation usage ratio λ is calculated; Then, the evaluation ratio λ is compared with 0.25: if λ<0.25, skip this training; if λ=0.25, use all samples in the experience memory pool to train the deep Q network; if λ>0.25, use the ten most recent samples in the experience memory pool to update the deep Q network every ten iterations; During the training process, the parameters of the neural network are adjusted through the back-propagation algorithm based on the difference between the prediction results of the deep Q network and the actual reward signal.
8. The electric shovel production scheduling method based on joint evolutionary exploration optimization according to claim 7 is characterized in that: The nonlinear activation function used in the hidden layer is the RelU function.
9. The electric shovel production scheduling method based on joint evolutionary exploration optimization according to claim 5 is characterized in that: In step 2.2.4, the specific process of Population1 selecting the environment includes: First, calculate the constraint violation CV(x) of each solution x. The calculation formula is: CV(x)=max{0,f2(x)-SC}; Where: f2(x) is the second objective function value, SC is the constraint value; Then, the dominance relationship between solutions is determined based on the constraint violation degree and the comparison of the objective function value, so that N excellent individuals are selected and retained in Population1 through non-dominated sorting and crowding distance; specifically, if two individuals x and y satisfy: CV(x) <CV(y); or Then x is said to dominate y, and non-dominated individuals are selected and N individuals are further screened out according to the crowding distance and retained in Population1.
10. The electric shovel production scheduling method based on joint evolutionary exploration optimization according to claim 5 is characterized in that: In step 2.1, the population individuals are initialized by randomly generating operation sequences and randomly allocating machines.
Citation Information
Cited By
Competitive evolution multi-task optimization method, system and equipment
CN121581338A