Flower unloading platform scheduling optimization system and method based on deep reinforcement learning and evolutionary algorithm

By constructing a flower unloading platform scheduling optimization model and using the DQN multi-strategy hybrid evolution algorithm, the problem of low efficiency of unloading platform scheduling in flower cold chain logistics is solved, and efficient resource utilization and operational costs are achieved.

CN120338425APending Publication Date: 2025-07-18KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510510398.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing technology lacks efficient unloading platform scheduling optimization methods in flower cold chain logistics, resulting in vehicles waiting in line, increasing platform idle time, and excessive resource consumption, affecting logistics efficiency and flower quality.

Method used

A flower unloading platform scheduling optimization model is constructed, a multi-strategy hybrid evolution algorithm based on DQN is adopted, and a deep reinforcement learning and evolution algorithm is combined to generate initial solutions and optimize solutions through multiple search operators to reduce vehicle queues and platform idle time and reduce resource consumption.

Benefits of technology

It effectively improves the efficiency of platform scheduling, reduces time waste and resource consumption in flower cold chain logistics, optimizes the logistics process, and reduces operating costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338425A_ABST
    Figure CN120338425A_ABST
Patent Text Reader

Abstract

The invention relates to a fresh flower unloading platform scheduling optimization system and method based on deep reinforcement learning and an evolutionary algorithm, and belongs to the field of operation planning optimization and intelligent scheduling. According to the method, a fresh flower unloading platform scheduling optimization model with minimization of the maximum completion time, minimization of platform idle time penalty cost and the like as targets is constructed, and a multi-strategy hybrid evolutionary algorithm based on DQN is provided for solving. The algorithm comprises the following steps of: firstly, generating diversified and high-quality initial solutions by adopting three initial solution generation methods, namely a random generation method, a first arrival first service scheduling rule and a minimum cost increment scheduling rule; in the initial stage of search, performing global search by using random reservation crossover operation and a two-point exchange operator; then, the DQN intelligent agent adaptively selects five search operators to perform local search so as to obtain a better solution; and finally, generating a new solution through reverse mutation. The platform scheduling efficiency can be effectively improved, and time waste and resource consumption in the flower cold-chain logistics process are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a fresh flower unloading platform scheduling optimization system and method based on deep reinforcement learning and evolutionary algorithm, belonging to the fields of operations research optimization and intelligent scheduling. Background Art

[0002] With the continuous transformation of people's consumption concepts, the fresh flower consumption market has grown rapidly. In the fresh flower cold chain logistics system, the fresh flower unloading platform scheduling, as an important link in the distribution center management, plays a key role in ensuring logistics efficiency and reducing operating costs. A reasonable platform scheduling arrangement can effectively reduce the waiting time of refrigerated vehicles, improve operation efficiency and resource utilization rate, thereby enhancing the overall operation efficiency of the logistics system. However, due to the limited resources of the unloading platform, especially during peak business hours, without scientific scheduling planning, it is easy to cause vehicles to queue for a long time, increase the idle time of the platform, reduce operation efficiency, and even affect the quality of fresh flowers, thus having an adverse impact on the operating efficiency of enterprises. Therefore, how to optimize the unloading platform scheduling plan under limited resource conditions to improve the overall operation efficiency of the fresh flower cold chain logistics has become one of the key research issues in the current logistics scheduling optimization field. The platform scheduling optimization problem can be modeled as an unrelated parallel machine scheduling problem, which belongs to the NP-hard problem and has characteristics such as a large problem scale, many constraint conditions, and high solution difficulty. For such problems, existing research mainly uses heuristic rule methods and mathematical optimization methods to solve them. Heuristic methods such as "first come, first served" are more common in practical applications because of their low computational complexity and simple implementation. However, their scheduling strategies rely on fixed rules and it is difficult to fully consider the comprehensive influence of multiple factors such as vehicle arrival time, operation duration, and resource constraints, resulting in the lack of globality in the optimization results. Mathematical optimization methods such as integer programming and mixed integer linear programming can obtain better solutions in problems of a certain scale, but when the scheduling scale expands, the computational complexity increases significantly, making it difficult to meet the requirements of scheduling efficiency and real-time performance in actual scenarios. In addition, the unloading platform scheduling problem in fresh flower cold chain logistics involves multiple factors such as vehicle arrival time, operation time, and platform resource allocation, further increasing the complexity of the problem. However, traditional methods have great limitations in dealing with such complex problems. Therefore, it is necessary to explore more efficient and practical solution methods. Summary of the Invention

[0003] The technical problem solved by the present invention: Provide a fresh flower unloading platform scheduling optimization system and method based on deep reinforcement learning and evolutionary algorithm, which can effectively reduce the queuing time of vehicles, the idle time of the platform, and resource consumption.

[0004] The technical solution adopted by the present invention is as follows: fully considering information such as the reservation time and operation duration of the vehicle, aiming to reasonably allocate platforms for the vehicle, and constructing an optimization model for the flower unloading platform scheduling that minimizes the maximum completion time, maximally reduces the penalty cost of platform idle time, operation delay time, as well as various costs such as refrigeration cost, carbon emission cost, and flower loss caused by delay. To solve this model, a multi-strategy hybrid evolutionary algorithm based on DQN is proposed for solution. This algorithm first uses three methods, namely random generation, first-come-first-served scheduling rule, and minimum cost increment scheduling rule, to generate initial solutions with diversity and high quality. Subsequently, through selection methods such as roulette wheel selection, elite selection, and best individual retention strategy, the individuals for the next round of iteration are screened. In the initial search stage, global search is carried out with the help of evolutionary algorithms such as random retention crossover operation and two-point exchange operator to expand the search space. On this basis, by introducing a DQN agent as the decision maker, five destruction and repair search operators are adaptively selected to further enhance the local search ability and find better solutions. Finally, the algorithm generates new solutions through reverse mutation to avoid premature convergence, thereby improving the platform scheduling efficiency and reducing time waste and resource consumption in the flower cold chain logistics process.

[0005] An optimization system for flower unloading platform scheduling based on deep reinforcement learning and evolutionary algorithm, comprising:

[0006] Flowers, the goods that need to complete the unloading operation on the platform;

[0007] Refrigerated warehouse, used to store the flowers harvested by refrigerated vehicles and transported back;

[0008] Refrigerated vehicle, used to go to the customer-specified location to harvest flowers and transport them to the logistics park for unloading operation;

[0009] Platform, a platform for refrigerated vehicles to dock and perform flower unloading operations.

[0010] Preferably, the flowers are perishable, prone to withering, and difficult to store for a long time.

[0011] Preferably, at the initial moment of scheduling, all platforms are in an available state.

[0012] Preferably, the time window reserved by the vehicle to the platform scheduling management system is greater than or equal to the time required for the vehicle to unload.

[0013] Preferably, once the flower unloading operation starts, it is not allowed to be interrupted midway.

[0014] Preferably, each platform can only provide unloading operation services for one refrigerated vehicle at a time.

[0015] An optimization method for the flower unloading platform scheduling based on deep reinforcement learning and evolutionary algorithm, the specific steps are as follows:

[0016] Step1: Use natural number encoding to represent each individual, where each individual corresponds to a solution. The structure of the solution includes two rows of data: the first row represents the vehicle sequence to be unloaded, and the second row corresponds to the operation platform assigned to each vehicle. Assuming the number of vehicles to be unloaded is n, the length of the solution is n;

[0017] Step2: Use three methods to generate the initial solutions, namely: randomly generate, first-come-first-served scheduling rule, and minimum cost increment scheduling rule. Among them, the random generation method is simple to implement, and the generated solutions have diversity. For the first-come-first-served scheduling rule, first consider the reserved operation time window of the vehicle, and give priority to allocating the vehicle with an earlier start time of the unloading operation time window and a shorter required operation time; for the allocation of the platform, give priority to choosing the platform with an earlier idle time; for the minimum cost increment scheduling rule, first sort the vehicles in ascending order according to the start time of the reserved operation time window of the vehicle, and give priority to allocating the vehicle with an earlier start time; subsequently, allocate the vehicle to the platform that minimizes the total cost increment. If the objective function value increments of multiple platforms are the same, randomly allocate the vehicle to the platforms with the same increment;

[0018] Step3: Comprehensively consider the makespan, platform idle time penalty cost, operation delay time penalty cost, refrigeration cost, carbon emission cost, and flower loss cost to construct an optimization objective function that minimizes the comprehensive cost. Take the reciprocal of this objective function as the fitness function of the algorithm. The larger the fitness value, the better the performance of the solution;

[0019] Step4: Use three selection strategies to select individuals for the iterative process of the next generation. In the roulette wheel selection strategy, the fitness value of an individual determines its probability of being selected. The higher the fitness value, the greater the probability of being selected for the next search operation; the elite selection strategy selects the top 10% of the individuals with the highest fitness values and directly participates in the next search operation to enhance the convergence and global search ability of the algorithm; the best individual retention strategy retains the individual with the highest fitness value and makes it directly enter the next generation without participating in the subsequent search operation to ensure the inheritance of high-quality solutions and improve the stability of the solutions;

[0020] Step 5: Use the random retention crossover operation and the two-point exchange operator based on vehicle numbers and platform numbers to perform a preliminary global search on the population individuals. Among them, the random retention crossover operation means selecting some encodings from the parent generation and directly retaining them in the offspring to generate new solutions. The two-point exchange search operator based on vehicle numbers means randomly selecting two encoding values at different positions in the vehicle encoding sequence for exchange to generate a new solution. If the fitness value of the new solution is better than the original solution, accept the new solution; otherwise, retain the original solution. The two-point exchange search operator based on platform numbers means randomly selecting two positions in the platform encoding sequence where the corresponding platform numbers are different, exchanging their values to generate a new solution. If the fitness value of the new solution is better, accept the new solution; otherwise, retain the original solution.

[0021] Step 6: Based on the solutions obtained in Step 5, use the DQN agent in deep reinforcement learning to adaptively select five search operators for local search to further optimize the quality of the solutions. The five search operators are disruption and repair operators, and the design of the five disruption and repair operators fully considers factors such as randomness, greedy strategy, vehicle appointment unloading time window, and platform resource utilization rate.

[0022] Step 7: Generate new solutions through reverse mutation operations to expand the diversity of the population and prevent the algorithm from falling into local optimal solutions. Among them, the reverse mutation operation means randomly generating two position indexes within the interval [1, n], and reversing the encoding sequences between these two positions in the first-row vehicle encoding and the second-row platform encoding of the solution respectively to obtain a new solution.

[0023] Furthermore, the specific operations of the five disruption and repair operators are as follows: The first operator randomly deletes a certain number of vehicles to disrupt the current solution and then repairs it by greedy insertion. The second operator deletes all vehicle operations on the platform with the largest load and its corresponding platform number, and then completes the repair by greedy insertion. The third operator sorts the vehicles according to their contribution degrees to the objective function value, deletes the vehicles with greater influence, and then restores them to the optimal position by greedy insertion. The fourth operator combines greediness and randomness, randomly deletes some vehicles from the contribution degree sorted list, and completes the insertion repair by using the greedy strategy. The fifth operator randomly deletes the vehicles that violate the appointment time window constraint and preferentially selects the platform with the earliest idle time to arrange the unloading task to optimize the insertion order and complete the repair. This method can effectively improve the search efficiency and the quality of the solutions.

[0024] The beneficial effects of the present invention are as follows: By fully considering key information such as the vehicle reservation operation time and the operation time required by the vehicle, an optimization model for platform scheduling is constructed with the goals of minimizing the makespan, maximizing the reduction of the penalty cost for platform idle time, reducing the penalty cost for operation delay time, and reducing the refrigeration cost, carbon emission cost, and fresh flower loss cost caused by the delay. Since the optimization goal involves multiple complex factors and the problem is difficult to solve, a multi-strategy hybrid evolutionary algorithm based on DQN is proposed. This algorithm has significant advantages in terms of solution quality, convergence speed, etc., and can effectively improve the solution efficiency and stability of complex optimization problems. Through the combination of this model and algorithm, the efficiency of platform scheduling can be significantly improved, time waste and resource consumption in the fresh flower cold chain logistics process can be reduced, thereby realizing the optimization of the logistics process, reducing the operation cost, and enhancing the economic benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a schematic diagram of the system model of the present invention;

[0026] Figure 2 is a schematic diagram of the vehicle and platform coding of the present invention;

[0027] Figure 3 is a schematic diagram of the random retention crossover operation of the present invention;

[0028] Figure 4 is a schematic diagram of the two-point exchange search operation based on vehicle coding of the present invention;

[0029] Figure 5 is a schematic diagram of the two-point exchange search operation based on platform coding of the present invention;

[0030] Figure 6 is a decision network diagram of platform scheduling based on DQN of the present invention;

[0031] Figure 7 is a schematic diagram of the reverse mutation operation of the present invention;

[0032] Figure 8 is a schematic diagram of the solution algorithm flow of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0033] The present invention will be further described below in conjunction with the drawings and specific embodiments.

[0034] Example 1: As Figure 1 shown, a fresh flower unloading platform scheduling optimization system based on deep reinforcement learning and evolutionary algorithm includes:

[0035] Fresh flowers, goods that need to complete the unloading operation on the platform;

[0036] A refrigerated warehouse for storing fresh flowers harvested and transported back by refrigerated vehicles;

[0037] Refrigerated vehicles for harvesting fresh flowers at customer-specified locations and transporting them to a logistics park for unloading operations;

[0038] A platform for refrigerated vehicles to dock and perform fresh flower unloading operations.

[0039] Furthermore, the fresh flowers are perishable, prone to withering, and difficult to preserve for a long time.

[0040] Furthermore, at the initial moment of scheduling, all platforms are available.

[0041] Furthermore, the time window reserved by the vehicle with the platform scheduling management system should be greater than or equal to the time required for vehicle unloading.

[0042] Furthermore, once the fresh flower unloading operation starts, it is not allowed to be interrupted midway.

[0043] Furthermore, each platform can only provide unloading operation services for one refrigerated vehicle at a time.

[0044] As Figure 2 shown, using natural number encoding, each individual represents a solution. The first row of the solution is the sequence of vehicles to be unloaded, and the second row is the corresponding platform number for each vehicle. Suppose there are 8 vehicles to be unloaded, the vehicle sequence is 8, 1, 5, 2, 4, 3, 6, 7, and the corresponding platform sequence is 2, 2, 1, 1, 3, 3, 2, 1. Therefore, the length of the solution is equal to the number of vehicles to be unloaded. If the number of vehicles to be unloaded is n, then the length of the solution is n.

[0045] As Figure 3As shown, the core idea of randomly retaining the crossover operation is to directly select some encodings from the parent generation and retain them in the offspring, thereby generating new solutions. The specific operation steps are as follows: First, select two individuals, denoted as Parent 1 and Parent 2. Assume the length of the solution is 8, and randomly generate two numbers i and j in the interval [1, 8], ensuring that i < j. Assume i = 3 and j = 5. Next, perform the crossover operation on the vehicle encoding: For the first row of vehicle encoding of Parent 1 and Parent 2, copy the encodings of Parent 1 and Parent 2 in the interval [3, 5] to the corresponding positions of Offspring 1 and Offspring 2 respectively, keeping the original order unchanged. Subsequently, delete the copied vehicle encodings from Parent 1 and Parent 2, and the remaining vehicle encodings will be used to fill the blank positions in Offspring 1 and Offspring 2. Specifically, the remaining vehicle encodings in Parent 1 fill the blank positions in Offspring 2, and the remaining vehicle encodings in Parent 2 fill the blank positions in Offspring 1. Subsequently, perform the crossover operation on the platform encoding: For the second row of platform encoding of Parent 1 and Parent 2, copy the platform encodings in the interval [3, 5] to Offspring 1 and Offspring 2 respectively, and keep the original order unchanged. Then, delete the copied platform encodings from Parent 1 and Parent 2, and the remaining platform encodings will fill the blank positions in Offspring 1 and Offspring 2 respectively. Specifically, the remaining platform encodings in Parent 1 fill the blank positions in Offspring 2, and the remaining platform encodings in Parent 2 fill the blank positions in Offspring 1.

[0046] As Figure 4 shown, the two-point exchange search operator based on vehicle numbers: Randomly select two different positions in the first row of vehicle encoding. Assume positions 6 and 7 in the first row of vehicle encoding are selected, and then exchange the values at these two positions to generate a new solution. If the fitness value of the new solution is better, accept this new solution; otherwise, retain the original solution.

[0047] As Figure 5 shown, the two-point exchange search operator based on platform numbers: Randomly select two different positions in the second row of platform encoding, and ensure that the values at the selected positions are different. Assume positions 6 and 7 in the second row of platform encoding are selected, and then exchange the values at these two positions to generate a new solution. If the fitness value of the new solution is better, accept this new solution; otherwise, retain the original solution.

[0048] As Figure 6 shown, the agent in DQN adaptively selects five disruption and repair operators, aiming to select search operators that help improve the quality of the population. DQN evaluates the optimization effect of each operator through a reward function: When the selection of an operator improves the fitness value of the optimal solution or the average fitness value of the population, a positive reward is given; otherwise, a negative reward is given. The specific operations of the five disruption and repair operators are as follows: The first operator: Based on Figure 2In the solution encoding method, a certain number of vehicle codes are randomly deleted to disrupt the current solution, and then it is repaired by the greedy insertion method. The second operator: Based on Figure 2 In the solution encoding method, all vehicle operations on the platform with the largest load and their corresponding platform numbers are deleted, and then the repair is completed by the greedy insertion method. The third operator: Based on Figure 2 In the solution encoding method, the vehicles are sorted according to their contribution degrees to the objective function value, the vehicles with greater influence are deleted, and then they are restored to the optimal position by the greedy insertion method. The fourth operator: Based on Figure 2 In the solution encoding method, combining greediness and randomness, some vehicles are randomly deleted from the contribution degree sorted list, and the greedy strategy is adopted to complete the insertion repair. Based on Figure 2 In the solution encoding method, the vehicles that violate the appointment time window constraint are randomly deleted, and the platform with the earliest idle time is preferentially selected to arrange the unloading task, so as to optimize the insertion order and complete the repair.

[0049] As Figure 7 shown, assuming the length of the solution is 8, two numbers i and j are randomly generated in the interval [1, 8] first, and it is ensured that i < j. Assume i = 3 and j = 6. Then, in the vehicle code of the first row of the solution, the coding sequence between positions 3 and 6 is reversed; at the same time, in the platform code of the second row of the solution, the coding between positions 3 and 6 is reversed.

[0050] As Figure 8 shown, three methods are first used to generate the initial solution, including: the random generation method, the first-come-first-served scheduling rule, and the minimum cost increment scheduling rule, to generate an initial solution with high quality and diversity. Subsequently, three selection strategies, namely roulette wheel selection, elite selection, and best individual retention, are used to screen individuals to ensure the diversity of the population while retaining high-quality individuals. In the initial stage of the search, global search is carried out through random retention crossover operations and the two-point exchange operator based on vehicle and platform numbers to expand the search space. Further, a DQN agent is introduced as a decision-making module to adaptively select five destruction and repair search operators to enhance the local search ability and seek better solutions. Finally, the reverse mutation operation is used to generate new individuals to improve the population diversity and avoid the algorithm falling into the local optimal solution.

[0051] A method for optimizing the fresh flower unloading platform scheduling based on deep reinforcement learning and evolutionary algorithm, the specific steps are as follows:

[0052] Step1: Use natural number encoding to represent each individual, where each individual corresponds to a solution. The structure of the solution includes two rows of data: the first row represents the sequence of vehicles to be unloaded, and the second row corresponds to the operation platform assigned to each vehicle. Assume the number of vehicles to be unloaded is n, then the length of the solution is n;

[0053] Step 2: Three methods are adopted to generate initial solutions, namely: random generation, first-come-first-served scheduling rule, and minimum cost increment scheduling rule. Among them, the random generation method is simple to implement, and the generated solutions are diverse. For the first-come-first-served scheduling rule, the reservation operation time window of the vehicle is considered first, and the vehicle with an earlier unloading operation time window start time and a shorter required operation time is preferentially allocated; for the platform allocation, the platform with an earlier idle time is preferentially selected; for the minimum cost increment scheduling rule, the vehicles are first sorted in ascending order according to the start time of the reservation operation time window, and the vehicles with an earlier start time are preferentially allocated; subsequently, the vehicle is allocated to the platform that minimizes the total cost increment. If the objective function value increments of multiple platforms are the same, the vehicle is randomly allocated to the platforms with the same increment;

[0054] Step 3: Considering the makespan, platform idle time penalty cost, operation delay time penalty cost, refrigeration cost, carbon emission cost, and flower loss cost comprehensively, an optimization objective function with the minimum comprehensive cost is constructed, and the reciprocal of this objective function is used as the fitness function of the algorithm. The larger the fitness value, the better the performance of the solution;

[0055] Step 4: Three selection strategies are adopted to select individuals for the iterative process of the next generation. In the roulette wheel selection strategy, the fitness value of an individual determines its probability of being selected. The higher the fitness value, the greater the probability of being selected for the next search operation; the elite selection strategy selects the top 10% of the individuals with the highest fitness values and directly participates in the next search operation to enhance the convergence and global search ability of the algorithm; the best individual retention strategy retains the individual with the highest fitness value, making it directly enter the next generation without participating in the subsequent search operation to ensure the inheritance of high-quality solutions and improve the stability of the solutions;

[0056] Step 5: Use random retention crossover operation and two-point exchange operator based on vehicle number and platform number to perform a preliminary global search on the population individuals; among them, the random retention crossover operation means selecting some codes from the parent generation and directly retaining them to the offspring to generate new solutions; the two-point exchange search operator based on vehicle number means randomly selecting two coding values at different positions in the vehicle coding sequence for exchange to generate a new solution. If the fitness value of this new solution is better than the original solution, the new solution is accepted, otherwise the original solution is retained; the two-point exchange search operator based on platform number means randomly selecting two positions in the platform coding sequence, the corresponding platform numbers of which are different, and exchanging their values to generate a new solution. If the fitness value of the new solution is better, the new solution is accepted, otherwise the original solution is retained;

[0057] Step6: Based on the solution obtained in Step5, use the DQN agent in deep reinforcement learning to adaptively select five search operators for local search to further optimize the quality of the solution; the five search operators are destruction and repair type operators, and the design of the five destruction and repair type operators fully considers factors such as randomness, greedy strategy, vehicle appointment unloading time window, and platform resource utilization rate;

[0058] Step7: Generate a new solution through reverse mutation operation to expand the diversity of the population and prevent the algorithm from falling into a local optimal solution; among them, the reverse mutation operation refers to randomly generating two position indexes within the interval [1, n], and reversing the coding sequences between these two positions in the vehicle coding of the first row and the platform coding of the second row of the solution respectively, so as to obtain a new solution.

[0059] The specific implementation manners of the present invention have been described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation manners. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can also be made without departing from the purpose of the present invention.

Claims

1. A fresh flower unloading platform scheduling optimization system based on deep reinforcement learning and evolutionary algorithms, characterized in that: Including: Fresh flowers, goods that need to complete the unloading operation on the platform; Refrigerated warehouse, used to store fresh flowers harvested by refrigerated vehicles and transported back; Refrigerated vehicle, used to go to the customer-specified location to harvest fresh flowers and transport them to the logistics park for unloading operation; Platform, a platform for refrigerated vehicles to dock and perform fresh flower unloading operations.

2. The optimized system for scheduling the fresh flower unloading platform based on deep reinforcement learning and evolutionary algorithm according to claim 1, wherein: The fresh flowers are perishable, prone to withering, and difficult to store for a long time.

3. The optimized system for the flower unloading platform scheduling based on deep reinforcement learning and evolutionary algorithm according to claim 1, wherein: At the initial moment of scheduling, all platforms are available.

4. The optimized system for the flower unloading platform scheduling based on deep reinforcement learning and evolutionary algorithm according to claim 1, characterized in that: The time window reserved by the vehicle for the platform scheduling management system should be greater than or equal to the time required for vehicle unloading.

5. The optimized system for the flower unloading platform scheduling based on deep reinforcement learning and evolutionary algorithm according to claim 1, characterized in that: Once the fresh flower unloading operation starts, it is not allowed to be interrupted midway.

6. The optimized system for scheduling the fresh flower unloading platform based on deep reinforcement learning and evolutionary algorithm according to claim 1, characterized in that: Each platform can only provide unloading operation services for one refrigerated vehicle at a time.

7. A method for optimizing the scheduling of a fresh flower unloading platform based on deep reinforcement learning and evolutionary algorithms, characterized in that: The specific steps are as follows: Step1: Represent each individual using natural number encoding, where each individual corresponds to a solution. The structure of the solution includes two rows of data: the first row represents the sequence of vehicles to be unloaded, and the second row corresponds to the operation platform assigned to each vehicle. Assuming the number of vehicles to be unloaded is n, the length of the solution is n; Step2: Use three methods to generate the initial solution, namely: random generation, first-come-first-served scheduling rule, and minimum cost increment scheduling rule. In the first-come-first-served scheduling rule, first consider the reserved operation time window of the vehicle, and give priority to allocating the vehicle with an earlier start time of the unloading operation time window and a shorter required operation time; for the allocation of the platform, give priority to choosing the platform with an earlier idle time. In the minimum cost increment scheduling rule, first sort the vehicles in ascending order according to the start time of their reserved operation time windows, and give priority to allocating the vehicles with an earlier start time; Subsequently, allocate the vehicle to the platform that minimizes the total cost increment. If the increment of the objective function value of multiple platforms is the same, randomly allocate the vehicle to the platforms with the same increment; Step3: Considering the makespan, platform idle time penalty cost, operation delay time penalty cost, refrigeration cost, carbon emission cost, and fresh flower loss cost comprehensively, construct an optimization objective function with the minimum comprehensive cost, and use the reciprocal of this objective function as the fitness function of the algorithm. The larger the fitness value, the better the performance of the solution; Step4: Use three selection strategies to select individuals for the iterative process of the next generation. In the roulette wheel selection strategy, the fitness value of the individual determines its probability of being selected. The higher the fitness value, the greater the probability of being selected for the next search operation. The elite selection strategy selects the top 10% of individuals with the highest fitness values and directly participates in the next search operation. The best individual retention strategy retains the individual with the highest fitness value and makes it directly enter the next generation without participating in the subsequent search operation; Step5: Conduct a preliminary global search on the population individuals using the random retention crossover operation and the two-point exchange operator based on vehicle numbers and platform numbers. Among them, the random retention crossover operation means selecting some encodings from the parent generation and directly retaining them in the offspring to generate a new solution. The two-point exchange search operator based on vehicle numbers means randomly selecting two encoding values at different positions in the vehicle encoding sequence and exchanging them to generate a new solution. If the fitness value of the new solution is better than that of the original solution, accept the new solution; otherwise, retain the original solution. The two-point exchange search operator based on platform numbers means randomly selecting two positions in the platform encoding sequence with different corresponding platform numbers, exchanging their values to generate a new solution. If the fitness value of the new solution is better, accept the new solution; otherwise, retain the original solution. Step6: Based on the solution obtained in Step5, use the DQN agent in deep reinforcement learning to adaptively select five search operators for local search. The five search operators are disruption and repair type operators. Step7: Generate new solutions through reverse mutation operations to expand the diversity of the population and prevent the algorithm from falling into local optimal solutions. Among them, the reverse mutation operation means randomly generating two position indexes within the interval [1, n], and reversing the encoding sequences between these two positions in the first-row vehicle encoding and the second-row platform encoding of the solution respectively to obtain a new solution.

8. An optimization method for the flower unloading platform scheduling based on deep reinforcement learning and evolutionary algorithm according to claim 7, characterized in that: The specific operations of the five disruption and repair operators are as follows: The first operator randomly deletes a certain number of vehicles to disrupt the current solution and then repairs it through greedy insertion. The second operator deletes all vehicle operations on the platform with the largest load and its corresponding platform number, and then completes the repair using the greedy insertion method. The third operator sorts the vehicles according to their contribution degrees to the objective function value, deletes the vehicles with greater influence, and then restores them to the optimal position through greedy insertion. The fourth operator combines greed and randomness, randomly deletes some vehicles from the contribution degree sorted list, and uses the greedy strategy to complete the insertion repair. The fifth operator randomly deletes the vehicles that violate the appointment time window constraint, and preferentially selects the platform with the earliest idle time to arrange the unloading task to optimize the insertion order and complete the repair.