DQN-based fruit freeze-drying production scheduling method, electronic equipment and storage medium

Through the DQN-based fruit freeze-dried production scheduling method, the problem of failure to effectively consider energy factors and long computing time in the prior art is solved, and fast and efficient production scheduling is achieved, reducing production costs and increasing profits.

CN120163376APending Publication Date: 2025-06-17HANSHAN NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510234489.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing fruit freeze-dried production scheduling methods fail to effectively consider energy factors, resulting in high production costs. The current production scheduling algorithm has a long operation time, making it difficult to meet the workshop's demand for rapid response on-site.

Method used

Using the fruit freeze-dried production scheduling method based on deep reinforcement learning (DQN) is used to construct a DQN model, and the optimal scheduling action is determined according to the status and electricity price changes of the fruit to be processed, production scheduling is optimized, and energy consumption is reduced.

Benefits of technology

It realizes rapid acquisition of scheduling solutions, high applicability, reduces computing time, significantly improves the efficiency and profit of production scheduling, and meets the rapid response needs of the workshop site.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163376A_ABST
    Figure CN120163376A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of intelligent fruit processing technologies, in particular to a DQN-based fruit freeze-drying production scheduling method, electronic equipment and a storage medium. The method comprises the following steps: S1, constructing a DQN model; s2, acquiring an environment state of a current processing batch, and determining an optimal scheduling action in the DQN model according to the environment state; s3, calculating the profit of the processing batch according to the optimal scheduling action, and taking the profit as an action reward; and S4, updating the state attributes of machines and workpieces of the next batch, returning the action rewards, and performing updating training on the DQN model. Scheduling schemes under various working conditions can be rapidly obtained through the DQN algorithm, the response speed is high, the applicability is high, compared with other reference algorithms, profits obtained through scheduling can be improved to different degrees, and the production scheduling effect is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of intelligent fruit processing technology, and in particular to a DQN-based fruit freeze-drying production scheduling method, an electronic device, and a storage medium. Background Art

[0002] The freeze-drying process is the development trend of fruit processing technology. Freeze-dried fruits are processed under low-temperature and vacuum conditions, which can better retain nutrients such as vitamins, minerals, fiber, and antioxidants in fruits. This method reduces the impact of heat and oxygen on fruits, thereby maximizing the preservation of the nutritional components, natural flavors, and tastes of fruits, extending their shelf life without the need to add preservatives, and is a typical representative of healthy foods.

[0003] Although the freeze-drying process has many advantages compared with other fruit processing methods, the freeze-drying process requires a long processing time and high energy consumption, resulting in a high price for freeze-dried products. Fruits generally need to be frozen to -30 - 50 °C and kept cold for a period of time before they are completely frozen. The power of freeze-drying machines is as high as dozens of kilowatts during the quick-freezing stage, and after the fruits are completely frozen, they also need to be kept warm in a vacuum environment for a long time to allow the water to completely sublimate. According to the characteristics that the energy consumption is high during the quick-freezing stage of the equipment, while the vacuum drying stage takes a long time, for this reason, time-of-use electricity price is an important factor that needs to be considered in production scheduling. Arranging the quick-freezing as much as possible during the valley time of the electricity price, and the valley-time electricity price can reduce the production cost. However, due to the tight-time constraints in fruit processing, delaying the processing task until the valley time may cause the fruits to deteriorate and increase the scrap cost, and increasing the idle waiting time of the equipment will lead to a decrease in the production capacity of the workshop. Therefore, production enterprises need to comprehensively consider the time-of-use electricity price and processing tight-time constraints, and perform reasonable scheduling to minimize the processing cost.

[0004] There are certain problems in the current scheduling methods for fruit freeze-drying preparation. For example, in terms of scheduling objectives, the existing scheduling methods usually consider the main objective of minimizing the total processing time, or minimizing the total tardiness of products, and rarely consider the energy factor, resulting in a high overall production cost and the inability to maximize the enterprise profit. At the same time, meta-heuristic algorithms represented by genetic algorithms can search for better scheduling results through iteration. However, when the scheduling scale is large, the meta-heuristic algorithms require a long operation time and it is difficult to meet the application requirements of rapid response in the workshop. Summary of the Invention

[0005] In order to solve the technical problems that the energy factor is not considered in the current fruit freeze-drying production scheduling process, the cost is high, and the operation time of the current production scheduling algorithm is long and it is difficult to meet the application requirements of the workshop, the embodiments of the present invention provide a DQN-based fruit freeze-drying production scheduling method, an electronic device, and a storage medium.

[0006] According to one aspect of the embodiments of the present invention, the present invention provides a fruit freeze-drying production scheduling method based on DQN, and the method includes:

[0007] S1. Taking the fruit slices to be processed as workpieces, arranging actions according to different sorting parameters of the workpieces, determining a plurality of scheduling actions to form a scheduling action space, determining a plurality of state features according to the state parameters of the workpieces to form a state space, and constructing a DQN model according to the scheduling action space and the state space;

[0008] S2. Obtaining the environmental state of the current processing batch, calculating the state features of the current processing batch according to the environmental state, and determining the optimal scheduling action in the DQN model according to the state features of the current processing batch, where the environmental state includes time-of-use electricity price, state attributes of machines and workpieces;

[0009] S3. Determining the workpieces to be processed in the current batch and the processing time period according to the optimal scheduling action, calculating the total value and total electricity cost of the workpieces to be processed in this processing batch, and obtaining the profit of this processing batch;

[0010] S4. Updating the state attributes of the machines and workpieces in the next batch, and using the profit of this processing batch as an action reward to update and train the DQN model.

[0011] In an alternative solution, in the step S1, the sorting parameters of the workpieces include size sorting, latest start time sorting, and value sorting; the scheduling actions further include a shutdown and waiting operation within a specific time interval.

[0012] In an alternative solution, the determining a plurality of state features according to the state parameters of the workpieces in the step S1 to form a state space specifically includes:

[0013] Defining the workpiece state parameters to be obtained;

[0014] Defining the workpieces as different workpiece sets according to the arrival time, latest start time, and scheduled situation in the workpiece state parameters;

[0015] Statistically analyzing the workpiece sets respectively, determining a plurality of state features of the workpieces, and forming a state space.

[0016] In an alternative solution, the step S2 specifically includes:

[0017] Obtaining the benchmark electricity cost of the current batch, the state attributes of the machines, and the state attributes of each workpiece;

[0018] Calculating the state attributes of the machines and workpieces according to the state space of the DQN model to obtain the state features of the current batch;

[0019] Determine the optimal scheduling action in the DQN model according to the status characteristics of the current batch.

[0020] In an alternative, the status attributes of the workpiece include the arrival time, the latest start time, the size, the processing time, and the scheduled situation of each workpiece.

[0021] In an alternative, the total electricity cost expenditure c of the processing batch b is:

[0022] c b =(PTP b1 *c p +PTF b1 *c f +PTO b1 *c o )*P1+(PTP b2 *c p +PTF b2 *c f +PTO b2 *c o )*P2;

[0023] Wherein, PTP b1 , PTF b1 and PTO b1 are the durations occupied by the freezing process of this batch at peak, normal, and valley electricity prices respectively; PTP b2 , PTF b2 and PTO b2 are the durations occupied by the drying process of this batch at peak, normal, and valley electricity prices respectively; c p , c f and c o are the peak electricity price, normal electricity price, and valley electricity price, P1 is the power of the freezing process, and P2 is the power of the drying process.

[0024] In an alternative, the action reward r of the current processing batch t is:

[0025] r t =V b -c b

[0026] Wherein, V b is the total value of the workpieces processed in the current processing batch.

[0027] In an alternative, the step S4 specifically includes:

[0028] Take the end time of the current processing batch as the planned start time for the next scheduling, and the unprocessed workpieces are returned as the unscheduled workpieces in the next batch and included in the environmental state of the next processing batch.

[0029] Return the input of the current processing batch, the optimal scheduling action, and the action reward r t to the experience replay pool for in-depth training of the DQN model.

[0030] According to the second aspect of the embodiments of the present invention, there is provided an electronic device, including a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus;

[0031] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations of the DQN-based fruit freeze-drying production scheduling method as described above.

[0032] According to the third aspect of the embodiments of the present invention, there is provided a storage medium, in which at least one executable instruction is stored. When the executable instruction runs on an electronic device, the electronic device is caused to execute the operations of the DQN-based fruit freeze-drying production scheduling method as described above.

[0033] The beneficial effects of the present invention are as follows:

[0034] 1. The DQN algorithm of the present invention is trained with random data and saves the trained deep Q-network model. In each subsequent scheduling, only the trained model needs to be imported to quickly obtain scheduling solutions under various working conditions, with high applicability, and solves the problem that the network model needs to be retrained when the number of workpieces to be scheduled changes in the past.

[0035] 2. In terms of operation time, when the scheduling scale is 500 workpieces, the operation time required by the DQN algorithm proposed by the present invention does not exceed 1 second, greatly saving the operation time, with a fast response speed, and can meet the real-time scheduling of the production workshop.

[0036] 3. In terms of scheduling results, compared with other benchmark algorithms, the scheduling method in the present invention can improve the profit obtained by scheduling to varying degrees under various machine capacities and various workshop processing requirements, and has a good scheduling effect. Description of the Drawings

[0037] The drawings are only used to illustrate the embodiments and are not considered to limit the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0038] Figure 1Shows the structural flowchart of the DQN-based fruit freeze-drying production scheduling method in an embodiment of the present invention.

[0039] Figure 2 Shows the overall framework of the DQN algorithm in an embodiment of the present invention

[0040] Figure 3 Shows the specific structural flowchart in step S1 in an embodiment of the present invention.

[0041] Figure 4 Shows the specific structural flowchart in step S2 in an embodiment of the present invention.

[0042] Figure 5 Shows the specific flowchart of step S4 in an embodiment of the present invention.

[0043] Figure 6 Shows the structural block diagram of an electronic device in an embodiment of the present invention. Detailed implementation manners

[0044] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein.

[0045] Embodiment 1:

[0046] Figure 1 Shows the structural flowchart of the DQN-based fruit freeze-drying production scheduling method provided in this embodiment. Figure 2 Shows the overall framework of the DQN algorithm in this embodiment.

[0047] Please refer to Figure 1 and Figure 2 , this embodiment provides a DQN-based fruit freeze-drying production scheduling method, which is applied to fruit freeze-drying production. The method includes:

[0048] S1. Taking the fruit slices to be processed as workpieces, arranging actions according to different sorting parameters of the workpieces, determining a plurality of scheduling actions to form a scheduling action space, determining a plurality of state features according to the state parameters of the workpieces to form a state space, and constructing a DQN model according to the scheduling action space and the state space.

[0049] Among them, the sorting parameters of the workpieces include but are not limited to the size, latest start time, and value of the workpieces. The state parameters of the workpieces include time-of-use electricity price, the state attributes of the machines, and the state attributes of the workpieces. Among them, the machine state attributes include machine capacity and idle time, and the workpiece state attributes include the arrival time, latest start time, size, processing time, and scheduling status of each workpiece. The information space dimensions included in these elements of the environmental state are not unified, and when there are many workpieces to be scheduled, the state space of the workpiece attributes will be too large, making it difficult for the deep Q-network to accurately identify each state. Therefore, these information of the environmental state are refined into a state space containing multiple state features.

[0050] S2. Obtain the environmental state of the current processing batch, calculate the state features of the current processing batch according to the environmental state, and determine the optimal scheduling action in the DQN model according to the state features of the current processing batch. The environmental state includes time-of-use electricity price, and the state attributes of the machines and workpieces.

[0051] Among them, the environmental state of the current processing batch refers to the time parameters, electricity price, and the state attributes of the machines and workpieces of the current scheduled processing batch. For example, if the start time of the current processing batch is 12 noon, during the peak electricity consumption period, and the latest start time of the workpiece is relatively long, and there are other conditions tending to stop and wait, after being determined by the DQN model that the optimal scheduling action is to stop and wait, then add a time gap, such as one hour, to the planned start time of this batch, and return an action reward value of 0.

[0052] S3. Determine the workpieces to be processed in the current batch and the processing time period according to the optimal scheduling action, calculate the total value and total electricity cost of the workpieces processed in this processing batch, and obtain the profit of this processing batch.

[0053] S4. Update the state attributes of the machines and workpieces in the next batch, and use the profit of this processing batch as the action reward to update and train the DQN model.

[0054] Among them, according to the current environmental state, determine the optimal scheduling action from the action space according to the deep Q-network, complete the scheduling of the current batch, update the next environmental state, and calculate the reward value of the current action. Return the relevant data of the current processing batch, the optimal scheduling action, and the reward value of the action to the experience replay pool for training the DQN model, so that the optimal batch forming strategy generated by the DQN model is more perfect.

[0055] In a specific embodiment, a specific application scenario is provided. That is, as the production basis, there are n workpieces in the potential processing orders of an enterprise. The workshop can reject the processing of any workpiece. If it chooses to process, it can obtain the profit corresponding to the value of the workpiece. The processing cost only considers the electricity cost generated by the equipment during the processing process, and fixed costs such as labor and equipment depreciation are ignored. The arrival time, latest start time, size, processing time, and value of each workpiece are different. Multiple workpieces can be grouped into a batch and processed simultaneously in the equipment as long as the total size of the workpieces does not exceed the equipment capacity. The start processing time of each batch cannot exceed the latest start time of all workpieces in the batch. Each batch needs to go through two stages of quick-freezing and drying in the same machine. The quick-freezing processing time and drying processing time of each batch are determined by the workpiece with the longest processing time in the batch. The power of the machine during quick-freezing and drying is different. Once a batch starts the freezing process, it must be continuously processed until drying is completed without interruption or insertion of other batches.

[0056] Since the ultimate goal of the enterprise is profit, only by combining energy conservation and profit can it meet the actual production application requirements. Therefore, the production scheduling method in the present invention fully considers the time-of-use electricity price, the tightness constraint of the workpiece, and the power change of the equipment, and the scheduling goal is to maximize the profit of n workpieces in the potential orders, that is, the total value of the selected processed workpieces minus the electricity cost.

[0057] In this scenario, the refinement scheme of step S1 is as follows.

[0058] As one of the preferences, the state space specifically includes 14 state features, which are respectively obtained by counting 4 different workpiece sets. Among them, the definitions of workpiece sets J1, J2, J3, and J4 are as follows.

[0059] J1: Un-scheduled workpieces that arrive before the machine idle time and whose latest start time is not less than the machine idle time.

[0060] J2: Workpieces that arrive during the postponed time period.

[0061] J3: Un-scheduled workpieces that cannot meet the latest start time due to postponed processing.

[0062] J4: Workpieces that arrive after the start processing time of the next batch plan.

[0063] As a preference, the sorting parameters of the workpieces include size sorting, latest start time sorting, and value sorting; the scheduling actions also include the operation of stopping and waiting within a specific time interval.

[0064] Among them, in terms of the action space design, at each decision-making moment, the agent needs to determine whether to start the next batch of processing at the original planned time or wait and postpone the planned start time of the next batch. If the processing starts at the original planned time, it is necessary to group the workpieces to be processed. When the total size of the arrived workpieces is greater than the machine capacity, it is necessary to determine which workpieces are given priority for processing according to different sorting parameters.

[0065] Specifically, grouping and batch processing according to the workpiece size helps to fill the machine as much as possible; grouping and batch processing according to the latest start time of the workpiece helps to give priority to the processing of urgent workpieces, thereby reducing the workpieces that cannot meet the latest start time; grouping and batch processing according to the workpiece value helps to give priority to the processing of high-value workpieces. These sorting and batch processing rules each have their own advantages, and it is necessary to select the batch processing rule according to different environmental states. The idle time of the machine is preset as the planned start time for the next batch of processing. Therefore, as a specific implementation manner, the action space consists of the following 4 actions.

[0066] Action 1: Group and process the workpieces in workpiece set J1 in descending order of size.

[0067] Action 2: Group and process the workpieces in workpiece set J1 in ascending order of the latest start time.

[0068] Action 3: Group and process the workpieces in workpiece set J1 in descending order of value.

[0069] Action 4: Stop and wait, and increase the planned start time of the next batch by a time interval. Preferably, a time interval is one hour. Of course, in actual applications, it can be adjusted to other durations according to needs.

[0070] Figure 3 The specific structural flowchart in step S1 is shown.

[0071] As one of the preferences, please refer to Figure 3 , determining multiple state features according to the state parameters of the workpieces in step S1 to form the state space specifically includes:

[0072] S11: Define the workpiece state parameters to be obtained;

[0073] S12: Define the workpieces as different workpiece sets according to the arrival time, the latest start time, and the scheduled situation in the workpiece state parameters;

[0074] S13: Statistically analyze the workpiece sets respectively, determine multiple state features of the workpieces, and form the state space.

[0075] Among them, the state information of the environment is defined, and state features are established. Subsequently, the rapid recognition of state features is achieved by inputting each state information. Specifically, the environmental state information includes time-of-use electricity price, the state attributes of the machine, and the state attributes of the workpiece, including but not limited to, the machine state attributes include machine capacity and idle time, and the workpiece state attributes include the arrival time, latest start time, size, processing time, and scheduling status of each workpiece.

[0076] To better understand this embodiment, a specific state space planning is provided. Among them, the state space includes 14 state features as shown below. The batch containing the workpiece with the longest processing time in the current batch is defined as the reference batch. The single-batch reference electricity cost is calculated by assuming that the reference batch starts processing at time zero.

[0077] State feature 1: The ratio of the total size of the workpieces in workpiece set J1 to the machine capacity.

[0078] State feature 2: The ratio of the maximum size in workpiece set J1 to the machine capacity.

[0079] State feature 3: The ratio of the standard deviation of the sizes in workpiece set J1 to the machine capacity.

[0080] State feature 4: The ratio of the minimum size in workpiece set J1 to the machine capacity.

[0081] State feature 5: The ratio of the standard deviation of the values of the workpieces in workpiece set J1 to the maximum value.

[0082] State feature 6: The ratio of the minimum value of the workpieces in workpiece set J1 to the maximum value.

[0083] State feature 7: The ratio of the result obtained by subtracting the planned start time of this batch from the maximum value of the latest start time of the workpieces in workpiece set J1 to the processing time of the reference batch.

[0084] State feature 8: The ratio of the result obtained by subtracting the machine idle time from the minimum value of the latest start time of the workpieces in workpiece set J1 to the processing time of the reference batch.

[0085] State feature 9: The ratio of the standard deviation of the latest start time of the workpieces in workpiece set J1 to the processing time of the reference batch.

[0086] State feature 10: The ratio of the total value of the workpieces in workpiece set J1 to the single-batch reference electricity cost.

[0087] State feature 11: The ratio of the standard deviation of the values of the workpieces in workpiece set J1 to the maximum value.

[0088] State feature 12: The ratio of the minimum value of the workpieces in workpiece set J1 to the maximum value.

[0089] Status feature 13: The ratio of the electricity cost saved by delaying the processing of a reference batch to the reference electricity cost.

[0090] Status feature 14: The ratio of the total size of the workpieces in workpiece set J4 to the machine capacity.

[0091] Figure 4 The specific structural flowchart in step S2 is shown.

[0092] As one of the preferences, please refer to Figure 4 , in this embodiment, step S2 specifically includes:

[0093] S21. Obtain the reference electricity cost of the current batch, the status attributes of the machine, and the status attributes of each workpiece;

[0094] S22. Calculate the status features of the current batch according to the state space of the DQN model for the status attributes of the machine and the workpiece;

[0095] S23. Determine the optimal scheduling action in the DQN model according to the status features of the current batch.

[0096] Among them, when the DQN model obtains the environmental state information of the current batch, calculates the status information of the current batch according to the definition of the status features, and compares and identifies it in the DQN model to determine the subsequent optimal scheduling action to maximize the processing profit of the enterprise.

[0097] Further preferably, in step S3, when determining the optimal scheduling action, the optimal scheduling action is executed as the processing plan for the current batch, and the current total profit is calculated and recorded and returned as the action reward. Specifically, the action reward r t for the current processing batch is:

[0098] r t = V b - c b

[0099] where V b is the total value of the processed workpieces in the current processing batch.

[0100] And the total electricity cost expenditure c b for this processing batch is:

[0101] c b = (PTP b1 * c p + PTF b1 * c f + PTO b1 * c o ) * P1 + (PTP b2 * cp +PTF b2 *c f +PTO b2 *c o )*P2;

[0102] Among them, PTP b1 , PTF b1 and PTO b1 are the durations occupied by the freezing process of this batch at peak, normal, and valley electricity prices respectively; PTP b2 , PTF b2 and PTO b2 are the durations occupied by the drying process of this batch at peak, normal, and valley electricity prices respectively; c p , c f and c o are the peak electricity price, normal electricity price, and valley electricity price respectively, P1 is the power of the freezing process, and P2 is the power of the drying process.

[0103] Figure 5 shows the specific flowchart of step S4 in this embodiment.

[0104] Preferably, please refer to Figure 5 , step S4 specifically includes:

[0105] S41. Take the end time of the current processing batch as the planned start time of the next scheduling, and the unprocessed workpieces are returned as the unscheduled workpieces of the next batch and included in the environmental state of the next processing batch.

[0106] S42. Return the input, optimal scheduling action, and action reward r t of the current processing batch to the experience replay pool for in-depth training of the DQN model.

[0107] Among them, for the calculation of the planned start time of the next scheduling, specifically, if action 4 is adopted, that is, stop and wait. The planned start time of the next scheduling is the planned start time of this scheduling batch plus a time gap. If action 1, action 2, or action 3 is adopted, then it is based on the planned start time of this scheduling batch plus the processing time of this batch in the freezing process and the drying process.

[0108] Embodiment 2:

[0109] Figure 6 shows the structural block diagram of the electronic device provided in this embodiment.

[0110] As shown in Figure 6As shown in the figure, this embodiment provides an electronic device, which includes: a processor 201, a memory 202, a communication interface 203, and a communication bus 204. Among them, the processor 201, the memory 202, and the communication interface 203 complete communication with each other through the communication bus 204;

[0111] The memory 202 is used to store at least one executable instruction, and the executable instruction causes the processor 201 to execute the operations of the DQN-based fruit freeze-drying production scheduling method in Embodiment 1.

[0112] Embodiment 3:

[0113] This embodiment provides a computer-readable storage medium, in which at least one executable instruction is stored. When the executable instruction runs on an electronic device, it causes the electronic device to execute the operations of the DQN-based fruit freeze-drying production scheduling method in Embodiment 1.

[0114] It should be noted that the above embodiments illustrate the present invention rather than limit the present invention, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.

Claims

1. A fruit freeze-dried production scheduling method based on DQN, characterized in that: The method comprises: S1. Take the fruit slices to be processed as workpieces, arrange actions according to different sorting parameters of the workpieces, determine multiple scheduling actions, form a scheduling action space, determine multiple state features according to the state parameters of the workpieces, form a state space, and build a DQN model based on the scheduling action space and the state space; S2. Obtain the environmental state of the current processing batch, calculate the state characteristics of the current processing batch according to the environmental state, and determine the optimal scheduling action in the DQN model according to the state characteristics of the current processing batch, wherein the environmental state includes the time-of-use electricity price, the state attributes of the machine and the workpiece; S3. Determine the workpieces to be processed in the current batch and the processing time period according to the optimal scheduling action, calculate the total value of the workpieces to be processed in the processing batch and the total electricity cost expenditure, and obtain the profit of the processing batch; S4. Update the state attributes of the machines and workpieces for the next batch, and use the profit of the processing batch as the action reward to update the DQN model.

2. The DQN-based fruit freeze-dried production scheduling method according to claim 1, characterized in that: In the step S1, the sorting parameters of the workpieces include size sorting, latest start time sorting and value sorting; the scheduling action also includes stopping and waiting for operation within a specific time interval.

3. The DQN-based fruit freeze-dried production scheduling method according to claim 1, characterized in that: Determining a plurality of state features according to the state parameters of the workpiece to form a state space in step S1 specifically includes: Define the artifact status parameters that need to be obtained; According to the arrival time, the latest start time and the scheduled status in the status parameters of the workpiece, the workpiece is defined as different workpiece sets; Statistics are taken on the workpiece sets respectively to determine multiple state characteristics of the workpieces and form a state space.

4. The DQN-based fruit freeze-dried production scheduling method according to claim 1, characterized in that: The step S2 specifically includes: Get the current batch's base electricity cost, the machine's status attributes, and each workpiece's status attributes; According to the state space of the DQN model, the state attributes of the machine and the workpiece are calculated to obtain the state characteristics of the current batch; Determine the optimal scheduling action in the DQN model based on the state characteristics of the current batch.

5. The DQN-based fruit freeze-dried production scheduling method according to claim 1, characterized in that: The state attributes of the workpieces include the arrival time, the latest start time, the size, the processing time and the scheduled status of each workpiece.

6. The DQN-based fruit freeze-dried production scheduling method according to claim 1, characterized in that: The total electricity cost of the processing batch c b for: c b =(PTP b1 *c p +PTF b1 *c f +PTO b1 *c o )*P1+(PTP b2 *c p +PTF b2 *c f +PTO b2 *c o )*P2; Among them, PTP b1 ,PTF b1 and PTO b1 The durations of the freezing process of the batch during peak, normal and valley electricity prices respectively; PTP b2 ,PTF b2 and PTO b2 is the duration of the drying process of the batch during peak, normal and valley electricity prices respectively; c p 、c f and c o are peak electricity prices, normal electricity prices and valley electricity prices, P1 is the power of the freezing process, and P2 is the power of the drying process.

7. The DQN-based fruit freeze-dried production scheduling method according to claim 1, characterized in that: The action reward r of the current processing batch t for: r t =V b -c b Among them, V b It is the total value of the workpieces processed in the current processing batch.

8. The DQN-based fruit freeze-dried production scheduling method according to claim 1, characterized in that: The step S4 specifically includes: The end time of the current processing batch is used as the planned start time of the next scheduling, and the unprocessed workpieces are returned as unscheduled workpieces of the next batch and included in the environmental status of the next processing batch; The input of the current processing batch, the optimal scheduling action and the action reward r t Returns the experience recycling pool for deep training of the DQN model.

9. An electronic device, characterized in that: It includes a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform the operations of the above-mentioned DQN-based fruit freeze-dried production scheduling method.

10. A storage medium, characterized in that: The storage medium stores at least one executable instruction. When the executable instruction is executed on the electronic device, the electronic device executes the operation of the DQN-based fruit freeze-dried production scheduling method.