Workshop production scheduling information determination method and apparatus

By using the trained target production decision model in the manufacturing workshop, and automatically determining the production scheduling information based on the production status data, the problem of poor reliability and availability of the workshop production scheduling strategy is solved, and highly intelligent production scheduling is achieved.

CN120069455APending Publication Date: 2025-05-30CHINA TOBACCO ZHEJIANG IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510236169.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The production process of modern manufacturing workshops is complex, the production methods are flexible and changeable, and are easily disturbed by dynamic events, resulting in poor reliability and availability of production scheduling strategies. It is difficult to consider long-term impacts based on decision-making models, which may lead to economic losses.

Method used

Through the trained target production decision model, production scheduling information is automatically determined based on the production status data of the production workshop, and a high-reliability and high-availability production scheduling strategy is generated. The model is trained by multiple training samples, which contain production status data, simulation scheduling information and predicted status data to ensure long-term impact considerations of decision-making.

Benefits of technology

The intelligence level of production scheduling is improved, the reliability and availability of production scheduling is ensured, and economic losses caused by decision-making errors are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069455A_ABST
    Figure CN120069455A_ABST
Patent Text Reader

Abstract

Embodiments of the invention provide a workshop production scheduling information determination method and apparatus. The method comprises the steps of obtaining production state data of a target production workshop; wherein the production state data comprises at least one of raw material inventory, equipment utilization rate of each production line, order quantity of at least one article category, order production plan information and order delivery time, and inputting the production state data into a pre-trained target production decision model, the production scheduling information of the target production workshop in the future preset time period is obtained, the production scheduling information can be automatically determined according to the production state data of the production workshop through the trained target production decision model, and a high-reliability and high-availability production scheduling strategy can be obtained, so that the intelligent level of production scheduling is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automatic control technology, and in particular, to a method and device for determining workshop production scheduling information. Background Art

[0002] Modern manufacturing has entered the era of meager profits, and the lean production management mode has replaced the traditional rough production management mode. Nowadays, the production process in the manufacturing workshop is long and complex, the production mode is flexible and changeable, and the production process is easily interfered by dynamic events, which poses challenges to the intelligent decision-making and rapid response capabilities of manufacturing enterprises.

[0003] Currently, using production management information tools or production decision models is an important way to achieve lean production. The production scheduling method based on production management information tools is usually: based on the method of backward scheduling according to the production task delivery date, the production scheduling result is output, and then the personnel arrangement, raw material configuration, and equipment shift of the entire workshop are overall scheduled manually according to the production scheduling result; the production scheduling method based on the production decision model is usually: an initial decision model is pre-trained based on a large number of training samples in advance. After the trained initial decision model is deployed and applied, through a large number of explorations, trial and errors, and optimizations, it finally converges to the optimal production scheduling strategy.

[0004] However, the production scheduling method based on production management information tools cannot output reliable production scheduling results under the production conditions of flexible and complex equipment shifts and mismatches of various production resources, resulting in poor reliability and availability of the production scheduling strategy; for the production scheduling method based on the production decision model, the trained initial decision model mainly focuses on immediate rewards or cumulative rewards within a limited number of steps, and insufficiently considers the long-term impact of decisions, which makes it difficult to obtain an initial decision model with good decision-making ability. In addition, in enterprise production scheduling, some exploratory scheduling decisions may bring irreversible consequences and cause greater economic losses. Summary of the Invention

[0005] Embodiments of the present invention provide a method and device for determining workshop production scheduling information, so that the trained target production decision model can automatically determine production scheduling information according to the production status data of the production workshop, obtain a production scheduling strategy with high reliability and high availability, and improve the intelligent level of production scheduling.

[0006] In a first aspect, embodiments of the present invention provide a method for determining workshop production scheduling information, and the method includes:

[0007] Obtain the production status data of the target production workshop; wherein, the production status data includes at least one of raw material inventory, equipment utilization rate of each production line, order quantity of at least one item category, order production plan information, and order delivery time;

[0008] Input the production status data into a pre-trained target production decision model to obtain the production scheduling information of the target production workshop within a preset time period in the future; wherein, the target production decision model is obtained by training with multiple training samples, and the training samples include first production status data, first simulated scheduling information corresponding to the first production status data, and second production status data. The first simulated scheduling information is obtained by processing the first production status data based on a production decision model to be processed, and the second production status data is the production status data deduced based on a state prediction model from the first production status data and the first simulated scheduling information. The first production status data is related to the deduced second production status data. The production scheduling information includes at least one of raw material procurement plan information, shift scheduling plan information for each production line, and raw material allocation information.

[0009] In a second aspect, an embodiment of the present invention further provides a device for determining workshop production scheduling information, and the device includes:

[0010] A status data acquisition module, configured to obtain the production status data of the target production workshop; wherein, the production status data includes at least one of raw material inventory, equipment utilization rate of each production line, order quantity of at least one item category, order production plan information, and order delivery time;

[0011] A scheduling information determination module, configured to input the production status data into a pre-trained target production decision model to obtain the production scheduling information of the target production workshop within a preset time period in the future; wherein, the target production decision model is obtained by training with multiple training samples, and the training samples include first production status data, first simulated scheduling information corresponding to the first production status data, and second production status data. The first simulated scheduling information is obtained by processing the first production status data based on a production decision model to be processed, and the second production status data is the production status data deduced based on a state prediction model from the first production status data and the first simulated scheduling information. The first production status data is related to the deduced second production status data. The production scheduling information includes at least one of raw material procurement plan information, shift scheduling plan information for each production line, and raw material allocation information.

[0012] In a third aspect, an embodiment of the present invention further provides an electronic device, and the electronic device includes:

[0013] One or more processors;

[0014] A storage device for storing one or more programs,

[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the workshop production scheduling information determination method according to any one of the embodiments of the present invention.

[0016] In a fourth aspect, an embodiment of the present invention further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the workshop production scheduling information determination method according to any one of the embodiments of the present invention when executed by a computer processor.

[0017] The technical solution of the embodiment of the present invention is to obtain production status data of a target production workshop, where the production status data includes at least one of raw material inventory, equipment utilization rate of each production line, order quantity of at least one item category, order production plan information, and order delivery time. Thus, the production status data is input into a pre-trained target production decision model to obtain production scheduling information of the target production workshop within a preset time period in the future. The target production decision model is obtained by training with multiple training samples. The training samples include first production status data, first simulated scheduling information corresponding to the first production status data, and second production status data. The first simulated scheduling information is obtained by processing the first production status data by a to-be-processed production decision model, and the second production status data is production status data deduced based on a status prediction model from the first production status data and the first simulated scheduling information. The first production status data is related to the deduced second production status data. The production scheduling information includes at least one of raw material procurement plan information, shift scheduling plan information for each production line, and raw material allocation information. The technical solution provided by the embodiment of the present invention can automatically determine production scheduling information according to the production status data of the production workshop through the trained target production decision model, and can obtain a production scheduling strategy with high reliability and high availability, thereby improving the intelligent level of production scheduling. Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the introduced drawings are only the drawings of a part of the embodiments to be described by the present invention, rather than all the drawings. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0019] Figure 1Schematic flowchart of a method for determining workshop production scheduling information provided by an embodiment of the present invention;

[0020] Figure 2 Schematic flowchart of another method for determining workshop production scheduling information provided by an embodiment of the present invention;

[0021] Figure 3 Schematic structural diagram of a device for determining workshop production scheduling information provided by an embodiment of the present invention;

[0022] Figure 4 Schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0023] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the accompanying drawings.

[0024] Embodiment 1

[0025] Figure 1 Schematic flowchart of a method for determining workshop production scheduling information provided by an embodiment of the present invention. This embodiment is applicable to situations where production scheduling and intelligent planning are required for a production workshop. This method can be executed by a device for determining workshop production scheduling information, and this device can be implemented in the form of software and / or hardware. The hardware can be an electronic device, such as a mobile terminal, a PC, or a server, etc.

[0026] As Figure 1 shown, the method for determining workshop production scheduling information includes:

[0027] S110. Obtain production status data of the target production workshop.

[0028] Among them, the target production workshop refers to an item production workshop for which production scheduling information needs to be determined for intelligent planning. For example, the target production workshop can be a tobacco product production workshop.

[0029] Among them, the production status data refers to the status information of the target production workshop in the current period. Specifically, the production status data may include at least one of the raw material inventory, the equipment utilization rate of each production line, the order quantity of at least one item category, the order production plan information, and the order delivery time. The raw material inventory refers to the inventory of materials required for producing items. The target production workshop can produce products of one or more item categories. For example, the target production workshop can produce products of item category A, item category B, and item category C. The target production workshop may include multiple production lines, and each production line can produce products of various item categories. Each production line is equipped with equipment for producing products. For example, taking a tobacco production workshop as an example, the equipment may include various types of equipment for tobacco production such as a cutting machine, a drying equipment, and a tipping equipment. The equipment utilization rate corresponding to a certain production line can be determined according to the working status information of the corresponding equipment of that production line within a period of time (the working status information may include: the duration of the running status, the duration of the idle status, the duration of shutdown for maintenance, the running efficiency, the duration of the fault status, the cumulative running time, and the equipment performance parameter information). The order quantity of a certain item category refers to the number of orders for which customers purchase products of that item category. The order production plan information includes information such as the planned start and end times of each process of products of a certain item category, the equipment usage arrangement, and the worker allocation. In specific applications, the order production plan can be represented concretely through a multi-dimensional array. The order delivery time refers to the order delivery date specified by the customer.

[0030] In this embodiment, the production status data of the target production workshop can be obtained periodically according to a preset scheduled task. For example, the scheduled task can be to obtain the production status data of the target production workshop every 24 hours to determine the production scheduling information of the target production workshop for the next day based on the production status data. It can also be that when it is detected that a preset trigger condition is met, the production status data of the target production workshop is obtained. For example, a server for executing the method for determining the workshop production scheduling information provided in this embodiment can be pre-configured in the control server of the target production workshop, and a preset control for enabling this method can be configured. When it is detected that this preset control is triggered, the production status data of the target production workshop can be obtained at this time.

[0031] S120. Input the production status data into a pre-trained target production decision model to obtain the production scheduling information of the target production workshop within a preset future time period.

[0032] Among them, the target production decision model refers to a reinforcement learning model that processes and analyzes the production status data and outputs the production scheduling information. In the specific application process, the target production decision model is a model that has been pre-trained. The preset future time period refers to the future time period with a preset duration from the current moment.

[0033] Among them, the production scheduling information refers to the information content on which the target production workshop relies when executing production activities. Specifically, the production scheduling information includes at least one of the raw material procurement plan information, the shift plan information of each production line, and the raw material allocation information corresponding to each production line. The raw material procurement plan information refers to the type and quantity information of raw materials to be procured. The raw material allocation information refers to the type and quantity information of raw material supplies to be allocated to each production line. The shift plan information refers to the shift information of a certain production line, the working hours of each shift, the staff of each shift, the product categories produced in each shift, the estimated production quantity information, etc.

[0034] Among them, the target production decision model is obtained by training with multiple training samples. The training samples include first production state data, first simulated scheduling information corresponding to the first production state data, and second production state data. The first simulated scheduling information is obtained by processing the first production state data based on the production decision model to be processed. The second production state data is the production state data deduced based on the state prediction model from the first production state data and the first simulated scheduling information. The first production state data is related to the deduced second production state data.

[0035] In this embodiment, the training generator for training to obtain the target production decision model includes the production decision model to be processed and the state prediction model. The production decision model to be processed is a decision model constructed based on prior knowledge and experience rules, and is used to generate decision actions according to the current state. The model parameters in the production decision model to be processed are default initial values. The state prediction model is a prediction model constructed based on prior knowledge and historical data, and is used to predict the next state according to the current state and decision actions. The state prediction model can be modeled based on statistical analysis and probability modeling of historical data, or can be modeled by extracting and learning from state data using a deep neural network model. The training generator is essentially an algorithm framework, which realizes the integration and combination of the production decision model to be trained and the state prediction model in the form of code and programs. By training the above training generator through multiple training samples in a loop, the production decision model to be trained determined at the end of the iteration is the target production decision model.

[0036] In this embodiment, the first production state data refers to the production state data generated by the target production workshop in the historical period. In the model training stage, inputting the first production state data into the production decision model to be processed can obtain the first simulated scheduling information; inputting the first simulated scheduling information and the first production state data into the state prediction model can obtain the second production state data.

[0037] It should be particularly noted that the state prediction model is mainly used to construct the second production state data during the training process.

[0038] Next, the training process of the target production decision model can be described. Each training sample includes first production status data. In the specific training process, the first production status data can be input into the production decision model to be processed. The production decision model to be processed can output first simulation scheduling information. Furthermore, the first simulation scheduling information and the first production status data are input into the state prediction model that has completed preliminary training to obtain second production status data. Furthermore, according to a preset reward function and the second production status data, the target reward value corresponding to the production decision model to be processed is determined. The parameters of the production decision model to be processed are adjusted according to the target reward value to obtain the current production decision model. Subsequently, the current production decision model is repeatedly trained in a loop by the remaining first production data in the training samples until the preset iteration end condition is met, and the obtained current production decision model is determined as the target production decision model.

[0039] In this embodiment, based on the production status data of the target production workshop obtained in the previous step, the production status data can be input into the pre-trained target production decision model. The target production decision model can output the production scheduling information of the target production workshop within a preset time period in the future by processing and analyzing the production status data. Exemplarily, the production scheduling information output by the target production decision model can be expressed as: {Date: 2025-02-12 / Purchase 1000 kg of Material A, 1500 kg of Material B, 1500 kg of Material C, and 2000 kg of Material D; Production Line 1: Morning Shift (8:00-16:00): Raw material allocation information, Worker A, Produce Brand A, Estimated Output: 50000 pieces; Night Shift (16:00-24:00): Raw material allocation information, Worker B, Produce Brand B, Estimated Output: 45000 pieces; Production Line 2: Raw material allocation information, Worker C, Morning Shift (8:00-16:00): Produce Brand B, Estimated Output: 48000 pieces; Night Shift (16:00-24:00): Equipment maintenance}.

[0040] The technical solution of the embodiment of the present invention obtains the production status data of the target production workshop, where the production status data includes at least one of raw material inventory, equipment utilization rate of each production line, order quantity of at least one item category, order production plan information, and order delivery time; inputs the production status data into a pre-trained target production decision model to obtain the production scheduling information of the target production workshop within a preset time period in the future. The target production decision model is obtained by training multiple training samples. The training samples include first production status data, first simulation scheduling information corresponding to the first production status data, and second production status data. The first simulation scheduling information is obtained by processing the first production status data based on a to-be-processed production decision model, and the second production status data is the production status data deduced based on a status prediction model from the first production status data and the first simulation scheduling information. The first production status data is related to the deduced second production status data. The production scheduling information includes at least one of raw material procurement plan information, shift scheduling plan information for each production line, and raw material allocation information. The technical solution provided by the embodiment of the present invention can automatically determine the production scheduling information according to the production status data of the production workshop through the trained target production decision model, and can obtain a production scheduling strategy with high reliability and high availability, thereby improving the intelligent level of production scheduling.

[0041] Embodiment 2

[0042] Figure 2 FIG. is a schematic flowchart of a method for determining workshop production scheduling information provided by Embodiment 2 of the present invention. On the basis of the foregoing embodiment, the training method of the target production decision model is further refined. Technical terms that are the same or corresponding to those in the above embodiment will not be described in detail here.

[0043] As Figure 2 shown, the method includes:

[0044] S210. Construct a first training sample set based on at least one set of first production status data of the target production workshop and a decision-making post-status data set corresponding to the at least one set of first production status data.

[0045] The decision-making post-status data set refers to the status data determined after the target production workshop completes the production scheduling task according to the decision determined by the first production status data. The first training sample set refers to the training sample set for training the status prediction model. The first training sample set includes multiple first training samples, and each first training sample includes a set of first production status data and the corresponding decision-making post-status data set. It should be particularly noted that the first production status data and the decision-making post-status data set are data contents that have been determined in the historical period and can be directly obtained in this step.

[0046] Next, the determination method of the decision-making state data set can be refined. In this embodiment, each set of first production state data has a corresponding decision-making state data set. The determination method of the decision-making state data set corresponding to any first production state data includes: for at least one set of first production state data, determining the first production state data as the first input data; inputting the first input data into a pre-trained preset decision model to obtain the historical production scheduling information corresponding to the target historical time period; obtaining the decision-making state data determined after the target production workshop executes the production task based on the historical production scheduling information; updating the decision-making state data as the first input data, and repeating the steps of determining the historical production scheduling information and the decision-making state data until the number of repetitions reaches the preset iteration number, then stopping the repetition; determining the decision-making state data determined by each iterative operation as the decision-making state data set.

[0047] In this embodiment, a certain first production state data can be used as the first input data; the first input data is input into a pre-trained preset decision model (here, the preset decision model refers to a pre-trained reinforcement model, not the production decision model to be processed and the target production decision model in this embodiment), and the historical production scheduling information corresponding to the target historical time period is obtained. In the historical period, the target production workshop can execute the production task according to the historical production scheduling information. Subsequently, the production state data of the target production workshop changes, and the production state data at this time is called the decision-making state data. Further, the decision-making state data can be updated as the first input data, and the steps of "inputting the first input data into the pre-trained preset decision model to obtain the historical production scheduling information corresponding to the target historical time period; obtaining the decision-making state data determined after the target production workshop executes the production task based on the historical production scheduling information" are repeated. When it is detected that the current number of repetitions reaches the preset iteration number, the repetition is stopped. Thus, the decision-making state data determined by each iterative operation can be determined as the decision-making state data set.

[0048] S220. Determine the first first training sample in the first training sample set as the current first training sample, and adjust the parameters of the state prediction model to be trained based on the current first training sample to obtain the current state prediction model.

[0049] In this embodiment, the state prediction model to be trained can be trained in a cyclic iteration manner according to each first training sample in the first training sample set, and the parameters of the state prediction model to be trained can be adjusted.

[0050] Each time the training is completed, a current state prediction model can be obtained.

[0051] Specifically, the training process for each first training sample is the same. Here, any one of the first training samples is taken as the current first training sample, and an exemplary description is given by taking the current first training sample as an example. Optionally, the specific implementation manner of adjusting the parameters of the to-be-trained state prediction model based on the current first training sample to obtain the current state prediction model may include:

[0052] S2201. Determine a predicted production state data set based on the current first production state data, the to-be-processed production decision model, and the to-be-trained state prediction model.

[0053] The predicted production state data set refers to the data set determined by the iterative processing of the to-be-processed production decision model and the to-be-trained state prediction model on the current first production state data.

[0054] In this embodiment, the specific implementation manner of determining the predicted production state data set based on the current first production state data, the to-be-processed production decision model, and the to-be-trained state prediction model may include: determining the first production state data in the current training sample as the current input data; inputting the current input data into the to-be-processed production decision model to obtain the first production scheduling information; inputting the first production scheduling information and the current input data into the to-be-trained state prediction model to obtain the predicted production state data; updating the predicted production state data as the current input data, and repeating the steps of determining the first production scheduling information and the predicted production state data until the number of repeated executions reaches the preset number of iterations, and then stopping the repetition; determining the predicted production state data determined by each iterative operation as the predicted production state data set.

[0055] In this embodiment, the current first production state data may be determined as the current input data; first, the current input data is input into the to-be-processed production decision model, and the to-be-processed production decision model processes and analyzes the current input data to obtain the first production scheduling information. Subsequently, the first production scheduling information and the current input data are input into the to-be-trained state prediction model, and the to-be-trained state prediction model can predict the production state data. Further, the predicted production state data may be updated as the current input data, and the steps of "inputting the current input data into the to-be-processed production decision model to obtain the first production scheduling information; inputting the first production scheduling information and the current input data into the to-be-trained state prediction model to obtain the predicted production state data" are repeated. When it is detected that the current number of repetitions reaches the preset number of iterations, the repetition is stopped. Thus, the predicted production state data determined by each iterative operation can be determined as the predicted production state data set.

[0056] S2202. Determine a first loss value based on the predicted production state data set and the decision-making posterior state data set in the current first training sample.

[0057] Among them, the first loss value refers to the loss value used to adjust the parameters of the state prediction model to be trained.

[0058] In this embodiment, the specific implementation manner of determining the first loss value based on the predicted production state data set and the decision-making state data set in the current first training sample may include: determining the target similarity between the predicted production state data set and the decision-making state data set, so as to determine the first loss value based on the target similarity.

[0059] Among them, the target similarity is used to characterize the similarity degree between the predicted production state data set and the decision-making state data set. The larger the value of the target similarity, the closer the predicted production state data set and the decision-making state data set are; the smaller the value of the target similarity, the less similar the predicted production state data set and the decision-making state data set are.

[0060] Optionally, the specific implementation manner of determining the target similarity between the predicted production state data set and the decision-making state data set may include:

[0061] S1. Based on at least one predicted raw material inventory in the predicted production state data set, construct a first raw material inventory sequence, based on at least one actual raw material inventory in the decision-making state data set, construct a second raw material inventory sequence, and perform Pearson correlation calculation on the first raw material inventory sequence and the second raw material inventory sequence to determine the raw material inventory similarity value.

[0062] In this embodiment, the predicted production state data set includes multiple predicted raw material inventories. These predicted raw material inventories can be constructed into a first raw material inventory sequence, and a second raw material inventory sequence can be constructed according to the multiple actual raw material inventories in the decision-making state data set. Thus, Pearson correlation calculation can be performed on the first raw material inventory sequence and the second raw material inventory sequence to obtain the raw material inventory similarity value.

[0063] Among them, the Pearson correlation coefficient is calculated using the following mathematical expression:

[0064]

[0065] In the formula, x i and y i respectively represent the i-th element of the first raw material inventory sequence and the second raw material inventory sequence, respectively represent the mean value of multiple actual raw material inventories and the mean value of multiple predicted raw material inventories.

[0066] S2. Based on at least one predicted production plan completion rate in the predicted production status dataset, construct a first production plan completion rate sequence. Based on at least one actual production plan completion rate in the decision - made status dataset, construct a second production plan completion rate sequence, and perform Pearson correlation calculation on the first production plan completion rate sequence and the second production plan completion rate sequence to determine the production plan completion rate similarity value.

[0067] In this embodiment, the predicted production status dataset includes multiple predicted production plan completion rates. These predicted production plan completion rates can be constructed into a first production plan completion rate sequence, and a second production plan completion rate sequence can be constructed according to the multiple actual production plan completion rates in the decision - made status dataset. Thus, Pearson correlation calculation can be performed on the first production plan completion rate sequence and the second production plan completion rate sequence to obtain the production plan completion rate similarity value.

[0068] S3. Based on at least one predicted equipment utilization rate in the predicted production status dataset, construct a first equipment utilization rate sequence. Based on at least one actual equipment utilization rate in the decision - made status dataset, construct a second equipment utilization rate sequence, and perform cosine similarity calculation on the first equipment utilization rate sequence and the second equipment utilization rate sequence to determine the equipment utilization rate similarity value.

[0069] In this embodiment, the predicted production status dataset includes multiple predicted equipment utilization rates. These predicted equipment utilization rates can be constructed into a first equipment utilization rate sequence, and a second equipment utilization rate sequence can be constructed according to the multiple actual equipment utilization rates in the decision - made status dataset. Thus, cosine similarity calculation can be performed on the first equipment utilization rate sequence and the second equipment utilization rate sequence to obtain the equipment utilization rate similarity value.

[0070] Among them, the cosine similarity is calculated using the following mathematical expression:

[0071]

[0072] In the formula, x i and y i respectively represent the i - th element of the first equipment utilization rate sequence and the second equipment utilization rate sequence.

[0073] S4. Based on at least one predicted order delivery timeliness rate in the predicted production status dataset, construct a first order delivery timeliness rate sequence. Based on at least one actual order delivery timeliness rate in the decision - made status dataset, construct a second order delivery timeliness rate sequence, and perform cosine similarity calculation on the first order delivery timeliness rate sequence and the second order delivery timeliness rate sequence to determine the order delivery timeliness rate similarity value.

[0074] In this embodiment, the predicted production status dataset includes multiple predicted order delivery timeliness rates. These predicted order delivery timeliness rates can be constructed into a first order delivery timeliness rate sequence, and a second order delivery timeliness rate sequence can be constructed based on the multiple actual order delivery timeliness rates in the decision - made status dataset. Thus, the cosine similarity between the first order delivery timeliness rate sequence and the second order delivery timeliness rate sequence can be calculated to obtain the order delivery timeliness rate similarity value.

[0075] In this embodiment, the Pearson correlation coefficient can measure the linear correlation between two continuous variables. Therefore, for the two continuous numerical indicators of raw material inventory and production plan completion rate, the Pearson correlation coefficient is used for measurement. The cosine similarity can measure the cosine value of the angle between two vectors and can well characterize the similarity characteristics of proportional indicators. Therefore, for the two proportional indicators of equipment utilization rate and order delivery timeliness rate, the cosine similarity is used for measurement.

[0076] S5. Determine the target similarity between the predicted production status dataset and the decision - made status dataset based on the raw material inventory similarity value, production plan completion rate similarity value, equipment utilization rate similarity value, and order delivery timeliness rate similarity value.

[0077] Specifically, the target similarity between the predicted production status dataset and the decision - made status dataset can be determined based on the weighted sum of the raw material inventory similarity value, production plan completion rate similarity value, equipment utilization rate similarity value, and order delivery timeliness rate similarity value.

[0078] In this embodiment, the weighted average similarity can comprehensively consider the similarity contributions of multiple indicators. By setting different weight coefficients, the importance of each indicator in the overall similarity evaluation can be flexibly adjusted. Based on this, the target similarity can be expressed as:

[0079] S = w 1 ·r 1 + w 2 ·r 2 + w 3 ·r 3 + r 4 ·cos 2

[0080] In the formula, r 1 and r 2 respectively represent the raw material inventory similarity value and the production plan completion rate similarity value, r 3 and r 4 respectively represent the equipment utilization rate similarity value and the order delivery timeliness rate similarity value, w 1 、w 2 、w 3 and w4 respectively represent the weights of the corresponding status data.

[0081] Correspondingly, the specific implementation manner of determining the first loss value based on the target similarity includes: determining the first loss value based on the target similarity and the first preset loss function; where the first preset loss function is:

[0082] L = 1 - exp(-α(1 - r) 2 )

[0083] In the formula, α represents a hyperparameter that controls the shape of the loss function, and r represents the target similarity.

[0084] In addition, the first loss value can also be determined based on the raw material inventory similarity value, the production plan completion rate similarity value, the equipment utilization rate similarity value, the order delivery timeliness rate similarity value, and the second preset loss function.

[0085] Among them, the second preset loss function is:

[0086] L' = v 1 ×[1 - exp(-α 1 (1 - r 1 ) 2 )] +

[0087] v 2 ×[1 - exp(-α 2 (1 - r 2 ) 2 )] +

[0088] v 3 ×[1 - exp(-α 3 (1 - r 3 ) 2 )] +

[0089] v 4 ×[1 - exp(-α 4 (1 - r 4 ) 2 )];

[0090] In the formula, v 1 , v 2 , v 3 and v 4 represent preset weight values, α 1 , α 2 , α 3 and α 4 represent hyperparameters that control the shape of the loss function, r 1 represents the raw material inventory similarity value, r 2 represents the production plan completion rate similarity value, r 3Represents the similarity value of equipment utilization rate, r 4 Represents the similarity value of on-time order delivery rate.

[0091] S2203. Adjust the parameters of the state prediction model to be trained based on the first loss value to obtain the current state prediction model.

[0092] In this embodiment, after obtaining the first loss value, the parameters of the state prediction model to be trained can be adjusted based on the first loss value to obtain the current state prediction model.

[0093] S230. Adjust the parameters of the production decision model to be processed based on the first production state data in the current first training sample, the first simulated scheduling information corresponding to the first production state data, and the second production state data determined by the current state prediction model to obtain the current production decision model.

[0094] Specifically, the specific implementation method for determining the current production decision model may include the following steps: input the first production state data into the production decision model to be processed to obtain the first simulated scheduling information; input the first production state data and the first simulated scheduling information into the current state prediction model to obtain the second production state data; based on the second production state data, determine the quantization index values of the production decision model to be processed under at least one evaluation dimension; based on the quantization index values under at least one evaluation dimension and a preset reward function, determine the initial reward value; perform scaling processing and normalization processing on the initial reward value to obtain the target reward value; adjust the parameters of the production decision model to be processed based on the target reward value to obtain the current production decision model.

[0095] Among them, at least one evaluation dimension includes the production efficiency evaluation dimension, the production cost evaluation dimension, the on-time order delivery rate evaluation dimension, and the inventory level evaluation dimension.

[0096] In this embodiment, the first production state data can be input into the production decision model to be processed, and the production decision model to be processed can output the first simulated scheduling information; furthermore, the first simulated scheduling information and the first production state data are input into the state prediction model that has completed preliminary training to obtain the second production state data. A preset reward function can be introduced as an evaluation and feedback mechanism for the production decision model to be processed. The preset reward function is used to measure the quality of the scheduling scheme generated by the decision model, and the calculation method of the reward value is designed from different optimization objectives and perspectives. This reward mechanism can effectively guide the decision model to learn and improve in a more optimized direction. The preset reward function is constructed by the following method:

[0097] Determine that the input variables of the preset reward function are the quantization index values of the production efficiency evaluation dimension, the production cost evaluation dimension, the on-time order delivery rate evaluation dimension, and the inventory level evaluation dimension;

[0098] Define a quantization index for each input variable, which is used to map the input variable into the interval [0, 1];

[0099] Use the weighted average method to construct the original reward function. The mathematical expression of the original reward function is:

[0100] R o = w E ·E + w C ·C + w T ·T + w I ·I

[0101] In the formula, w E , w C , w T and w I respectively represent the weights of production efficiency, production cost, delivery time, and inventory level. E, C, T, and I respectively represent the quantization indexes of production efficiency, production cost, delivery time, and inventory level; in this embodiment, the weights of production efficiency, production cost, delivery time, and inventory level can be determined by combining expert experience, data analysis, and sensitivity analysis.

[0102] After constructing the original reward function, the method for constructing the reward function further includes:

[0103] Perform scaling processing and normalization processing on the reward function. The scaling processing uses the following mathematical expression:

[0104]

[0105] In the formula, R o represents the original reward function, R min and R max respectively represent the minimum value and the maximum value of the original reward function. Among them, the determination methods of R min and R max can first obtain the preliminary range through theoretical analysis, and then adjust it based on historical data and simulation experiments, or be given by algorithm experts based on experience.

[0106] After the scaling processing, the value range of the reward function is adjusted to the sensitive area of the Sigmoid function used for the normalization processing, improving the effect of the normalization processing.

[0107] The normalization processing uses the following mathematical expression:

[0108]

[0109] In the formula, R s represents the reward function after the scaling processing, R nRepresents the normalized reward function.

[0110] In this embodiment, the reward function is normalized by the Sigmoid function, effectively improving the stability, discrimination ability, and interpretability of the reward function. First, the value range of the normalized reward function is restricted to the interval [0, 1], improving the function stability. Second, the non-linear characteristic of the Sigmoid function can enhance the discrimination ability of the reward function for different decision-making behaviors. When the original reward value is small, the change of the normalized reward value is relatively gentle, and when the original reward value is large, the change of the normalized reward value is relatively steep. In addition, the normalized reward value can be interpreted as the quality level of the decision-making behavior. When the R n value is close to 1, it indicates that the decision-making behavior is extremely excellent, while when the R n value is close to 0, it indicates that the decision-making behavior is extremely poor.

[0111] After determining the target reward value based on the above method, the parameter of the production decision model to be processed can be adjusted by the target reward value to obtain the current production decision model.

[0112] S240. Update the next first training sample of the current first training sample to the current first training sample, update the current state prediction model to the state prediction model to be trained, update the current production decision model to the production decision model to be processed, and repeat the steps of determining the current state prediction model and the current production decision model.

[0113] S250. When the preset iteration end condition is met, determine the obtained current production decision model as the target production decision model.

[0114] In this embodiment, for the state prediction model, the preset iteration end condition can be that the prediction accuracy meets the accuracy threshold (MAE or RMSE is lower than a specific threshold); for the production decision model, the preset iteration end condition can be set through decision quality (percentage value of production efficiency improvement or percentage value of cost reduction), decision response time, constraint satisfaction, robustness, operation stability, etc. When it is detected that the preset iteration end condition is met, the obtained current production decision model can be determined as the target production decision model at this time.

[0115] Optionally, in order to introduce an exploration mechanism in the training process of the production decision model, enable the production decision model to try new decisions to a certain extent and discover potentially optimal strategies, use the preset reward function to evaluate the decision-making behavior of the production decision model and optimize it, specifically including:

[0116] When generating decision-making behaviors for the production decision-making model, an exploration mechanism is introduced to randomly select an action with an exploration probability ε; a preset reward function is used to evaluate the decision-making behaviors generated by the production decision-making model, and the immediate reward corresponding to the selected decision-making action is calculated; based on the immediate reward, a reinforcement learning algorithm is used to optimize the production decision-making model and update the parameters of the production decision-making model.

[0117] Among them, randomly selecting an action with an exploration probability ε specifically includes: generating a random number r between 0 and 1; if r < ε, randomly select an action from the action space as the decision result; if r ≥ ε, use the optimal action given by the production decision-making model in the next round as the decision result.

[0118] Specifically, the initial value ε 0 and the decay rate α of the exploration probability ε can be set, and the exploration probability is updated according to the mathematical expression ε = ε 0 ·α t in each iteration, where t represents the iteration number of the current step. In this way, in continuous training iterations, as the performance of the production decision-making model improves, the value of the exploration probability is gradually reduced, so that the production decision-making model gradually transitions from exploration to using the optimal strategy.

[0119] S260. Obtain the production status data of the target production workshop; among them, the production status data includes the raw material inventory, the equipment utilization rate of each production line, the order quantity of at least one item category, the order production plan information, and the order delivery time.

[0120] S270. Input the production status data into the pre-trained target production decision-making model to obtain the production scheduling information of the target production workshop within a preset future time period.

[0121] Among them, the production scheduling information includes at least one of the raw material procurement plan information, the shift scheduling plan information of each production line, and the raw material allocation information.

[0122] The training method of this embodiment introduces a state prediction model, which can model the non-Markov property of production state data and predict the long-term impact of scheduling decisions, helping the production decision-making model to weigh short-term benefits and long-term interests and make a more globally optimal scheduling decision (since the state data set generated by the state prediction model includes both the immediate changes in the system state after the decision and reflects the impact of the current decision on multiple future time points. In this way, the cumulative effect of the decision can be evaluated, and the production decision-making model can learn the short-term and long-term consequences of the decision through interaction with the state prediction model). Moreover, the state prediction model can, to a certain extent, replace the actual production environment, provide a simulated future state for the production decision-making model, and reduce the exploration requirements of the production decision-making model in actual production. This method trains the production decision-making model through an iterative optimization method, introduces a closed-loop feedback mechanism throughout the training process, and optimizes the state prediction model and the production decision-making model by continuously comparing the differences between the predicted state and the actual state.

[0123] Generally speaking, the training method of the production decision-making model in this embodiment can improve the generalization ability and robustness of the production decision-making model, accelerate the training convergence speed of the production decision-making model, and enhance the exploration ability and optimization efficiency of the production decision-making model. Compared with the existing training methods, the training method of this embodiment can dynamically generate simulated state data by introducing a state prediction model, and train the production decision-making model based on these data, enabling the production decision-making model to access a more diverse state space and learn more generalized decision-making strategies. Moreover, due to the introduction of the state prediction model, it is possible to predict and simulate future states in advance, providing more abundant training data for the production decision-making model. Through the feedback optimization of similarity evaluation and reward functions, the training direction of the production decision-making model can be more efficiently guided, thus accelerating the training convergence speed. Taking the production scheduling production decision-making model of a tobacco production enterprise as an example, the production decision-making model trained by this training method can more accurately predict various changes and risks in the production process by continuously optimizing the state prediction model, and accordingly adjust the production decision-making model, enabling the production decision-making model to make an optimal scheduling decision according to the real-time production state, thereby improving the intelligent level of production scheduling.

[0124] In the technical solution of this embodiment, during the process of training the target production decision model, a first training sample set is constructed based on at least one set of first production status data of the target production workshop and a decision-making post-status data set corresponding to the at least one set of first production status data. Subsequently, the first first training sample in the first training sample set is determined as the current first training sample, and the parameter of the to-be-trained status prediction model is adjusted based on the current first training sample to obtain the current status prediction model. Furthermore, based on the first production status data in the current first training sample, the first simulation scheduling information corresponding to the first production status data, and the second production status data determined by the current status prediction model, the parameter of the to-be-processed production decision model is adjusted to obtain the current production decision model. Further, the next first training sample of the current first training sample is updated as the current first training sample, the current status prediction model is updated as the to-be-trained status prediction model, and the current production decision model is updated as the to-be-processed production decision model, and the steps of determining the current status prediction model and the current production decision model are repeatedly executed. Finally, when the preset iteration end condition is satisfied, the obtained current production decision model is determined as the target production decision model. By introducing a status prediction model, the training method of this embodiment can dynamically generate simulated status data, and train the production decision model based on these data, enabling the production decision model to access a more diverse state space and learn more generalized decision-making strategies. Moreover, due to the introduction of the status prediction model, the future status can be predicted and simulated in advance, providing richer training data for the decision model. In addition, through the feedback optimization of similarity evaluation and reward function, the training direction of the decision model can be guided more efficiently, thereby accelerating the training convergence speed.

[0125] Embodiment III

[0126] Figure 3 is a schematic structural diagram of a device for determining workshop production scheduling information provided by an embodiment of the present invention, as Figure 3 shown, the device includes: a status data acquisition module 310 and a scheduling information determination module 320.

[0127] Among them, the status data acquisition module 310 is configured to acquire the production status data of the target production workshop; wherein, the production status data includes at least one of the raw material inventory, the equipment utilization rate of each production line, the order quantity of at least one item category, the order production plan information, and the order delivery time;

[0128] A scheduling information determination module 320, configured to input the production status data into a pre-trained target production decision model to obtain production scheduling information of the target production workshop within a preset time period in the future; wherein, the target production decision model is obtained by training with a plurality of training samples, and the training samples include first production status data, first simulated scheduling information corresponding to the first production status data, and second production status data. The first simulated scheduling information is obtained by processing the first production status data based on a production decision model to be processed, and the second production status data is production status data deduced based on a status prediction model for the first production status data and the first simulated scheduling information. The first production status data is related to the deduced second production status data, and the production scheduling information includes at least one of raw material procurement plan information, shift plan information for each production line, and raw material allocation information.

[0129] Based on the above embodiments, optionally, the workshop production scheduling information determination device further includes: a model training module; the model training module includes:

[0130] A sample set determination unit, configured to construct a first training sample set based on at least one set of first production status data of the target production workshop and a decision-making post-state data set corresponding to the at least one set of first production status data;

[0131] A current prediction model determination unit, configured to determine the first training sample at the head of the first training sample set as the current first training sample, and adjust the parameters of the production status prediction model to be trained based on the current first training sample to obtain the current status prediction model;

[0132] A current decision model determination unit, configured to adjust the parameters of the production decision model to be processed based on the first production status data in the current first training sample, the first simulated scheduling information corresponding to the first production status data, and the second production status data determined by the current status prediction model to obtain the current production decision model;

[0133] A loop iteration processing unit, configured to update the next first training sample of the current first training sample as the current first training sample, update the current status prediction model as the production status prediction model to be trained, update the current production decision model as the production decision model to be processed, and repeat the steps of determining the current status prediction model and the current production decision model;

[0134] A target model determination unit, configured to determine the obtained current production decision model as the target production decision model when a preset iteration end condition is satisfied.

[0135] Based on the above embodiments, optionally, the sample set determination unit further includes: a post - decision state data set determination subunit; the post - decision state data set determination subunit is specifically configured to, for the at least one set of first production state data, determine the first production state data as the first input data; input the first input data into a pre - trained preset decision model to obtain the historical production scheduling information corresponding to the target historical time period; obtain the post - decision state data determined after the target production workshop executes the production task based on the historical production scheduling information; update the post - decision state data as the first input data, and repeat the steps of determining the historical production scheduling information and the post - decision state data until the number of repetitions reaches a preset number of iterations, then stop repeating; determine the post - decision state data determined in each iteration operation as the post - decision state data set.

[0136] Based on the above embodiments, optionally, the current prediction model determination unit includes:

[0137] A prediction data set determination subunit, configured to determine a predicted production state data set based on the current first production state data, the production decision model to be processed, and the state prediction model to be trained;

[0138] A loss value determination subunit, configured to determine a first loss value based on the predicted production state data set and the post - decision state data set in the current first training sample;

[0139] A prediction model determination subunit, configured to adjust the parameters of the state prediction model to be trained based on the first loss value to obtain the current state prediction model.

[0140] Based on the above embodiments, optionally, the prediction data set determination subunit is specifically configured to determine the first production state data in the current training sample as the current input data; input the current input data into the production decision model to be processed to obtain the first production scheduling information; input the first production scheduling information and the current input data into the state prediction model to be trained to obtain the predicted production state data;

[0141] update the predicted production state data as the current input data, and repeat the steps of determining the first production scheduling information and the predicted production state data until the number of repetitions reaches a preset number of iterations, then stop repeating; determine the predicted production state data determined in each iteration operation as the predicted production state data set.

[0142] Based on the above embodiments, optionally, the loss value determination subunit is specifically configured to determine the target similarity between the predicted production state data set and the post - decision state data set, and determine the first loss value based on the target similarity.

[0143] Based on the above embodiments, optionally, the loss value determination subunit is specifically further configured to construct a first raw material inventory quantity sequence based on at least one predicted raw material inventory quantity in the predicted production status dataset, construct a second raw material inventory quantity sequence based on at least one actual raw material inventory quantity in the decision-making post-status dataset, and perform Pearson correlation calculation on the first raw material inventory quantity sequence and the second raw material inventory quantity sequence to determine the raw material inventory similarity value; construct a first production plan completion rate sequence based on at least one predicted production plan completion rate in the predicted production status dataset, construct a second production plan completion rate sequence based on at least one actual production plan completion rate in the decision-making post-status dataset, and perform Pearson correlation calculation on the first production plan completion rate sequence and the second production plan completion rate sequence to determine the production plan completion rate similarity value; construct a first equipment utilization rate sequence based on at least one predicted equipment utilization rate in the predicted production status dataset, construct a second equipment utilization rate sequence based on at least one actual equipment utilization rate in the decision-making post-status dataset, and perform cosine similarity calculation on the first equipment utilization rate sequence and the second equipment utilization rate sequence to determine the equipment utilization rate similarity value; construct a first order delivery timeliness rate sequence based on at least one predicted order delivery timeliness rate in the predicted production status dataset, construct a second order delivery timeliness rate sequence based on at least one actual order delivery timeliness rate in the decision-making post-status dataset, and perform cosine similarity calculation on the first order delivery timeliness rate sequence and the second order delivery timeliness rate sequence to determine the order delivery timeliness rate similarity value; determine the target similarity between the predicted production status dataset and the decision-making post-status dataset based on the raw material inventory similarity value, the production plan completion rate similarity value, the equipment utilization rate similarity value, and the order delivery timeliness rate similarity value.

[0144] Based on the above embodiments, optionally, the loss value determination subunit is specifically further configured to determine the target similarity between the predicted production status dataset and the decision-making post-status dataset based on the weighted sum of the raw material inventory similarity value, the production plan completion rate similarity value, the equipment utilization rate similarity value, and the order delivery timeliness rate similarity value;

[0145] Correspondingly, the determining the first loss value based on the target similarity includes:

[0146] Determining a first loss value based on the target similarity and a first preset loss function; wherein, the first preset loss function is:

[0147] L = 1 - exp(-α(1 - r) 2 );

[0148] In the formula, α represents a hyperparameter for controlling the shape of the loss function, and r represents the target similarity;

[0149] Or; determining a first loss value based on the raw material inventory similarity value, the production plan completion rate similarity value, the equipment utilization rate similarity value, the order delivery timeliness rate similarity value, and a second preset loss function; wherein, the second preset loss function is:

[0150] L' = v 1 ×[1 - exp(-α 1 (1 - r 1 ) 2 )] +

[0151] v 2 ×[1 - exp(-α 2 (1 - r 2 ) 2 )] +

[0152] v 3 ×[1 - exp(-α 3 (1 - r 3 ) 2 )] +

[0153] v 4 ×[1 - exp(-α 4 (1 - r 4 ) 2 )];

[0154] In the formula, v 1 , v 2 , v 3 and v 4 represent preset weight values, α 1 , α 2 , α 3 and α 4 represent hyperparameters for controlling the shape of the loss function, r 1 represents the raw material inventory similarity value, r 2 represents the production plan completion rate similarity value, r 3 represents the equipment utilization rate similarity value, r 4 represents the order delivery timeliness rate similarity value.

[0155] Based on the above - mentioned embodiments, optionally, the current decision - making model determination unit includes:

[0156] A simulation scheduling information determination subunit, configured to input the first production status data into the to - be - processed production decision model to obtain first simulation scheduling information;

[0157] A second state data determination subunit, configured to input the first production state data and the first simulated scheduling information into the current state prediction model to obtain second production state data;

[0158] An evaluation quantization value determination subunit, configured to determine quantization index values of the to-be-processed production decision model under at least one evaluation dimension based on the second production state data; wherein, the at least one evaluation dimension includes a production efficiency evaluation dimension, a production cost evaluation dimension, an order delivery timeliness rate evaluation dimension, and an inventory level evaluation dimension;

[0159] An initial reward value determination subunit, configured to determine an initial reward value based on the quantization index values under the at least one evaluation dimension and a preset reward function;

[0160] A target reward value determination subunit, configured to perform scaling processing and normalization processing on the initial reward value to obtain a target reward value;

[0161] A decision model determination subunit, configured to perform parameter adjustment on the to-be-processed production decision model based on the target reward value to obtain a current production decision model.

[0162] The technical solution of the embodiment of the present invention, by obtaining the production state data of the target production workshop, wherein the production state data includes at least one of raw material inventory, equipment utilization rate of each production line, order quantity of at least one item category, order production plan information, and order delivery time; inputting the production state data into a pre-trained target production decision model to obtain production scheduling information of the target production workshop within a preset future time period, wherein the target production decision model is obtained by training with a plurality of training samples, and the training samples include first production state data, first simulated scheduling information corresponding to the first production state data, and second production state data, the first simulated scheduling information is obtained by processing the first production state data based on the to-be-processed production decision model, the second production state data is production state data deduced based on the state prediction model for the first production state data and the first simulated scheduling information, the first production state data is related to the deduced second production state data, and the production scheduling information includes at least one of raw material procurement plan information, shift scheduling plan information of each production line, and raw material distribution information. The technical solution provided by the embodiment of the present invention can automatically determine production scheduling information according to the production state data of the production workshop through the trained target production decision model, and can obtain a production scheduling strategy with high reliability and high availability, thereby improving the intelligent level of production scheduling.

[0163] The workshop production scheduling information determination device provided by the embodiment of the present invention can execute the workshop production scheduling information determination method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0164] It should be noted that the various units and modules included in the above device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present invention.

[0165] Embodiment 4

[0166] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Figure 4 It shows a block diagram of an exemplary electronic device 40 suitable for implementing the implementation manner of the embodiment of the present invention. Figure 4 The displayed electronic device 40 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.

[0167] As Figure 4 shown, the electronic device 40 is presented in the form of a general-purpose computing device. The components of the electronic device 40 may include but are not limited to: one or more processors or processing units 401, a system memory 402, and a bus 403 connecting different system components (including the system memory 402 and the processing unit 401).

[0168] The bus 403 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any bus structure in a variety of bus structures. For example, these architectures include but are not limited to Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0169] The electronic device 40 typically includes a variety of computer system-readable media. These media can be any available media accessible by the electronic device 40, including volatile and non-volatile media, removable and non-removable media.

[0170] The system memory 402 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 404 and / or cache memory 405. The electronic device 40 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 406 can be used to read and write non-removable, non-volatile magnetic media ( Figure 4 not shown, commonly referred to as a "hard disk drive"). Although Figure 4Not shown in the figure, a disk drive for reading and writing a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM or other optical medium) can be provided. In these cases, each drive can be connected to the bus 403 through one or more data medium interfaces. The memory 402 can include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present invention.

[0171] A program / utility 408 having a set (at least one) of program modules 407 can be stored in, for example, the memory 402. Such program modules 407 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 407 generally perform the functions and / or methods in the embodiments described in the present invention.

[0172] The electronic device 40 can also communicate with one or more external devices 409 (such as a keyboard, a pointing device, a display 410, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 40, and / or communicate with any device that enables the electronic device 40 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 411. And, the electronic device 40 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 412. As shown in the figure, the network adapter 412 communicates with other modules of the electronic device 40 through the bus 403. It should be understood that although Figure 4 not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 40, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0173] The processing unit 401 executes various functional applications and page processing by running the programs stored in the system memory 402, such as implementing the workshop production scheduling information determination method provided by the embodiments of the present invention.

[0174] Embodiment 5

[0175] The embodiments of the present invention also provide a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the workshop production scheduling information determination method when executed by a computer processor. The method includes:

[0176] Obtain the production status data of the target production workshop; wherein, the production status data includes at least one of raw material inventory, equipment utilization rate of each production line, order quantity of at least one item category, order production plan information, and order delivery time.

[0177] Input the production status data into a pre-trained target production decision model to obtain the production scheduling information of the target production workshop within a preset time period in the future; wherein, the target production decision model is obtained by training with multiple training samples, and the training samples include first production status data, first simulated scheduling information corresponding to the first production status data, and second production status data. The first simulated scheduling information is obtained by processing the first production status data based on a production decision model to be processed, and the second production status data is the production status data deduced from the first production status data and the first simulated scheduling information based on a status prediction model. The first production status data is related to the deduced second production status data. The production scheduling information includes at least one of raw material procurement plan information, shift scheduling plan information of each production line, and raw material distribution information.

[0178] The computer storage medium of the embodiments of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, apparatus, or device.

[0179] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device.

[0180] The program code contained on a computer-readable medium can be transmitted using any suitable medium, including - but not limited to - wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0181] The computer program code for performing the operations of the embodiments of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages - such as Java, Smalltalk, C++; and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0182] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments here, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A method for determining workshop production scheduling information, characterized in that: include: Acquire production status data of the target production workshop; wherein the production status data includes at least one of raw material inventory, equipment utilization rate of each production line, order quantity of at least one item category, order production plan information, and order delivery time; The production status data is input into a pre-trained target production decision model to obtain production scheduling information of the target production workshop within a preset time period in the future; wherein the target production decision model is trained by multiple training samples, and the training samples include first production status data, first simulated scheduling information corresponding to the first production status data, and second production status data, the first simulated scheduling information is obtained by processing the first production status data based on the production decision model to be processed, the second production status data is production status data deduced from the first production status data and the first simulated scheduling information based on a status prediction model, the first production status data and the deduced second production status data are related, and the production scheduling information includes: raw material procurement plan information, shift scheduling information for each of the production lines, and at least one of raw material allocation information.

2. The method according to claim 1, characterized in that: The method further comprises: Constructing a first training sample set based on at least one set of first production status data of the target production workshop and a post-decision status data set corresponding to the at least one set of first production status data; Determine the first first training sample in the first training sample set as the current first training sample, and adjust the parameters of the state prediction model to be trained based on the current first training sample to obtain the current state prediction model; Adjusting parameters of a production decision model to be processed based on the first production status data in the current first training sample, the first simulation scheduling information corresponding to the first production status data, and the second production status data determined by the current status prediction model to obtain a current production decision model; Updating the next first training sample of the current first training sample to the current first training sample, updating the current state prediction model to the state prediction model to be trained, updating the current production decision model to the production decision model to be processed, and repeating the steps of determining the current state prediction model and the current production decision model; When the preset iteration end condition is met, the obtained current production decision model is determined as the target production decision model.

3. The method according to claim 2, characterized in that The method for determining the post-decision status data set corresponding to the at least one set of first production status data includes: For the at least one set of first production status data, determining the first production status data as first input data; Inputting the first input data into a pre-set decision model that has been trained to obtain historical production scheduling information corresponding to a target historical time period; Acquire post-decision state data determined by the target production workshop after executing the production task based on the historical production scheduling information; The post-decision state data is updated to the first input data, and the steps of determining the historical production scheduling information and the post-decision state data are repeatedly executed until the number of repeated executions reaches a preset number of iterations, and then the repetition is stopped; The post-decision state data determined by each iterative operation is determined as a post-decision state data set.

4. The method according to claim 2, characterized in that: The step of adjusting the parameters of the state prediction model to be trained based on the current first training sample to obtain the current state prediction model includes: Determine a predicted production status data set based on the current first production status data, the production decision model to be processed, and the status prediction model to be trained; Determine a first loss value based on the predicted production status data set and the post-decision status data set in the current first training sample; Based on the first loss value, parameters of the state prediction model to be trained are adjusted to obtain a current state prediction model.

5. The method according to claim 4, characterized in that The step of determining a predicted production status data set based on the current first production status data, the production decision model to be processed, and the status prediction model to be trained includes: Determine the first production status data in the current training sample as the current input data; Inputting the current input data into the production decision model to be processed to obtain first production scheduling information; Inputting the first production scheduling information and the current input data into the state prediction model to be trained to obtain predicted production state data; The predicted production status data is updated to the current input data, and the steps of determining the first production scheduling information and predicting the production status data are repeatedly executed until the number of repeated executions reaches a preset number of iterations, and then the repetition is stopped; The predicted production status data determined by each iterative operation is determined as a predicted production status data set.

6. The method according to claim 4, characterized in that The determining of a first loss value based on the predicted production status data set and the post-decision status data set in the current first training sample comprises: A target similarity between the predicted production status dataset and the post-decision status dataset is determined to determine a first loss value based on the target similarity.

7. The method according to claim 6, characterized in that The determining of the target similarity between the predicted production status dataset and the post-decision status dataset comprises: Based on at least one predicted raw material inventory in the predicted production status data set, a first raw material inventory sequence is constructed, based on at least one actual raw material inventory in the post-decision status data set, a second raw material inventory sequence is constructed, and a Pearson correlation calculation is performed on the first raw material inventory sequence and the second raw material inventory sequence to determine a raw material inventory similarity value; Based on at least one predicted production plan completion rate in the predicted production status data set, a first production plan completion rate sequence is constructed; based on at least one actual production plan completion rate in the post-decision status data set, a second production plan completion rate sequence is constructed; and a Pearson correlation calculation is performed on the first production plan completion rate sequence and the second production plan completion rate sequence to determine a production plan completion rate similarity value; Based on at least one predicted equipment utilization in the predicted production status data set, a first equipment utilization sequence is constructed; based on at least one actual equipment utilization in the post-decision status data set, a second equipment utilization sequence is constructed; and cosine similarity calculation is performed on the first equipment utilization sequence and the second equipment utilization sequence to determine an equipment utilization similarity value; Based on at least one predicted order delivery timeliness rate in the predicted production status data set, a first order delivery timeliness rate sequence is constructed; based on at least one actual order delivery timeliness rate in the post-decision status data set, a second order delivery timeliness rate sequence is constructed; and cosine similarity calculation is performed on the first order delivery timeliness rate sequence and the second order delivery timeliness rate sequence to determine an order delivery timeliness rate similarity value; Based on the raw material inventory similarity value, the production plan completion rate similarity value, the equipment utilization rate similarity value and the order delivery timeliness rate similarity value, the target similarity between the predicted production status data set and the post-decision status data set is determined.

8. The method according to claim 7, characterized in that The determining of the target similarity between the predicted production status dataset and the post-decision status dataset based on the raw material inventory similarity value, the production plan completion rate similarity value, the equipment utilization rate similarity value, and the order delivery timeliness rate similarity value includes: Determine the target similarity between the predicted production status dataset and the post-decision status dataset based on a weighted sum of the raw material inventory similarity value, the production plan completion rate similarity value, the equipment utilization rate similarity value, and the order delivery timeliness rate similarity value; Accordingly, determining the first loss value based on the target similarity includes: A first loss value is determined based on the target similarity and a first preset loss function; wherein the first preset loss function is: L=1-exp(-α(1-r) 2 ); In the formula, α represents the hyperparameter that controls the shape of the loss function, and r represents the target similarity; Or; based on the raw material inventory similarity value, the production plan completion rate similarity value, the equipment utilization rate similarity value, the order delivery timeliness rate similarity value and a second preset loss function, determine the first loss value; wherein the second preset loss function is: L'=v1×[1-exp(-α1(1-r1) 2 )]+ v2×[1-exp(-α2(1-r2) 2 )]+ v3×[1-exp(-α3(1-r3) 2 )]+ v4×[1-exp(-α4(1-r4) 2 )];; Wherein, v1, v2, v3 and v4 represent preset weight values, α1, α2, α3 and α4 represent hyperparameters that control the shape of the loss function, r1 represents the similarity value of raw material inventory, r2 represents the similarity value of production plan completion rate, r3 represents the similarity value of equipment utilization rate, and r4 represents the similarity value of order delivery timeliness rate.

9. The method according to claim 2, characterized in that: The step of adjusting parameters of a production decision model to be processed based on the first production status data in the current first training sample, the first simulation scheduling information corresponding to the first production status data, and the second production status data determined by the current status prediction model to obtain a current production decision model includes: Inputting the first production status data into the production decision model to be processed to obtain first simulation scheduling information; Inputting the first production status data and the first simulation scheduling information into the current status prediction model to obtain second production status data; Based on the second production status data, determine the quantitative index value of the to-be-processed production decision model under at least one evaluation dimension; wherein the at least one evaluation dimension includes a production efficiency evaluation dimension, a production cost evaluation dimension, an order delivery timeliness evaluation dimension, and an inventory level evaluation dimension; Determining an initial reward value based on the quantitative indicator value under the at least one evaluation dimension and a preset reward function; Scaling and normalizing the initial reward value to obtain a target reward value; The parameters of the production decision model to be processed are adjusted based on the target reward value to obtain a current production decision model.

10. A device for determining workshop production scheduling information, characterized in that: include: A status data acquisition module, used to acquire production status data of a target production workshop; wherein the production status data includes at least one of raw material inventory, equipment utilization rate of each production line, order quantity of at least one item category, order production plan information, and order delivery time; A scheduling information determination module is used to input the production status data into a pre-trained target production decision model to obtain the production scheduling information of the target production workshop within a preset time period in the future; wherein the target production decision model is obtained by training with multiple training samples, and the training samples include first production status data, first simulated scheduling information corresponding to the first production status data, and second production status data, the first simulated scheduling information is obtained by processing the first production status data based on the production decision model to be processed, the second production status data is the production status data deduced from the first production status data and the first simulated scheduling information based on a status prediction model, the first production status data is related to the deduced second production status data, and the production scheduling information includes: at least one of raw material procurement plan information, shift scheduling information for each of the production lines, and raw material allocation information.