AI production operation management method based on large model

By adopting a large model-based AI method in production operation management, integrating multi-source data and abstracting it into Markov decision-making process, the problems of lagging equipment status judgment and unscientific order priority decision-making under traditional management methods are solved, and the optimal allocation of production resources and the improvement of corporate profits are achieved.

CN120069489APending Publication Date: 2025-05-30SHANGHAI KEZHI ELECTRIC AUTOMATION CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510553759.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional production operation management methods are difficult to efficiently process multi-source data such as equipment operation, order management and supply chain, and cannot provide timely and accurately support decisions, resulting in lagging equipment status judgments, unscientific order priority decisions, and unreasonable allocation of production resources, which affects order delivery and corporate profits.

Method used

Using AI production operation management methods based on large models, by building an operation management database, integrating equipment operation data, order management data and supply chain data, abstract production scheduling is a Markov decision-making process, defining state space, action space, transfer probability and reward functions, and using artificial intelligence to generate dynamic scheduling instructions to optimize production resource allocation.

Benefits of technology

Real-time and accurate reflection of the production environment is achieved, and strong support is provided for decision-making. By comprehensively considering multiple factors, optimizing production resource allocation, improving the on-time delivery rate of orders and corporate profits, and enhancing the ability to respond to market changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069489A_ABST
    Figure CN120069489A_ABST
Patent Text Reader

Abstract

The invention discloses an AI production operation management method based on a large model, particularly relates to the field of production operation management, and aims to solve the problems of data isolation, low decision-making efficiency, poor dynamic adaptability and the like in traditional production scheduling. The method comprises the following steps: establishing a multi-source database, abstracting production scheduling into a Markov decision process, and defining a state space containing production operation management 11 production operation management characteristics such as equipment load, order priority, material inventory, safety coefficient and the like; designing five action spaces of process adjustment, batch division, shutdown decision, safe inventory adjustment and resource allocation, calculating a transition probability in combination with historical data, and quantifying an action value through a reward function. According to the scheme, an operation task is decomposed into an equipment management and order allocation double-agent system, an equipment management agent optimizes startup and shutdown and an equipment allocation strategy based on a real-time load and a fault state, and an order allocation agent realizes accurate resource allocation through dynamic priority and safe inventory adjustment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of production and operation management. More specifically, the present invention relates to an AI production and operation management method based on a large model. Background Art

[0002] In the current field of production and operation management, with the increasingly fierce market competition and diverse customer demands, traditional production and operation management methods face numerous challenges. Enterprises need to process a large amount of production data, including equipment operation status, order information, material flow, etc. The complexity and dynamics of this data make it extremely difficult to accurately grasp the production and operation status.

[0003] Traditional methods mainly rely on the independent operation of multiple systems and empirical decision-making models: Internet of Things sensors and device logs are used to collect equipment load and fault data, ERP / MES production and operation management systems manage order priorities and production plans, and WMS production and operation management systems monitor material inventory and flow. Traditional methods handle equipment failures through manual inspections, determine order priorities based on a single dimension such as delivery date or profit, and allocate material inventory according to fixed rules.

[0004] However, in actual use, there are still some drawbacks. For example, traditional management methods are difficult to efficiently process multi-source data such as equipment operation, order management, and supply chain, and cannot provide timely and accurate support for decision-making. For example, when analyzing equipment load data, it is impossible to quickly extract key information from a large amount of time-series data, resulting in a lag in the judgment of equipment status. At the same time, when traditional methods determine order priorities and arrange production sequences, they cannot comprehensively consider factors such as contract profit and remaining delivery time, and often make decisions based on experience, resulting in unreasonable allocation of production resources, affecting order delivery and corporate profits. Finally, there are obstacles to data circulation and collaborative work among different departments in traditional methods, and links such as equipment management, order allocation, and material management are disjointed, unable to form an efficient production and operation system, reducing the enterprise's response speed to market changes. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides an AI production and operation management method based on a large model, through the following solutions to solve the problems raised in the above background art.

[0006] To achieve the above object, the present invention provides the following technical solution: An AI production and operation management method based on a large model, including: S1: Database construction: Construct an operation management database based on equipment operation data sets, order management data sets, and supply chain data sets; S2: Production and operation modeling: Abstract production scheduling as a Markov decision process, and define the state space S, action space A, transition probability P, and reward function R; The state space S is a set of vectors that describe the specific situation of the production system at a certain moment. Suppose there are n devices, m orders, and k types of materials in the production system. The state s ∈ S can be represented as a vector; The action space A is an abstract set of decisions that can be taken during production operations. These decisions will affect the state of the production system. The action space is divided into process adjustment actions, order batch division actions, shutdown decisions, safety inventory adjustment actions, and resource allocation actions; The transition probability P(s′∣s,a) represents the probability of transitioning to state s′ after taking action a (a ∈ A) in state s; The reward function R is used to evaluate the quality of taking action a in state s; S3: Data collection: Obtain operation data from the operation management database based on the state space S, action space A, transition probability P, and reward function R defined in step S2. After preprocessing the obtained operation data, a Markov decision data set is obtained; S4: Operation task decomposition judgment: Construct an operation scheduling task based on the Markov decision data set, and decompose the operation scheduling task into a device management task and an order allocation task. Based on the allocation, it is judged in real time whether the actions in the action space of the device management task and the order allocation task are executed; The device management task includes the state space S eq 、the action space A eq and the agent π eq ; The order allocation task includes the state space S ord 、the action space A ord and the agent π ord ; S5: Intelligent action execution: Based on the judgment result in step S4, generate dynamic scheduling instructions through artificial intelligence to execute the actions in the action space.

[0007] Preferably, the device operation data set includes the unique device number, the device workload time series data, and the current device failure status. The device operation data set is constructed based on the data collected by the Internet of Things sensor devices and the device production logs; The order management data set includes order priority data, delivery date data, material requirement data, and production plan data. The order management data set is constructed based on the ERP system database, the MES system database, and the enterprise production logs; The supply chain data set includes material inventory data, material inbound and outbound record data, and purchase record data. The supply chain data set is constructed based on the WMS system database.

[0008] Preferably, the state s is specifically: s = [l1 , l 2 ,..., l n , p 1 , p 2 ,..., p m , q 1 , q 2 ,..., q k , b 1 , b 2 ,..., b k , t 1 , t 2 ,..., t m , f 1 , f 2 ,..., f n , where: l i (i = 1, 2,..., n) represents the load of the i-th device, expressed by the proportion of the time currently occupied by the device, and the calculation method is: l i = (remaining time of the current task + total time of tasks in the queue) / total available time of the device, and the value range is between [0, 1], where 0 means the device is idle and 1 means the device is running at full load.

[0009] p j (j = 1, 2,..., m) represents the priority of the j-th order, and the calculation method is: p j = 0.6 × contract profit / max(profit of all orders) + 0.4 × remaining delivery time / min(remaining time of all orders); q l (l = 1, 2,..., k) represents the inventory quantity of the l-th material. The inventory quantity affects the continuity of production. If the inventory of a certain material is insufficient, it may lead to production interruption, and the calculation method is: q l = current inventory - allocated but unused quantity + quantity of purchases in transit; b l (l = 1, 2,..., k) represents the safety stock coefficient of the l-th material; t j (j = 1, 2,..., m) represents the remaining delivery time of the j-th order, that is, the remaining time from the current moment to the order delivery deadline; f i (i = 1, 2,..., n) represents the failure status of the i-th device, represented by a binary value, where 0 means the device is running normally and 1 means the device has a failure.

[0010] Preferably, the process adjustment action refers to adjusting the process sequence of a certain order to one of all possible permutations and combinations. Assuming that order j has oj process, and the process adjustment action can be represented as any one of all permutations and combinations of these j processes {σ 1 , σ 2 , …, σ oj}; The order batch division action mentioned above refers to combining orders with a priority > 1.5 into one production batch, Batch({orders j}, priority threshold = 1.5); The shutdown decision refers to shutting down a faulty device, PowerOnOff(device i , status ∈ {0, 1}); The safety stock adjustment action refers to adjusting the safety stock coefficient of the kth material. The adjustment strategy is that all order demand quantities = available inventory × (1 - min(safety stock coefficient)), that is, the current safety stock coefficient value just meets the order demand quantity that can be allocated; The resource allocation action refers to the action of allocating a certain device and material to a certain order or process based on an allocation strategy, including device allocation and material allocation; The device allocation refers to allocating device i to the lth process of order j for processing. The allocation strategy is device selection = argmin i (l i × (1 + f i ))), that is, preferentially select a device with a low load and normal operation; The material allocation refers to allocating a certain quantity of material k to the production of order j. The allocation strategy is the allocation quantity = min(order demand quantity, available inventory × (1 - safety stock coefficient)).

[0011] Preferably, the specific calculation method of the transition probability P(s′∣s,a) is as follows: , where N(s,a,s′) represents the actual occurrence times of directly transitioning from state s to state s′ after executing action a in historical data, and ∑ s′′ N(s,a,s′′) refers to the total number of times of transitioning from state s to all possible subsequent states s′′ after executing action a in historical data.

[0012] Preferably, the reward function is defined as: R(s,a) = α × production efficiency improvement - β × cost increase - γ × delivery time delay + δ × equipment utilization rate improvement, where α, β, γ, and δ are weight coefficients used to balance the importance of different indicators. The production efficiency improvement is measured by the increase in the number of orders completed per unit time; the cost increase includes equipment usage cost, material cost, and labor cost; the delivery time delay is represented by the difference between the actual delivery time and the order delivery time; the equipment utilization rate improvement is measured by the proportional change in the actual operating time of the equipment to the total available time.

[0013] Preferably, the state space S eq includes the equipment load status and the equipment failure status, and the action space A eq includes the shutdown decision and the equipment allocation in the resource allocation action. The π eq refers to calculating the actual reward R based on the state space S eq data in combination with the reward function R and the transition probability P(s′∣s,a). act Judge whether to retain the action PowerOnOff (equipment i , status ∈ {0,1}) and argmin i (l i ×(1 + f i )); The actual reward R act = P(S′ eq ∣S eq , A eq ) × R. When the actual reward is greater than the preset value, execute the action in the current action space; if the reward is less than the preset value, do not execute.

[0014] Preferably, the state space S ord includes the order priority, inventory quantity, safety stock coefficient, and remaining delivery time. The action space A ord includes the process adjustment action, order batch division action, safety stock adjustment action, and material allocation in the resource allocation action. The π ord refers to calculating the actual reward R based on the state space S ord data in combination with the reward function R and the transition probability P(s′∣s,a). act Judge whether to retain the action; The actual reward R act = P(S′ ord ∣S ord , A ord ) × R. When the actual reward is greater than the preset value, execute the action in the current action space; if the reward is less than the preset value, do not execute. If the actual rewards of multiple actions are all greater than the preset value, retain the action with the maximum actual reward.

[0015] Preferably, based on the actions retained in step S4, the artificial intelligence generates dynamic scheduling instructions to adjust the order processes according to real-time priorities and device statuses, combines orders with high priorities according to the real-time order priorities and preferentially assigns them to devices that meet the requirements in step S4 for production. At the same time, it allocates materials according to the allocated devices and order requirements to adjust the safety inventory coefficient. When the safety inventory coefficient does not meet the production requirements, it sends a material procurement requirement to the administrator device.

[0016] Technical effects and advantages of the present invention: By constructing an operation management database and integrating device operation data sets, order management data sets, and supply chain data sets, the present invention can efficiently store and process massive production data. Combining with large model technology, it can deeply mine and analyze the data, accurately reflect the status of the production environment in real time, and provide strong support for decision-making; The present invention abstracts production scheduling as a Markov decision process, defines the state space, action space, transition probability, and reward function, making the decision-making more scientific and reasonable. When determining the order priority and production sequence, it comprehensively considers multiple factors such as contract profit and remaining delivery time, evaluates the impact of different decisions on production efficiency, cost, and delivery time through the reward function, realizes the optimal allocation of production resources, and improves the order on-time delivery rate and enterprise profit; The present invention decomposes the operation scheduling task into a device management task and an order allocation task, clarifies the state space, action space, and agents of each task, realizes the collaborative work of links such as device management, order allocation, and material management, generates dynamic scheduling instructions through artificial intelligence, adjusts order processes, allocates devices and materials, and adjusts the safety inventory coefficient according to the real-time production situation, improves the overall efficiency of production operation, and enhances the enterprise's response ability to market changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic diagram of the overall structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] As shown in the appended Figure 1 An AI production operation management method based on a large model includes: S1: Database construction: Construct an operation management database based on device operation data sets, order management data sets, and supply chain data sets; Specifically, the device operation data set includes the unique device number, the time series data of the device workload, and the current fault status of the device. The device operation data set is constructed based on the data collected by the Internet of Things sensor devices and the device production logs; The order management data set includes order priority data, delivery date data, material requirement data, and production plan data. The order management data set is constructed based on the ERP system database, the MES system database, and the enterprise production logs; The supply chain data set includes material inventory data, material inbound and outbound record data, and purchase record data. The supply chain data set is constructed based on the WMS system database; S2: Production operation modeling: Abstract the production scheduling as a Markov decision process, and define the state space S, the action space A, the transition probability P, and the reward function R; Specifically, the state space S is a vector set that describes the specific situation of the production system at a certain moment. It contains characteristic information in multiple dimensions, and this information can comprehensively reflect the real-time state of the production environment. Specifically, assume that there are n devices, m orders, and k types of materials in the production system. The state s ∈ S can be represented as a vector. The specific state s is: s = [l 1 , l 2 ,..., l n , p 1 , p 2 ,..., p m , q 1 , q 2 ,..., q k , b 1 , b 2 ,..., b k , t 1 , t 2 ,..., t m , f 1 , f 2 ,..., f n , where: l i (i = 1, 2,..., n) represents the load of the i-th device, which is represented by the proportion of the time currently occupied by the device. The calculation method is: l i = (the remaining time of the current task + the total time of the tasks in the queue) / the total available time of the device, and the value range is between [0, 1]. 0 means the device is idle, and 1 means the device is running at full load. For example, if the device is available for 24 hours per day, the remaining time of the current task of device 1 is 3 hours, and the total time of the queue tasks is 2 hours, then l 1=(3 + 2) / 24 ≈ 0.208; The remaining time for the current task of Device 2 is 5 hours, and the total time of the queued tasks is 1 hour, then l 2 =(5 + 1) / 24 = 0.25; The remaining time for the current task of Device 3 is 1 hour, and the total time of the queued tasks is 0 hour, then l 3 =1 / 24 ≈ 0.042.

[0020] p j p(j = 1, 2,..., m) represents the priority of the j-th order. The calculation method is: p j = 0.6×contract profit / max(all order profits) + 0.4×remaining delivery time / min(all order remaining times). The larger the value, the higher the priority. For example, the profit of Order 1 is 5000 yuan and the remaining delivery time is 5 days; the profit of Order 2 is 3000 yuan and the remaining delivery time is 3 days; the profit of Order 3 is 4000 yuan and the remaining delivery time is 4 days. The maximum profit of all orders is 5000 yuan and the minimum remaining delivery time is 3 days, then p 1 = 0.6×(5000 / 5000) + 0.4×(5 / 3) ≈ 1.27, p 2 = 0.6×(3000 / 5000) + 0.4×(3 / 3) = 0.76, p 3 = 0.6×(4000 / 5000) + 0.4×(4 / 3) ≈ 1.01; q l q(l = 1, 2,..., k) represents the inventory quantity of the l-th material. The inventory quantity affects the continuity of production. If the inventory of a certain material is insufficient, it may lead to production interruption. The calculation method is: q l = current inventory - allocated but unused quantity + in-transit purchase quantity. For example, the current inventory of Material 1 is 100 pieces, the allocated but unused quantity is 20 pieces, and the in-transit purchase quantity is 30 pieces, then q 1 = 100 - 20 + 30 = 110; The current inventory of Material 2 is 80 pieces, the allocated but unused quantity is 10 pieces, and the in-transit purchase quantity is 20 pieces, then q 2 = 80 - 10 + 20 = 90; The current inventory of Material 3 is 120 pieces, the allocated but unused quantity is 30 pieces, and the in-transit purchase quantity is 40 pieces, then q 3 = 120 - 30 + 40 = 130; b l b(l = 1, 2,..., k) represents the safety stock coefficient of the l-th material. The safety stock coefficient is set by the enterprise based on historical data and industry experience, representing the inventory ratio reserved by the enterprise to cope with uncertainties such as demand fluctuations and supply delays; t j(j = 1, 2, ..., m) represents the remaining delivery time of the j-th order, that is, the remaining time from the current moment to the order delivery deadline. For example, assume that the delivery date of order 1 is 5 days away from the current date, t 1 = 5; the delivery date of order 2 is 3 days away from the current date, t 2 = 3; the delivery date of order 3 is 4 days away from the current date, t 3 = 4; f i (i = 1, 2, ..., n) represents the failure status of the i-th device, which is represented by a binary value. 0 indicates that the device is operating normally, and 1 indicates that the device has a failure.

[0021] The action space A is an abstract set of decisions that can be taken in the production operation process. These decisions will affect the state of the production system. The action space is divided into process adjustment actions, order batch division actions, shutdown decisions, safety inventory adjustment actions, and resource allocation actions; The process adjustment action refers to adjusting the process sequence of a certain order to one of all possible permutations and combinations. Assume that order j has o j processes. The process adjustment action can be represented as any one of all permutations and combinations {σ j , σ 1 , …, σ 2} of these o oj processes. For example, if the original process sequence of order j is [1, 2, 3], after the process adjustment, it can become [1, 3, 2], [2, 1, 3], [2, 3, 1], [3, 1, 2], [3, 2, 1].

[0022] The order batch division action refers to merging orders with a priority > 1.5 into one production batch, Batch({order j}, priority threshold = 1.5); The shutdown decision refers to shutting down the faulty device, PowerOnOff(device i , status ∈ {0, 1}); The safety inventory adjustment action refers to adjusting the safety inventory coefficient of the k-th material. The adjustment strategy is that all order demand quantities = available inventory × (1 - min(safety inventory coefficient)), that is, the current safety inventory coefficient value just meets the requirement that the distributable inventory can meet the order demand quantities; The resource allocation action refers to the action of allocating a certain device and material to a certain order or process based on the allocation strategy, including device allocation and material allocation; The device allocation refers to allocating device i to the l-th process of order j for processing. The allocation strategy is device selection = argmin i (l i×(1 + f i ))), that is, preferentially select the equipment with low load and normal operation; The material distribution mentioned above refers to allocating a certain quantity of material k to the production of order j, and the distribution strategy is: allocation quantity = min(order demand, available inventory × (1 - safety stock coefficient)).

[0023] The transition probability P(s′∣s,a) represents the probability of transferring to state s′ after taking action a (a ∈ A) in state s; It should be further noted that this probability is calculated based on historical data. The specific calculation method of the transition probability P(s′∣s,a) is as follows: , where N(s,a,s′) represents the actual occurrence times of directly transferring from state s to state s′ after executing action a in historical data, and ∑ s′′ N(s,a,s′′) refers to the total number of times of transferring from state s to all possible subsequent states s′′ after executing action a in historical data.

[0024] The reward function R is used to evaluate the quality of taking action a in state s. The reward function is related to production efficiency, cost, and delivery time. The reward function is defined as: R(s,a) = α × production efficiency improvement - β × cost increase - γ × delivery time delay + δ × equipment utilization improvement, where α, β, γ, and δ are weight coefficients used to balance the importance of different indicators. The production efficiency improvement is measured by the increase in the number of orders completed per unit time; the cost increase includes equipment usage cost, material cost, and labor cost; the delivery time delay is represented by the difference between the actual delivery time and the order delivery time; the equipment utilization improvement is measured by the change in the ratio of the actual operating time of the equipment to the total available time.

[0025] It should be further noted that α, β, γ, and δ are used to balance the importance of different indicators and are adaptively adjusted according to the production and operation goals of the enterprise. For example, when the enterprise needs to increase the consideration of equipment utilization, the value of δ can be increased, and α + β + γ + δ = 1.

[0026] S3: Data collection: Obtain operation data from the operation management database based on the state space S, action space A, transition probability P, and reward function R defined in step S2. After performing preprocessing operations on the obtained operation data, a Markov decision data set is obtained; Specifically, the preprocessing operations include missing value filling, duplicate value deletion, and normalization operations; The formula for the normalization operation is as follows: l inormalized = (l i - l imin ) / (l imax - limin ) Among them, l inormalized is the normalized value, and l i is the original value, and l imin and l imax are the minimum value and the maximum value respectively; S4: Operation task decomposition judgment: Based on the Markov decision data set, construct an operation scheduling task, and decompose the operation scheduling task into a device management task and an order allocation task. Based on the allocation, judge whether the actions in the action spaces of the device management task and the order allocation task are executed in real time; The device management task includes a state space S eq , an action space A eq , and an agent π eq ; The order allocation task includes a state space S ord , an action space A ord , and an agent π ord ; It should be further noted that the state space S eq includes the device load state and the device failure state. The action space A eq includes the shutdown decision and the device allocation in the resource allocation action. The π eq refers to calculating the actual reward R eq by artificial intelligence based on the data in the state space S act in combination with the reward function R and the transition probability P(s′∣s,a), and judging whether the actions PowerOnOff(device i , state ∈ {0,1}) and argmin i (l i ×(1 + f i )) should be retained; Specifically, the actual reward R act = P(S′ eq ∣S eq , A eq ) × R. When the actual reward is greater than the preset value, execute the action in the current action space. If the reward is less than the preset value, do not execute; The state space S ord includes the order priority, inventory quantity, safety stock coefficient, and remaining delivery time. The action space A ord includes the process adjustment action, order batch division action, safety stock adjustment action, and material allocation in the resource allocation action. The π ord refers to calculating the actual reward R ord by artificial intelligence based on the data in the state space S act and judging whether the action should be retained; Specifically, the actual reward Ract =P(S′ ord ∣S ord ,A ord ) × R. When the actual reward is greater than the preset value, execute the action space action of the current action. If the reward is less than the preset value, do not execute. If the actual rewards of multiple actions are all greater than the preset value, retain the action with the largest actual reward; It should be specifically noted that the preset value is determined by the enterprise based on experimental data and historical data, and is used to measure the impact of the current action on the production efficiency improvement, cost reduction, delivery efficiency, and equipment utilization rate improvement in the production and operation process; S5: Intelligent action execution: Based on the judgment result in step S4, generate a dynamic scheduling instruction through artificial intelligence to execute the action space action; It should be specifically noted that the artificial intelligence generates a dynamic scheduling instruction to adjust the order process based on the actions retained in step S4 according to the real-time priority and equipment status, and combines and preferentially allocates the orders with high priority to the equipment that meets the requirements in step S4 for production according to the real-time order priority. At the same time, allocate materials according to the allocated equipment and order requirements to adjust the safety inventory coefficient. When the safety inventory coefficient does not meet the production demand, send a material procurement requirement to the administrator's device; Secondly: In the attached drawings of the disclosed embodiments of the present invention, only the structures related to the disclosed embodiments of the present disclosure are involved. Other structures can refer to the general design. Without conflict, the same embodiment and different embodiments of the present invention can be combined with each other; Finally: The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A large-model-based AI production operation management method, characterized in that: include: S1: Database construction: Build an operation management database based on the equipment operation data set, order management data set, and supply chain data set; S2: Production operation modeling: Abstract production scheduling as a Markov decision process, define the state space S, action space A, transition probability P and reward function R; The state space S is a set of vectors that describe the specific status of the production system at a certain moment. Suppose there are n devices, m orders, and k materials in the production system. The state s∈S can be represented as a vector. The action space A is an abstract set of decisions that can be taken during the production operation process. These decisions will affect the state of the production system. The action space is divided into process adjustment actions, order batch division actions, shutdown decisions, safety stock adjustment actions, and resource allocation actions; The transition probability P(s′|s,a) represents the probability of transitioning to state s′ after taking action a (a∈A) in state s; The reward function R is used to evaluate the quality of taking action a in state s; S3: Data collection: Based on the state space S, action space A, transition probability P and reward function R defined in step S2, the operation data is obtained from the operation management database, and the obtained operation data is preprocessed to obtain a Markov decision data set; S4: Operation task decomposition judgment: construct the operation scheduling task based on the Markov decision data set, and decompose the operation scheduling task into the equipment management task and the order allocation task, and judge whether the action space action of the equipment management task and the order allocation task is executed in real time based on the allocation; The device management task includes a state space S eq , action space A eq and agent π eq ; The order allocation task includes the state space S ord , action space A ord and agent π ord ; S5: Intelligent action execution: Based on the judgment result in step S4, dynamic scheduling instructions are generated through artificial intelligence to execute action space actions.

2. According to the big model-based AI production operation management method of claim 1, it is characterized by: The equipment operation data set includes the equipment unique number, equipment workload time series data and the current fault status of the equipment, and the equipment operation data set is constructed based on the IoT sensor equipment collection data and equipment production log; The order management data set includes order priority data, delivery date data, material requirement data and production plan data, and the order management data set is constructed based on the ERP system database, the MES system database and the enterprise production log; The supply chain data set includes material inventory data, material in-and-out record data, and purchase record data, and the supply chain data set is constructed based on the WMS system database.

3. According to the big model-based AI production operation management method of claim 1, it is characterized by: The state s is specifically: s=[l1,l2,...,l n ,p1,p2,...,p m ,q1,q2,...,q k ,b1,b2,...,b k ,t1,t2,...,t m ,f1,f2,...,f n ],in: l i (i=1,2,...,n) represents the load of the ith device, which is expressed by the proportion of time currently occupied by the device. The calculation method is: l i = (remaining time of current task + total time of tasks in queue) / total available time of device, the value range is [0,1], 0 means the device is idle, 1 means the device is running at full capacity; p j (j=1,2,...,m) represents the priority of the jth order, which is calculated as follows: j =0.6×contract profit / max(profit of all orders)+0.4×remaining delivery time / min(remaining time of all orders); q l (l=1,2,...,k) represents the inventory quantity of the lth material. The inventory quantity will affect the continuity of production. If the inventory of a certain material is insufficient, it may cause production interruption. The calculation method is: q l = Current inventory - allocated but unused quantity + in-transit purchase quantity; b l (l=1,2,...,k) represents the safety stock factor of the lth material; t j (j=1,2,...,m) represents the remaining delivery time of the jth order, that is, the remaining time from the current moment to the order delivery deadline; f i (i=1,2,...,n) indicates the fault status of the ith device, expressed in binary values, where 0 indicates that the device is operating normally and 1 indicates that the device is faulty.

4. The AI ​​production operation management method based on a large model according to claim 1 is characterized in that: The process adjustment action refers to adjusting the process sequence of a certain order to one of all possible permutations and combinations. Assume that order j has o j The process adjustment action can be expressed as this o j All permutations and combinations of the process {σ1,σ2,…,σ oj Any one of}; The order batching action refers to merging orders with a priority level greater than 1.5 into one production batch. j }, priority threshold = 1.5); The shutdown decision refers to shutting down the faulty device, PowerOnOff (device i , state∈{0,1}); The safety stock adjustment action refers to adjusting the safety stock coefficient of the kth material. The adjustment strategy is that all order requirements = available inventory × (1-min (safety stock coefficient)), that is, the current safety stock coefficient value just satisfies the allocable inventory to meet the order requirements; The resource allocation action refers to the action of allocating certain equipment and materials to a certain order or process based on the allocation strategy, including equipment allocation and material allocation; The equipment allocation refers to allocating equipment i to the lth process of order j for processing. The allocation strategy is equipment selection = argmin i (l i ×(1+f i )), that is, giving priority to equipment with low load and normal operation; The material allocation refers to allocating a certain amount of material k to the production of order j, and the allocation strategy is allocation quantity = min (order demand, inventory available quantity × (1-safety stock coefficient)).

5. The AI ​​production operation management method based on a large model according to claim 1 is characterized in that: The reward function is defined as: R(s,a)=α×production efficiency improvement-β×cost increase-γ×delay in delivery+δ×improvement in equipment utilization, where α, β, γ, and δ are weight coefficients used to balance the importance of different indicators. The improvement in production efficiency is measured by the increase in the number of orders completed per unit time; the cost increase includes equipment usage cost, material cost, and labor cost; the delay in delivery is represented by the difference between the actual delivery time and the order delivery time; the improvement in equipment utilization is measured by the change in the ratio of the actual operating time of the equipment to the total available time.

6. The AI ​​production operation management method based on a large model according to claim 1 is characterized in that: The state space S eq Including equipment load status and equipment failure status, the action space A eq Including shutdown decisions and resource allocation actions in device allocation, the π eq Refers to artificial intelligence based on the state space S eq The data is combined with the reward function R and the transition probability P(s′|s,a) to calculate the actual reward R act Determine the action PowerOnOff (device i , state ∈ {0,1}) and argmin i (l i ×(1+f i )) whether to retain; The actual reward R act =P(S′ eq ∣S eq ,A eq )×R, when the actual reward is greater than the preset value, the current action space action is executed, and if the reward is less than the preset value, it is not executed.

7. The AI ​​production operation management method based on a large model according to claim 1 is characterized in that: State space S ord Including order priority, inventory quantity, safety stock factor and remaining delivery time, the action space A ord Including process adjustment actions, order batch division actions, safety stock adjustment actions, and material allocation in resource allocation actions. ord Refers to artificial intelligence based on the state space S ord The data is combined with the reward function R and the transition probability P(s′|s,a) to calculate the actual reward R act Determine whether the action is retained; The actual reward R act =P(S′ ord ∣S ord ,A ord )×R, when the actual reward is greater than the preset value, the current action space action is executed. If the reward is less than the preset value, it is not executed. If the actual rewards of multiple actions are greater than the preset value, the action with the largest actual reward is retained.

Citation Information

Patent Citations

  • Casting enterprise raw material inventory management control method based on Markov decision theory

    CN111126905A

  • MTO enterprise order processing method and system

    CN113592240A

  • Dynamic scheduling method and system for unstable hybrid job shop with randomly arrived orders

    CN118780540A

  • Intelligent production scheduling method and system for chemical fiber products based on reinforcement learning

    CN119378948A