Aerospace product hot processing workshop dynamic decision-making method, device and equipment

By employing a neural network-based intelligent decision-making method in aerospace thermal processing workshops, multimodal states are dynamically identified and production parameters are adjusted in real time, solving the problem of inefficient decision-making in existing technologies and achieving efficient production process control and quality assurance.

CN119761884BActive Publication Date: 2025-11-21WUHAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411700211.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-11-21
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

Existing intelligent optimization algorithms in aerospace thermal processing workshops struggle to dynamically identify multimodal states, making it difficult to achieve efficient intelligent collaborative scheduling of quality and energy efficiency during production. This is especially true when dynamic events such as emergency order insertions and machine failures occur, making it difficult to make efficient decisions.

Method used

An intelligent decision-making method based on a neural network model is adopted. By acquiring production data from the hot processing workshop, abstracting state information, using online Q-networks and target Q-networks for value assessment, generating control action information in the action space, and optimizing the decision-making process through a reward function, production parameters are adjusted in real time.

Benefits of technology

It enables dynamic identification of multimodal states in aerospace thermal processing workshops, rapid response to emergencies, ensuring processing quality and improving production efficiency, thereby enhancing the intelligent control capabilities of the production process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119761884B_ABST
    Figure CN119761884B_ABST
Patent Text Reader

Abstract

The application discloses an aerospace product hot processing workshop dynamic decision method, device and equipment, and belongs to the technical field of hot processing production. The method comprises the following steps: acquiring first production data in the process of producing a target workpiece in a hot processing workshop; abstracting the first production data in a state space to obtain first workshop state information of the hot processing workshop; inputting the first workshop state information into an intelligent decision model to obtain first control action information of an action space output by the intelligent decision model, the intelligent decision model being constructed based on a neural network model; and generating a production decision scheme for producing the target workpiece in the hot processing workshop according to the first control action information. The method can dynamically identify multi-modal states of the workshop and realize efficient decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of thermal processing production technology, and in particular relates to a dynamic decision-making method, device and equipment for thermal processing workshops of aerospace products. Background Technology

[0002] The production mode of aerospace thermally processed products (such as aircraft main landing gear and rocket engine heads) is mainly single-piece, small-batch production, characterized by long processes, slow cycle times, high energy consumption, and extremely high quality requirements. The thermal processing workshop is the main carrier of the aerospace thermally processed product production process, and its management and control methods directly determine the quality and efficiency of the thermally processed product production process.

[0003] Currently, the management and control of aerospace product thermal processing workshops mainly relies on manual experience, such as manual production scheduling, process design, and manual troubleshooting, which is primarily a low-efficiency manual management approach. The thermal processing production process is subject to numerous quality and energy consumption disturbances. When dynamic events such as urgent order insertions or machine malfunctions occur in the thermal processing workshop, efficient intelligent collaborative scheduling and decision-making for quality and energy efficiency become difficult. Automated management and control is the development trend of thermal processing workshop management, but the current integrated management and control system technology for aerospace thermal processing workshops is developing slowly and has limited functionality, only offering data acquisition, transmission, visualization, and a limited amount of scheduling optimization using genetic algorithms. Integrating advanced intelligent optimization algorithms into the aerospace thermal processing workshop management and control system is the underlying key technology for transforming inefficient manual management into highly efficient intelligent management.

[0004] However, current mainstream intelligent optimization algorithms are typically metaheuristic methods (such as evolutionary algorithms, swarm intelligence-based algorithms, and physics-based algorithms). These methods are mostly static, and their fixed search strategies, numerous iteration requirements, static fitness functions, lack of real-time data fusion, and fixed parameter settings make it difficult to dynamically identify multimodal states in the workshop and make efficient decisions. Therefore, improving the dynamic perception and real-time efficient decision-making capabilities of decision-making algorithms has become a key challenge that urgently needs to be addressed in the intelligent control system technology for the production process of aerospace product thermal processing workshops. Summary of the Invention

[0005] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes a dynamic decision-making method, apparatus, and equipment for aerospace product thermal processing workshops, which can dynamically identify multimodal states of the workshop and achieve efficient decision-making.

[0006] Firstly, this application provides a dynamic decision-making method for aerospace product thermal processing workshops, the method comprising:

[0007] Obtain the first production data during the production of the target workpiece in the heat treatment workshop;

[0008] The first production data is abstracted in the state space to obtain the first workshop state information of the heat treatment workshop;

[0009] The first workshop status information is input into the intelligent decision-making model to obtain the first control action information of the action space output by the intelligent decision-making model. The intelligent decision-making model is constructed based on a neural network model.

[0010] Based on the first control action information, a production decision plan for producing the target workpiece in the heat treatment workshop is generated.

[0011] According to one embodiment of this application, the first workshop state information of the state space includes the product processing quality, order status information, machine utilization rate, and overall machine energy consumption average of the target workpiece during the production process.

[0012] According to one embodiment of this application, the expression for the product processing quality is:

[0013]

[0014] Wherein, Q(t) represents the product processing quality; α, β, γ, δ, and ζ are the weights of the state parameters that determine the quality; T represents the state parameter corresponding to the forging temperature; P represents the state parameter corresponding to the forging pressure; S micro-structure The state parameters corresponding to the microstructure; G represents the state parameter corresponding to strain; G represents the state parameter corresponding to geometric dimensions.

[0015] According to one embodiment of this application, the order status information is characterized as a set of current orders and urgent orders for the target workpiece:

[0016] W = {p i ,d i ,w i )|i=1,2,3…,n}∪{(p E ,d E ,w E )}

[0017] Where W represents the set of current orders and urgent orders; i is the order number, p i Indicates the processing time of order i; d i Indicates the delivery deadline for order i; ω i p represents the weight of order i; E Indicates the processing time for urgent orders; d E Indicates the delivery deadline for urgent orders; ω E Indicates the weight of urgent orders; ω E Greater than ωi .

[0018] According to one embodiment of this application, the neural network model includes an online Q-network and a target Q-network. The online Q-network is used to predict and output the first control action information based on the first workshop state information, and the online Q-network is used to update the parameters of the target Q-network.

[0019] According to one embodiment of this application, the step of inputting the first workshop state information into the intelligent decision-making model to obtain the first control action information of the action space output by the intelligent decision-making model includes:

[0020] The first workshop status information is input into the online Q network. The online Q network is used to evaluate the value of multiple possible actions corresponding to the first workshop status information. Among the multiple possible actions, the action combination with the highest Q value is determined as the first control action information.

[0021] The first control action information is evaluated using the target Q-network to obtain a target Q-value for the first control action information, and the target Q-value is used to update the online Q-network.

[0022] According to one embodiment of this application, after generating a production decision scheme for the hot working workshop to produce the target workpiece based on the first control action information, the method further includes:

[0023] When the production decision plan is executed in the heat treatment workshop, the second production data corresponding to the production decision plan is obtained;

[0024] The second production data is abstracted in the state space to obtain the second workshop state information of the heat treatment workshop;

[0025] The combination of the first workshop status information, the first control action information, the reward value corresponding to the first control action information, and the second workshop status information is used as a training sample, and the training sample is stored in the experience replay pool.

[0026] The training samples are extracted from the experience replay pool to update the parameters of the intelligent decision-making model.

[0027] According to one embodiment of this application, the reward value is obtained based on a reward function, wherein the reward function R is:

[0028] R=ω1·R geo +ω2·R mech +ω3·R def +ω4·R eff -P viol

[0029] Among them, R geo R represents the bonus value for the size and shape deviation of the target workpiece. mech R represents the bonus value for the mechanical performance deviation of the target workpiece. def R represents the reward for defect detection of the target workpiece. eff R represents the reward for the production efficiency of the target workpiece. viol This indicates that actions in the first control action information that may lead to quality problems will be given a negative reward, and ω1, ω2, ω3, and ω4 are weight parameters that characterize the relative importance.

[0030] Secondly, this application provides a dynamic decision-making device for aerospace product thermal processing workshops, the device comprising:

[0031] The acquisition module is used to acquire the first production data during the production of the target workpiece in the hot working workshop;

[0032] The first processing module is used to abstract the first production data in the state space to obtain the first workshop state information of the heat processing workshop;

[0033] The second processing module is used to input the first workshop status information into the intelligent decision model to obtain the first control action information of the action space output by the intelligent decision model, wherein the intelligent decision model is constructed based on a neural network model;

[0034] The third processing module is used to generate a production decision plan for the hot processing workshop to produce the target workpiece based on the first control action information.

[0035] Thirdly, this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the dynamic decision-making method for aerospace product thermal processing workshops as described in the first aspect above.

[0036] Fourthly, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the dynamic decision-making method for aerospace product thermal processing workshops as described in the first aspect above.

[0037] Fifthly, this application provides a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the dynamic decision-making method for aerospace product thermal processing workshops as described in the first aspect.

[0038] In a sixth aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the dynamic decision-making method for aerospace product thermal processing workshops as described in the first aspect above.

[0039] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application.

[0040] The present invention provides a dynamic decision-making method, apparatus, and equipment for aerospace product thermal processing workshops, which has the following advantages over the prior art:

[0041] (1) By abstracting the first production data monitored in real time, the first workshop status information is obtained as the input of the intelligent decision-making model, and a decision-making production plan with dynamic adjustment of production parameters is obtained. This ensures that key parameters such as temperature, pressure, stress and strain in the production process are always kept within the optimal range. In the production process of the aerospace product thermal processing workshop, the multimodal state of the workshop is dynamically identified to achieve efficient decision-making.

[0042] (2) It can respond quickly when dynamic events occur in the thermal processing workshop of aerospace products, perceive the status of the workshop in real time, make intelligent decisions on workshop actions based on status characteristics, and make real-time dynamic feedback adjustments based on the workshop actions after the decision, so as to ensure processing quality and improve production efficiency. It has practical application value for the development of production process control system for thermal processing workshop of aerospace products, and can achieve efficient decision-making and intelligent control. Attached Figure Description

[0043] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0044] Figure 1 This is one of the flowcharts illustrating the dynamic decision-making method for aerospace product thermal processing workshops provided in this application embodiment;

[0045] Figure 2 This is a schematic diagram illustrating the mapping relationship between the state space and the action space provided in the embodiments of this application;

[0046] Figure 3 This is the second flowchart illustrating the dynamic decision-making method for aerospace product thermal processing workshops provided in this application embodiment;

[0047] Figure 4 This is a schematic diagram of the structure of the dynamic decision-making device for the thermal processing workshop of aerospace products provided in the embodiments of this application;

[0048] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0049] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0050] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0051] In recent years, deep reinforcement learning methods, which integrate deep learning and reinforcement learning, have proven effective in solving real-time dynamic scheduling problems for various types of workshops, including job workshops, flexible job workshops, and assembly line workshops, in the field of dynamic workshop scheduling. However, the multimodal states, complex production processes, stringent quality indicators, and energy consumption indicators in aerospace thermal processing pose challenges to the application of deep reinforcement learning in scheduling decisions for aerospace thermal processing workshops. Therefore, there is an urgent need to establish a general scheduling decision-making model and method applicable to the field of aerospace product thermal processing.

[0052] The following description, in conjunction with the accompanying drawings, details the dynamic decision-making method, dynamic decision-making device, electronic equipment, and readable storage medium for aerospace product thermal processing workshops provided in this application, through specific embodiments and application scenarios.

[0053] Among them, the dynamic decision-making method for the thermal processing workshop of aerospace products can be applied to the terminal, specifically executed by the hardware or software in the terminal.

[0054] The terminal includes, but is not limited to, portable communication devices such as mobile phones or tablets with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). It should also be understood that, in some embodiments, the terminal may not be a portable communication device, but rather a desktop computer with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads).

[0055] The following embodiments describe a terminal including a display and a touch-sensitive surface. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, mouse, and joystick.

[0056] The dynamic decision-making method for aerospace product thermal processing workshops provided in this application embodiment can be executed by an electronic device or a functional module or entity in an electronic device that can realize the dynamic decision-making method for aerospace product thermal processing workshops. The electronic devices mentioned in this application embodiment include, but are not limited to, mobile phones, tablets, computers, cameras, and wearable devices. The following uses an electronic device as the execution subject to illustrate the dynamic decision-making method for aerospace product thermal processing workshops provided in this application embodiment.

[0057] like Figure 1 As shown, the dynamic decision-making method for the aerospace product thermal processing workshop includes:

[0058] Step 110: Obtain the first production data during the production of the target workpiece in the heat treatment workshop;

[0059] The heat treatment workshop is a production area where metals or other materials are heat-treated and used to produce aerospace products.

[0060] The target workpiece can be an aerospace product.

[0061] The primary production data can include output data, equipment utilization rate, energy consumption data, and quality data, which can reflect the production efficiency, resource utilization, and product quality of the heat processing workshop.

[0062] Step 120: Abstract the first production data in the state space to obtain the first workshop state information s of the heat treatment workshop;

[0063] The first workshop state information s in the state space can include temperature, pressure, stress and strain, geometric dimensions, microstructure, order queue, emergency orders, machine utilization rate, and overall machine energy consumption average during the production process of aerospace thermal processing products.

[0064] In this step, the first production data is preprocessed by data cleaning and standardization, and key indicators such as production efficiency, equipment utilization rate, and pass rate are calculated on the preprocessed first production data to obtain the first workshop status information s.

[0065] Step 130: Input the first workshop state information s into the intelligent decision model to obtain the first control action information a in the action space output by the intelligent decision model. The intelligent decision model is constructed based on a neural network model.

[0066] Among them, the first control action information 'a' is designed based on the action space and may include: forging temperature adjustment, forging pressure adjustment, forging speed control, product heat treatment parameter adjustment, insertion of emergency orders, adjustment of production sequence, machine task balancing allocation, and high and low power machine task balancing allocation.

[0067] like Figure 2 As shown, there is a mapping relationship between the workshop state information in the state space and the control action information in the action space, and the control action information can affect the corresponding workshop state information.

[0068] Workshop status information includes: temperature, pressure, stress and strain, geometry, microstructure, current order queue, priority and arrival time of urgent orders, average machine utilization, and average overall machine energy consumption for aerospace thermal processing products.

[0069] Control action information includes: forging temperature adjustment, forging pressure adjustment, forging speed control, product heat treatment parameter adjustment, insertion of emergency orders, adjustment of production sequence, machine task balancing allocation, and high and low power machine task balancing allocation.

[0070] In this step, the first workshop state information s is input into the intelligent decision model. The intelligent decision model extracts features from the first workshop state information s and generates and outputs the first control action information a in the action space based on the mapping relationship between the workshop state and the control action.

[0071] Step 140: Based on the first control action information, generate a production decision plan for the hot processing workshop to produce the target workpiece.

[0072] The production decision-making scheme includes production task allocation, equipment scheduling plan, material requirements list and risk assessment, etc., which are used to control the equipment in the heat processing workshop to control the equipment to execute the first control action information.

[0073] In this step, the first control action information is analyzed to determine the content of the control action information, including equipment status, current production status, process parameters, etc., and to clarify the requirements such as the specifications, quantity and delivery time of the target workpiece.

[0074] Based on the characteristics of the target workpiece, design the corresponding hot working process flow, including the sequence of each process and the operation requirements.

[0075] Based on the first control action information, the optimal process parameters such as temperature, time, and pressure are set to ensure the quality of the target workpiece. Simulation tools are used to predict the production effect under different parameters, optimize the process parameters, and form a complete production decision plan.

[0076] According to the dynamic decision-making method for aerospace product thermal processing workshops provided in this application embodiment, by abstracting the first production data monitored in real time, the first workshop status information is obtained as the input of the intelligent decision-making model, and a decision-making production plan for dynamically adjusting production parameters is obtained. This ensures that key parameters such as temperature, pressure, stress and strain during the production process are always kept within the optimal range. In the production process of aerospace product thermal processing workshops, the multimodal states of the workshop are dynamically identified to achieve efficient decision-making.

[0077] In some embodiments, the first workshop state information of the state space includes the product processing quality, order status information, machine utilization rate, and average overall machine energy consumption of the target workpiece during the production process.

[0078] In some embodiments, the expression for the product processing quality is:

[0079]

[0080] Wherein, Q(t) represents the product processing quality; α, β, γ, δ, and ζ are the weights of the state parameters that determine the quality; T represents the state parameter corresponding to the forging temperature; P represents the state parameter corresponding to the forging pressure; S micro-structure The state parameters corresponding to the microstructure; G represents the state parameter corresponding to strain; G represents the state parameter corresponding to geometric dimensions.

[0081] In some embodiments, order status information is represented as a set of current orders and urgent orders for the target workpiece:

[0082] W = {p i ,d i ,w i )|i=1,2,3…,n}∪{(p E ,d E ,w E )}

[0083] Where W represents the set of current orders and urgent orders; i is the order number, p i Indicates the processing time of order i; d i Indicates the delivery deadline for order i; ω i p represents the weight of order i; E Indicates the processing time for urgent orders; d E Indicates the delivery deadline for urgent orders; ω E ω represents the weight of urgent orders. E Greater than ω i .

[0084] In actual execution, the state information of the first workshop is a = {Q(t), W, U}. std (t), U ave-E (t)}, specifically as follows:

[0085]

[0086] In the formula, Q(t) represents the product processing quality; α, β, γ, δ, and ζ are the weights of the various state parameters that determine the quality; T represents the forging temperature; P represents the forging pressure; and S... micro-structure Indicates microstructure (such as average grain size); G represents strain; p represents geometric dimension; p represents strain. i Indicates the processing time of order i; d i Indicates the delivery deadline for order i; ω i p represents the weight (priority) of order i; E Indicates the processing time for urgent orders; d E Indicates the delivery deadline for urgent orders; ω E The weight of the emergency order (ω) E Greater than ω i W represents the set of current orders and urgent orders; U std (t) represents the standard deviation of the machine's average utilization rate; U ave (t) represents the average machine utilization rate; U energy (t) represents the total energy consumption of the workshop under the current state; U ave-E (t) represents the average energy consumption of the entire machine; n represents the total number of workpieces; m represents the total number of machines; i represents the workpiece serial number; j represents the machine serial number; h i This represents the total number of operations on a workpiece (o∈(1,2,...,h)). i ); o represents the process identifier (o∈(1,2,...,h) i );t ioj This represents the processing time of process o on machine j for workpiece i; This represents the power consumption of machine j during startup; This represents the power consumption of machine j when it is off. This represents the processing power of machine j per unit time; ST represents the waiting power of machine j per unit time; ioj This indicates the start time of process o for workpiece i on machine j; FT ioj This represents the end time of process o for workpiece i on machine j; x ioj This indicates that the value of operation o for workpiece i is 1 during processing by machine j (otherwise it is 0); y ioj The value x indicates that machine j is in a powered-off state during operation o of processing workpiece i (otherwise it is 0);ioi‘o’j This indicates that the process o of workpiece i is preceded by the process o′ of workpiece i', and is 1 (otherwise 0) during processing by machine j.

[0087] In some embodiments, the neural network model includes an online Q-network and a target Q-network. The online Q-network is used to predict and output the first control action information a based on the first workshop state information s, and the online Q-network is used to update the parameters of the target Q-network.

[0088] In this embodiment, the neural network model is a Double Deep Q-Network (DDQN), and the intelligent decision-making model can be a decision agent based on the DDQN algorithm. It has been reconstructed by combining aerospace thermal processing status, actions, and rewards. The computer test environment for simulation is: implemented based on Python language and Tensorflow 2.0 framework, Intel(R) Core(TM) i5-10400F CPU@2.90GHz, RAM32GB.

[0089] In some embodiments, inputting the first workshop state information s into the intelligent decision model to obtain the first control action information a in the action space output by the intelligent decision model includes:

[0090] The first workshop status information s is input to the online Q network. The online Q network is used to evaluate the value of multiple possible actions corresponding to the first workshop status information s. Among the multiple possible actions, the action combination with the highest Q value is determined as the first control action information a.

[0091] The first control action information a is evaluated using the target Q network to obtain the target Q value of the first control action information, and the target Q value is used to update the online Q network.

[0092] In actual execution, the first control action information is represented as:

[0093] a = [a temp ,a press ,a speed ,a heat_treat ,a prod_order ,a task_bal ,a power_bal ]

[0094] In the formula, 'a' represents the action vector, where each element represents a specific action, and:

[0095] Forging temperature regulation action: a temp ∈[T min ,Tmax ], T min and T max Indicates the minimum and maximum permitted temperatures;

[0096] Forging pressure adjustment action: a press ∈[P min ,P max ], P min and P max Indicates the minimum and maximum permitted pressure;

[0097] Forging speed control action: a speed ∈[S min ,S max ], S min and S max Indicates the minimum and maximum permitted speeds;

[0098] Heat treatment parameter (heat treatment furnace) adjustment actions: a heat_treat ∈[H min H max ], H min and H max Indicates the minimum and maximum permitted heat treatment temperatures;

[0099] Production sequencing action: a prod_order ∈{all possible production orders};

[0100] Machine task balancing and allocation actions: a task_bal ∈{all possible task assignments};

[0101] High and low power machine task balancing and allocation action: a power_bal ∈{Task balance of all possible high and low power machines}.

[0102] In some embodiments, after generating a production decision scheme for the hot working workshop to produce the target workpiece based on the first control action information, the method further includes:

[0103] When the production decision plan is executed in the heat treatment workshop, the second production data corresponding to the production decision plan is obtained;

[0104] The second production data is abstracted in the state space to obtain the second workshop state information of the heat treatment workshop;

[0105] The combination of the first workshop status information, the first control action information, the reward value corresponding to the first control action information, and the second workshop status information s' is used as a training sample, and the training sample is stored in the experience replay pool.

[0106] The training samples are extracted from the experience replay pool to update the parameters of the intelligent decision-making model.

[0107] like Figure 3 As shown, the decision agent includes an online Q-Network, a target Q-Network, a loss function, and an experience replay pool.

[0108] The current state s (first workshop state information) of the aerospace thermal processing workshop is input into an online Q-Network. The online Q-Network evaluates the value of multiple possible actions corresponding to the first workshop state information s. Among the multiple possible actions, the action combination with the highest Q value is determined as the first control action information a, and the evaluation Q(s, a; θ) and periodic replication parameters are obtained. The online Q-Network can also input the periodic replication parameters into a target Q-Network. The evaluation Q(s, a; θ) includes the first workshop state information s, the first control action information a, and the current network weight θ of the online Q-Network.

[0109] The first control action information 'a' is evaluated using the target Q-Network, and the target Q-Network outputs the target Q-value, max. a′ Q(s', a'; θ') is the target Q value, which includes the second workshop state information s', the second control action information a', and the current network weight θ' of the target Q-Network.

[0110] The loss function is used to evaluate Q(s, a; θ) and the target max. a′ Q(s', a'; θ') is calculated to update the network parameters of the online Q-Network.

[0111] After executing the production decision plan in the hot processing workshop, the reward value r corresponding to the first control action information a is obtained through the reward function, and the second production data of the hot processing workshop is collected. The second production data is abstracted in the state space to obtain the second workshop state information s' of the hot processing workshop. The combination (s, a, r, s') of the first workshop state information s, the first control action information a, the reward value r corresponding to the first control action information a, and the second workshop state information s' is used as training samples and stored in the experience playback pool.

[0112] Input (s, a) from the experience replay pool into the online Q-Network, and input s' from the experience replay pool into the target Q-Network.

[0113] The steps for training and applying the intelligent decision-making model are shown in Table 1.

[0114] Table 1. Steps for training and applying intelligent decision-making models

[0115]

[0116]

[0117] Initialize the online Q-Network and the target Q-Network, store training samples using the experience replay buffer, improve training efficiency and stability through random sampling, and then reduce the bias of Q-value estimation and improve training stability through the DDQN method; train in a simulated environment, collect samples of state, action, reward and next state, update network parameters, etc., and the model training is completed.

[0118] Double DQN is used to reduce the bias in Q-value estimation. A fixed target network is introduced, and its parameters are updated periodically to stabilize the training process.

[0119] The output of the online Q-network is a vector of length equal to the size of the action space. For each state s, the online Q-network will output every possible combination of actions a. i Q value Q(s,a) i ), select the action combination with the highest Q value (argmax) ai Q(s,a i DDQN uses an experience replay buffer to store the agent's experience, i.e., samples (s, a, r, s′) of states, actions, rewards, and the next state. DDQN uses the target network to compute the Q-value: y = R + γQ. target (s′,argmax a′ Q(a,s′)) is updated using the minimum loss function:

[0120] In this embodiment, the system can respond quickly to dynamic events in the aerospace product thermal processing workshop, perceive the workshop status in real time, make intelligent decisions on workshop actions based on status characteristics, and make real-time dynamic feedback adjustments based on the workshop actions after the decisions, thereby ensuring processing quality and improving production efficiency. This system has practical application value for the development of a production process control system for aerospace product thermal processing workshops, and can achieve efficient decision-making and intelligent control.

[0121] In some embodiments, the reward value is obtained based on a reward function R, which is:

[0122] R=ω1·R geo +ω2·R mech +ω3·R def +ω4·R eff -P viol

[0123] Among them, R geo R represents the bonus value for the size and shape deviation of the target workpiece. mech R represents the bonus value for the mechanical performance deviation of the target workpiece. def R represents the reward for defect detection of the target workpiece. eff P represents the reward for the production efficiency of the target workpiece. viol This indicates that actions in the first control action information a that may lead to quality problems will be given a negative reward, and ω1, ω2, ω3, and ω4 are weight parameters that characterize relative importance.

[0124] In practice, the reward function is designed as a comprehensive reward function that balances quality indicators and efficiency.

[0125] Among them, quality indicators include geometric accuracy, mechanical properties, material uniformity and defect detection, etc. Efficiency balance is to introduce rewards related to order delivery time while ensuring quality, so as to ensure that urgent orders are completed on time, while taking into account production efficiency, reducing production time and energy consumption.

[0126] R=ω1·R geo +ω2·R mech +ω3·R def +ω4·R eff -P viol

[0127] In the formula, R geo R represents the bonus value for the size and shape deviation of the part. mech R represents the bonus value for mechanical performance deviation. def R represents the reward for defect detection. eff Rewards representing productivity, while ensuring quality, take productivity into account, such as reducing time and energy consumption, etc. P viol This indicates that actions that violate quality standards will be punished, while actions that may lead to quality problems will be negatively rewarded, such as excessively high or low temperatures or pressures exceeding safe limits. ω1, ω2, ω3, and ω4 are weighting parameters that adjust the relative importance of each part.

[0128] In this embodiment, by setting a reward function, the learning behavior of the intelligent decision-making model can be effectively guided to make better decisions in the complex environment of the hot processing workshop, thereby improving the overall performance of the intelligent decision-making model.

[0129] After training the model in a simulated environment, it is integrated into the existing MES (Manufacturing Execution System) in the control workshop for communication and application. This invention can ensure the quality of aerospace thermal processing products and improve production efficiency by dynamically adjusting production parameters in real time.

[0130] Before testing the model in the actual aerospace product thermal processing workshop, the algorithm model is trained, and a communication protocol interface is built between the algorithm and the MES system of the control workshop. This interface can read the status of the thermal processing workshop and issue action commands to ensure that it can cope with dynamic events such as emergency order insertion and complete offline verification.

[0131] The algorithm performance is tested using standard sets MK01 to MK10. Since these test sets typically do not include quality factors, in this embodiment, the machine-ordered individuals in the state vector space and action vector space are retained, while other individuals are set to 0. The first workshop state information is represented as a state space vector, and the first control action information is represented as an action space vector. The state space vector s and action space vector a are:

[0132] s = [0,0,0,0,W,0,0]

[0133] a = [0,0,0,0,a] prod_order ,0,0]

[0134] The results are shown in Table 2, demonstrating that the algorithm of this invention is superior to other algorithms. For example, the optimal completion time is better than that of the genetic algorithm in all 10 problem cases.

[0135] Table 2 shows the minimum completion time for different algorithms on the standard test set.

[0136]

[0137]

[0138] Furthermore, according to the statistics of the factory's existing manual event management methods, as shown in Table 3, manual scheduling and process design usually take 2 to 3 days, emergency orders require about 6 hours to reschedule, and manual scheduling during fault handling takes 0.5 hours. However, the intelligent management method based on the present invention can significantly improve management efficiency.

[0139] Table 3. Efficiency Comparison of Different Control Methods in the Dynamic Management of Aerospace Product Thermal Processing Workshop

[0140] event Manual control Intelligent control based on deep reinforcement learning Existing production schedule and process design 2-3 days 0.5 to 1 hour Urgent orders arrive and production is rescheduled 6 hours 1 to 5 minutes Scheduling during fault handling 0.5 to 1 hour 5-10 minutes

[0141] In this embodiment, by precisely controlling the actions, the microstructure during the forging process is optimized, reducing production defects and improving the mechanical properties and geometric accuracy of the parts. By introducing an emergency order insertion mechanism, changes in production demand are quickly responded to, the production sequence is dynamically adjusted, order waiting time is reduced, production bottlenecks and resource waste are minimized, and the production system's responsiveness to emergency orders and unforeseen circumstances is enhanced, ensuring timely delivery. This improves overall production efficiency, reduces energy consumption, and achieves green manufacturing. Optimizing production efficiency while maintaining high quality enhances the company's competitiveness.

[0142] The dynamic decision-making method for aerospace product thermal processing workshops provided in this application can be executed by a dynamic decision-making device for aerospace product thermal processing workshops. This application uses the example of a dynamic decision-making device for aerospace product thermal processing workshops executing the dynamic decision-making method for aerospace product thermal processing workshops to illustrate the dynamic decision-making device for aerospace product thermal processing workshops provided in this application.

[0143] This application also provides a dynamic decision-making device for aerospace product thermal processing workshops.

[0144] like Figure 4 As shown, the dynamic decision-making device for the aerospace product thermal processing workshop includes:

[0145] The acquisition module 410 is used to acquire the first production data during the production of the target workpiece in the hot working workshop;

[0146] The first processing module 420 is used to abstract the first production data in the state space to obtain the first workshop state information of the heat processing workshop;

[0147] The second processing module 430 is used to input the first workshop status information into the intelligent decision model to obtain the first control action information of the action space output by the intelligent decision model, wherein the intelligent decision model is constructed based on a neural network model;

[0148] The third processing module 440 is used to generate a production decision plan for the hot processing workshop to produce the target workpiece based on the first control action information.

[0149] According to the embodiment of this application, the dynamic decision-making device for aerospace product thermal processing workshops abstracts the first production data monitored in real time to obtain the first workshop status information as input to the intelligent decision-making model, thereby obtaining a decision-making production plan that dynamically adjusts production parameters. This ensures that key parameters such as temperature, pressure, stress, and strain during the production process are always kept within the optimal range. During the production process of aerospace product thermal processing workshops, the device dynamically identifies the multimodal state of the workshop and achieves efficient decision-making.

[0150] The dynamic decision-making device for the aerospace product thermal processing workshop in this application embodiment can be an electronic device or a component of an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a tablet computer, laptop computer, handheld computer, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc., and this application embodiment does not specifically limit it.

[0151] The dynamic decision-making device for the aerospace product thermal processing workshop in this embodiment can be a device with an operating system. This operating system can be Linux, Windows, or other possible operating systems; this embodiment does not specifically limit its use.

[0152] The dynamic decision-making device for the thermal processing workshop of aerospace products provided in this application embodiment can realize all the processes implemented in the above-described embodiments of the dynamic decision-making method for the thermal processing workshop of aerospace products. To avoid repetition, these processes will not be described again here.

[0153] In some embodiments, such as Figure 5 As shown, this application embodiment also provides an electronic device 500, including a processor 501, a memory 502, and a computer program stored in the memory 502 and executable on the processor 501. When the program is executed by the processor 501, it implements the various processes of the above-described embodiment of the dynamic decision-making method for the thermal processing workshop of aerospace products and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0154] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0155] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described embodiment of the dynamic decision-making method for the thermal processing workshop of aerospace products and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0156] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0157] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned dynamic decision-making method for the thermal processing workshop of aerospace products.

[0158] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0159] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described embodiment of the dynamic decision-making method for the thermal processing workshop of aerospace products, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0160] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0161] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0162] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the dynamic decision-making method for the aerospace product thermal processing workshop of the various embodiments of this application.

[0163] In the description of this application, "first feature" and "second feature" may include one or more of the features.

[0164] In the description of this application, "multiple" means two or more.

[0165] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0166] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0167] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.

Claims

1. A dynamic decision-making method for aerospace product thermal processing workshops, characterized in that, include: Obtain the first production data during the production of the target workpiece in the heat treatment workshop; The first production data is abstracted in the state space to obtain the first workshop state information of the heat processing workshop; the first workshop state information includes the product processing quality of the target workpiece in the production process, order status information, machine utilization rate and overall machine energy consumption average value. The intelligent decision-making model is built on a neural network model. The intelligent decision-making model includes an online Q-network, a target Q-network, a loss function, and an experience replay pool. The first workshop state information is input into the intelligent decision-making model to obtain the first control action information of the action space output by the intelligent decision-making model. The first workshop state information is input into the online Q-network, and the online Q-network evaluates the value of multiple possible actions corresponding to the first workshop state information. Among the multiple possible actions, the action combination with the highest Q value is determined as the first control action information. Based on the first control action information, a production decision plan for producing the target workpiece in the heat treatment workshop is generated.

2. The dynamic decision-making method for aerospace product thermal processing workshops according to claim 1, characterized in that, The expression for the product processing quality is: Wherein, Q(t) represents the product processing quality; α, β, γ, δ, and ζ are the weights of the state parameters that determine the quality; T represents the state parameter corresponding to the forging temperature; P represents the state parameter corresponding to the forging pressure; S micro-structure The state parameters corresponding to the microstructure; G represents the state parameter corresponding to strain; G represents the state parameter corresponding to geometric dimensions.

3. The dynamic decision-making method for aerospace product thermal processing workshops according to claim 1, characterized in that, The order status information is represented as a set of current orders and urgent orders for the target workpiece: W={(p i ,d i ,w i )|i=1,2,3…,n}∪{(p E ,d E ,w E )} Where W represents the set of current orders and urgent orders; i is the order number, p i Indicates the processing time of order i; d i Indicates the delivery deadline for order i; ω i p represents the weight of order i; E Indicates the processing time for urgent orders; d E Indicates the delivery deadline for urgent orders; ω E Indicates the weight of urgent orders; ω E Greater than ω i .

4. The dynamic decision-making method for aerospace product thermal processing workshops according to claim 1, characterized in that, The neural network model includes an online Q-network and a target Q-network. The online Q-network is used to predict and output the first control action information based on the first workshop state information. The online Q-network is used to update the parameters of the target Q-network. The target Q-network is used to evaluate the value of the first control action information to obtain the target Q-value of the first control action information. The target Q-value is used to update the online Q-network.

5. The dynamic decision-making method for aerospace product thermal processing workshops according to claim 4, characterized in that, The online Q-network is used to predict and output the first control action information based on the first workshop status information, including: The first workshop status information is input into the online Q network. The online Q network is used to evaluate the value of multiple possible actions corresponding to the first workshop status information. Among the multiple possible actions, the action combination with the highest Q value is determined as the first control action information.

6. The dynamic decision-making method for aerospace product thermal processing workshops according to claim 1, characterized in that, After generating a production decision plan for the target workpiece in the heat treatment workshop based on the first control action information, the method further includes: When the production decision plan is executed in the heat treatment workshop, the second production data corresponding to the production decision plan is obtained; The second production data is abstracted in the state space to obtain the second workshop state information of the heat treatment workshop; The combination of the first workshop status information, the first control action information, the reward value corresponding to the first control action information, and the second workshop status information is used as a training sample, and the training sample is stored in the experience replay pool. The training samples are extracted from the experience replay pool to update the parameters of the intelligent decision-making model.

7. The dynamic decision-making method for aerospace product thermal processing workshops according to claim 6, characterized in that, The reward value is obtained based on a reward function, R, which is: R=ω1·R geo +ω2·R mech +ω3·R def +ω4·R eff -P viol Among them, R geo R represents the bonus value for the size and shape deviation of the target workpiece. mech R represents the bonus value for the mechanical performance deviation of the target workpiece. def R represents the reward for defect detection of the target workpiece. eff P represents the reward for the production efficiency of the target workpiece. viol This indicates that actions in the first control action information that may lead to quality problems will be given a negative reward, and ω1, ω2, ω3, and ω4 are weight parameters that characterize the relative importance.

8. A dynamic decision-making device for a thermal processing workshop of aerospace products, characterized in that, include: The acquisition module is used to acquire the first production data during the production of the target workpiece in the hot working workshop; The first processing module is used to abstract the first production data in the state space to obtain the first workshop state information of the heat processing workshop; The second processing module constructs an intelligent decision-making model based on a neural network. The intelligent decision-making model includes an online Q-network, a target Q-network, a loss function, and an experience replay pool. The first workshop state information is input into the intelligent decision-making model to obtain the first control action information of the action space output by the intelligent decision-making model. That is, the first workshop state information is input into the online Q network, and the value of multiple possible actions corresponding to the first workshop state information is evaluated by the online Q network. Among the multiple possible actions, the action combination with the highest Q value is determined as the first control action information. The third processing module is used to generate a production decision plan for the hot processing workshop to produce the target workpiece based on the first control action information.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the dynamic decision-making method for the thermal processing workshop of aerospace products as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Self-adaptive scheduling method for intelligent workshop

    CN111199272A

  • Process parameter optimization method based on reinforcement learning

    CN116048028A