Equipment parameter error simulation and optimization method, device, equipment and medium

By combining equipment parameter error simulation and optimization methods of LSTM and DDQN, digital twin technology is used to achieve high-precision simulation and real-time optimization of equipment errors, solving the problem of insufficient time dependence and real-time in the existing technology, improving the consistency of manufacturing quality, and is suitable for high-precision manufacturing, smart factories and aerospace fields.

CN120493491APending Publication Date: 2025-08-15CHINA ORDNANCE EQUIP GRP AUTOMATION RES INST CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510486732.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the field of intelligent manufacturing and industrial Internet of Things, the simulation and optimization methods of equipment errors have problems such as insufficient time-dependent processing, poor real-time performance, limited computing resources and lack of online learning capabilities, and it is difficult to adapt to dynamic changes in the manufacturing environment, resulting in insufficient error prediction accuracy and untimely optimization of process parameters.

Method used

The equipment parameter error simulation and optimization method based on Markov decision-making process is adopted, combined with the time dependence of LSTM network capture errors, online parameter optimization is used using DDQN, and the model is updated in real time and parallel computing is realized through digital twin technology, supporting efficient learning of edge devices.

Benefits of technology

It realizes high-precision simulation and real-time optimization of manufacturing equipment errors, improves manufacturing quality consistency, adapts to dynamic changes in the manufacturing environment, solves the shortcomings of traditional methods in time dependence and real-time, and has important theoretical value and practical application significance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493491A_ABST
    Figure CN120493491A_ABST
Patent Text Reader

Abstract

The invention discloses an equipment parameter error simulation and optimization method, device, equipment and medium, and relates to the technical field of intelligent manufacturing and industrial Internet of Things, and the method realizes high-precision simulation and real-time optimization of manufacturing equipment errors by combining DDQN and LSTM. The LSTM layer effectively captures the dynamic characteristics of errors changing along with time, and the defect of time-dependent processing is overcome; the DDQN realizes online optimization of error prediction network parameters through cooperative work of an evaluation network and a target network, solves the problem of insufficient adaptability of a pre-training model to specific manufacturing equipment, and improves the error prediction accuracy; the integration of the digital twinning technology supports online learning and parallel model retraining, solves the problem of calculation power limitation of edge equipment, dynamically adapts to manufacturing environment changes, remarkably improves the consistency of manufacturing quality, can be widely applied to the fields of high-precision manufacturing, intelligent factories, aerospace and the like, and provides a brand new solution for error prediction and process optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent manufacturing and industrial Internet of Things, and in particular to a method, device, equipment and medium for simulating and optimizing equipment parameter errors based on reinforcement learning and LSTM. Background Art

[0002] In the fields of intelligent manufacturing and the Industrial Internet of Things (IIoT), the simulation and optimization of equipment errors are core technologies for improving manufacturing precision and efficiency. Error prediction and optimization methods based on traditional machine learning or deep learning have been widely used. For example, neural network-based error prediction methods, by fitting the relationship between manufacturing errors and process parameters, can predict errors to a certain extent and recommend parameter adjustment solutions.

[0003] However, existing technologies face key challenges in practical applications: First, manufacturing errors often change dynamically over time. In particular, factors such as equipment aging and wear can significantly affect the accumulation of errors. Traditional neural networks (such as MLP or CNN) cannot effectively capture this time dependence, resulting in insufficient error prediction accuracy; second, existing methods mainly rely on offline training and static models, which are difficult to adapt to the dynamic changes of the manufacturing environment and cannot achieve real-time process parameter optimization; in addition, in complex manufacturing scenarios, the sources of errors are diverse and mutually coupled. It is difficult to optimize multiple parameters simultaneously using simple mathematical models and machine learning methods. The manufacturing system itself has limited computing power resources and cannot support deep learning network calculations of large-scale parameters; finally, existing technologies lack efficient online learning mechanisms and cannot dynamically update models based on real-time data, making it difficult to meet the real-time requirements of the manufacturing process.

[0004] The patent document with patent number CN202210297055.8 in the prior art discloses "A method for sensitivity analysis of manufacturing errors of cable net antennas based on a proxy model". This invention provides a method for sensitivity analysis of manufacturing errors based on a Kriging proxy model, including: classifying and merging manufacturing error variables according to structural symmetry characteristics, performing initial sampling of manufacturing errors and calculating the corresponding antenna accuracy, constructing a Kriging proxy model based on the current sample points to simulate the functional relationship between manufacturing errors and antenna accuracy, calculating the sensitivity index of each type of error factor using a global sensitivity analysis strategy, and adding new sample points through a dynamic point addition strategy to update the proxy model, and finally judging whether the analysis process is complete by using a weighted relative error convergence criterion. This invention proposes a mechanism for online insertion, updating, and convergence judgment, and because the sensitivity index of each type of error factor is calculated independently, interference from non-important factors can be avoided.

[0005] However, this invention does not take into account the dynamic characteristics of manufacturing errors over time (such as equipment aging, wear, etc.), lacks real-time optimization and online learning capabilities, and the construction and updating of the Kriging agent model requires a large amount of computing resources, making it difficult to run efficiently on manufacturing equipment with limited computing power. Summary of the Invention

[0006] In view of the above problems, the present invention provides a method, device, apparatus, and medium for simulating and optimizing equipment parameter errors to overcome or at least partially solve the above problems. The method solves the problem of dynamic prediction of equipment errors and process parameter optimization in the manufacturing process.

[0007] The present invention provides the following solutions:

[0008] A method for simulating and optimizing equipment parameter errors, comprising:

[0009] The equipment error simulation and parameter optimization process is modeled as a Markov decision process, and an error prediction action benefit calculation module and a parameter optimization benefit calculation module are determined. The error prediction action benefit calculation module combines the absolute value of the error and the threshold constraint as a penalty term to evaluate the accuracy of the error prediction. The parameter optimization benefit calculation module combines the mean square error between the actual parameters and the ideal parameters to evaluate the actual application effect of the parameter optimization and neural network model.

[0010] Obtain historical manufacturing error data and ideal manufacturing parameters of equipment;

[0011] Inputting the historical manufacturing error data into a pre-trained equipment error prediction and parameter optimization model, the equipment error prediction and parameter optimization model comprising an error prediction network and a parameter optimization network; the error prediction network is configured to calculate and output a predicted error vector using the input historical manufacturing error data; and the parameter optimization network is configured to calculate and output optimized manufacturing parameters using the input error vector and the ideal manufacturing parameters;

[0012] Setting the current manufacturing parameters of the manufacturing equipment to the optimized manufacturing parameters through the digital twin system and monitoring the actual manufacturing parameters of the manufacturing equipment collected by the sensor;

[0013] Utilizing the error prediction action benefit calculation module and the parameter optimization benefit calculation module in combination with the actual manufacturing parameters to calculate and obtain the error prediction benefit and the parameter optimization benefit;

[0014] The error prediction benefit and the historical manufacturing error data are aggregated using an information fusion module to obtain updated historical error data; the updated historical error data, the error vector, and samples from the error prediction network pre-training process are stored in an error prediction experience buffer; the parameter optimization benefit, the optimized manufacturing parameters, and samples from the parameter optimization network pre-training process are stored in a parameter optimization experience buffer;

[0015] Determining that the current error prediction network has performed a target number of error predictions, then determining to use the data in the error prediction experience buffer to online update the network parameters of the error prediction network using a dual deep reinforcement learning network;

[0016] If it is determined that the number of parameter optimizations currently performed by the parameter optimization network is greater than the target number, the equipment error prediction and parameter optimization model is retrained and updated using the data model retraining tool in the parameter optimization experience buffer.

[0017] Preferably, the error prediction action benefit calculation module is represented by the following formula:

[0018]

[0019] Where: r t,1 represents the error prediction return, represents the error of the nth manufacturing parameter in the tth manufacturing, represents the error prediction result of the nth manufacturing parameter in the tth manufacturing, α1 and β1 represent the prediction error penalty factors, The threshold value indicating the tolerable deviation between the prediction error and the actual error;

[0020] The parameter optimization benefit calculation module is represented by the following formula:

[0021]

[0022] Where: r t,2 represents the error prediction return, represents the nth actual manufacturing parameter actually monitored in the tth manufacturing, represents the nth ideal manufacturing parameter given in the tth manufacturing.

[0023] Preferably, the error prediction network includes two long short-term memory network layers and two fully connected layers; the long short-term memory network layers are used to capture the time dependency of manufacturing errors, and the fully connected layers are used to fuse the coupled error relationships of multiple parameters.

[0024] Preferably: the parameter optimization network includes an input layer, a hidden layer and an output layer, the number of neurons in the input layer is equal to the dimension of the manufacturing parameters; the hidden layer includes multiple fully connected layers and a Dropout layer, the fully connected layer uses ReLU=max(·,0) as the activation function, and the Dropout layer is used to prevent overfitting; the number of neurons in the output layer is equal to the dimension of the optimal setting parameters, and a linear activation function is used to obtain the regression analysis results of the optimized parameters.

[0025] Preferably, the equipment error prediction and parameter optimization model adopts a staged training process of first pre-training the hyperparameters of the parameter optimization network based on random search, and then pre-training the error prediction network using a dual deep reinforcement learning method.

[0026] Preferably, the pre-training method of the parameter optimization network includes:

[0027] Step 1: Define the number of hidden layers, number of neurons, Dropout ratio, learning rate, and optimizer value range;

[0028] Step 2: Set the number of random search iterations, the value range of which is 50-100 times;

[0029] Step 3: Randomly sample hyperparameter combinations. In each iteration, randomly sample a set of hyperparameter values from the hyperparameter space.

[0030] Step 4: Train the model using the sampled hyperparameter combination and evaluate the loss, accuracy, and F1 score on the validation set.

[0031] Step 5: After all iterations are completed, select the hyperparameter combination with the best performance on the validation set as the final result.

[0032] Preferably, the pre-training method of the error prediction network includes:

[0033] Step 1: Randomly sample x training Mini-batchE from the training set of sample set ε1 x sample;

[0034] Step 2: From Mini-batchE x Select sample e t =(s t ,a t ,r t,1 ,s t+1 );

[0035] Step 3: Use the Q learning policy network to calculate the Q value Q(s) of the current state t ,a t :θ), and the next state s t+1 All actions in at+1 The Q value Q(s t+1 ,a t+1 :θ), select the next action with the maximum Q value

[0036] Step 4: Calculate the next state s t+1 The target network Q value

[0037] Step 5: Calculate learning objectives γ is the discount factor;

[0038] Step 6: Calculate the loss function L(θ)=(yQ(s t ,a t :θ)) 2 , obtain the gradient through back propagation of the policy network

[0039] Step 7: Update the policy network parameters using gradient descent Where ρ is the learning rate;

[0040] Step 8: Update the target network parameters θ with weighted hyperparameters - ←τθ+(1-τ)θ - ;

[0041] Step 9: Repeat steps 2 to 8 until Mini-batch E is traversed x If all samples in the training set are included, then one round of training is completed and the average error prediction gain is calculated on the validation set;

[0042] Step 10: Adjust the learning rate ρ and the number of batch training samples x, and return to step 1;

[0043] Step 11: In all training rounds, the policy network parameters with the best performance are selected as the pre-trained error prediction network parameters.

[0044] An equipment parameter error simulation and optimization device, used to perform the above-mentioned equipment parameter error simulation and optimization method, the device comprising:

[0045] A simulation and optimization modeling unit is used to model the equipment error simulation and parameter optimization process as a Markov decision process, and determine an error prediction action benefit calculation module and a parameter optimization benefit calculation module; the error prediction action benefit calculation module combines the absolute value of the error and the threshold constraint as a penalty term to evaluate the accuracy of the error prediction; the parameter optimization benefit calculation module combines the mean square error of the actual parameters and the ideal parameters to evaluate the actual application effect of the parameter optimization and neural network model;

[0046] A data acquisition unit, used to obtain historical manufacturing error data and ideal manufacturing parameters of the equipment;

[0047] a prediction and optimization unit, configured to input the historical manufacturing error data into a pre-trained equipment error prediction and parameter optimization model, wherein the equipment error prediction and parameter optimization model includes an error prediction network and a parameter optimization network; the error prediction network is configured to calculate and output a predicted error vector using the input historical manufacturing error data; and the parameter optimization network is configured to calculate and output optimized manufacturing parameters using the input error vector and the ideal manufacturing parameters;

[0048] An actual manufacturing parameter acquisition unit, configured to set the current manufacturing parameters of the manufacturing equipment to the optimized manufacturing parameters through the digital twin system and to monitor and acquire the actual manufacturing parameters of the manufacturing equipment acquired by the sensor;

[0049] A benefit calculation unit, configured to calculate an error prediction benefit and a parameter optimization benefit by using the error prediction action benefit calculation module and the parameter optimization benefit calculation module in combination with the actual manufacturing parameters;

[0050] An experience cache unit is configured to aggregate the error prediction benefit and the historical manufacturing error data using an information fusion module to obtain updated historical error data; store the updated historical error data, the error vector, and samples from the error prediction network pre-training process into an error prediction experience cache; and store the parameter optimization benefit, the optimized manufacturing parameters, and samples from the parameter optimization network pre-training process into a parameter optimization experience cache;

[0051] An experience replay unit, configured to determine whether the current error prediction network has performed a target number of error predictions, and then determine whether to use the data in the error prediction experience buffer to perform online updating of network parameters of the error prediction network using a dual deep reinforcement learning network;

[0052] The periodic retraining unit is used to determine that the number of parameter optimizations currently performed by the parameter optimization network is greater than the target number, and then use the data model retraining tool in the parameter optimization experience buffer area to retrain and update the equipment error prediction and parameter optimization model.

[0053] An equipment parameter error simulation and optimization device, the device comprising a processor and a memory:

[0054] The memory is used to store program code and transmit the program code to the processor;

[0055] The processor is used to execute the above-mentioned equipment parameter error simulation and optimization method according to the instructions in the program code.

[0056] A computer-readable storage medium is used to store program code, and the program code is used to execute the above-mentioned equipment parameter error simulation and optimization method.

[0057] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0058] The embodiments of the present application provide a method, device, equipment, and medium for simulating and optimizing equipment parameter errors. This method achieves high-precision simulation and real-time optimization of manufacturing equipment errors by combining DDQN and LSTM. The LSTM layer effectively captures the dynamic characteristics of errors changing over time, solving the shortcomings of time-dependent processing; DDQN achieves online optimization of error prediction network parameters through the collaborative work of the evaluation network and the target network, solving the problem of insufficient adaptability of pre-trained models to specific manufacturing equipment and improving error prediction accuracy; the integration of digital twin technology supports online learning and parallel model retraining, solves the problem of edge device computing power limitations, dynamically adapts to changes in the manufacturing environment, and significantly improves manufacturing quality consistency. It can be widely used in high-precision manufacturing, smart factories, aerospace, and other fields, providing a new solution for error prediction and process optimization, with important theoretical value and practical application significance.

[0059] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be derived from these drawings without inventive effort.

[0061] Figure 1 This is a flow chart of a method for simulating and optimizing equipment parameter errors provided by an embodiment of the present invention;

[0062] Figure 2 This is a structural diagram of a neural network model for equipment error simulation and optimization based on LSTM provided in an embodiment of the present invention;

[0063] Figure 3 Schematic diagram of the error prediction network pre-training process based on DDQN provided by an embodiment of the present invention;

[0064] Figure 4This is a flowchart of equipment error prediction and parameter optimization model execution and online learning in the digital twin system provided by an embodiment of the present invention;

[0065] Figure 5 This is a structural diagram of a digital twin model of a robotic arm for an assembly process provided by an embodiment of the present invention;

[0066] Figure 6 Schematic diagram of robot arm error prediction and parameter optimization recommendation results provided by an embodiment of the present invention;

[0067] Figure 7 Schematic diagram of an equipment parameter error simulation and optimization device provided by an embodiment of the present invention;

[0068] Figure 8 It is a schematic diagram of an equipment parameter error simulation and optimization device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0069] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of the present invention.

[0070] See also Figure 1 , is a method for simulating and optimizing equipment parameter errors provided by an embodiment of the present invention, such as Figure 1 As shown, the method may include:

[0071] S101: Model the equipment error simulation and parameter optimization process as a Markov decision process, and determine an error prediction action benefit calculation module and a parameter optimization benefit calculation module; the error prediction action benefit calculation module combines the absolute value of the error and the threshold constraint as a penalty term to evaluate the accuracy of the error prediction; the parameter optimization benefit calculation module combines the mean square error of the actual parameters and the ideal parameters to evaluate the actual application effect of the parameter optimization and neural network model; in specific implementation, the embodiment of the present application can provide that the error prediction action benefit calculation module is represented by the following formula:

[0072]

[0073] Where: r t,1 represents the error prediction return, represents the error of the nth manufacturing parameter in the tth manufacturing, represents the error prediction result of the nth manufacturing parameter in the tth manufacturing, α1 and β1 represent the prediction error penalty factors, The threshold value indicating the tolerable deviation between the prediction error and the actual error;

[0074] The parameter optimization benefit calculation module is represented by the following formula:

[0075]

[0076] Where: r t,2 represents the error prediction return, represents the nth actual manufacturing parameter actually monitored in the tth manufacturing, represents the nth ideal manufacturing parameter given in the tth manufacturing.

[0077] S102: Obtain historical manufacturing error data and ideal manufacturing parameters of the equipment;

[0078] S103: Input the historical manufacturing error data into a pre-trained equipment error prediction and parameter optimization model, wherein the equipment error prediction and parameter optimization model includes an error prediction network and a parameter optimization network; the error prediction network is used to calculate and output a predicted error vector using the input historical manufacturing error data; the parameter optimization network is used to calculate and output optimized manufacturing parameters using the input error vector and the ideal manufacturing parameters; in specific implementation, the embodiment of the present application can provide that the error prediction network includes two layers of long short-term memory network layers and two layers of fully connected layers; the long short-term memory network layer is used to capture the time dependence of the manufacturing error, and the fully connected layer is used to fuse the coupled error relationship of multiple parameters.

[0079] The parameter optimization network includes an input layer, a hidden layer, and an output layer. The number of neurons in the input layer is equal to the dimension of the manufacturing parameters. The hidden layer includes multiple fully connected layers and a Dropout layer. The fully connected layer uses ReLU=max(·,0) as the activation function, and the Dropout layer is used to prevent overfitting. The number of neurons in the output layer is equal to the dimension of the optimal setting parameters, and a linear activation function is used to obtain the regression analysis results of the optimized parameters.

[0080] The equipment error prediction and parameter optimization model adopts a staged training process of first pre-training the hyperparameters of the parameter optimization network based on random search, and then pre-training the error prediction network using a dual deep reinforcement learning method.

[0081] The pre-training method of the parameter optimization network includes:

[0082] Step 1: Define the number of hidden layers, number of neurons, Dropout ratio, learning rate, and optimizer value range;

[0083] Step 2: Set the number of random search iterations, the value range of which is 50-100 times;

[0084] Step 3: Randomly sample hyperparameter combinations. In each iteration, randomly sample a set of hyperparameter values from the hyperparameter space.

[0085] Step 4: Train the model using the sampled hyperparameter combination and evaluate the loss, accuracy, and F1 score on the validation set.

[0086] Step 5: After all iterations are completed, select the hyperparameter combination with the best performance on the validation set as the final result.

[0087] The pre-training method of the error prediction network includes:

[0088] Step 1: Randomly sample x training Mini-batchE from the training set of sample set ε1 x sample;

[0089] Step 2: From Mini-batchE x Select sample e t =(s t ,a t ,r t,1 ,s t+1 );

[0090] Step 3: Use the Q learning policy network to calculate the Q value Q(s) of the current state t ,a t :θ), and the next state s t+1 All actions in a t+1 The Q value Q(s t+1 ,a t+1 :θ), select the next action with the maximum Q value

[0091] Step 4: Calculate the next state s t+1 The target network Q value

[0092] Step 5: Calculate learning objectives γ is the discount factor;

[0093] Step 6: Calculate the loss function L(θ)=(yQ(s t ,a t :θ)) 2 , obtain the gradient through back propagation of the policy network

[0094] Step 7: Update the policy network parameters using gradient descent Where ρ is the learning rate;

[0095] Step 8: Update the target network parameters θ with weighted hyperparameters - ←τθ+(1-τ)θ- ;

[0096] Step 9: Repeat steps 2 to 8 until Mini-batch E is traversed x If all samples in the training set are included, then one round of training is completed and the average error prediction gain is calculated on the validation set;

[0097] Step 10: Adjust the learning rate ρ and the number of batch training samples x, and return to step 1;

[0098] Step 11: In all training rounds, the policy network parameters with the best performance are selected as the pre-trained error prediction network parameters.

[0099] S104: Setting the current manufacturing parameters of the manufacturing equipment to the optimized manufacturing parameters through the digital twin system and monitoring the actual manufacturing parameters of the manufacturing equipment collected by the sensor;

[0100] S105: Calculating error prediction benefit and parameter optimization benefit by using the error prediction action benefit calculation module and the parameter optimization benefit calculation module in combination with the actual manufacturing parameters;

[0101] S106: Aggregating the error prediction benefit and the historical manufacturing error data using an information fusion module to obtain updated historical error data; storing the updated historical error data, the error vector, and samples from the error prediction network pre-training process into an error prediction experience buffer; storing the parameter optimization benefit, the optimized manufacturing parameters, and samples from the parameter optimization network pre-training process into a parameter optimization experience buffer;

[0102] S107: determining that the current error prediction network has performed a target number of error predictions, then determining to use the data in the error prediction experience buffer to online update the network parameters of the error prediction network using a dual deep reinforcement learning network;

[0103] S108: If it is determined that the number of parameter optimizations currently performed by the parameter optimization network is greater than the target number, the equipment error prediction and parameter optimization model is retrained and updated using the data model retraining tool in the parameter optimization experience buffer.

[0104] In the specific implementation of playback and online update, the present invention provides an equipment error prediction and parameter optimization model execution and online learning method in a digital twin system, such as Figure 4 The specific description is as follows.

[0105] Step 1: Initialize model deployment. Deploy the pre-trained equipment error prediction and parameter optimization model (including the error prediction network and parameter optimization network) into the digital twin system of the manufacturing equipment.

[0106] Step 2: Initialize the equipment status. If the manufacturing equipment has a manufacturing record before being connected to the digital twin system, its historical manufacturing data record is stored in the experience buffer of the digital twin system, and the number of manufacturing times t = the number of historical manufacturing times + 1.

[0107] Step 3: Preprocessing of historical manufacturing data. In the t-th manufacturing, if t≤b, that is, the number of times the current manufacturing equipment is used is less than the number of historical manufacturing data required, then let s t =(0,0,0,…,s t-1,error ,s t-1,time ,s t,time ), that is, fill zeros in front of the historical manufacturing data sequence, so that s t If t>b, the historical manufacturing data meets the model input requirements and can be used directly.

[0108] Step 4: Error prediction and parameter optimization. The historical manufacturing error data s of this (tth) manufacturing t and the desired manufacturing parameters (ideal manufacturing parameters) t,ideal Input the equipment error prediction and parameter optimization model in the digital twin system, calculate through the neural network model, and output the error prediction result a t and the recommended manufacturing parameter setting vector o t,set .

[0109] Step 5: Manufacturing execution and monitoring. The digital twin system sets the manufacturing parameters of the actual manufacturing equipment to the actual manufacturing equipment through virtual-real interaction. t,set And monitor the actual manufacturing parameters of manufacturing equipment in real time t,real .

[0110] Step 6: Benefit calculation and experience caching. The benefit calculation module of the digital twin system obtains the error prediction benefit r through formula (1) and formula (2). t,1 And the parameter optimization benefit r t,2 , and summarize through the information fusion module to obtain the updated historical error data s t+1 , the sample e t =(s t ,a t ,r t,1 ,s t+1 ) and g t =(o t,ideal ,a t ,o t,set ,r t,2) are respectively stored in the error prediction experience cache E1 and the parameter optimization experience cache E2 of the digital twin system. The size of the experience cache E1 is the number of batch training samples x optimized during pre-training. It adopts an automatic queue cache method to automatically delete timed data. The experience cache E2 has no size limit and its size is controlled by manual cleaning (during regular retraining).

[0111] Step 7: Experience playback. x represents the number of batch training samples of the error prediction network, which is equivalent to the online learning cycle. That is, the current error prediction network has made x error predictions and needs experience playback to update the network parameters. Specifically, a sample e is randomly drawn from the experience buffer E1. t =(s t ,a t ,r t,1 ,s t+1 ), adopt the dual deep Q reinforcement learning method to update the error prediction network parameters, specifically refer to 3-2) step 3-8, until all experience samples in the experience buffer area E1 are traversed, and the error prediction network parameter update is completed.

[0112] Step 8: Retrain regularly. N1 represents the training cycle of the parameter optimization network. This means the current parameter optimization network has already undergone N1 parameter optimizations. Since the number of training samples required for machine learning models exceeds the number of samples required for reinforcement learning, N1>>x. Embed a parallel computing model training and packaging tool within the digital twin system. Divide the experience buffer E2 into a training set and a validation set at an 8:2 ratio. Retrain the parameter optimization network by referring to steps 1-5 in 3-1). Update the hyperparameters of the parameter optimization network through model releases, and clear the experience buffer E2 (retaining recent experience data in a certain ratio).

[0113] The equipment parameter error simulation and optimization method provided in the embodiments of the present application captures the time-varying characteristics of equipment errors through an LSTM network. This method can efficiently simulate the impact of factors such as equipment aging and wear on manufacturing accuracy. Combined with a manufacturing parameter optimization neural network, it optimizes process parameters in real time to achieve ideal manufacturing results.

[0114] Reinforcement learning and random search methods based on dual deep Q networks are used to achieve centralized pre-training and online learning of error prediction networks and parameter optimization networks to eliminate the influence of manufacturing equipment heterogeneity.

[0115] By leveraging the virtual-real synchronization and parallel computing technologies of digital twin systems to achieve regular parameter updates and synchronized execution of models, this approach can be widely applied in fields such as high-precision manufacturing, smart factories, aerospace, semiconductor, and medical device manufacturing, and is suitable for a variety of scenarios such as equipment fault diagnosis, health management, and error correction. Compared with traditional methods, this approach can effectively handle complex error sources, time dependencies, and real-time requirements. It enables online learning of virtual systems and virtual-real-prediction manufacturing system parameter optimization through digital twins, demonstrating significant practical value and innovation, providing a new solution for error control and process optimization in manufacturing equipment.

[0116] The following is a detailed introduction to the equipment parameter error simulation and optimization method provided in the embodiments of the present application.

[0117] The method provided in this application, by combining the time series modeling capabilities of the long short-term memory network (LSTM) and the online optimization capabilities of the dual deep Q network (DDQN) reinforcement learning, can effectively address the shortcomings of existing technologies in time-dependent processing, real-time optimization decision-making and online learning capabilities.

[0118] At the same time, considering the computing power limitations of edge manufacturing equipment, digital twin technology is further integrated into the system. As a virtual mirror, the digital twin system can synchronize the status of physical equipment in real time, use the actual monitored process parameters for online learning and regular retraining of neural network models, and share computing power pressure through cloud or edge computing resources. This ensures high-precision prediction of manufacturing errors and precise optimization of manufacturing parameters while ensuring computational efficiency. This proposed method can not only significantly improve the consistency of manufacturing quality, but also provide a new technical means for the field of intelligent manufacturing, with important theoretical value and practical application significance.

[0119] The embodiment of the present application provides an equipment error simulation and optimization method based on reinforcement learning and LSTM, which can be applied to error simulation prediction and control parameter optimization of various types of intelligent manufacturing, robot autonomous operation, equipment operation and maintenance, etc. The algorithm is applied to the digital twin system, and the parameter errors of the actual equipment are predicted through digital twin virtual simulation calculation to form an equipment error simulation and optimization system based on reinforcement learning and LSTM.

[0120] The embodiment of the present application provides a Markov-based equipment error simulation and optimization problem modeling method. Specifically, the equipment error simulation and parameter optimization process is modeled as a Markov decision process, represented as a four-tuple<S,A,r,γ> , where S represents the state space, A represents the action space, r represents the return, and γ represents the long-term discount factor.

[0121] Each manufacturing equipment is equivalent to an intelligent agent, which makes action decisions based on the real-time observed status data during the manufacturing process. Specifically, in the t-th manufacturing, the state information of the intelligent agent is:

[0122] s t =

[0123] [s t-b,error ,s t-b,time ,…,s t-b+i,error ,s t-b+i,time ,…,s t-1,error ,s t-1,time ,s t,time ]

[0124] s t Represents historical manufacturing error records, s t-b+i,error ,s t-b+i,time Parameter error and cumulative working time corresponding to the t-b+i-th manufacturing, b is the effective step length of historical data.

[0125] The error of the nth manufacturing parameter in the t-b+ith manufacturing (the average value of the difference between the real-time monitoring value and the set value during the manufacturing process); s t-b+i,time It represents the cumulative working time of the manufacturing equipment at the beginning of the t-b+i-th manufacturing (accumulated from the time the manufacturing equipment is put into use, t=0, s t,time =0).

[0126] The agent uses the state information s t , select action Indicates the prediction error (a positive number indicates that the actual manufacturing parameters are too large and the manufacturing parameter settings need to be reduced; a negative number indicates that the actual manufacturing parameters are too small and the manufacturing parameter settings need to be increased). t , based on ideal manufacturing parameters Adjusting manufacturing parameters Complete this manufacturing, based on the actual manufacturing parameters monitored Get the error prediction benefit r t,1 , expressed as:

[0127]

[0128] The former represents the sum of the difference between the prediction error and the actual error, the latter is to constrain the deviation of the prediction error not to exceed the threshold, α1 and β1 represent the prediction error penalty factors; and the parameter optimization benefit r t,2 , expressed as:

[0129]

[0130] is the mean square error between the actual manufacturing parameters and the ideal manufacturing parameters. Subsequently, the agent updates its state information s t+1 =[s t-b+1,error ,s t-b+1,time ,…,s t,error ,s t,time ,s t+1,time ].

[0131] In order to control manufacturing errors and improve equipment manufacturing quality, it is necessary to find the optimal error prediction strategy π1(s,a:θ1) and parameter optimization strategy π2(a,o:θ2) so that r t,1 and r t,2 Minimum. It is found that the manufacturing error prediction process is a typical Markov process, while the parameter optimization process is a multi-parameter linear optimization process. Therefore, it is necessary to design a two-layer policy network, which is divided into two parts: the error prediction policy network and the parameter optimization policy network. The former is updated online through reinforcement learning, and the latter is optimized through pre- / retraining.

[0132] The embodiment of the present application provides an LSTM-based equipment error simulation and optimization neural network model, the network model structure is as follows Figure 2 As shown. It includes two parts: an LSTM-based error prediction network and a parameter optimization network. The information fusion and benefit calculation complexity are low, and there is a clear calculation method. It uses online mathematical calculations and is not within the scope of neural network modeling. Specifically:

[0133] Error prediction network: It consists of two layers of LSTM (Long Short-Term Memory) and two fully connected layers. The LSTM layer is used to capture the time dependency of manufacturing errors. The number of LSTM units is determined by the complexity of the time series (16 / 32 / 64 / 128), and the time step is determined by the time span of the historical data (historical data effective step b). The fully connected layer is used to integrate the coupled error relationship of multiple parameters. The number of units is determined by the dimension of the manufacturing parameters. The input of the error prediction network is the historical error data s t , through neural network operation, the output is the predicted error vector a t .

[0134] Parameter Optimization Network: Consists of an input layer, a hidden layer, and an output layer. It is used to optimize the setting parameters of the manufacturing equipment based on the prediction error and the ideal manufacturing parameters, so that the error between the actual manufacturing parameters and the ideal manufacturing parameters is as small as possible. The number of neurons in the input layer is equal to the dimension of the manufacturing parameters; the hidden layer is composed of multiple fully connected layers and Dropout layers. The fully connected layer uses ReLU=max(·,0) as the activation function, and the Dropout layer is used to prevent overfitting; the number of neurons in the output layer is equal to the dimension of the optimal setting parameters, and a linear activation function is used to obtain the regression analysis results of the optimized parameters. The input of the parameter optimization network is the predicted error vector at and ideal manufacturing parameters o t,ideal , calculated through neural network, the output is the optimized manufacturing parameters o t,set .

[0135] The embodiment of the present application provides a segmented neural network pre-training method that combines reinforcement learning and machine learning, which pre-trains the error prediction network and the parameter optimization network respectively.

[0136] All historical manufacturing data of the same type of manufacturing equipment are collected to form sample sets ε1 and ε2, which correspond to error prediction sample e and parameter optimization sample e, respectively. They are divided into training set, test set and validation set in the ratio of 7:2:1.

[0137] First, the parameter optimization network is pre-trained based on the sample set ε2. The random search method is used on the training set to optimize the hyperparameters of the parameter optimization network (number of hidden layers, number of neurons, dropout ratio, learning rate and optimizer). The specific steps are as follows:

[0138] Step 1: Define the number of hidden layers, number of neurons, dropout ratio, learning rate, and optimizer value range.

[0139] Step 2: Set the number of random search iterations, which determines the breadth of the search and ranges from 50 to 100 times.

[0140] Step 3: Randomly sample hyperparameter combinations. In each iteration, a set of hyperparameter values is randomly sampled from the hyperparameter space.

[0141] Step 4: Train the model using the sampled hyperparameter combination and evaluate the loss, accuracy, and F1 score on the validation set.

[0142] Step 5: After all iterations are completed, select the hyperparameter combination with the best performance on the validation set as the final result.

[0143] Then, based on the trained parameter optimization network, the error prediction network is pre-trained on the sample set ε1, and the dual deep Q reinforcement learning method is used to construct the strategy network and the target network respectively. The structure is as follows Figure 3 The specific steps are as follows:

[0144] Step 1: Sample Mini-batch. Randomly sample a batch training Mini-batchE from the training set of sample set ε1 x (Contains x samples).

[0145] Step 2: Sample extraction. From Mini-batch E x Select sample e t =(s t ,at ,r t,1 ,s t+1 ).

[0146] Step 3: Calculate the Q value of the policy network. Use the Q learning policy network to calculate the Q value Q(s) of the current state t ,a t :θ), and the next state s t+1 All possible actions a t+1 The Q value Q(s t+1 ,a t+1 :θ), select the next action with the maximum Q value

[0147] Step 4: Calculate the target network Q value. Calculate the next state s t+1 The target network Q value

[0148] Step 5: Calculate learning objectives γ is the discount factor.

[0149] Step 6: Calculate the loss function L(θ)=(yQ(s t ,a t :θ)) 2 , obtain the gradient through back propagation of the policy network

[0150] Step 7: Update policy network parameters. Use gradient descent method to update policy network parameters. Where ρ is the learning rate (the learning rates of the error prediction network and the parameter optimization network are independent).

[0151] Step 8: Target network parameter update. Update the target network parameter θ with the weighted hyperparameter - ←τθ+(1-τ)θ - .

[0152] Step 9: Repeat steps 2 to 8 until Mini-batch E is traversed b If all samples in the training set are included, a round of training is completed and the average error prediction gain is calculated on the validation set.

[0153] Step 10: Adjust the learning rate ρ and the number of batch training samples x, and return to step 1.

[0154] Step 11: In all training rounds, select the policy network parameters with the best performance as the pre-trained error prediction network parameters, and the error prediction network pre-training is completed.

[0155] It can be seen that the equipment parameter error simulation and optimization method provided in the embodiment of the present application is based on the error prediction process modeling of manufacturing equipment based on the Markov decision process, with historical manufacturing error data as the state, the predicted error as the action, and the difference between the predicted error and the actual error as the benefit. Two benefit calculation methods are designed, including the error prediction benefit (Formula 1) - a penalty term that combines the absolute value of the error with the threshold constraint, which is used to evaluate the accuracy of the error prediction; and the parameter optimization benefit (Formula 2) - the mean square error between the actual parameter and the ideal parameter, which is used to evaluate the actual application effect of parameter optimization and the entire neural network model.

[0156] The error prediction network is composed of an LSTM network, and the parameter optimization network is composed of multiple fully connected layers (including dropout layers). The two layers work together to simulate and optimize the error. The training process is staged, first pre-training the parameter optimization network (using random search to optimize hyperparameters), then pre-training the error prediction network (based on DDQN).

[0157] A technical solution that achieves virtual-reality synchronization, online data collection, and dynamic model updates through a digital twin system. An experience replay mechanism is used to cache real-time manufacturing data and periodically update the error prediction network parameters based on the DDQN network. Periodic retraining supported by parallel computing is used to update the hyperparameters of the parameter optimization network. When the experience replay cycle overlaps with the retraining cycle, the parameter optimization network and the error prediction network are updated synchronously. In environments with limited computing power, this technology achieves efficient operation through phased model updates (high-frequency updates for the error prediction network and low-frequency updates for the parameter optimization network) and cache management (automatic cleanup of old data).

[0158] The specific application of the method provided in the embodiment of the present application is explained below using a robotic arm as an example.

[0159] The implementation object is a grabbing robot arm in the assembly process, and the parameters that need to be controlled and monitored are the pulling force, clamping force and material weight.

[0160] During operation, the robotic arm must adjust its lifting and clamping forces based on the weight of the material to prevent it from falling and deforming. Table 1 lists the types of materials that the robotic arm can grip.

[0161] The ideal clamping force is calculated as follows:

[0162] F 夹紧 =μ×m×g×α

[0163] μ is the friction coefficient, which depends on the smoothness of the clamped material, m is the mass of the material (in kg), and g is the acceleration due to gravity (9.81 m / s 2 ), α=1.5 is the safety factor.

[0164] The ideal pull-out force is calculated as follows:

[0165] F 拔起 =m×g×α

[0166] Table 1 Types of materials gripped by the robotic arm

[0167] Material Type Material A Material B Material C Material D Weight (g) 515±5 420±5 1018±7 826±5 Friction coefficient 0.3 0.25 0.56 0.37

[0168] The implementation process of the equipment error simulation and optimization system based on digital twins is as follows.

[0169] (1) Construct a digital twin model of the assembly process, including dynamic models, static models, behavioral models, rule models, and set models. The upper computer of the production line PLC control system collects the pulling force, clamping force (setting parameters) and material weight of the robot arm. By deploying micro pressure sensors, the pulling force and clamping force (actual parameters) actually applied by the robot arm are obtained. Digital twin models such as Figure 5 shown.

[0170] (2) The historical manufacturing data of the robotic arm is imported into the digital twin system, its error simulation and optimization neural network model is pre-trained, and the model retraining tool and the online learning DDQN network are deployed in parallel to support subsequent online updates.

[0171] (3) The robotic arm is started through the digital twin system, and the actual production line production status is monitored in real time. The current material weight is obtained as g = 515g, and the material type is A (friction coefficient is 0.3); the ideal clamping force is calculated to be 2.27N, and the ideal pulling force is calculated to be 7.58.

[0172] (4) Call the historical manufacturing data of the robot arm (historical clamping force error, pulling force error, continuous working time), input the historical manufacturing data and the ideal clamping force and pulling force into the error simulation and optimization neural network model, and output the error prediction results and parameter optimization recommendation results, such as Figure 6 shown.

[0173] (5) The digital twin system controls the PLC system to set the clamping force and pulling force of the robot arm, and the actual clamping force and pulling force are obtained through sensor feedback (the average value within the monitoring period, and the monitoring period in this embodiment is from the beginning to the end of a grasping operation). The deviations from the ideal value and the set value are calculated respectively. The deviations from the actual value and the ideal value correspond to the benefits of the parameter optimization network, and the deviations from the actual value and the set value correspond to the benefits of the error prediction network. The calculation method refers to formulas (1) and (2);

[0174] (6) The grasping data of the robot arm is processed and stored in the error prediction experience buffer and the parameter optimization experience buffer. When the number of data in the error prediction experience buffer is greater than 100, the earliest front-end sequence data is automatically cleared.

[0175] (7) Every 100 captures, the digital twin system calls the experience data in the error prediction experience buffer and performs online updates of the error prediction network parameters based on the DDQN network.

[0176] (8) Every 1,000 captures, the digital twin system calls the experience data in the parameter optimization experience cache and retrains the entire error simulation and optimization neural network model through the model retraining tool. The training process adopts a two-stage pre-training method (first optimize the parameter optimization network hyperparameters, and then optimize the error prediction network parameters). The error simulation and optimization neural network model is updated by publishing the retrained model.

[0177] It is understandable that since both online learning and retraining of the model use parallel computing, if a manufacturing task occurs during the calculation period, error prediction and parameter optimization are performed according to the current (unupdated) neural network model to avoid confusion in the model update process and affect manufacturing quality.

[0178] In summary, the equipment parameter error simulation and optimization method provided in this application achieves high-precision simulation and real-time optimization of manufacturing equipment errors by combining DDQN and LSTM. The LSTM layer effectively captures the dynamic characteristics of errors changing over time, solving the shortcomings of time-dependent processing; DDQN achieves online optimization of error prediction network parameters through the collaborative work of the evaluation network and the target network, solving the problem of insufficient adaptability of pre-trained models to specific manufacturing equipment and improving error prediction accuracy; the integration of digital twin technology supports online learning and parallel model retraining, solves the problem of edge device computing power limitations, dynamically adapts to changes in the manufacturing environment, and significantly improves manufacturing quality consistency. It can be widely used in high-precision manufacturing, smart factories, aerospace and other fields, providing a new solution for error prediction and process optimization, with important theoretical value and practical application significance.

[0179] See also Figure 7 , the embodiment of the present application can also provide an equipment parameter error simulation and optimization device, such as Figure 7 As shown, the device for executing the above-mentioned equipment parameter error simulation and optimization method may include:

[0180] The simulation optimization modeling unit 701 is used to model the equipment error simulation and parameter optimization process as a Markov decision process, and determine an error prediction action benefit calculation module and a parameter optimization benefit calculation module. The error prediction action benefit calculation module combines the absolute value of the error and the threshold constraint as a penalty term to evaluate the accuracy of the error prediction. The parameter optimization benefit calculation module combines the mean square error between the actual parameters and the ideal parameters to evaluate the actual application effect of the parameter optimization and neural network model.

[0181] A data acquisition unit 702 is used to acquire historical manufacturing error data and ideal manufacturing parameters of the equipment;

[0182] A prediction and optimization unit 703 is configured to input the historical manufacturing error data into a pre-trained equipment error prediction and parameter optimization model, wherein the equipment error prediction and parameter optimization model includes an error prediction network and a parameter optimization network; the error prediction network is configured to calculate and output a predicted error vector using the input historical manufacturing error data; and the parameter optimization network is configured to calculate and output optimized manufacturing parameters using the input error vector and the ideal manufacturing parameters.

[0183] The actual manufacturing parameter acquisition unit 704 is used to set the current manufacturing parameters of the manufacturing equipment to the optimized manufacturing parameters through the digital twin system and monitor and obtain the actual manufacturing parameters of the manufacturing equipment acquired by the sensor;

[0184] A benefit calculation unit 705 is configured to calculate an error prediction benefit and a parameter optimization benefit by using the error prediction action benefit calculation module and the parameter optimization benefit calculation module in combination with the actual manufacturing parameters;

[0185] The experience cache unit 706 is configured to aggregate the error prediction benefit and the historical manufacturing error data using an information fusion module to obtain updated historical error data; store the updated historical error data, the error vector, and samples from the error prediction network pre-training process into an error prediction experience cache; and store the parameter optimization benefit, the optimized manufacturing parameters, and samples from the parameter optimization network pre-training process into a parameter optimization experience cache;

[0186] An experience replay unit 707 is configured to determine whether the current error prediction network has performed a target number of error predictions, and then determine whether to use the data in the error prediction experience buffer to perform online updating of network parameters of the error prediction network using a dual deep reinforcement learning network;

[0187] The periodic retraining unit 708 is used to determine that the number of parameter optimizations currently performed by the parameter optimization network is greater than the target number, and then use the data model retraining tool in the parameter optimization experience buffer area to retrain and update the equipment error prediction and parameter optimization model.

[0188] The present application also provides an equipment parameter error simulation and optimization device, the device comprising a processor and a memory:

[0189] The memory is used to store program code and transmit the program code to the processor;

[0190] The processor is used to execute the steps of the above-mentioned equipment parameter error simulation and optimization method according to the instructions in the program code.

[0191] like Figure 8 As shown, an equipment parameter error simulation and optimization device provided by an embodiment of the present application may include: a processor 10, a memory 11, a communication interface 12, and a communication bus 13. The processor 10, the memory 11, and the communication interface 12 all communicate with each other via the communication bus 13.

[0192] In the embodiment of the present application, the processor 10 can be a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit, a digital signal processor, a field programmable gate array or other programmable logic devices, etc.

[0193] The processor 10 may call a program stored in the memory 11 . Specifically, the processor 10 may execute operations in an embodiment of the equipment parameter error simulation and optimization method.

[0194] The memory 11 is used to store one or more programs. The program may include program code, and the program code includes computer operating instructions. In the embodiment of the present application, the memory 11 stores at least a program for implementing the following functions:

[0195] The equipment error simulation and parameter optimization process is modeled as a Markov decision process, and an error prediction action benefit calculation module and a parameter optimization benefit calculation module are determined. The error prediction action benefit calculation module combines the absolute value of the error and the threshold constraint as a penalty term to evaluate the accuracy of the error prediction. The parameter optimization benefit calculation module combines the mean square error between the actual parameters and the ideal parameters to evaluate the actual application effect of the parameter optimization and neural network model.

[0196] Obtain historical manufacturing error data and ideal manufacturing parameters of equipment;

[0197] Inputting the historical manufacturing error data into a pre-trained equipment error prediction and parameter optimization model, the equipment error prediction and parameter optimization model comprising an error prediction network and a parameter optimization network; the error prediction network is configured to calculate and output a predicted error vector using the input historical manufacturing error data; and the parameter optimization network is configured to calculate and output optimized manufacturing parameters using the input error vector and the ideal manufacturing parameters;

[0198] Setting the current manufacturing parameters of the manufacturing equipment to the optimized manufacturing parameters through the digital twin system and monitoring the actual manufacturing parameters of the manufacturing equipment collected by the sensor;

[0199] Utilizing the error prediction action benefit calculation module and the parameter optimization benefit calculation module in combination with the actual manufacturing parameters to calculate and obtain the error prediction benefit and the parameter optimization benefit;

[0200] The error prediction benefit and the historical manufacturing error data are aggregated using an information fusion module to obtain updated historical error data; the updated historical error data, the error vector, and samples from the error prediction network pre-training process are stored in an error prediction experience buffer; the parameter optimization benefit, the optimized manufacturing parameters, and samples from the parameter optimization network pre-training process are stored in a parameter optimization experience buffer;

[0201] Determining that the current error prediction network has performed a target number of error predictions, then determining to use the data in the error prediction experience buffer to online update the network parameters of the error prediction network using a dual deep reinforcement learning network;

[0202] If it is determined that the number of parameter optimizations currently performed by the parameter optimization network is greater than the target number, the equipment error prediction and parameter optimization model is retrained and updated using the data model retraining tool in the parameter optimization experience buffer.

[0203] In one possible implementation, the memory 11 may include a program storage area and a data storage area, wherein the program storage area can store an operating system and application programs required for at least one function (such as a file creation function, a data reading and writing function), etc.; the data storage area can store data created during use, such as initialization data, etc.

[0204] In addition, the memory 11 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device or other volatile solid-state storage device.

[0205] The communication interface 12 may be an interface of a communication model, used for connecting to other devices or systems.

[0206] Of course, it needs to be explained that Figure 8 The structure shown does not constitute a limitation on the equipment parameter error simulation and optimization device in the embodiment of the present application. In actual applications, the equipment parameter error simulation and optimization device may include Figure 8 More or fewer components than shown, or combinations of certain components.

[0207] An embodiment of the present application may also provide a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to execute the steps of the above-mentioned equipment parameter error simulation and optimization method.

[0208] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0209] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or certain parts of the embodiments of the present application.

[0210] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0211] The above description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention are included in the scope of protection of the present invention.

Claims

1. A method for simulating and optimizing equipment parameter errors, characterized in that: include: The equipment error simulation and parameter optimization process is modeled as a Markov decision process, and an error prediction action benefit calculation module and a parameter optimization benefit calculation module are determined. The error prediction action benefit calculation module combines the absolute value of the error and the threshold constraint as a penalty term to evaluate the accuracy of the error prediction. The parameter optimization benefit calculation module combines the mean square error between the actual parameters and the ideal parameters to evaluate the actual application effect of the parameter optimization and neural network model. Obtain historical manufacturing error data and ideal manufacturing parameters of equipment; Inputting the historical manufacturing error data into a pre-trained equipment error prediction and parameter optimization model, the equipment error prediction and parameter optimization model comprising an error prediction network and a parameter optimization network; the error prediction network is configured to calculate and output a predicted error vector using the input historical manufacturing error data; The parameter optimization network is used to calculate and output optimized manufacturing parameters using the input error vector and the ideal manufacturing parameters; Setting the current manufacturing parameters of the manufacturing equipment to the optimized manufacturing parameters through the digital twin system and monitoring the actual manufacturing parameters of the manufacturing equipment collected by the sensor; Utilizing the error prediction action benefit calculation module and the parameter optimization benefit calculation module in combination with the actual manufacturing parameters to calculate and obtain the error prediction benefit and the parameter optimization benefit; Aggregating the error prediction benefit and the historical manufacturing error data using an information fusion module to obtain updated historical error data; storing the updated historical error data, the error vector, and samples from the error prediction network pre-training process into an error prediction experience buffer; Storing the parameter optimization benefit, the optimized manufacturing parameters, and samples of the parameter optimization network pre-training process in a parameter optimization experience buffer; Determining that the current error prediction network has performed a target number of error predictions, then determining to use the data in the error prediction experience buffer to online update the network parameters of the error prediction network using a dual deep reinforcement learning network; If it is determined that the number of parameter optimizations currently performed by the parameter optimization network is greater than the target number, the equipment error prediction and parameter optimization model is retrained and updated using the data model retraining tool in the parameter optimization experience buffer.

2. The equipment parameter error simulation and optimization method according to claim 1, characterized in that: The error prediction action benefit calculation module is represented by the following formula: Where: r t,1 represents the error prediction return, represents the error of the nth manufacturing parameter in the tth manufacturing, represents the error prediction result of the nth manufacturing parameter in the tth manufacturing, α1 and β1 represent the prediction error penalty factors, The threshold value indicating the tolerable deviation between the prediction error and the actual error; The parameter optimization benefit calculation module is represented by the following formula: Where: r t,2 represents the error prediction return, represents the nth actual manufacturing parameter actually monitored in the tth manufacturing, represents the nth ideal manufacturing parameter given in the tth manufacturing.

3. The equipment parameter error simulation and optimization method according to claim 1, characterized in that: The error prediction network includes two long short-term memory network layers and two fully connected layers; the long short-term memory network layers are used to capture the time dependency of manufacturing errors, and the fully connected layers are used to fuse the coupled error relationships of multiple parameters.

4. The equipment parameter error simulation and optimization method according to claim 1, characterized in that: The parameter optimization network includes an input layer, a hidden layer, and an output layer. The number of neurons in the input layer is equal to the dimension of the manufacturing parameters. The hidden layer includes multiple fully connected layers and a Dropout layer. The fully connected layer uses ReLU=max(·,0) as the activation function, and the Dropout layer is used to prevent overfitting. The number of neurons in the output layer is equal to the dimension of the optimal setting parameters, and a linear activation function is used to obtain the regression analysis results of the optimized parameters.

5. The equipment parameter error simulation and optimization method according to claim 1, characterized in that: The equipment error prediction and parameter optimization model adopts a staged training process of first pre-training the hyperparameters of the parameter optimization network based on random search, and then pre-training the error prediction network using a dual deep reinforcement learning method.

6. The equipment parameter error simulation and optimization method according to claim 5, characterized in that: The pre-training method of the parameter optimization network includes: Step 1: Define the number of hidden layers, number of neurons, Dropout ratio, learning rate, and optimizer value range; Step 2: Set the number of random search iterations, the value range of which is 50-100 times; Step 3: Randomly sample hyperparameter combinations. In each iteration, randomly sample a set of hyperparameter values from the hyperparameter space. Step 4: Train the model using the sampled hyperparameter combination and evaluate the loss, accuracy, and F1 score on the validation set. Step 5: After all iterations are completed, select the hyperparameter combination with the best performance on the validation set as the final result.

7. The equipment parameter error simulation and optimization method according to claim 6, characterized in that: The pre-training method of the error prediction network includes: Step 1: Randomly sample x training Mini-batchE from the training set of sample set ε1 x sample; Step 2: From Mini-batchE x Select sample e t =(s t ,a t ,r t,1 ,s t+1 ); Step 3: Use the Q learning policy network to calculate the Q value Q(s) of the current state t ,a t :θ), and the next state s t+1 All actions in a t+1 The Q value Q(s t+1 ,a t+1 :θ), select the next action with the maximum Q value Step 4: Calculate the next state s t+1 The target network Q value Step 5: Calculate learning objectives γ is the discount factor; Step 6: Calculate the loss function L(θ)=(yQ(s t ,a t :θ)) 2 , obtain the gradient through back propagation of the policy network Step 7: Update the policy network parameters using gradient descent Where ρ is the learning rate; Step 8: Update the target network parameters θ with weighted hyperparameters - ←τθ+(1-τ)θ - ; Step 9: Repeat steps 2 to 8 until all mini-batch E are traversed. x If all samples in the training set are included, then one round of training is completed and the average error prediction gain is calculated on the validation set; Step 10: Adjust the learning rate ρ and the number of batch training samples x, and return to step 1; Step 11: In all training rounds, the policy network parameters with the best performance are selected as the pre-trained error prediction network parameters.

8. An equipment parameter error simulation and optimization device, characterized in that: The device is used to execute the equipment parameter error simulation and optimization method according to any one of claims 1 to 7, comprising: A simulation and optimization modeling unit is used to model the equipment error simulation and parameter optimization process as a Markov decision process, and determine an error prediction action benefit calculation module and a parameter optimization benefit calculation module; the error prediction action benefit calculation module combines the absolute value of the error and the threshold constraint as a penalty term to evaluate the accuracy of the error prediction; the parameter optimization benefit calculation module combines the mean square error of the actual parameters and the ideal parameters to evaluate the actual application effect of the parameter optimization and neural network model; A data acquisition unit, used to obtain historical manufacturing error data and ideal manufacturing parameters of the equipment; a prediction and optimization unit, configured to input the historical manufacturing error data into a pre-trained equipment error prediction and parameter optimization model, wherein the equipment error prediction and parameter optimization model includes an error prediction network and a parameter optimization network; the error prediction network is configured to calculate and output a predicted error vector using the input historical manufacturing error data; and the parameter optimization network is configured to calculate and output optimized manufacturing parameters using the input error vector and the ideal manufacturing parameters; An actual manufacturing parameter acquisition unit, configured to set the current manufacturing parameters of the manufacturing equipment to the optimized manufacturing parameters through the digital twin system and to monitor and acquire the actual manufacturing parameters of the manufacturing equipment acquired by the sensor; A benefit calculation unit, configured to calculate an error prediction benefit and a parameter optimization benefit by using the error prediction action benefit calculation module and the parameter optimization benefit calculation module in combination with the actual manufacturing parameters; An experience cache unit is configured to aggregate the error prediction benefit and the historical manufacturing error data using an information fusion module to obtain updated historical error data; store the updated historical error data, the error vector, and samples from the error prediction network pre-training process into an error prediction experience cache; and store the parameter optimization benefit, the optimized manufacturing parameters, and samples from the parameter optimization network pre-training process into a parameter optimization experience cache; An experience replay unit, configured to determine whether the current error prediction network has performed a target number of error predictions, and then determine whether to use the data in the error prediction experience buffer to perform online updating of network parameters of the error prediction network using a dual deep reinforcement learning network; The periodic retraining unit is used to determine that the number of parameter optimizations currently performed by the parameter optimization network is greater than the target number, and then use the data model retraining tool in the parameter optimization experience buffer area to retrain and update the equipment error prediction and parameter optimization model.

9. An equipment parameter error simulation and optimization device, characterized in that: The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the equipment parameter error simulation and optimization method according to any one of claims 1 to 7 according to the instructions in the program code.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store program code, and the program code is used to execute the equipment parameter error simulation and optimization method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • A method for sensitivity analysis of cable net antenna manufacturing errors based on surrogate model

    CN114707116B